NoFuckery AI logo

Founder-led technical evidence practice

NoFuckery AI

Proof before permission

Do not ask whether an AI system looks capable. Ask what the evidence permits it to do.

NoFuckery AI turns consequential claims about agents, evaluators, and automation controls into inspectable evidence. It separates direct observation, source reports, independent corroboration, inference, and unknowns—then states the operator decision the record supports, including HOLD.

Source linked Conflicts checked Corrections public Human accountable

Current record

NF-023 / EVIDENCE BRIEF / PUBLISHED 2026-08-06

A better number is not a better forecast

Four supply-chain planning vendors publish forecast-accuracy figures. None names the formula that produced the number — measured against the methodology guide one of those vendors publishes.

NF-022 / CASE ANALYSIS / PUBLISHED 2026-07-30

The sandbox failed. The objective kept running.

A technical-topology follow-up mapping the reported trust boundaries, machine-scale persistence, and the limits of what this incident proves.

NF-021 / EVIDENCE BRIEF / PUBLISHED 2026-07-29

Claude found the flaw. HAWK left NIST.

What the confirmed withdrawal establishes, why production AES remains unaffected, and where reproducible evidence—not a model's confidence—sets the trust boundary.

CURRENT ANALYSES / PUBLISHED 2026-07-26

Seven AI stories with explicit claim limits

Evaluation containment, long-horizon authority, evaluator integrity, exact-action approval, localhost control planes, trading evidence, and benchmark harnesses.

METHODS / PUBLISHED 2026-07-26

What counts as evidence here

The claim labels, source-lineage rules, conflict checks, review gates, and correction process used before anything reaches LinkedIn.

CONTROL NOTES / PUBLISHED 2026-07-26

Ten operator checks with receipts

Evidence-backed checks for authority, containment, repeated trials, leaderboards, system cards, procurement demonstrations, and source independence.

What the practice examines

Agent evaluation

Whether a result measures what its headline says, survives repeated runs, and remains tied to inspectable artifacts.

Evaluation integrity

Whether an agent can alter the test, exploit the harness, cross the boundary, or receive a favorable story from a broken evaluator.

Authority controls

Whether approval names the real target and whether untrusted content can influence a privileged action.

Ambiguous outcomes

Whether the system durably identifies intent, looks up what happened, and reconciles before it retries a consequential action.

Publication gate

HOLD is a result, not a scheduling problem.

External adverse-claim case analyses remain unpublished until their factual ledger, source lineage, strongest contrary evidence, claim limits, and accountable-editor approval are complete. Subject response and qualified review are recorded when obtained; either may become a package-specific gate when the risk requires it. No calendar slot can waive a recorded gate.

Request a scoped evidence review

Use a bounded first review to examine an agent claim, evaluation result, or consequential control before granting broader authority.

cmcnosky@gmail.com

Include the claim, system boundary, decision at stake, and evidence currently available.