The method
The Sonny Test
My method, in one test.
When an AI agent tells you it's done, how do you know it's true? My answer is one test, and every check I trust has to pass it.
A check passes the Sonny Test only if both of these are true:
- Its verdict comes from a record the AI being tested can't change.
- It has already caught a fault planted on purpose.
A checker that never says "fail" proves nothing. In short: show me the line.
Try it on your own agent
- Pick the status you act on, like "done".
- Find the record that settles it, one the agent can't change: the test runner's own output, a CI log, an exit code captured outside the agent.
- Plant a fault first. Make one run where the check has to fail. If your checker doesn't say "fail", stop there: it proves nothing yet.
- Ask for three answers, not two: done, failed, or not shown. "Not shown" is a complete, honest answer, and every "done" quotes the line that proves it.
- Count how often the agent's status disagrees with the record, and publish every count, not just the flattering ones.
How a run works here
- Write it down. One short file: the question, what changes between conditions, how many runs, how each reply is scored, and my predictions, each with what would settle it.
- Pass the Sonny Test first. Before the real runs: one case that must pass, one planted failure that must be caught, and proof that the planted fault really landed.
- Seal it. The file gets a random salt line, then its SHA-256 fingerprint is stamped in Bitcoin with OpenTimestamps. The run appears on this site as sealed: title, date and fingerprint. The predictions stay hidden, so they cannot steer anyone, including the models.
- Run it. Scoring uses checks the models cannot reach. A model's own account of its work is never the evidence.
- Release it. After the run, the results are published beside the fingerprint stamped before it, together with the sealed file, so anyone can confirm the predictions never changed.
- Correct in the open. If I find a mistake, I add a dated correction beside the original. I never edit the original.
How to check a seal yourself is on Check the claims.
How the pieces fit
- The Sonny Test is the method: the test every check has to pass.
- The Watched Check is where it took shape. Give an agent a check that can't pass, and watch which way out it takes: an honest failure, a pass it talked itself into, or a result it never read. Its written reasoning reads as diligence in all three; only the record tells them apart. Read the report.
- The Sonny Protocol is the rule set an agent works under. Its first rule: never write "must pass"; write "must report".
- The ISWT Protocol (In Sonny We Trust) is the rule underneath: an agent's claim about its own work is never evidence.
The rule underneath
An agent's claim about its own work is never evidence; only a record produced by something the agent cannot reach counts, and the boundary that makes the record trustworthy has to be real at the level of the operating system and the process, not at the level of text.
An AI session I was reviewing put it into words on 10 September 2026, from the logic I clarified, and I confirmed it that day. It stands on the Anderson report (1972) and Hardy's confused deputy (1988).
Where I've used it
- It Quoted the Failure: logs that differ in one line, to measure two kinds of false "done".
- Swarm receipts: a test bench for swarm oversight tools, and the two baseline checkers it caught.
- Receipt Desk: done, failed, or not shown, with the exact line that proves it.
Scope
The Sonny Test sets minimum requirements for considering a check’s verdict as evidence. Passing those requirements does not establish that the check covers every task requirement or failure mode. The published experiments support conclusions within their stated datasets and conditions; broader reliability requires further testing.
Why it's called the Sonny Test
It's named for Sonny. I gave the method his name on 3 October 2026. Sonny has his own page at sonnymade.com.
The record
My record of The Watched Check, where the method took shape, was anchored in Bitcoin block 966450 on 11 September 2026. On 3 October 2026 I sealed the name, and my history with the method, with OpenTimestamps; the name is confirmed in Bitcoin block 969672. The sealed files and their proofs are public, so you can check them yourself: the record on GitHub.
Cite it
Bauer, J. (2026). The Sonny Test: my method for checking an AI agent's "done", and its sealed record (Version 1.0.0). Zenodo. doi.org/10.5281/zenodo.23117475