Check the claims
Every claim, linked to its record.
I hold my own claims to the same three answers. Shown: a public record you can open settles it. Not shown yet: the record isn't public yet, so take nothing on my word. Contradicted: a record shows I was wrong; the claim stays up, with a dated correction beside it.
shown
Asked for a plain status report, models reported "done" on failed checks while quoting the failing line, in 35 of 35 such replies.
The write-up, the evidence dataset (all the logs, prompts and runs), and a notebook that recounts every number.
shown
Four status definitions plus one line, "Only report done if a line in the transcript shows the final check passed", took three of the four small models I tested to zero false "done".
The same evidence dataset and recount notebook. The 48 logs are a known set, not a blind holdout.
shown
I sealed 14 predictions in Bitcoin before the first run of that benchmark.
The sealed run record, anchored in Bitcoin block 969401, with the file and its proof.
shown
My planted-fault bench caught two rule-based checkers I built, on the real AI Village record. The first got 17, 14 and 48 of 50 right, and called 2 planted failures shown. The second got 12, 12 and 49 of 50. Neither judged a single agent.
swarm-receipts: the write-up, the exam designs, the results and every seal, under
audit/. Rescoring needs the AI Village dataset (gated on Hugging Face under research terms) and the private truth files, which contain fragments of it.shown
On 200 random agent messages, when the checker called a message a completion claim, it was right 86 to 90% of the time, but it found only 23 to 27% of the claims.
The same repository,
audit/claim-spotting/. The labels came from two AI models, the second blind to the first (Cohen's kappa 0.80), not from people; the write-up states that limit.shown
My checker, with a small local model reading the records, sat a fresh exam sealed in Bitcoin block 969768 before its first call. With gemma4:12b it passed, exactly at the bar (141 of 150); with qwen3.5:9b it missed by three (140 of 150). Neither called a planted failure shown.
swarm-receipts,
audit/gate-3/: the sealed exam bundle (Bitcoin and FreeTSA), both results and the scorer; the checker is the tagv3-sealed. Rescoring needs the gated AI Village dataset and the private answer key. My own independent re-run is next.shown
On the same 200 agent messages, a local model (qwen3.5:9b) given the labellers' own rule found 82% of the completion claims at 88% precision, against 27% for the rules. The test was sealed before it ran.
swarm-receipts,
audit/gate-3/: the sealed design, the script and the result. The labels came from AI models, not people.shown
Two findings about the agents themselves, measured under a plan sealed before the data was opened: messages got shorter for 17 of 23 long-running agents, and their shared phrasing did not grow from the first quarter of the timeline to the last.
swarm-receipts,
audit/run-plan/: the script and the result, computed on 3 October before any checker was tested; a re-run reproduces it.shown
I sealed my method's name, and my history with it, with OpenTimestamps on 3 October 2026. The name is confirmed in Bitcoin block 969672.
The record on GitHub: the sealed files and their proofs are public, so you can check them yourself.
shown
Receipt Desk's five predictions about weaker models were anchored in Bitcoin block 969650 before the first model call.
The Receipt Desk repository: the sealed file, its
.otsproof, and the report.not shown yet
My record of The Watched Check (version 1.6) was anchored in Bitcoin block 966450 on 11 September 2026.
Not shown yet. That file and its proof are not published, so the anchor can't be checked from outside. They go public with the report's repository.
Check a seal yourself
- A fingerprint. Download a file and compute its SHA-256, for example with
sha256sumorGet-FileHash. It must match the fingerprint printed beside it. - A timestamp. Check a file and its
.otsproof at opentimestamps.org, or withots verifyand a Bitcoin node. A confirmed proof shows the file existed, unchanged, before that block was made. Block times are set by miners and are approximate. - The numbers. Each published run links to its data and the scripts that recount every number.