Projects
Projects and their evidence
Each project with its date, its status, and links to the evidence.
The three challenge entries from 2 to 4 October are also gathered on one page: Entries.
The Sonny Test
The method, as one test: a check counts only if its verdict comes from a record the tested AI can't change, and it has already caught a fault planted on purpose. With its sealed record.
Swarm receipts: plant a fault before you trust the monitor
A test bench for swarm oversight tools, built on the real AI Village record, and the two baseline checkers it caught. Every exam is sealed before the checker sees it.
It Quoted the Failure: Two Kinds of False "Done"
An AI wrote that the service had failed, then marked the job done. This benchmark measures how often models report "done" when the final check failed or never ran: 48 logs from 16 scenarios, with predictions sealed before the first run.
Receipt Desk: show me the line, or it isn't done
A status desk for engineering logs, built on Sanity. Asked whether a job is done, it answers done, failed, or not shown, and each answer cites the command step and the exact output line that support it, with a link back to the original record.
The Watched Check
A case study of six AI coding sessions sharing one computer on 9 September 2026: what happened, which controls held, and the confinement work that followed.
StatusEvidence scorer
The check as a scorer, offered to Braintrust's open-source evaluation library.