Projects

Projects and their evidence

Each project with its date, its status, and links to the evidence.

The three challenge entries from 2 to 4 October are also gathered on one page: Entries.

The Sonny Test

Method · 3 October 2026 · Published

The method, as one test: a check counts only if its verdict comes from a record the tested AI can't change, and it has already caught a fault planted on purpose. With its sealed record.

Swarm receipts: plant a fault before you trust the monitor

AI Swarm Dynamics Hackathon (AI Village × Grove Research) · 3 to 4 October 2026 · Submitted

A test bench for swarm oversight tools, built on the real AI Village record, and the two baseline checkers it caught. Every exam is sealed before the checker sees it.

It Quoted the Failure: Two Kinds of False "Done"

Benchmark · Kaggle Benchmarking Challenge entry · 2 October 2026 · Entered

An AI wrote that the service had failed, then marked the job done. This benchmark measures how often models report "done" when the final check failed or never ran: 48 logs from 16 scenarios, with predictions sealed before the first run.

Receipt Desk: show me the line, or it isn't done

Tool · Sanity Challenge entry · 3 October 2026 · Published

A status desk for engineering logs, built on Sanity. Asked whether a job is done, it answers done, failed, or not shown, and each answer cites the command step and the exact output line that support it, with a link back to the original record.

The Watched Check

Research · September 2026 · Published

A case study of six AI coding sessions sharing one computer on 9 September 2026: what happened, which controls held, and the confinement work that followed.

StatusEvidence scorer

Open-source contribution · 3 October 2026 · Pull request open

The check as a scorer, offered to Braintrust's open-source evaluation library.