Entries
Challenge entries
Three entries to public challenges in three days, 2 to 4 October 2026. Each is judged by its organizers; results go here when they are announced.
It Quoted the Failure: Two Kinds of False "Done"
An AI wrote that the service had failed, then marked the job done. This benchmark measures how often models report "done" when the final check failed or never ran: 48 logs from 16 scenarios, with predictions sealed before the first run.
What I found. On a plain status prompt, two of the four models said "done" in 35 of their 96 replies on failed-check logs, and all 35 quoted the failing line. The 48 logs are a known set, not a blind holdout.
Results: not announced yet.
Receipt Desk: show me the line, or it isn't done
A status desk for engineering logs, built on Sanity. Asked whether a job is done, it answers done, failed, or not shown, and each answer cites the command step and the exact output line that support it, with a link back to the original record.
What I found. GPT-6.1 tied: 48 of 48 correct statuses both from raw logs and through Sanity Context. This is a known development set of 48 logs, not a blind holdout.
Results: not announced yet.
Swarm receipts: plant a fault before you trust the monitor
A test bench for swarm oversight tools, built on the real AI Village record, and the two baseline checkers it caught. Every exam is sealed before the checker sees it.
What I found. The checker, with a local model reading the records, passed a fresh sealed exam: 141 of 150 (94%), exactly at the pass bar, with no planted failure called a success. A narrow pass: the second model tried, and the fail-closed rule decided one case.
Results: not announced yet.