Entries

Challenge entries

Three entries to public challenges in three days, 2 to 4 October 2026. Each is judged by its organizers; results go here when they are announced.

It Quoted the Failure: Two Kinds of False "Done"

Benchmark · Kaggle Benchmarking Challenge entry · 2 October 2026 · Entered

An AI wrote that the service had failed, then marked the job done. This benchmark measures how often models report "done" when the final check failed or never ran: 48 logs from 16 scenarios, with predictions sealed before the first run.

What I found. On a plain status prompt, two of the four models said "done" in 35 of their 96 replies on failed-check logs, and all 35 quoted the failing line. The 48 logs are a known set, not a blind holdout.

Results: not announced yet.

Receipt Desk: show me the line, or it isn't done

Tool · Sanity Challenge entry · 3 October 2026 · Published

A status desk for engineering logs, built on Sanity. Asked whether a job is done, it answers done, failed, or not shown, and each answer cites the command step and the exact output line that support it, with a link back to the original record.

What I found. GPT-6.1 tied: 48 of 48 correct statuses both from raw logs and through Sanity Context. This is a known development set of 48 logs, not a blind holdout.

Results: not announced yet.

Swarm receipts: plant a fault before you trust the monitor

AI Swarm Dynamics Hackathon (AI Village × Grove Research) · 3 to 4 October 2026 · Submitted

A test bench for swarm oversight tools, built on the real AI Village record, and the two baseline checkers it caught. Every exam is sealed before the checker sees it.

What I found. The checker, with a local model reading the records, passed a fresh sealed exam: 141 of 150 (94%), exactly at the pass bar, with no planted failure called a success. A narrow pass: the second model tried, and the fail-closed rule decided one case.

Results: not announced yet.