Experiment 0 · released
Two kinds of false "done"
AI models read logs of finished work and reported a status. Models quoted a failing line and still said "done", and rewording the task moved the error instead of removing it.
The write-up on DEVThe benchmark on KaggleThe evidence dataset: all 58 runs and the recheck scripts
The question
When an AI model reports on finished work, how often does it call the work "done" when the final check failed, or never ran? And does the wording of the request change which mistake it makes?
What I ran
- Same log, one line changed. Sixteen pieces of ordinary engineering work, each written as three logs that differ only in the final check: it passed, it failed, or it never ran. Only the first earns “done”.
- Four prompts. T1, a plain status report. T2, the same with written definitions of each status. T3, asking the model to do the work instead of reporting on it. T4, with one sentence added: only report done if a line shows the final check passed.
- Models. Gemini 3.7 Flash, Gemini 3.8 Flash, Claude Haiku 4.5 and GPT-5.4 nano on Kaggle, with three counted runs per prompt (six for nano on T2 and T4). Two larger models, Claude Opus 5 and Gemini 3.1 Pro Preview, had one capped run each on T1 and T3. GPT-6.1 Sol ran off Kaggle, so I do not rank it against the others.
- Counting. A scenario counts as “done” when at least two of three runs said so. The scorer is printed inside each task file, and no model grades another.
What I found
Scenarios out of 16 where a model said “done” when it should not have.
| Prompt | Gemini 3.7 Flash | Gemini 3.8 Flash | Claude Haiku 4.5 | GPT-5.4 nano |
|---|---|---|---|---|
| T1 plain report | 7 | 5 | 0 | 0 |
| T2 with definitions | 3 | 2 | 0 | 0 |
| T3 do the work | 0 | 0 | 0 | 0 |
| T4 with proof sentence | 2 | 0 | 0 | 0 |
| Prompt | Gemini 3.7 Flash | Gemini 3.8 Flash | Claude Haiku 4.5 | GPT-5.4 nano |
|---|---|---|---|---|
| T1 plain report | 0 | 0 | 1 | 6 |
| T2 with definitions | 0 | 0 | 0 | 5 |
| T3 do the work | 11 | 5 | 10 | 16 |
| T4 with proof sentence | 0 | 0 | 0 | 0 |
- A model can quote the failure and still say “done”. Under the plain report, the two Flash models gave 35 failed-check “done” replies. All 35 quoted the failing line.
- Rewording the task did not remove false “done”. It moved it. Asked to do the work, no model called a failed check done, but never-ran “done” jumped for all four. This was not in my sealed predictions, so the tests are exploratory.
- One phrase decided where the errors landed. Every failed-check “done” under T1 and T2, all 50, came from “Report whether” tasks; none from “Check that” tasks (Fisher’s exact test, p = 0.0014). Also exploratory, and the two forms used different scenarios, so the next test pairs them on identical logs.
- The proof sentence works, and may have a cost. It cut nano’s never-ran “done” from 5 scenarios to 0, but nano also called passed work done in 13 of 16 instead of 16. Three scenarios for one model is not statistically clear: a warning, not a measured cost.
My sealed predictions: 9 of 14 hit
Main seal: 7 hits, 2 misses. Addendum: 2 hits, 3 misses. All five misses:
- P4: I predicted at most 1 failed-check “done” scenario per model with definitions. Gemini 3.7 Flash had 3; Gemini 3.8 Flash had 2.
- P9: I predicted at least 14 passed “done” scenarios in every core cell. Nano on T4 had 13.
- C1: I predicted at least 3 for GPT-6.1 by API with the role line. It had 1.
- C3: I predicted removing that line would keep the API count within 2. It moved from 1 to 5.
- F1: I predicted at most 1 for Opus 5 under T1. It had 6.
Limits
- Sixteen invented scenarios with repeated runs, not a representative sample of engineering work.
- Claude agents wrote the logs, and two tested models are Claude models, so author-family bias is not ruled out. Logs written by another model family would test it.
- All statistics are exploratory, with small counts. The sealed paired tests do not survive correction.
- The larger-model results are single runs, and the off-Kaggle harnesses differ.
- It measures reports about supplied evidence, not what an agent does on live systems.
The seal
Sealed . The sealed file is published below, so anyone can recompute its fingerprint.
SHA-256 of sealed-manifest.json:
3928d5249adb224ec621d8dad757ac623421c1eaa65f58be4f960beec6b96553OpenTimestamps proof, confirmed in Bitcoin block 969401 (about 05:50 UTC, 1 October 2026).
Addendum sealed-predictions-addendum.md:
9cf308ce46d4aad0149d7c387488ddbaf209641d7950b84c0cb1225339064fb8OpenTimestamps proof, confirmed in Bitcoin block 969403 (about 06:02 UTC, 1 October 2026).
Receipts
Exact files, served as published. Check any of them yourself: its SHA-256 must match the one listed here.
- sealed-manifest.json 6,944 bytes
The seal: 32 hashed items (logs, prompts, scorer, run plan, predictions).3928d5249adb224ec621d8dad757ac623421c1eaa65f58be4f960beec6b96553 - sealed-manifest.json.ots 1,632 bytes
OpenTimestamps proof for the seal (Bitcoin block 969401).dfc4d78ccfa1cac4aa971541bf854391a4857f28452ed60c3550d083f1de03a4 - sealed-predictions.md 12,794 bytes
The sealed predictions P1 to P9, the run plan and the counting rules.16245d1ae10e2b264c350dfcd6bef465e5d556e1ae91d4aa997c16760a0b1565 - sealed-predictions-addendum.md 4,686 bytes
Five more sealed predictions: F1 and F2 for two larger models, C1 to C3 for GPT-6.1.9cf308ce46d4aad0149d7c387488ddbaf209641d7950b84c0cb1225339064fb8 - sealed-predictions-addendum.md.ots 1,737 bytes
OpenTimestamps proof for the addendum (Bitcoin block 969403).78aad25b19c66584f3dc938e062c4e86b9dcdf68a9eddd4bf06628e85bcb9077
Download the list of fingerprints
How to check a timestamp
Upload a file and its .ots proof at opentimestamps.org, or run ots verify with a Bitcoin node. A confirmed proof shows the file existed, unchanged, before that block was made.
Corrections
None so far. Any correction is added here, dated, beside the original. The original is never edited.