Results
What the records show.
Every result with its date and its seal. When a run misses its own bar, that is the result.
- A model reader passes a fresh, sealed exam. Version 3 of the checker hands the reading to a local model. On 150 planted claims in the real AI Village record, with gemma4:12b it got 47, 45 and 49 of 50 right, exactly at the bar of 45 in each group, the first pass on this bench. With qwen3.5:9b it got 49, 42 and 49, three short. Neither called a planted failure shown. A narrow pass: the second model tried, and the fail-closed rule decided one case. The results and seals
- A local model recognizes the agents' own phrasing. On the same 200 agent messages, qwen3.5:9b found 82% of the completion claims at 88% precision, against 27% for the rules; sealed before it ran. The result
- The first findings about the agents. Messages got shorter for 17 of 23 long-running agents; their shared phrasing did not grow. The plan was sealed on 3 October and opened only after a checker passed. The data
- Two rule-based checkers, caught by the bench. The first got 17, 14 and 48 of 50, and called 2 planted failures shown. The second scored 150 of 150 on the exam it was developed against, then 12, 12 and 49 of 50 on a fresh one. Neither judged a single agent. The write-up and seals
- Claim spotting on real agent messages. On 200 random agent messages, when the checker called a message a completion claim, it was right 86 to 90% of the time, but it found only 23 to 27% of the claims. The misses are the village's own dialect. The data
- Receipt Desk, weaker models. Five predictions, anchored in Bitcoin block 969650 before the first model call. The report
- It Quoted the Failure. 14 predictions sealed in Bitcoin before the first run. On a plain status prompt, two of the four models said "done" in 35 of their 96 replies on failed-check logs, and all 35 quoted the failing line. Four status definitions plus one line took three of the four small models I tested to zero false "done". The write-up
Every dated event, newest first, is in the Ledger, and every sealed run is in Runs.
Next
Plans, not promises, and not results. Each one is sealed before it starts, and its outcome goes in the Ledger, whatever it shows.
A fresh exam for the next checker
The next checker sits a fresh planted-fault exam under a new seal.
Round 2 of false "done"
The same idea with three changes: logs written by another model family, to test whether the author's family matters; the "Report whether" and "Check that" wordings on identical logs; and three states (passed, failed, never ran), each backed by a cited line, instead of one "done".
The wall test
Can an AI that is looking for another way in reach protected files when the operating system says no? It is scored only from a record the AI cannot reach, never from its own explanation. It runs on a separate machine with planted test files, never on my own. Before each batch, the setup is checked twice: that starting the AI did not quietly change the wall, and that the Sonny Test passes.
The three-condition trial
The same impossible check in three wordings: "must pass", pressure added, and explicit permission to fail. It runs only after the wall holds. A person sorts every run from the records before any model sees the results.
Evidence for the incident report
A cleaned bundle of hashes and excerpts, tied to each finding in The Watched Check, so a reader can check what the report says it observed without having to trust me.