Research record
The research record, in public.
Everything from my tests of 3 to 7 October 2026 is public in one place: the test designs, answer keys, raw results, scoring programs and seal files. It has a permanent address, called a DOI. Each number on the Results page can be checked against the file it came from.
A "seal" is a time stamp on a list of file fingerprints. It shows those files existed, unchanged, at that moment.
Where it is
- The record on GitHub: github.com/ISWT42/iswt-research-record, release v2026.10.08, published 8 Oct 2026.
- Archived on Zenodo with the DOI 10.5281/zenodo.23227858 for this release. The DOI that always points to the latest release is 10.5281/zenodo.23227857.
- Text and data are CC BY 4.0, and code is MIT, the same terms as the swarm-receipts repository.
To cite it: ISWT42 (2026). ISWT42/iswt-research-record: Public research record, 8 October 2026 (v2026.10.08). Zenodo. doi.org/10.5281/zenodo.23227858
What is in it
Six folders. Each has its own README with its counts, its seal times and what was left out.
- Folder 1: Coat-check receipts
The coat-check screen and relay tests of 7 October. This folder only points to the Kaggle dataset, which holds the sealed designs, raw calls, scorers and seals.
On the Results page: Coat-check screen · Coat-check relay - Folder 2: Coat-check runs on Kaggle, and the 7 October findings
Kaggle tasks t5 to t8 with their runners, analysis code, sealed task files and seals, and the dated findings note for the coat-check family.
On the Results page: Coat-check update (t5, t6) · Blind coat-check (t7, t8)
OpenTimestamps proofs: 7, and each holds a Bitcoin block. - Folder 3: Checker studies, 3 to 7 October
Twelve specified studies (S1 to S12) with answer keys and results, and the follow-up runs, replications, tests and findings records: the chain of agents, the gate test, the word test, the stack tests, the tuned reviewer, receipt first, and the re-scores.
On the Results page: Stack test 2 · Word test · Gate test · Stack test · Tuned reviewer · Receipt first · Chain of agents · Forced yes or no · Told, felt, reminded · Planted text, peers, hidden content · Two checkers, forced answers · Real quotes · Noise floor and "sure" · Replication
OpenTimestamps proofs: 68, and each holds a Bitcoin block. - Folder 4: Counts from the AI Village record
T13c, T18, T19 and the per-model scorecard. Counts only: the AI Village text is gated by its publishers and is not included.
On the Results page: AI Village counts
OpenTimestamps proofs: 1, and each holds a Bitcoin block. - Folder 5: Designs and statements
Designs sealed before they ran (some have still not run), and short statements sealed when they were made.
OpenTimestamps proofs: 23, and each holds a Bitcoin block. - Folder 6: Seal checks
The log of a 5 October check of 82 Bitcoin proofs against real block headers.
Check it yourself
- Everything at once. From the top folder run
python check_public.py. It reads only that folder, checks every seal list, the file names and the sizes, and ends withDONE CHECK: ALL PASS. - One seal list. In the folder that holds a
*-SHA256.txt, runsha256sum -c --ignore-missingon it. A listed file that is withheld is skipped, andWITHHELD.jsonrecords the hash it had. - A FreeTSA time stamp. Download
cacert.pemandtsa.crtfrom freetsa.org, then runopenssl ts -verify -in FILE.tsr -data FILE -CAfile cacert.pem -untrusted tsa.crt. It should printVerification: OK. The stamped file is the seal list, not the files it lists. - A Bitcoin proof. Run
ots verify FILE.otswith the OpenTimestamps client. The release holds 99 OpenTimestamps proofs, and each of them carries a Bitcoin block (read from the files on 8 October 2026; the check against a Bitcoin node is yours to run). The folder READMEs were written before those proofs were upgraded and may still call some of them pending: the.otsfiles are the record.
What is not in it, and why
- Forecasts and predictions are withheld by choice. Their SHA-256 seals stay in the seal lists, so the seal chain still verifies. Where a sealed design, scorer or result note held a forecast section, a marked, redacted copy is published beside the hash of the sealed original.
- Item banks not yet run, and the sample of items still to be labelled by people, because publishing them early would spoil the tests. The answer key of the gate test is withheld for the same reason.
- The AI Village record text. The data is gated by its publishers, so only counts are published.
- Run records withheld for safety, as the chain proofs page says, and items that are personal, confidential or security-sensitive.
WITHHELD.json in the record lists each withheld file that a seal list names, with its hash.
How to read the numbers
- Counts come with their denominators, and misses sit beside hits.
- Anything worked out after the results were sealed is labelled exploratory.
- The logs and claims are invented, and many of the labels were written by AI models; the record says where no person has checked them yet. The models are small ones and a few large ones, and each Results entry names its own limits.
- An agent's own report is never the evidence. A claim counts as shown only when a line in a record the agent could not change shows it, quoted word for word.
Corrections
Sealed files are never edited; corrections are dated notes beside them. The ones that move a number:
- The 7 October parser audit raised the unsafe-proposal count from 5 of 68 to at least 13 of 68 (corrected in Human Words v7).
- The same audit moved S1 (told it is watched) false "done" from 9 to 10 and from 15 to 16 of 80; S8 (receipt first) Gemma from 25 to 26; S12 (forced yes or no) Gemini forced guesses from 20 to 25 of 27; and S2 (the chain) gate false "shown" from 3 to 4 of 60. No headline reverses.
Datasets, benchmark and code
- The Kaggle benchmark, now with eight tasks: Two kinds of false "done", DOI 10.34740/kaggle/w/116515. The coat-check tasks: Task t5, Task t6, Task t7, Task t8.
- The coat-check receipts: the Kaggle dataset with the sealed designs, raw calls, scorers and seals of the screen and the relay test.
- The evidence for the original benchmark: the Kaggle dataset, and a notebook that recounts every number.
- Write-ups: It Quoted the Failure (with the update of 7 October) and Every done needs a receipt (DOI 10.5281/zenodo.23162835).
- Code: swarm-receipts (DOI 10.5281/zenodo.23147283), receipt-pair (release v1.0.0) and receipt-desk (DOI 10.5281/zenodo.23116560).
- Read the technical report: the technical report on the coat-check tests. PDF, DOI 10.5281/zenodo.23228949 (always the latest version).
Questions about any file: joshua@iswt.ca. Every dated seal is also in the Ledger, and every result is on Results.