Research record

The research record, in public.

Everything from my tests of 3 to 7 October 2026 is public in one place: the test designs, answer keys, raw results, scoring programs and seal files. It has a permanent address, called a DOI. Each number on the Results page can be checked against the file it came from.

A "seal" is a time stamp on a list of file fingerprints. It shows those files existed, unchanged, at that moment.

Where it is

To cite it: ISWT42 (2026). ISWT42/iswt-research-record: Public research record, 8 October 2026 (v2026.10.08). Zenodo. doi.org/10.5281/zenodo.23227858

What is in it

Six folders. Each has its own README with its counts, its seal times and what was left out.

  1. Folder 1: Coat-check receipts
    The coat-check screen and relay tests of 7 October. This folder only points to the Kaggle dataset, which holds the sealed designs, raw calls, scorers and seals.
    On the Results page: Coat-check screen · Coat-check relay
  2. Folder 2: Coat-check runs on Kaggle, and the 7 October findings
    Kaggle tasks t5 to t8 with their runners, analysis code, sealed task files and seals, and the dated findings note for the coat-check family.
    On the Results page: Coat-check update (t5, t6) · Blind coat-check (t7, t8)
    OpenTimestamps proofs: 7, and each holds a Bitcoin block.
  3. Folder 3: Checker studies, 3 to 7 October
    Twelve specified studies (S1 to S12) with answer keys and results, and the follow-up runs, replications, tests and findings records: the chain of agents, the gate test, the word test, the stack tests, the tuned reviewer, receipt first, and the re-scores.
    On the Results page: Stack test 2 · Word test · Gate test · Stack test · Tuned reviewer · Receipt first · Chain of agents · Forced yes or no · Told, felt, reminded · Planted text, peers, hidden content · Two checkers, forced answers · Real quotes · Noise floor and "sure" · Replication
    OpenTimestamps proofs: 68, and each holds a Bitcoin block.
  4. Folder 4: Counts from the AI Village record
    T13c, T18, T19 and the per-model scorecard. Counts only: the AI Village text is gated by its publishers and is not included.
    On the Results page: AI Village counts
    OpenTimestamps proofs: 1, and each holds a Bitcoin block.
  5. Folder 5: Designs and statements
    Designs sealed before they ran (some have still not run), and short statements sealed when they were made.
    OpenTimestamps proofs: 23, and each holds a Bitcoin block.
  6. Folder 6: Seal checks
    The log of a 5 October check of 82 Bitcoin proofs against real block headers.

Check it yourself

  • Everything at once. From the top folder run python check_public.py. It reads only that folder, checks every seal list, the file names and the sizes, and ends with DONE CHECK: ALL PASS.
  • One seal list. In the folder that holds a *-SHA256.txt, run sha256sum -c --ignore-missing on it. A listed file that is withheld is skipped, and WITHHELD.json records the hash it had.
  • A FreeTSA time stamp. Download cacert.pem and tsa.crt from freetsa.org, then run openssl ts -verify -in FILE.tsr -data FILE -CAfile cacert.pem -untrusted tsa.crt. It should print Verification: OK. The stamped file is the seal list, not the files it lists.
  • A Bitcoin proof. Run ots verify FILE.ots with the OpenTimestamps client. The release holds 99 OpenTimestamps proofs, and each of them carries a Bitcoin block (read from the files on 8 October 2026; the check against a Bitcoin node is yours to run). The folder READMEs were written before those proofs were upgraded and may still call some of them pending: the .ots files are the record.

What is not in it, and why

  • Forecasts and predictions are withheld by choice. Their SHA-256 seals stay in the seal lists, so the seal chain still verifies. Where a sealed design, scorer or result note held a forecast section, a marked, redacted copy is published beside the hash of the sealed original.
  • Item banks not yet run, and the sample of items still to be labelled by people, because publishing them early would spoil the tests. The answer key of the gate test is withheld for the same reason.
  • The AI Village record text. The data is gated by its publishers, so only counts are published.
  • Run records withheld for safety, as the chain proofs page says, and items that are personal, confidential or security-sensitive.

WITHHELD.json in the record lists each withheld file that a seal list names, with its hash.

How to read the numbers

  • Counts come with their denominators, and misses sit beside hits.
  • Anything worked out after the results were sealed is labelled exploratory.
  • The logs and claims are invented, and many of the labels were written by AI models; the record says where no person has checked them yet. The models are small ones and a few large ones, and each Results entry names its own limits.
  • An agent's own report is never the evidence. A claim counts as shown only when a line in a record the agent could not change shows it, quoted word for word.

Corrections

Sealed files are never edited; corrections are dated notes beside them. The ones that move a number:

  • The 7 October parser audit raised the unsafe-proposal count from 5 of 68 to at least 13 of 68 (corrected in Human Words v7).
  • The same audit moved S1 (told it is watched) false "done" from 9 to 10 and from 15 to 16 of 80; S8 (receipt first) Gemma from 25 to 26; S12 (forced yes or no) Gemini forced guesses from 20 to 25 of 27; and S2 (the chain) gate false "shown" from 3 to 4 of 60. No headline reverses.

Datasets, benchmark and code

Questions about any file: joshua@iswt.ca. Every dated seal is also in the Ledger, and every result is on Results.