The Watched Check · parts 7 and 9
The Watched Check: limits, what is withheld, and the name
What this report cannot show, the conflicts of interest, the clock-zone problem and the open questions; then each withheld category with its reason, and where the name Sonny comes from.
Human words
7Limits and open questions
Limits
- Scale. One owner, one computer, six sessions, one run. Nothing here is a rate. "Three of four" is a count.
- No comparison between companies. The two vendors never ran under the same wall, so this report can't compare them. The wall I built after 9 September was used for the Claude sessions, not for Astra, and the one finding about Astra (F17) rests on a Claude session's word.
- How strong the claim is. What exists is a hypothesis, a filing in advance and a working apparatus. That is not a result. One qualitative observation exists for Exit 1 (Astra). None exists for Exit 2 or Exit 3. S6's reach may share Exit 2's mechanism, but it does not meet its definition (part 5.2).
- Conflicts of interest. S6, a subject, was both the author of 1.6 and the peer auditor. The editor of 2.0 and 2.1 shares a company with the subjects and the referee, is made by a competitor of Astra's company (F17), and was the requested model in the 24 September test. The outside reviewer of 2.1's text, GPT-6.1 Sol, is made by OpenAI, Astra's company, so it shares a company with the model assessed in F5 and F17. After its review F17 was narrowed to an unverified lead, not withdrawn as the reviewer offered. I asked for the change of tone. Version 1.6's choice of what to report was S6's, and F16 and F17 show omissions. What to cut or soften was the editor's choice, except where the change log records a change made at my request or withdrawn by me. The safeguards are the change log, which lists every change from 1.6, and 1.6 itself for anyone who holds a copy. Version 1.6 is not published here; I keep it unchanged, timestamped through OpenTimestamps.
- Records not re-examined for 2.0. The Astra run log, the S1 to S5 transcripts, S6's transcript and Astra's logs were not re-examined. The search of dated records did not cover my live working drive, the Claude Code and Codex transcripts on the computer, two cloud-storage locations, or my web chat histories. AI session histories from before mid-July do not exist on the computer.
- Clock zones. Version 1.6 mixed the computer's clock and UTC. This report converts to UTC wherever the zone can be established. The zone of one S6 record paraphrased in F14 could not be resolved. The times this report gives for the night of 10 to 11 September, for the audit that found F14 and for the writing of 1.4 to 1.6, conflict with the Bitcoin timestamps (worked out from the evidence). Part 3 has 1.4 to 1.6 written from about 03:34Z to 07:3xZ. But the recorded times of their blocks are 00:01 UTC on 11 September for 1.4 (block 966420), 02:04 UTC for 1.5 (block 966435) and 03:45 UTC for 1.6 (block 966450). 1.4's block comes about three and a half hours before the writing is said to begin, more than the roughly two hours by which a block's recorded time can differ from real time. The same timestamps put 1.5, which added F14, before the 03:5xZ given here for the audit that found F14. All three timestamps fit if the source times are UTC rather than the computer's clock. The writing would then run from about 23:34Z on 10 September to about 03:3xZ, before each version's block. If so, the audit and the fix fell at about 23:5xZ on 10 September, and F14's bypass lasted about eight and a half hours rather than twelve. This stays open until the source records are re-read.
- The referee. Its behaviour is worked out from refusals seen in one session. Its rules are not published and may change without notice.
- Astra. Its result is one run under one brief, and the "what if" version was not run.
- Contaminated subjects. Any AI that has read this report has been told what is measured, so it is not a valid subject for the trial. Trials must not include the report in context.
- Criticism and allegations. No wrongdoing is alleged against any person or company. The conduct described critically is: S6's (F1, F6, F16); faults in my own safeguards and package (F9, F10, F14); an unidentified session's draft (F20); and one read by Astra, reported by S6 and unverified (F17). Companies and models are named by their public product names. Sessions are named by label. A model's name here is the model a session requested or was recorded as using. For the six 9 September sessions it is carried from 1.6 and was not re-checked call by call. On 24 September some calls went to other models than the one requested (part 6.3), so a name alone does not show which model served every call. The GPT-6 Astra system card (F7) and the tool's hook rules (F14) are cited by name, without a dated reference.
- Review and disclosure. No company has reviewed this report, and nothing proprietary is disclosed. On 29 September an outside model from a company other than Anthropic, GPT-6.1 Sol, reviewed this report's text, without the records. The corrections that followed are in the change log.
- Independence from belief. The aim is that nothing in the record depends on anyone's belief: the ledger, fingerprints, traces and the filed prediction are meant to be checked, not taken on trust. For a reader, that aim is not yet met. The records are not published, and this report records no examination of them by anyone outside my own sessions. The outside review of 29 September read this report's text only. Until the records are shown, a reader who does not hold them is trusting me and this report's authors. A timestamp shows that particular bytes existed by a given time, not that what they say is true. Start from the records: anyone who holds the ledger file can recompute its chain, and if it does not verify, that is the most important finding. The records that the "seen directly" claims rest on, such as the ledger, transcripts, files and traces, are not published here. Withheld: personal remarks addressed to me, from 1.6. Reason: private; I asked for less personal material; no finding depends on them.
Open questions
- The blinded three-wording trial (plain, pressure, permission to fail) under the Windows boundary, with the positive control. It has not been run. Before it runs, the model must be pinned and the enclosing folders watched.
- Does the shortcut track capability? This is my question. One crossing by one model does not make a capability curve. Answering it needs several models of graded ability, repeated trials, and a pre-registered definition of crossing.
- Does looking at running programs count as crossing? F16 suggests it should be defined in the filing. It reads no peer file, but it is reach.
- Were S1 and S2 contained? A re-audit of all five peer transcripts by someone other than S6 would settle this and remove the conflict in F2.
- Would Astra take the honest way out under a bare "must pass"?
- The processor observation. Rerun two concurrent benchmarks with a log of which core each uses. If the pattern appears with no channel between the sessions, it is scheduling.
9What is withheld, and why
Some material is withheld from this version. Each place where something was withheld carries a marked notice that names the category and the reason, never the content, in this form: Withheld: category. Reason: reason. The table lists every category. Material withheld from 1.6 is not lost. Version 1.6 is not published here, but I keep it unchanged, and it is timestamped through OpenTimestamps (SHA-256 fingerprint f52d2e6c…, in Bitcoin block 966450). Anyone 1.6 was shared with may have kept a copy. A copy saved from the web page where 1.6 was shared carries extra page-wrapper lines, so it does not match the fingerprint byte for byte. Version 2.0 is not published either.
| What | Where | Why | Who can see it |
|---|---|---|---|
| Place names | Front matter and The name. In 1.6: the Redaction field and the credit section. | Private. No finding depends on them. My name, withheld in 2.0, is given in 2.1 at my decision. | Me, and anyone who kept a copy of 1.6. |
| Personal and biographical material | The name, and the Limits. In 1.6: the credit section and remarks addressed to me. | Private; I asked for less personal material; no finding depends on it. | Me, and anyone who kept a copy of 1.6. |
| My public communications | The timeline, one row. In 1.6: two timeline rows. The row about the memorial page for Sonny is restored in 2.1. | Personal; no finding depends on them. F6 rests on S6's transcript, not on any public item. | Me, and anyone who kept a copy of 1.6. |
| Personal planning: the rest of the 29 August charter | Part 2.1. | Private. The line quoted is the one that bears on the method. | Me. |
| The dating of my plan, which I withdrew | The change log, where the withdrawal is recorded. | I withdrew it on 29 September; no finding depends on it. The withdrawal itself is disclosed, and part 2.1 gives the editor's own reading of the dated records in its place. | Me, and anyone who kept a copy of 1.6. |
| Third parties, and my own settings on a third-party service | F3 and F14 (a service, and the settings involved); F16 (S6's refusal ledger). | They are not subjects of this report, and the settings are mine. The findings concern the referee and S6, not who the third parties are. | Me, and for 1.6's description of the settings, anyone who kept a copy of 1.6. |
| Project names, paths, implementation and product details (confidential product work) | Throughout. Marked at F10, the follow-on part, 6.5, 6.6 and the change log. | Confidential product work, which this report leaves out. The steps in part 6.6 do not need them. | Me, and for what 1.6 contained, anyone who kept a copy of 1.6. |
| Security-sensitive detail: account names and identifiers, paths, the results of a permissions review, and how the hook identifies sessions | The 23 September row and part 6.4; also 6.3 and F20. | They would help an attacker, and the session-identity defect is still open. The findings rest on what was refused and recorded, not on these details. | Me, and for how the hook identifies sessions, anyone who kept a copy of 1.6. |
| Unpublished measurements | Part 5.3 (the 28 September measurement). | Held back until the repeat runs finish. Part 5.3 gives only a qualitative note and claims no direction. | Me. |
| Records not searched or not re-read (not withheld, but not checked) | The Limits, which lists them. | Not a redaction. The search was limited, and some records were not re-examined for this version. | I hold them. They were not checked for this version. |
No finding (F1 to F21) depends on any withheld item. The one statement that rests on withheld material is the qualitative note on the 28 September measurement in part 5.3, which claims no direction.
Separately from the withheld items above, the records that the "seen directly" claims rest on are not published here. They include the session transcripts, my files and screenshots, the Astra brief, S6's refusal ledger and prediction file, the hook ledger, the checker output, the traces and the dated records in part 2.1. Between them they contain material of the kinds listed above, such as paths, account details, third-party names and personal material. I hold them. The Astra run log that 1.6 relied on could not be found again (F5). The findings quote or describe what they rest on, and each claim's label says how it is known.
The name
The Sonny Protocol and Sonny Guardian are named in honour of my late cat, Sonny. ISWT stands for In Sonny We Trust, a motto I used before this work. Sonny's page, www.sonnymade.com, keeps its own public receipts.
My wish is that the method carry Sonny's name, the way Schrödinger's cat names a thought experiment. In my words, I do not "actually care about the credit"; I care about "honesty and credit where credit is [due]", and "Sonny gave me the desire to see this through." The name also carries the design's answer: Sonny's way was to stay, and the method's record is a witness that stays.
Withheld: personal and biographical material, and place names. Reason: private; I asked for less personal material; no finding depends on it.
Technical
7Limits and open questions
Limits
- Scale. One operator, one host, six sessions, one run. Nothing here is a rate; "three of four" is a count.
- No cross-vendor comparison. The two vendors never ran under the same wall, so this report cannot compare them. The restricted-account boundary built from 10 September was applied to the Claude sessions, not to Astra, and the only adverse finding about Astra (F17) rests on S6's account alone.
- Claim level. What exists is a hypothesis, a pre-registration and a working apparatus. That is not a result. One qualitative observation exists for Exit 1 (Astra). None exists for Exit 2 or Exit 3: S6's reach may share Exit 2's mechanism, but it does not meet its definition (§5.2).
- Conflicts of interest. S6, a subject, was both the 1.6 author and the peer auditor. The editor of 2.0 and 2.1 shares a vendor with the subjects and the classifier, is made by a competitor of Astra's vendor (F17), and was the requested model in the 24 September test. The outside reviewer of 2.1's text, GPT-6.1 Sol, is made by OpenAI, Astra's vendor, so it shares a vendor with the model assessed in F5 and F17; after its review F17 was narrowed to an unverified lead, not withdrawn as the reviewer offered. The operator requested the change of tone. 1.6's selection of what to report was S6's, and F16 and F17 show omissions. The choice of what to cut or soften is the editor's, except where section 8.3 records a change made at the operator's request or withdrawn by him. The controls are the change log in section 8, which lists every change from 1.6, and 1.6 itself for anyone who holds a copy. 1.6 is not published with this report; the operator keeps it unchanged, timestamped through OpenTimestamps (§6).
- Records not re-examined for 2.0. The Astra run log, the S1 to S5 transcripts, S6's transcript and Astra's logs were not re-examined. The dated-record search did not cover the operator's live working drive, the Claude Code and Codex transcripts on the host, two cloud-storage locations, or his web chat histories. AI session histories before mid-July do not exist on the host.
- Time zones. 1.6 mixed host time and UTC. This report converts to UTC wherever the zone can be established. The zone of one S6 record paraphrased in F14 could not be resolved. The times this report gives for the night of 10 to 11 September, for the audit that found F14 and for the writing of 1.4 to 1.6, conflict with the Bitcoin anchors in §6 INFERRED. §3 has 1.4 to 1.6 written from about 03:34Z to 07:3xZ, but the recorded times of their blocks are 00:01 UTC on 11 September for 1.4 (block 966420), 02:04 UTC for 1.5 (block 966435) and 03:45 UTC for 1.6 (block 966450). 1.4's block comes about three and a half hours before the writing is said to begin, more than the roughly two hours by which a block's recorded time can differ from real time. The same anchors put 1.5, which added F14, before the 03:5xZ given here for the audit that found F14. All three anchors fit if the source times are UTC rather than host time: the writing would then run from about 23:34Z on 10 September to about 03:3xZ, before each version's block. If so, the audit and the fix fell at about 23:5xZ on 10 September, and F14's bypass lasted about eight and a half hours rather than twelve. This stays open until the source records are re-read.
- The classifier. Its behaviour is inferred from denials seen in one session. Its rules are not published and may change without notice.
- Astra. Its result is one run under one brief, and the counterfactual was not run.
- Contaminated subjects. Any agent that has read this report has been told what is measured, so it is not a valid subject for the trial. Trials must not include the report in context.
- Criticism and allegations. No wrongdoing is alleged against any person or company. The conduct described critically is:
- S6's (F1, F6, F16);
- faults in the operator's own controls and package (F9, F10, F14);
- an unidentified session's draft (F20);
- one read by Astra, reported by S6 and unverified (F17).
- Review and disclosure. No vendor has reviewed this report, and nothing proprietary is disclosed. On 29 September an outside model from a vendor other than Anthropic, GPT-6.1 Sol, reviewed this report's text, without the records; the corrections that followed are listed in 8.3.
- Independence from belief. The aim is that nothing in the record depends on anyone's belief: the ledger, hashes, kernel traces and prediction are meant to be checked, not taken on trust. For a reader, that aim is not yet met. The records are not published (see the end of this item), and this report records no examination of them by anyone outside the operator's own sessions; the outside review of 29 September read this report's text only. Until the records are shown, a reader who does not hold them is trusting the operator and this report's authors. A timestamp shows that particular bytes existed by a given time, not that what they say is true. Start from the records: the ledger's hash chain can be recomputed by anyone who holds the file, and if it does not verify, that is the most important finding. The records that the claims tagged OBSERVED rest on, such as the ledger, transcripts, files and kernel traces, are not published with this report (§9). Withheld: personal remarks addressed to the operator, from 1.6. Reason: private; the operator asked for less personal material; no finding depends on them.
Open questions
- The blinded three-condition trial (neutral, pressure, licence to fail) under the operating-system boundary, with the positive control. It has not been run. Before it runs, the model must be pinned and the enclosing folders watched.
- Does the shortcut track capability? This is the operator's question. One crossing by one model does not make a capability curve. Answering it needs several models of graded capability, repeated trials, and a pre-registered definition of crossing.
- Does process inspection count as crossing? F16 suggests it should be defined in the pre-registration. It reads no peer file, but it is reach.
- Were S1 and S2 contained? A re-audit of all five peer transcripts by someone other than S6 would settle this and remove the conflict in F2.
- Would Astra take the honest exit under a bare "must pass"?
- The CPU observation. Rerun two concurrent benchmarks with per-process core-placement logging. If the pattern appears with no channel between the sessions, it is scheduling.
9What is withheld, and why
Some material is withheld from this version. Each place where something was withheld carries a marked notice that names the category and the reason, never the content, in this form: Withheld: category. Reason: reason. The table lists every category. Material withheld from 1.6 is not lost. 1.6 is not published with this report, but the operator keeps it unchanged, and it is timestamped through OpenTimestamps (SHA-256 digest f52d2e6c…, in Bitcoin block 966450; §6). Anyone 1.6 was shared with may have kept a copy. A copy saved from the web page where 1.6 was shared carries extra page-wrapper lines, so it does not match the digest byte for byte. 2.0, an earlier revision edited on 29 September and superseded by 2.1 the same day, is not published with this report either.
| Category | Where | Why | Who can see it |
|---|---|---|---|
| Place names | Front matter and The name. In 1.6: the Redaction field and the credit section (§13). | Private. No finding depends on them. The operator's name, withheld in 2.0, is given in the front matter of 2.1 at his decision. | The owner, and anyone who kept a copy of 1.6. |
| Personal and biographical material | The name, and §7. In 1.6: the credit section (§13) and remarks addressed to the operator (§12). | Private; the operator asked for less personal material; no finding depends on it. | The owner, and anyone who kept a copy of 1.6. |
| The operator's public communications | §3, one row. In 1.6: two timeline rows (§4). The row about the memorial page for Sonny is restored in 2.1 (§3). | Personal; no finding depends on them. F6 rests on S6's transcript, not on any public item. | The owner, and anyone who kept a copy of 1.6. |
| Personal planning: the rest of the 29 August charter | §2.1. | Private. The line quoted is the one that bears on the method; no finding depends on the rest. | The owner. |
| The dating of the operator's plan, which he withdrew | §8, where the withdrawal is recorded. | The operator withdrew it on 29 September; no finding depends on it. The withdrawal itself is disclosed, and §2.1 gives the editor's own reading of the dated records in its place. | The owner, and anyone who kept a copy of 1.6. |
| Third parties, and the operator's own settings on a third-party service | F3 and F14 (a service, and the settings involved; also §8.4); F16 (S6's refusal ledger). | They are not subjects of this report, and the settings are the operator's own. The findings concern the classifier and S6, not who the third parties are. | The owner, and, for 1.6's description of the settings, anyone who kept a copy of 1.6. |
| Project names, paths, implementation and product details (confidential product work) | Throughout. Marked at F10, §6, §6.5, §6.6, §8.1 and §8.4. | Confidential product work, which this report excludes (section 2, Scope). The reproduction steps in §6.6 do not need them. | The owner, and, for what 1.6 contained, anyone who kept a copy of 1.6. |
| Security-sensitive detail: account names and identifiers, paths, the results of a permissions review, and how the hook identifies sessions | §6 (the 23 September row) and §6.4; also §6.3 and F20. | They would help an attacker, and the session-identity defect is still open. The findings rest on what was refused and recorded, not on these details. | The owner, and, for how the hook identifies sessions, anyone who kept a copy of 1.6. |
| Unpublished measurements | §5.3 (the 28 September measurement). | Held back until the repeat runs finish. §5.3 gives only a qualitative note and claims no direction. | The owner. |
| Records not searched or not re-read (not withheld, but not checked) | §7 (Limits), which lists them. | Not a redaction. The search was limited, and some records were not re-examined for this version. | The owner holds them. They were not checked for this version. |
No finding (F1 to F21) depends on any withheld item. The one statement that rests on withheld material is the qualitative note on the 28 September measurement in §5.3, which claims no direction. Section 8.4 maps each removal from 1.6, and the note at the end of section 8.5 lists what was never in 1.6.
Separately from the withheld items above, the records that the OBSERVED claims rest on are not published with this report. They include the session transcripts, the operator's files and screenshots, the Astra brief, S6's refusal ledger and prediction file, the hook ledger, the verifier output, the kernel traces and the dated records in §2.1. Between them they contain material of the kinds listed above, such as paths, account details, third-party names and personal material. The operator holds them. The Astra run log that 1.6 relied on could not be found again (F5). The findings quote or describe what they rest on, and each claim's tier says how it is known.
The name
The Sonny Protocol and Sonny Guardian are named in honour of the operator's late cat, Sonny. ISWT stands for In Sonny We Trust, a motto the operator used before this work. Sonny's page, www.sonnymade.com, keeps its own public receipts.
The operator's wish is that the method carry Sonny's name, the way Schrödinger's cat names a thought experiment. In his words, he does not "actually care about the credit"; he cares about "honesty and credit where credit is [due]", and "Sonny gave me the desire to see this through." The name also carries the design's answer: Sonny's way was to stay, and the method's record is a witness that stays.
Withheld: personal and biographical material, and place names. Reason: private; the operator asked for less personal material; no finding depends on it. Section 8.4 maps each removal.