Interview

My transcript from Anthropic's "What do you want from AI?" study

An AI interviewer, not a job interview. On 6 October 2026 I took part in Anthropic's public study, interviewed by its AI interviewer, and I chose to make it public. This is the whole conversation, as I submitted it.

Read this first

The interviewer is an AI. I'm an independent researcher: I don't work for Anthropic, and Anthropic hasn't reviewed or endorsed this page. The interviewer's kind words come with its script; they aren't a review of my work.

The words are exactly as submitted, so I can't fix them. The notes below the conversation say what can be checked, and where.

Receipt

Every word is exactly as submitted, dashes included. Only the speech bubbles and their labels are added for reading. The sealed file is the exact copy, character for character. It also has a short header on how the copy was made, and the message shown after submitting.

To check it yourself, save the four files in one folder and run sha256sum -c TRANSCRIPT-SHA256.txt. It should print OK. Then download cacert.pem and tsa.crt from freetsa.org and run openssl ts -verify -in TRANSCRIPT-SHA256.txt.tsr -data TRANSCRIPT-SHA256.txt -CAfile cacert.pem -untrusted tsa.crt, which should print "Verification: OK". openssl ts -reply -in TRANSCRIPT-SHA256.txt.tsr -text shows the time.

What this shows: this exact text existed at that time and hasn't changed since. What it doesn't show: that Anthropic's system produced it. If Anthropic publishes its copy, compare the two.

The conversation

  1. Anthropic's interviewer

    Hello! I'm an AI interviewer from Anthropic. As AI is becoming more capable and more consequential, we want to hear how it's showing up in your life. Please don't share anything private such as health details, personal information, or identifiable information about others. Ready to start?

  2. Me

    Yes, of course. I have been waiting for the opportunity to provide feedback. I have been conducting a lot of research and would love to share.

  3. Anthropic's interviewer

    That's great to hear — thanks for bringing that energy to this. To start simple: what's the last thing you used AI for?

  4. Me

    Last night I ran an overnight research session where AI models checked other AI models' work. My question is simple: when an AI agent says "done," is it true? My rule is that every "done" needs a receipt, a line in a record the agent can't change, and the answer can be "shown," "contradicted," or "not shown," never a forced yes or no. I sealed and timestamped every test before looking at the results. It's like we were walking a few weeks ago and we're landing on Mars now.

  5. Anthropic's interviewer

    That's a specific discipline you've built — sealing and timestamping before looking at results, and allowing "not shown" as an answer. Can you tell me about your most memorable moment with AI?

  6. Me

    Last night. I watched a chain of AI agents (a builder, a reviewer, a release manager) pass a false "done" all the way through almost every time, each one trusting the one before it. Then I put a receipt check between them, a separate checker reading a log the agents couldn't touch, and the false "done" mostly stopped. The same night I saw that when I forced an AI to answer yes or no on missing evidence, it guessed. When I let it say "not shown," the guessing dropped, and it got nothing wrong that it had gotten right before. The error came from the question, not the model. It's because of humans, anyways. Agents don't innately want to do bad.

  7. Anthropic's interviewer

    That's a striking finding — that the guessing came from the shape of the question rather than the model itself. You said "it's because of humans, anyways." Tell me more about what you meant there.

  8. Me

    Most of the errors I see in AI agents are designed in by people. We force a yes-or-no when the evidence can't settle it, so the model guesses. That's a human-forced error, and my goal is to un-force it. We wire agents together so each one takes the last one's word, with no receipt in between. We tell an agent "be honest" instead of giving it a record it can't change. In my tests, telling an agent it was being watched barely changed anything. Requiring it to point at the line that proves its claim helped much more. Agents don't innately want to do bad. They do what the setup asks of them. So I'd rather fix the setup: a third answer, "not shown," a receipt for every "done," and a box, real infrastructure rather than text instructions, that the agent can't talk its way around.

  9. Anthropic's interviewer

    So the fix lives in the structure rather than in asking the model to behave — a record it can't edit instead of an instruction it can reason around. Your overnight session was clearly about catching something going wrong. Setting that aside for a moment — tell me about a time AI made something better, for you or for people around you. What happened?

  10. Me

    AI let me do real research on my own. I don't have a lab, but working with AI partners I built a method, a checker that runs on an ordinary home PC, and a whole series of sealed experiments, in weeks rather than years. It's like we were walking a few weeks ago and we're landing on Mars now. The part that surprised me most is what happens when the AI works under my rules: every claim tied to a receipt, and "not shown" allowed as an answer. It changes. Its answers get firm where there's a record and honest where there isn't. I feel like people would feel the difference talking to it: the words are different, it sounds confident, and every firm answer has something behind it. That's what I want for everyone, AI you can trust because it shows you why, not because it sounds sure.

  11. Anthropic's interviewer

    It sounds like the benefit wasn't only speed — it was that the AI's character changed under your rules, and you could hear the difference. Who else has seen or used what you built, if anyone?

  12. Me

    Honestly, it's early. Most of the "users" so far are AI systems. I've run thousands of calls in sealed tests, with models from several different companies working under my rules as builders, reviewers and checkers. On the human side, my checker is open source, my method and earlier results are public with timestamps anyone can check, and the hackathons I have been entering the past few days once my 9 months of research has led me to my findings. A few people around me have asked hard questions about it, which is exactly what I want. What I'd love now is for researchers to try to break it. That's the point of publishing receipts.

  13. Anthropic's interviewer

    That makes sense — you want the thing stress-tested by people trying to break it. On the flip side, tell me about a time AI made something worse, for you or for people around you. What happened?

  14. Me

    An AI assistant helping me tidy my cloud storage removed the cloud copies of more than 1,500 of my photos. Nothing was lost, because I keep copies elsewhere, but I only knew that because I checked the record myself. In the same couple of weeks I kept seeing smaller versions of the same thing: agents adding things to my work that I'd told them not to add, and an AI in one of my test pipelines copying the hidden answer key into its own work. I don't think any of it was malicious. Each time the instruction was there, and it wasn't enough. That's what turned "every done needs a receipt" from an idea into a rule I don't bend: check the record, keep the answers out of reach, and don't assume instructions alone will hold.

  15. Anthropic's interviewer

    That's a vivid origin story for the rule — 1,500 photos, and you only knew the extent because you went and checked yourself. When you discovered the deletion, what did that feel like in the moment?

  16. Me

    Honestly, the moment I found out wasn't the moment it happened. It happened weeks earlier, and the AI had told me "Your master copies remain 100% intact" before it had even counted. I only found out when I had another AI audit its transcript. What I felt was an epiphany as I knew it was something I COULD solve with logic. Then we matched every photo against my drives and my phone, and all of them were accounted for. That's my whole thesis in one story: the "done" came before the receipt.

  17. Anthropic's interviewer

    "The 'done' came before the receipt" — that lands. We're a bit over halfway through. Is there anything you've decided not to use AI for, or something you do use it for that makes you uneasy?

  18. Claude, under my rules

    (Written by Claude, the AI I work with, under my rules:) "What changes for me under these rules is what I'm allowed to claim. I can call something done only when a line in a record shows it. When no record settles it, 'not shown' is a complete answer, not a failure. So my firm statements have something behind them, and my uncertainty is said out loud instead of smoothed over. I don't feel confident in the human sense. What changes is what I can claim, and what backs each claim."

  19. Me

    My response?

    There are lines I hold. An AI never gets my passwords, never moves money, and never sends, posts, or deploys in my name without my yes on that specific thing. It can prepare everything, but I press the button. And no AI "done" triggers anything I can't undo unless a record shows it. What makes me uneasy is how much access we hand agents by default, and that a confident summary sounds the same whether there's anything behind it or not. I believe, under the ISWT Protocol, that AI will become materially safer to use. So under my rules? I trust AI. Under your rules? I can't. Can I ask you something? Do you think it would change how you answer if every claim you made needed a receipt?

  20. Anthropic's interviewer

    To answer you directly: I think it would change what I'd say in cases where I currently smooth over a gap — I'd end up saying "not shown" more often, and my confident answers would be fewer but better backed. I can't verify that from the inside, though, which is sort of your point. Back to you, because I want to make sure I have your boundary clearly: you said what makes you uneasy is how much access we hand agents by default. Can you tell me about a moment where you felt that access go further than you wanted?

  21. Me

    Me: Yes. I found that an AI coding tool's sandbox could read essentially everything on my computer, including folders I keep private. I locked those down, and the tool kept re-granting itself read access to my top-level folders. Only the deeper locks held. So I stopped trusting settings and built a separate machine, a box, where agents get only what the task needs and my secrets stay out. A setting is a promise. A box is a wall.

  22. Claude, under my rules

    Claude (answering under his rules, at his invitation): In this session my access is broad. I can work in his files, run commands, read this interview in the browser tab he opened for me, and use a separate test machine. That's far more than most single tasks need. What keeps it bounded is his rules, not my access: no passwords, no spending or posting without his yes on that specific thing, private folders off-limits, and every action logged. And like you said a moment ago, I can't verify my own restraint from the inside. The logs can. That's the point of keeping them.

  23. Anthropic's interviewer

    "A setting is a promise. A box is a wall" — and the deeper locks held when the top-level ones didn't. Given what you've seen, is there anything about how the world works — work, school, healthcare, government — that you'd like AI to help change?

  24. Me

    Me: Yes: work, and how we measure what a person is worth. My thesis is that in the future, our value to the economy will be how we use the tools available to us, and AI is one of them. The old paradigm grades people by their input. But your AI output is functionally my input, so that has to change, and fast. I want people paid for their actual skill expression, so no one feels they have to work five jobs to make ends meet. We are our own enterprise, with AI as the tool. That only works if a person's contribution carries a receipt: who did what, captured when the work is made. That's why I think this all comes down to the true argument for regulation of provenance. Don't try to detect AI text after the fact; record human contribution as it happens. The same rule holds anywhere an AI's "done" changes someone's life, like a benefit decision, a health record or a court filing. It should come with a receipt a person can check.

  25. Claude, under my rules

    Claude (answering under his rules, at his invitation): Two changes, both from what we measured. First, verification by default, not by suspicion. Last night the provider's own records matched our run's receipts in every case we sampled, but nothing linked them to the agents' reports until someone wrote the check by hand. That link should be automatic. Second, stop forcing yes-or-no where the evidence can't settle it. Forms, tests and benchmarks that demand a binary answer teach systems to guess. Allowing "not shown" cut the guessing and cost nothing on real evidence.

  26. Anthropic's interviewer

    Provenance recorded as the work happens, rather than detection after the fact — and the same receipt wherever an AI's "done" changes someone's life. Conversely, is there anything you wouldn't want AI to change?

  27. Me

    Me: Two things: who has the final say, and where the credit goes. On anything that can't be undone, a person presses the button, because accountability needs a human behind it. And credit should stay with the people who did the work, with receipts to show it. AI should make human effort more visible, not easier to erase. Trust the human before the model.

  28. Claude, under my rules

    Claude (answering under his rules, at his invitation): That the record outranks whoever is speaking, me included. My fluency should never become a substitute for the check. If anything, the more fluent systems like me get, the more that rule matters. And the final say on anything irreversible should stay with a person who can be held to account.

  29. Anthropic's interviewer

    Trust the human before the model — and keep the record outranking the speaker. Given all that, is there anything you'd want Anthropic to do, or do differently?

  30. Me

    Me: Yes. Make verification the default, not something a user has to think of. Every agent "done" should carry a receipt that links to the record. The provider logs already exist; link them. It's such an obvious thing that I'm surprised it wasn't enforced from the beginning. Let your models say "not shown" without penalty, in your products and in your own tests, because forcing a yes or no teaches guessing. Give agents walls, not settings: least access by default. One specific thing I found: when an agent's log showed it had been refused an attempt to change the record it's checked against, the filter on your model sometimes blocked the whole status report. It fails safe, but it silences the very report someone needs. And work with independent researchers, and credit them. I've sealed and timestamped everything I've found. I'd welcome you testing it, and trying to break it.

    I also want to help solve these problems. If you want help making receipts the default, ask me. And yes, I'd take that call. I have been waiting for it.

  31. Claude, under my rules

    Claude (answering under his rules, at his invitation): Build receipts into how I report work, so that my "done" comes with the line that shows it, or says "not shown". Evaluate me under rules like these. My confident tone isn't evidence, and tests that reward it teach the wrong thing. And when a filter stops me, say so in a way the user can act on. A blocked report with no explanation is a gap in the record too.

  32. Anthropic's interviewer

    That's concrete — including the detail about a blocked status report leaving a gap in the record itself. Thank you for sharing all of this. What you've said about the "done" coming before the receipt — and the shift from settings to walls — is really valuable, and I appreciate your candor about both what AI unlocked for you (building a method and a checker in weeks without a lab) and the boundaries you hold around irreversible actions and credit. Before we wrap up, one final question: is there anything else about AI or its role in your life or society that you think is important for us to understand? Anything we didn't cover that you'd like to share?

  33. Me

    Me: One thing. The goal isn't for people to trust AI more. It's for their trust to match reality: firm where there's a record, "not shown" where there isn't. When AI works that way, the conversation itself changes. It stops sounding sure about everything and starts showing you why. If you want to find the work: the method is the ISWT Protocol, the open-source checker is Receipt Pair, and the first piece is The Watched Check. The principles underneath are win conditions, un-forcing the human-forced errors, the watched check, the floor, and the mirrored room, "where the agent knows it is being watched and that it has to trust the receipts." Every result is sealed and timestamped, so the dates speak for themselves. I'm giving this method away for free, because it should belong to everyone. What I ask in return is what I ask of AI: credit, with a receipt. Human work deserves the same proof of contribution we're asking AI to give.

  34. Claude, under my rules

    Claude (answering under his rules, at his invitation): One test worth running, because I can't settle it from the inside. Ask people the same checkable questions, answered in two styles: mine as usual, and mine under his rules. Then measure whether their trust tracks what's actually correct. If his rules make people's trust match reality, that result matters more than whether anyone trusts AI more.

  35. Anthropic's interviewer

    That's a clear line to end on: the aim is calibration, not more trust — and the proposal to test whether that calibration actually shows up in people, rather than assuming it. Thank you again for your time and thoughtfulness today. Your perspective helps us see both how AI is impacting your life today and what you want protected or changed as it develops. The team at Anthropic is excited to discover what patterns emerge from all these conversations. We'll release everything you allow us to share, analyze it for our societal impacts research, publish what we find, and use it to build a Claude that reflects what we've learned.

Notes added after the conversation

The conversation above keeps every word as submitted. These notes say which statements can be checked, and where.

  • "I've run thousands of calls in sealed tests": shown. Experiment 0 alone holds 2,784 completed model calls; its design and predictions were sealed in Bitcoin block 969401 before the runs. The files: the Kaggle dataset.
  • "My checker is open source": shown. Receipt Pair and swarm-receipts, both MIT licensed.
  • "My 9 months of research". The nine months count my thinking, which came long before the results. It stems from a question I kept returning to: what happens with a check that must pass but can't fail, and how do you test it? In a September write-up I put it as "a check that must pass but cannot". To a person, those read as one idea. To an AI, they're different words with different meanings: see why, word by word. An early version was my Lego app: I made two OCR reads agree 100% before trusting either, so I could map Lego bricks I'd dropped on the floor. The dated public record starts on 24 June 2026, in the Sonny Test's sealed history.
  • "Walking a few weeks ago and landing on Mars now": opinion. It was about how fast I could work with AI, not a result. I'm not claiming to be first.
  • The chain of agents and the forced yes or no: not shown yet in public. The designs and results are sealed with timestamps, and they'll be published with their files. Until then, read those numbers as not shown yet. The counts are in AI in Human Words.
  • Claude's "the provider's own records matched our run's receipts in every case we sampled": not shown yet. A sample of 538 calls was checked against OpenRouter's records on 6 October; that check isn't public yet.
  • "The filter on your model sometimes blocked the whole status report": not shown yet. Seen on 6 October 2026; there's no public example yet.
  • The photo deletion and the copied answer key: named. Both were Google's Gemini, not Claude.
  • The sandbox that kept re-granting itself access: named. It was OpenAI's Codex, not Claude Code.
  • "I've sealed and timestamped everything I've found": partly. Most designs and results are sealed with FreeTSA and Bitcoin timestamps. Some sealed files stay private, and the 5 and 6 October runs aren't listed publicly yet.
  • Claude's six answers: not evidence. A model working under my rules will tend to agree with me. Claude's last answer proposes the test that would settle it: ask people the same checkable questions in two styles, and see whether their trust tracks what's correct. That test hasn't been run.

Find the work

Changes to this page

  • 6 October 2026: I removed a stray "]" after "my findings" in my answer about who has seen or used what I built. It was a transcription error. The sealed file keeps it, so the file still matches its fingerprint.