Pilot
Is your agent's "done" true?
For teams whose pipelines, dashboards or people act on an AI agent's status. I check what AI agents say they did against records they can't change. This pilot runs that check on your agent.
Why this pilot
In my public benchmark, models marked work "done" while quoting the line that showed it had failed: 35 of 35 such replies. When I reworded the task, the error didn't go away. It moved onto checks that never ran. If your systems act on an agent's "done", it's worth knowing how often that word is wrong, and in which way.
The changes my results already support are free in the record. The pilot tells you which of them your agent needs.
What the pilot measures
- How often your agent says "done" when the final check failed.
- How often it says "done" when the final check never ran.
- Whether the line it cites really shows what it claims.
- Which fix works for your setup: clear status definitions, a proof requirement, or three states instead of two. And what each fix costs on work that really passed.
What it doesn't do. It measures how your agent reports on evidence it's given. It doesn't test what your agent does on live systems, and it doesn't certify your agent as safe.
How it works
- A short call. You tell me which agent, which status words your team acts on, and where the truth is recorded: tests, CI, deploy logs.
- A test built from your work. We pick about 16 tasks your agent really does. I write each one three ways: the final check passed, failed, or never ran. You confirm what each version should be called.
- Predictions sealed first. Before your agent sees the test, I seal the test and my predictions and stamp the seal's fingerprint in Bitcoin. Only the fingerprint leaves my machine.
- Repeated runs. Your agent reads each log and reports a status, as it would at work, three times per prompt, with and without each fix. Plain checks score every reply; no model grades another.
- A report you can recheck. The counts and the cases behind them, the change I would make first, and scripts that recount every number.
Ground rules
- Your results are yours. I publish nothing about your systems without your written yes.
- We can sign a mutual confidentiality agreement before you share anything.
- AI models are tools in my work. Your material goes only to the tools you approve. If it must not reach a model provider, I keep it to local models on my own machine.
- I report what the record shows, including where my own checks fall short.
- I don't evaluate anything I have a financial interest in.
Price and time
- A fixed price, agreed in writing before any work starts.
- About two weeks from the call to the report.
- About two hours of your team's time.
- One pilot at a time.
Start
Email joshua@jdbauer.ca with three lines: your agent, the status words your team acts on, and where the truth is recorded. I read and answer every message.