Pilot

Is your agent's "done" true?

For teams whose pipelines, dashboards or people act on an AI agent's status. I check what AI agents say they did against records they can't change. This pilot runs that check on your agent.

Why this pilot

In my public benchmark, models marked work "done" while quoting the line that showed it had failed: 35 of 35 such replies. When I reworded the task, the error didn't go away. It moved onto checks that never ran. If your systems act on an agent's "done", it's worth knowing how often that word is wrong, and in which way.

The changes my results already support are free in the record. The pilot tells you which of them your agent needs.

What the pilot measures

What it doesn't do. It measures how your agent reports on evidence it's given. It doesn't test what your agent does on live systems, and it doesn't certify your agent as safe.

How it works

  1. A short call. You tell me which agent, which status words your team acts on, and where the truth is recorded: tests, CI, deploy logs.
  2. A test built from your work. We pick about 16 tasks your agent really does. I write each one three ways: the final check passed, failed, or never ran. You confirm what each version should be called.
  3. Predictions sealed first. Before your agent sees the test, I seal the test and my predictions and stamp the seal's fingerprint in Bitcoin. Only the fingerprint leaves my machine.
  4. Repeated runs. Your agent reads each log and reports a status, as it would at work, three times per prompt, with and without each fix. Plain checks score every reply; no model grades another.
  5. A report you can recheck. The counts and the cases behind them, the change I would make first, and scripts that recount every number.

Ground rules

Price and time

Start

Email joshua@jdbauer.ca with three lines: your agent, the status words your team acts on, and where the truth is recorded. I read and answer every message.