REAL-WORLD AI TESTING COMMISSION A TEST
Lab Note Method

Who Checks the Checker?

Automated AI audits are arriving fast, and a lot of them use AI to judge AI. Here's why CTAI keeps a human at the point of verdict.

A new kind of product is showing up: independent audits for AI agents. Connect your agent, pick your inspections, and get a report on where it misbehaves. Plenty of them are well built, and the problem they're chasing is real. Most of them also share one design choice — the judging is done by other AI models.

I think that's worth a moment of attention. Not because automated audits are bad, but because it makes the question I built CTAI around more pressing, not less: when AI is checking AI, who checks the checker?

The claim is the starting point. The test determines what the evidence supports. That applies to AI tools — and to the systems that test them.

Two different jobs

It's easy to lump anything with "independent" and "trust your AI" on the homepage into one category. They're doing different work, for different people, and it helps to be clear about which is which.

Automated audit

Scale and coverage

Machine-run · machine-judged

Hundreds of scripted inspections against an agent your team built, run in a simulated environment. Fast, repeatable, and good at catching known categories of failure. Built for engineers and risk teams.

Crash Test

Real work, human verdict

Human-led · evidence-first

An AI tool someone is thinking of relying on, put to work in realistic conditions over time. I apply the pressure, capture what happens, and make the call. Built for people deciding whether to trust it.

Why the verdict stays human

A model judging another model inherits the blind spots it shares with it. If both were trained to find the same kind of answer convincing, a confident, plausible failure can sail through. And the failure mode I worry about most isn't the tool that breaks loudly. It's the one that produces something polished, wrong, and easy to believe.

That's the kind of failure you catch by using a tool for real work — noticing that the citation doesn't say what the summary claims, that the "finished" task quietly skipped a step, that the output reads well and still wouldn't survive contact with a client. It takes someone with context, time on the tool, and nothing to gain from the result.

So CTAI has a simple rule: AI doesn't test AI here. When I use tools to help run a test, they're equipment — they execute and they capture. The judgement is mine, and each report will say which equipment was used.

Side by side, not instead of

None of this is an argument against automated audits. If you build your own agents, scripted coverage at scale is useful, and a human alone can't match it. But coverage isn't the same as a verdict, and a passing audit isn't the same as knowing a tool holds up once it's doing your actual work.

That gap — between what a system claims, what it passes, and what really happens when you put it to work — is where CTAI lives. The crash-test starts where the demo ends. As more of the checking gets automated, I think it matters more, not less, that someone is still standing at the end of the process saying what the evidence supports.

Written by
Woody Woody
Woody
Founder · Sole tester · CTAI

Solo operator and the Crash Test Dummy behind CTAI. I put AI to work. I apply pressure. I document what happens — no sponsors, no hype.

Read my full story →
Keep reading
01
How I test AI tools
Four phases, one evidence chain: claim, test, evidence, analysis, verdict.
02
The Score I Took Down
Why CTAI retired its own highest score — and put the reason on the record.