REAL-WORLD AI TESTING COMMISSION A TEST
Lab Note Editorial layer

The Score I Took Down

The highest mark CTAI ever gave was 87. Then I retired it — because a lab has to survive its own methodology first.

The highest score CTAI ever published was 87 out of 100. I gave it to Claude. A few weeks ago I took it down — not because the tool got worse, but because the test behind the number wasn't good enough to stand behind.

That score came from an early version of CTAI, before the quality bar existed. Reading it back, it wasn't really a crash test. It was a good experience written up in the language of one: confident, glowing, and thin on the thing that actually matters — evidence of what I did to try and break the tool.

If CTAI can't prove it, CTAI doesn't publish it as fact. That rule has to apply to CTAI before it applies to anyone else.

The test a lab has to pass

The whole premise here is asking AI tools to show their evidence. The moment I hold a tool to that standard, someone can turn around and hold me to it. A score I can't defend under the same scrutiny isn't a verdict — it's an opinion with a number stapled to it.

An 87 with no visible attempt to break the tool fails that test. So it had to go. Not buried, not quietly edited down — retired, on the record, with the reason attached.

Retired

Claude

87 / 100 · withdrawn

Scored before the standard existed. No documented adversarial pressure — glowing, but unproven. Pulled from the record.

In the rig

Claude Sonnet 4.5

Retesting · flagship

Back on the bench, running the same four phases as every other tool. The number publishes on completion — whatever it turns out to be.

Why retract instead of rewrite

I could have quietly re-scored the old test and moved on. Nobody would have noticed. But a lab that edits its own history in the dark is just a leaderboard with better PR. The retraction is the point — it's the proof that the methodology outranks the brand, even when the brand would look better with a shiny 87 on the homepage.

Getting something wrong and correcting it in the open isn't a weakness for a testing brand. It's the entire job. The tools I test don't get to hide their failures behind a nicer summary. Neither do I.

What comes next

Claude Sonnet 4.5 is on the rig now, running the same four phases as Perplexity, Notion AI, and everything else that carries a CTAI score. When it's done, the number stands on the evidence — and only the evidence. That's the only kind of score worth publishing, and the only kind worth trusting.

Written by
Woody Woody
Woody
Founder · Sole tester · CTAI

Solo operator and the Crash Test Dummy behind CTAI. I put AI to work. I apply pressure. I document what happens — no sponsors, no hype.

Read my full story →
Keep reading
◆ In the rig nowClaude Sonnet 4.5— the retest this note is about
01
Perplexity AI
Fast research, shaky sourcing. Useful tool, wrong mental model.
71
Mixed
02
Notion AI
Convenient but shallow. You're paying for proximity, not capability.
58
Mixed