REAL-WORLD AI TESTING COMMISSION A TEST
Lab Note Editorial layer

One Platform, Two Layers

Most review sites pick a lane — data or opinion. CTAI runs both, on purpose.

CTAI was never designed as two things. It started as one — a place to publish test results. But the results on their own were incomplete. A score tells you what happened. It does not tell you why it matters, what I noticed along the way, or what changed in how I think about the tool category after two weeks of daily use.

That second layer — the thinking, the founder perspective, the part that is harder to structure — turned out to be the part that made the testing credible. Not because opinion validates data, but because transparency about process is what separates a lab from a leaderboard.

The two layers

Layer 1 Structure

Crash Tests

Structured evaluations. Four-phase methodology, scoring rubric, repeatable and auditable — built to surface friction under pressure. The evidence layer: what happened, measured against a standard.

Layer 2 Perspective

Lab Notes

Lighter, readable, founder-led. The context behind the tests — what I learned, what patterns keep repeating, what the scores alone cannot capture. The editorial layer.

Layer 1 is engineered. It follows rules. Every crash test runs the same phases, uses the same rubric, produces the same artefact structure. That rigidity is the whole point — it is what makes the results comparable and the methodology defensible.

Layer 2 is lighter. It follows instinct. A lab note starts with an observation, not a hypothesis. The writing is personal because the perspective is personal. Nobody else is running these tests this way, so nobody else can write this context.

What breaks without one

Strip the lab notes and CTAI becomes a database. Useful, maybe. But the reader has no reason to trust the person behind the scores. Every AI review site has numbers. The question is always: who produced them, and do they understand what they measured?

Strip the crash tests and CTAI becomes a blog. Interesting, maybe. But the writing has no structural authority underneath it. Opinions without methodology are just content, and the internet has enough of that.

Testing without editorial is a spreadsheet. Editorial without testing is a blog. The two layers are what make this an operating system.

The feedback loop

The layers are not independent. They pressure each other. Writing lab notes forces me to articulate what I actually learned from a crash test — not just the score, but what shifted in my understanding. If I cannot write clearly about what a test revealed, the test was not thorough enough.

Running crash tests forces the editorial to stay grounded. I cannot drift into abstract thinking about AI when I have two weeks of structured data sitting next to the essay. The data keeps the writing honest; the writing keeps the data meaningful.

Every test shapes the next test

This is still how both sides evolve. The test produces the evidence. The note decodes what it meant. And the learning shapes what gets tested next. The two layers grow together — each one raising the standard the other has to meet.

Written by
Woody Woody
Woody
Founder · Sole tester · CTAI

Solo operator and the Crash Test Dummy behind CTAI. I put AI to work. I apply pressure. I document what happens — no sponsors, no hype.

Read my full story →
Keep reading
◆ In the rig nowClaude Sonnet 4.5— flagship retest in progress
01
Perplexity AI
Fast research, shaky sourcing. Useful tool, wrong mental model.
71
Mixed
02
Notion AI
Convenient but shallow. You're paying for proximity, not capability.
58
Mixed