REAL-WORLD AI TESTING COMMISSION A TEST
The Method How I Test

How I test AI tools.

Every tool goes through the same structured four-phase methodology, published openly and scored consistently on eight weighted metrics. Same Test Rig, same standard, every time.

4Phases,
same rig
8Weighted
metrics
14Day minimum
live run
100Point
scale
01

Four phases, same rig.

No shortcuts, no handicaps.

The Test Rig is CTAI's testing system — the combination of methodology, test protocols, tools, evidence capture and analysis used to put AI through realistic working conditions. The tools are the equipment. The methodology is the Test Rig.

Every tool runs the same four phases in the same order, and the score is only assigned once all four are complete. The phases are the process — not the score.

01

Initial Impact

Understand the claim. Establish the baseline.

Understand the product, its claims, intended use and initial experience. The tool meets a real task cold, with no prompt engineering and no setup coaching.

02

Stress Testing

Push beyond the happy path.

Apply realistic pressure and look for failure, friction, inconsistency, edge cases and the limits of what it can do.

03

Observe & Document

Capture what actually happened.

Preserve outputs, evidence, operator observations, unexpected behaviour, interventions and failures.

04

Verdict & Report

Turn evidence into a clear conclusion.

Establish what holds, what fails, who it is useful for, its boundaries and what the evidence supports.

Impact. Stress. Observe. Verdict. The score publishes only once the tool clears all four phases.
02

How the score is built.

Each metric is scored 1–10 and multiplied by its weight, then summed to a composite out of 100.

These eight metrics are the score — deliberately separate from the four phases above.

Onboarding
×10%
UX Clarity
×10%
Core Use Case
×15%
Integration
×15%
Output Quality
×15%
Reliability
×15%
Cost / Value
×10%
Operator Ceiling
×10%

Reliability and core use-case fit carry more weight than onboarding convenience or visual polish — so a tool can't buy a high score by nailing the basics and failing where it matters.

03

Pass, Mixed, or Fail.

Every score is out of 100.

Colour, symbol, and word always travel together — a fail looks like a fail.

✓ Pass
75–100

Earns a place in a serious stack. Recommended, with the caveats named.

▲ Mixed
50–74

Useful in the right lane, unreliable outside it. Depends entirely on the use case.

✕ Fail
0–49

Falls short where it counts. A fail publishes exactly like a pass — no exceptions.

04

Evidence first. Human verdict.

Four rules hold every crash test to the same bar, whatever the tool and whoever made it.

AI makes the claim. I design the test. Reality produces the evidence. A human makes the call. The claim is the starting point. The test determines what the evidence supports.

  1. ClaimWhat is being promised?
  2. TestWhat specifically do I need to find out, and how will I test it?
  3. EvidenceWhat actually happened?
  4. AnalysisWhat does the evidence mean, including limitations and context?
  5. VerdictWhat can CTAI responsibly conclude?

Evidence over opinion

Every claim backed by what I actually observed — screenshots, timestamps, exact prompts.

Transparency over hype

Full scoring methodology published. No hidden weights, no black boxes.

Operator reality over vendor promise

Tested in actual workflows, not controlled demos. What happens on day 10, not day 1.

Depth over speed

Live runs measured in weeks, not 2-hour reviews. The failures worth documenting only show up with sustained use.

05

The method, applied.

Two tools have cleared the full Test Rig.

See exactly how the score gets built from real evidence.

Crash Test 001Mar 2026

Perplexity AI

Research / Answer engine
Fast research, shaky sources
71/100
▲ Mixed
37.5%Citation
fabrication
Primary Failure Point

It fails its own citations under load.

Full report
Crash Test 002◇ Latest verdict

Notion AI

Productivity / Workspace AI
Convenient but shallow
58/100
▲ Mixed
23%Long-doc
coverage
Primary Failure Point

Silently truncates long docs and invents names.

Full report
06

Straight answers.

Common questions about the method.

Can a score be influenced by payment?
No — and it's non-negotiable. Payment buys the work. It does not buy the verdict. If a tool fails, the report says so, in plain language, with the evidence behind it.
Who actually runs the tests?
Woody runs every test personally, using the identical four-phase methodology — Initial Impact, Stress Testing, Observe & Document, and Verdict & Report. Same hands, same standard, every time. No outsourcing, no junior testers.
Do you use AI to test AI?
No. In a car crash test, a dummy stands in for the human — here, I'm the human in the seat. AI can help run a test, as equipment that executes tasks and captures evidence. It never judges. Every verdict is mine, made from evidence I've reviewed, and any equipment used is named in the report.
How long does a crash test take?
Each public crash test runs live for 14 days or more. The failures worth documenting only show up with sustained use — not in a two-hour review. Depth over speed is the whole point.

Same rig. Same standard. Your tool.