Every tool goes through the same structured four-phase methodology, published openly and scored consistently on eight weighted metrics. Same Test Rig, same standard, every time.
No shortcuts, no handicaps.
The Test Rig is CTAI's testing system — the combination of methodology, test protocols, tools, evidence capture and analysis used to put AI through realistic working conditions. The tools are the equipment. The methodology is the Test Rig.
Every tool runs the same four phases in the same order, and the score is only assigned once all four are complete. The phases are the process — not the score.
Understand the product, its claims, intended use and initial experience. The tool meets a real task cold, with no prompt engineering and no setup coaching.
Apply realistic pressure and look for failure, friction, inconsistency, edge cases and the limits of what it can do.
Preserve outputs, evidence, operator observations, unexpected behaviour, interventions and failures.
Establish what holds, what fails, who it is useful for, its boundaries and what the evidence supports.
Each metric is scored 1–10 and multiplied by its weight, then summed to a composite out of 100.
These eight metrics are the score — deliberately separate from the four phases above.
Reliability and core use-case fit carry more weight than onboarding convenience or visual polish — so a tool can't buy a high score by nailing the basics and failing where it matters.
Every score is out of 100.
Colour, symbol, and word always travel together — a fail looks like a fail.
Earns a place in a serious stack. Recommended, with the caveats named.
Useful in the right lane, unreliable outside it. Depends entirely on the use case.
Falls short where it counts. A fail publishes exactly like a pass — no exceptions.
Four rules hold every crash test to the same bar, whatever the tool and whoever made it.
AI makes the claim. I design the test. Reality produces the evidence. A human makes the call. The claim is the starting point. The test determines what the evidence supports.
Every claim backed by what I actually observed — screenshots, timestamps, exact prompts.
Full scoring methodology published. No hidden weights, no black boxes.
Tested in actual workflows, not controlled demos. What happens on day 10, not day 1.
Live runs measured in weeks, not 2-hour reviews. The failures worth documenting only show up with sustained use.
Two tools have cleared the full Test Rig.
See exactly how the score gets built from real evidence.
It fails its own citations under load.
Silently truncates long docs and invents names.
Common questions about the method.