REAL-WORLD AI TESTING COMMISSION A TEST
Real-world AI testing
Crash Test Report CT003
Specimen 003 LLM / Reasoning · Flagship Retest

Claude Sonnet 4.5

In the rig · 0 / 4 phases complete
Crash Test Score
—/100
8 weighted metrics · 4 phases · 0 logged events
Verdict◆ In the rig
Confidence—
Risk—

Verdict pending.

Status In the rig

The retest.

The earlier Claude score was retired as inflated, so this flagship run re-evaluates Claude Sonnet 4.5 from scratch against the current standard. Nothing on this page is a published result.

Decision Snapshot Best Use

Verdict pending.

The recommendation lands here once the full testing period is logged and scored.

—
Key Signal Pending
Instrument readouts populate once the run is complete.
01 Weighted Breakdown

How the score will be built.

Each metric is scored 1–10 and multiplied by its weight. Scores populate once all four phases are complete — no partial scores.

Onboarding ×10%
—
UX Clarity ×10%
—
Core Use Case ×15%
—
Integration ×15%
—
Output Quality ×15%
—
Reliability ×15%
—
Cost / Value ×10%
—
Operator Ceiling ×10%
—
Pass 75–100Mixed 50–74Fail 0–49
Composite · — / 100 · Pending
02 Phase Findings

Four phases, same rig.

The same four-phase gauntlet every tool runs. Findings drop into each row as the phase is completed and logged.

01 Initial Impact

First contact

Awaiting run

First-contact quality, setup friction, and trust calibration under real onboarding conditions.

02 Stress Testing

Under pressure

Awaiting run

Edge cases, failure modes, and reliability under real-world adversarial pressure.

03 Observe & Document

Real workflow

Awaiting run

Findings drop in as the phase is completed and logged.

04 Verdict & Report

The call

Awaiting run

The score, band, and recommendation publish here once all four phases are complete and logged. No partial verdicts, no early calls — that's the whole point of the rig.

03 Instrument Readout

What the data will show.

Every number here will trace to a logged, timestamped observation — nothing reconstructed after the fact.

—
Pending

Instrument readouts populate once the run is complete.

Woody
Founder
Sole tester · CTAI
Woody
Operator Verdict

Verdict pending. The score, band, and recommendation publish here once all four phases are complete and logged. No partial verdicts, no early calls — that's the whole point of the rig.

Evidence Log CTAI+ 0 events · the log fills as the run happens

No observations logged yet — this test is in the rig.

Tested by
Woody
Woody
Founder · Sole tester · CTAI

Solo operator and the Crash Test Dummy behind CTAI. I put AI to work. I apply pressure. I document what happens — no sponsors, no hype.

About the operator →