Claude Sonnet 4.5
Verdict pending.
The retest.
The earlier Claude score was retired as inflated, so this flagship run re-evaluates Claude Sonnet 4.5 from scratch against the current standard. Nothing on this page is a published result.
Verdict pending.
The recommendation lands here once the full testing period is logged and scored.
How the score will be built.
Each metric is scored 1–10 and multiplied by its weight. Scores populate once all four phases are complete — no partial scores.
Four phases, same rig.
The same four-phase gauntlet every tool runs. Findings drop into each row as the phase is completed and logged.
First-contact quality, setup friction, and trust calibration under real onboarding conditions.
Edge cases, failure modes, and reliability under real-world adversarial pressure.
Findings drop in as the phase is completed and logged.
The score, band, and recommendation publish here once all four phases are complete and logged. No partial verdicts, no early calls — that's the whole point of the rig.
What the data will show.
Every number here will trace to a logged, timestamped observation — nothing reconstructed after the fact.
Instrument readouts populate once the run is complete.
Sole tester · CTAI
Verdict pending. The score, band, and recommendation publish here once all four phases are complete and logged. No partial verdicts, no early calls — that's the whole point of the rig.