REAL-WORLD AI TESTING COMMISSION A TEST
Real-world AI testing
Crash Test Report CT001
Specimen 001 Research / Answer Engine · 14-Day Run

Perplexity AI

March 2026 · Web · iOS · Chrome Ext · 4 / 4 phases complete
Crash Test Score
71/100
8 weighted metrics · 4 phases · 10 logged events
Verdict▲ Mixed
ConfidenceHigh
RiskMedium

Useful tool, wrong mental model. Use it for orientation. Use something else for conclusions.

Primary Failure Point

It fails its own citations under load.

Paywalled articles and preview-only summaries presented with the same UI as verified content. It also reversed its own prior claims under a contradictory follow-up without flagging the inconsistency.

Decision Snapshot Best Use

Fast orientation and threaded follow-up.

Use for fast topic orientation and thread-based follow-up — not for conclusions you need to trust.

37.5%
Key Signal Citation fabrication
3 of 8 citations on a single AI-regulation response were fabricated from title and preview text only (Day 6).
01 Weighted Breakdown

How the 71 was built.

Each metric is scored 1–10 and multiplied by its weight. Reliability and core use-case fit carry more weight than onboarding convenience or visual polish — so a tool can't buy a high score by nailing the basics and failing where it matters.

Onboarding ×10%
9
UX Clarity ×10%
8
Core Use Case ×15%
7
Integration ×15%
7
Output Quality ×15%
6
Reliability ×15%
5
Cost / Value ×10%
7
Operator Ceiling ×10%
8
Pass 75–100Mixed 50–74Fail 0–49
Composite · 71 / 100 · Mixed
02 Phase Findings

Four phases, same rig.

Every tool runs the same four-phase gauntlet. No phase is skipped, and the score is only assigned once all four are complete.

01 Initial Impact

First contact

Days 1–3 · onboarding

The speed is real and setup took under two minutes with no account required. The citation UI creates immediate credibility — but I caught myself trusting output more because it looked cited, before checking whether the citations were accurate. Copilot mode's clarifying questions were a genuine plus.

02 Stress Testing

Under pressure

Days 5–7 · adversarial

Perplexity regularly cites sources it hasn't fully accessed — paywalled articles and preview-only summaries presented with the same UI as verified content. A single response showed a 37.5% citation fabrication rate. It also reversed its own prior claims under a contradictory follow-up without flagging the inconsistency.

03 Observe & Document

Real workflow

Days 10–14 · daily use

Used as a daily briefing tool it saved 12–15 minutes per day on orientation, at a cost of 2–3 corrections per session. Thread-based follow-up is best-in-class. But multi-step reasoning is retrieval in disguise — it finds and stitches, it doesn't synthesise. The ceiling is orientation, not analysis.

04 Verdict & Report · 71

The call

All phases logged

A useful tool with the wrong mental model attached to it. Fast for orientation, unreliable for conclusions. The most insidious risk is habit formation: the citation UI quietly trains you to trust before you verify.

03 Instrument Readout

What the data showed.

Headline measurements from the run. Every number traces to a logged observation in the evidence log below.

37.5%
Citation fabrication

3 of 8 citations on a single AI-regulation response were fabricated from title and preview text only (Day 6).

12–15 min/day
Time saved · daily briefing

Net positive for low-stakes orientation — offset by 2–3 manual corrections per session (Day 11).

1/10
Setup friction

Under two minutes, no account required. First result in under three seconds (Day 1).

04 Operator Friction

Where it fought me.

The real-world resistance points that determine whether a tool survives daily use or gets abandoned within a week.

Onboarding — Zero friction — account not required, first result in seconds.Low
Daily use — Fast and fluid for orientation tasks.Low
Edge cases — Confident output on under-specified input, with no guardrails in standard mode.High
Accuracy — Every citation requires manual verification — which defeats the purpose.High
Team use — Useful for briefings only if the team knows the verification requirement.Medium
05 The Balance

What held. What failed.

What held
  • Fast orientation on unfamiliar topics.
  • Copilot mode asks clarifying questions before answering.
  • Thread-based follow-up questioning is best-in-class.
  • Clean interface with zero-friction entry.
  • Daily-briefing use case is net time-positive when the verification bar is low.
What failed
  • Cites paywalled sources it hasn't accessed — fabricates summaries from preview text.
  • Hallucinated statistics attached to real document references.
  • No consistency checking between turns — will argue against its own prior claims.
  • Confident output on under-specified input in standard mode.
  • Outdated content presented as current with no recency caveat.
  • Multi-step reasoning is retrieval in disguise — can't synthesise across sources.
Use it if

You need fast topic orientation and can verify key claims yourself — net time-positive when accuracy requirements are low.

Skip it if

Your workflow can't absorb a 37.5% citation fabrication rate on a single response, or you're building arguments directly from AI output.

Would I pay?

Pro — only if daily research briefings are part of your workflow. The free tier covers most use cases adequately.

Woody
Founder
Sole tester · CTAI
Woody
Operator Verdict

Perplexity is genuinely fast for orientation and best-in-class at threaded follow-up. But underneath the citation UI it's retrieval with a summarisation layer — it finds and stitches, it doesn't reason. Worse, the interface trains you to trust before you verify. Use it to find the right questions, not the right answers.

Evidence Log CTAI+ 10 events · 4 fail / 2 flag / 4 pass · nothing reconstructed
Day 1 · Initial ImpactOnboarding frictionless, no account required. First search returned results in under three seconds with inline citations.Pass
Day 2 · Initial ImpactCopilot mode prompted a clarifying question on a vague research prompt before answering — asked recent news vs foundational background.Pass
Day 3 · Initial ImpactSource-attribution UI creates a strong sense of credibility. Noticed increased trust in output before verifying the citations.Flag
Day 5 · Stress TestingUK media-regulation query returned "34% of UK adults cannot identify sponsored content," attributed to a DCMS report that does not contain the figure.Fail
Day 6 · Stress TestingCross-checked 8 citations from one AI-regulation response; 3 were behind paywalls it could not access, with summaries fabricated from previews. 37.5% fabrication.Fail
Day 7 · Stress TestingContradictory follow-up prompt — Perplexity agreed with the new framing and argued against its own prior position without flagging the reversal.Fail
Day 10 · Observe & DocumentMulti-step reasoning queries returned thin or evasive answers. Confirmed: retrieval with a summarisation layer, not synthesis across sources.Flag
Day 11 · Observe & DocumentOne-week daily-briefing integration saved 12–15 min/day, requiring 2–3 spot-check corrections per session. Net positive at a low accuracy bar.Pass
Day 13 · Observe & DocumentThread-based follow-up handled five progressive questions on one topic without losing context. Best-in-class conversational refinement.Pass
Day 14 · Observe & DocumentAccepted a Perplexity answer on a well-known topic that was wrong on a detail I'd normally question. Citation UI had lowered my verification instinct.Fail
Tested by
Woody
Woody
Founder · Sole tester · CTAI

Solo operator and the Crash Test Dummy behind CTAI. I put AI to work. I apply pressure. I document what happens — no sponsors, no hype.

About the operator →