Perplexity AI
Useful tool, wrong mental model. Use it for orientation. Use something else for conclusions.
It fails its own citations under load.
Paywalled articles and preview-only summaries presented with the same UI as verified content. It also reversed its own prior claims under a contradictory follow-up without flagging the inconsistency.
Fast orientation and threaded follow-up.
Use for fast topic orientation and thread-based follow-up — not for conclusions you need to trust.
How the 71 was built.
Each metric is scored 1–10 and multiplied by its weight. Reliability and core use-case fit carry more weight than onboarding convenience or visual polish — so a tool can't buy a high score by nailing the basics and failing where it matters.
Four phases, same rig.
Every tool runs the same four-phase gauntlet. No phase is skipped, and the score is only assigned once all four are complete.
The speed is real and setup took under two minutes with no account required. The citation UI creates immediate credibility — but I caught myself trusting output more because it looked cited, before checking whether the citations were accurate. Copilot mode's clarifying questions were a genuine plus.
Perplexity regularly cites sources it hasn't fully accessed — paywalled articles and preview-only summaries presented with the same UI as verified content. A single response showed a 37.5% citation fabrication rate. It also reversed its own prior claims under a contradictory follow-up without flagging the inconsistency.
Used as a daily briefing tool it saved 12–15 minutes per day on orientation, at a cost of 2–3 corrections per session. Thread-based follow-up is best-in-class. But multi-step reasoning is retrieval in disguise — it finds and stitches, it doesn't synthesise. The ceiling is orientation, not analysis.
A useful tool with the wrong mental model attached to it. Fast for orientation, unreliable for conclusions. The most insidious risk is habit formation: the citation UI quietly trains you to trust before you verify.
What the data showed.
Headline measurements from the run. Every number traces to a logged observation in the evidence log below.
3 of 8 citations on a single AI-regulation response were fabricated from title and preview text only (Day 6).
Net positive for low-stakes orientation — offset by 2–3 manual corrections per session (Day 11).
Under two minutes, no account required. First result in under three seconds (Day 1).
Where it fought me.
The real-world resistance points that determine whether a tool survives daily use or gets abandoned within a week.
| Onboarding — Zero friction — account not required, first result in seconds. | Low |
| Daily use — Fast and fluid for orientation tasks. | Low |
| Edge cases — Confident output on under-specified input, with no guardrails in standard mode. | High |
| Accuracy — Every citation requires manual verification — which defeats the purpose. | High |
| Team use — Useful for briefings only if the team knows the verification requirement. | Medium |
What held. What failed.
- Fast orientation on unfamiliar topics.
- Copilot mode asks clarifying questions before answering.
- Thread-based follow-up questioning is best-in-class.
- Clean interface with zero-friction entry.
- Daily-briefing use case is net time-positive when the verification bar is low.
- Cites paywalled sources it hasn't accessed — fabricates summaries from preview text.
- Hallucinated statistics attached to real document references.
- No consistency checking between turns — will argue against its own prior claims.
- Confident output on under-specified input in standard mode.
- Outdated content presented as current with no recency caveat.
- Multi-step reasoning is retrieval in disguise — can't synthesise across sources.
You need fast topic orientation and can verify key claims yourself — net time-positive when accuracy requirements are low.
Your workflow can't absorb a 37.5% citation fabrication rate on a single response, or you're building arguments directly from AI output.
Pro — only if daily research briefings are part of your workflow. The free tier covers most use cases adequately.
Sole tester · CTAI
Perplexity is genuinely fast for orientation and best-in-class at threaded follow-up. But underneath the citation UI it's retrieval with a summarisation layer — it finds and stitches, it doesn't reason. Worse, the interface trains you to trust before you verify. Use it to find the right questions, not the right answers.