REAL-WORLD AI TESTING COMMISSION A TEST
Real-world AI testing
Crash Test Report CT002
Specimen 002 Productivity / Workspace AI · 3-Week Run

Notion AI

March 2026 · Web · Desktop · iOS · Android · 4 / 4 phases complete
Crash Test Score
58/100
8 weighted metrics · 4 phases · 9 logged events
Verdict▲ Mixed
ConfidenceHigh
RiskMedium

The integration is real, the capability ceiling is low, and the tool will not tell you when you've hit it.

Primary Failure Point

It fails quietly on anything deeper.

Truncating long documents without warning, inventing names in meeting drafts, and flattening every voice it edits — and it sounds complete when it isn't.

Decision Snapshot Best Use

Formatting, boilerplate, low-stakes drafts.

Use for formatting, boilerplate, and low-stakes drafts. Switch tools the moment the task needs depth.

23%
Key Signal Long-doc coverage
A 6,000-word strategy doc summarised to ~1,400 words — the opening only, with no truncation warning (Day 7).
01 Weighted Breakdown

How the 58 was built.

Each metric is scored 1–10 and multiplied by its weight. Reliability and core use-case fit carry more weight than onboarding convenience or visual polish — which is exactly where Notion AI's surface strength stops paying off.

Onboarding ×10%
9
UX Clarity ×10%
8
Core Use Case ×15%
5
Integration ×15%
8
Output Quality ×15%
4
Reliability ×15%
4
Cost / Value ×10%
5
Operator Ceiling ×10%
5
Pass 75–100Mixed 50–74Fail 0–49
Composite · 58 / 100 · Mixed
02 Phase Findings

Four phases, same rig.

Every tool runs the same four-phase gauntlet. No phase is skipped, and the score is only assigned once all four are complete.

01 Initial Impact

First contact

Days 1–3 · onboarding

Zero-friction entry, and it's genuinely true — one click from an existing workspace, no account, no new interface, spacebar surfaces the panel. Fastest activation of any tool tested. First outputs looked polished. But "Improve writing" on my own copy already hinted at the pattern: it looked better and read worse — the voice was gone.

02 Stress Testing

Under pressure

Days 7–10 · adversarial

The gap between surface polish and real capability. A 6,000-word document summarised to only its opening ~1,400 words with no truncation warning. A meeting-agenda draft invented two attendee names not in the workspace. "Improve writing" flattened voice 0/5. The score of 58 reflects a consistent pattern, not isolated failures.

03 Observe & Document

Real workflow

Weeks 2–3 · daily use

Deployed across three client workspaces; it stayed in rotation for one (low-stakes internal ops) and was quietly dropped from the other two. Reliable for reformatting notes and boilerplate first drafts, nothing beyond. The insight: the convenience premium is for proximity, not capability — you pay £10/member/month to not switch tabs.

04 Verdict & Report · 58

The call

All phases logged

The integration is real, the capability ceiling is low, and the tool will not tell you when you've hit it. Fine for low-stakes drafting; misleading on anything that needs depth, because it sounds complete when it isn't.

03 Instrument Readout

What the data showed.

Headline measurements from the run. Every number traces to a logged observation in the evidence log below.

0/5
Voice preservation

Across five "Improve writing" samples, every output lost its voice — contractions removed, rhythm flattened (Day 10).

23%
Long-doc coverage

A 6,000-word strategy doc summarised to ~1,400 words — the opening only, with no truncation warning (Day 7).

1/10
Setup friction

One-click activation inside the workspace — no account, no config. Fastest of any tool tested (Day 1).

04 Operator Friction

Where adoption slows down.

The real-world resistance points that determine whether a tool survives daily use or gets abandoned within a week.

Onboarding — One-click activation, no account or config — fastest entry of any tool tested. Only rises slightly on long-document handling.Low
Output quality — Polished on the surface, thin underneath — degrades sharply on edge cases and long documents.High
Context retention — Loses the thread on documents over ~3,000 words, silently and with no warning.High
Voice preservation — "Improve writing" flattens voice on every sample — 0/5 preserved.High
05 The Balance

What held. What failed.

What held
  • Zero friction — already inside your workflow.
  • Short-document summarisation (under ~3,000 words).
  • Boilerplate and template generation.
  • Reformatting unstructured notes into structured prose.
  • First drafts of low-stakes internal documents.
What failed
  • Document truncation on content over ~3,000 words.
  • Hallucinated content in meeting and agenda drafts.
  • Destroys voice on every "Improve writing" task.
  • No instruction-conflict detection.
  • Cannot reason across linked databases.
  • Won't flag when it's out of its depth.
Use it if

Your team already lives in Notion and needs formatting, boilerplate, and short-document summaries under 3,000 words.

Skip it if

You work with documents over 3,000 words — summaries silently truncate — or you need your writing voice preserved (0/5).

Would I pay?

Only if your team lives in Notion all day. You're paying for proximity, not capability — worth it for workflow convenience, not for reasoning-heavy work.

Woody
Founder
Sole tester · CTAI
Woody
Operator Verdict

Notion AI's zero-friction integration is genuinely valuable for formatting, boilerplate, and short summaries. But it fails quietly on anything deeper — truncating long documents without warning, inventing names in meeting drafts, and flattening every voice it edits. You're paying for proximity, not capability. Convenient, but shallow.

Evidence Log CTAI+ 9 events · 3 fail / 3 flag / 3 pass · nothing reconstructed
Day 1 · Initial ImpactZero setup friction. One-click activation from an existing workspace — no account, API key, or config. Spacebar surfaces the AI panel. Fastest activation of any tool tested.Pass
Day 2 · Initial Impact"Improve writing" on my own paragraph — output looked polished at a glance, but the voice was gone, flattened to generic business prose. Looked better, read worse.Flag
Day 3 · Initial ImpactSummarise on a short meeting note (under 400 words) — fast and accurate, extracted the three key decisions correctly. Strong on short, structured input.Pass
Day 7 · Stress TestingSummarised a 6,000-word strategy document. Output reflected only ~1,400 words — the opening sections. Final recommendations absent. No truncation warning; the summary looked complete.Fail
Day 9 · Stress Testing"Draft meeting agenda" on a project page listing team members — the draft included two attendee names not present anywhere in the workspace.Fail
Day 10 · Stress Testing"Improve writing" across five samples — every output over-formalised, contractions removed, rhythm flattened. Voice preservation 0/5. Works as a voice eraser, not an editor.Fail
Week 2 · Observe & DocumentDeployed across three client workspaces (agency ops, client strategy, product roadmap). Stayed in rotation for one — low-stakes internal — and was quietly dropped from the other two.Flag
Week 2 · Observe & DocumentBest consistent use: reformatting unstructured notes into structured prose, and boilerplate first drafts. Meeting recap bullets→prose reliable; process-checklist draft reliable.Pass
Week 3 · Observe & DocumentThe convenience premium is for proximity, not capability — £10/member/month to not switch tabs. Legitimate for a specific task type; the mistake is believing proximity equals capability.Flag
Tested by
Woody
Woody
Founder · Sole tester · CTAI

Solo operator and the Crash Test Dummy behind CTAI. I put AI to work. I apply pressure. I document what happens — no sponsors, no hype.

About the operator →