REAL-WORLD AI TESTING COMMISSION A TEST
Real-world AI testing
Crash Test Report CT003
Specimen 003 AI Browser Assistant · 6-Month Live Use

Dia Browser

April–October 2026 · macOS · Pre-Rig retrospective
Crash Test Score
63/100
8 weighted metrics · 6 months live use · 16 recorded events
Verdict▲ Mixed
Would I pay?No
Run6 months

Good for research, discovery and brainstorming. Not at the output standard my work requires.

Pre-Rig retrospective: six months of live use, April–October 2026 — not a controlled Test Rig run. Method and disclosure below.

Primary Failure Point

Build work needed a second tool.

It couldn't reach my repo, run the build or see a render, and it was read-only. Nothing carried between sessions. I became the relay — and every handoff was another thing to check.

Decision Snapshot Best Use

In-browser research, discovery and brainstorming.

Use it for research, discovery, rebuilding context from your tabs, brainstorming and honest critique — and remain the quality gate yourself.

7/10
Key Signal Operator Ceiling
How far it takes the operator before the human has to take over: through research and critique, then a handoff to the repo, the render and a second tool.
01 Weighted Breakdown

How the 63 was built.

Each metric is scored 1–10 and multiplied by its weight. It scores highest where it's easiest to start and where research goes deepest. Reliability and output quality pull it down, because of the hallucination and the slow responses.

Onboarding ×10%
9
UX Clarity ×10%
5
Core Use Case ×15%
7
Integration ×15%
6
Output Quality ×15%
6
Reliability ×15%
5
Cost / Value ×10%
6
Operator Ceiling ×10%
7
Pass 75–100Mixed 50–74Fail 0–49
Composite · 63 / 100 · Mixed
02 Phase Findings

Six months, four phases.

This wasn't run phase by phase. Six months of live use came first; the retrospective sets out what happened against the same four phases every crash test uses.

01 Initial Impact

First contact

13 April 2026 · day one

The claim I started from was Dia's own: on day one it described memory across sessions as a feature, and warned me, “You need to remain the quality gate.” Setup cost almost nothing — one prompt from an open tab, a URL or a brief returned structured, specific work.

02 Stress Testing

Under pressure

Apr–Oct · real build work

Real build work on a hands-on Eleventy site. It couldn't reach my repo or run the build — git and localhost pages were blocked — and images reached it as text descriptions, not pixels. It could advise but not act, and usage ceilings cut some passes short.

03 Observe & Document

Real workflow

Apr–Oct · daily use

Context held for about a week inside a thread; by July Dia said plainly that nothing carried between sessions, so I kept a Notion memory stack. Hallucination and slow responses wore me down. What held up: it rebuilt the full project state from my tabs and mapped 774 commits to a restore point.

04 Verdict & Report · 63

The call

October 2026 · retrospective

Strong in the browser, but short of the output standard I need for build work. Cost-to-value sits in the middle: I got real value from it, but not enough that I'd pay for it.

03 Instrument Readout

What the data showed.

Headline numbers from six months of live use. Each traces to the session transcripts and artifacts behind this retrospective — there was no Rig evidence log.

1 week
Context held · in-thread

Within a long chat thread. Between separate sessions nothing carried — by July Dia said so plainly.

774 commits
Mapped to a restore point

Across 5 branches and 11 tags, to find a restore point from 19 June.

3 docs
Memory kept outside Dia

Master Context, Session Primer and AI Handoff in Notion, maintained purely to hand back the context it dropped.

04 Operator Friction

Where it fought me.

The real-world resistance points that determine whether a tool survives daily use or gets abandoned within a week.

Onboarding — Cost almost nothing — an open tab, a URL or a pasted brief was enough to start.
Memory — Nothing carried between separate sessions; I maintained a Notion stack to hand context back.
Build access — Couldn't reach my repo or run the build; git and localhost pages were blocked.
Visual work — Images reached it as text descriptions, not pixels, so every visual call came back to my eye.
Accuracy — Confident answers I had to check and correct.
Speed — Slow responses — the defining frustration across daily use.
05 The Balance

What held. What failed.

What held
  • Near-zero setup — structured, specific work from one prompt.
  • Research, discovery and everyday brainstorming.
  • Rebuilt the full project state from my open tabs (4 Sep).
  • Corrected itself when I challenged a detail it had wrong.
  • Mapped 774 commits to find a restore point (19 Jun).
  • Pushed back on another tool's uncritically positive feedback (9 May).
  • Design critique stayed useful without seeing the design.
What failed
  • Nothing carried between sessions, despite describing memory as a feature.
  • No access to my repo or the build; git and localhost were blocked.
  • Images reached it as text descriptions, not pixels.
  • Read-only, and usage ceilings cut some passes short.
  • Hallucination: confident answers I had to check and correct.
  • Response delay — the defining frustration across daily use.
Use it if

You want in-browser research, discovery, brainstorming and honest critique — paired with a coding agent and your own memory layer.

Skip it if

You need hands-on build work, visual QA, or anything that needs it to remember last week.

Would I pay?

No. I got real value from it, but not enough to pay for at the standard my work requires.

06 Method & Disclosure

How this test was run.

Scope — Dia's built-in AI assistant, used daily as my research, discovery and build partner while I rebuilt crashtestingai.com.
Run — Six months of live use, April–October 2026. A pre-Rig retrospective, not a controlled Test Rig run — like CT-001 and CT-002, it predates the Rig.
Evidence — Session transcripts and artifacts from that period. Hallucination and response delay are my observations across daily use, not logged session by session.
Equipment — Dia retrieved the transcripts, artifacts and quotes, and drafted a self-graded report, which I set aside. Claude prepared the scoring sheet and laid out this report.
Verdict — Every score, quote and the verdict are mine.
Woody
Founder
Sole tester · CTAI
Woody
Operator Verdict

Dia is a good AI browser. For research, discovery and everyday brainstorming it did a good job. But it hallucinated too often, slow responses were the defining frustration, and it couldn't reach my repo, see a render or take any action — so build work always needed a second tool. Strong in the browser, but short of the output standard I need for build work.

Share
Evidence Log CTAI+ 16 events · 6 fail / 5 flag / 5 pass · reconstructed from transcripts and artifacts
13 Apr · Initial ImpactDay one. Asked to assess its fit for my workflow, Dia described memory across sessions as a feature, and said: “I'll always produce something that looks polished… You need to remain the quality gate.”Flag
Apr · Initial ImpactPointed at an open tab, a URL or a pasted brief, it returned structured, specific work in one prompt, with no setup.Pass
7 May · Stress TestingRead-only: “I can't set reminders or create calendar events — I only have read access.” It could advise but not act.Fail
9 May · Observe & DocumentPushed back on another tool's uncritically positive feedback rather than agreeing with it.Pass
12 May · Stress TestingImage attachments reached it as text descriptions, not pixels: “I can read the content and structure… but I can't see the visual styling…” Every visual call came back to my eye.Fail
Apr–Oct · Stress TestingCouldn't reach my repo or run the build. Git was blocked, localhost pages were security-blocked and files on my Desktop were out of reach.Fail
Apr–Oct · Stress TestingManual relay to show it a render: screenshot, resize with sips, drop the file into its work folder.Flag
Apr–Oct · Stress TestingUsage ceilings cut some passes short mid-edit.Flag
Jul · Observe & DocumentDia said plainly that nothing carried between sessions — after describing memory across sessions as a feature in April. Within long threads it had held context for about a week.Fail
Apr–Oct · Observe & DocumentMaintained a Notion stack (Master Context, Session Primer, AI Handoff) purely to hand Dia back the context it dropped.Flag
Apr–Oct · Observe & DocumentAnything in the repo needed a second agent: Cursor, later Claude Code.Flag
Apr–Oct · Observe & DocumentHallucination: confident answers I had to check and correct. My observation across daily use; not logged session by session.Fail
Apr–Oct · Observe & DocumentResponse delay: the defining frustration across daily use. Not logged session by session.Fail
4 Sep · Observe & DocumentAfter a holiday break, rebuilt the full project state from my open tabs, and corrected itself when I challenged a detail it had wrong.Pass
Apr–Oct · Observe & DocumentMapped 774 commits across 5 branches and 11 tags to find a restore point (19 Jun).Pass
Apr–Oct · Observe & DocumentDesign critique stayed useful even though it couldn't see the design.Pass