Dia Browser
Good for research, discovery and brainstorming. Not at the output standard my work requires.
Pre-Rig retrospective: six months of live use, April–October 2026 — not a controlled Test Rig run. Method and disclosure below.
Build work needed a second tool.
It couldn't reach my repo, run the build or see a render, and it was read-only. Nothing carried between sessions. I became the relay — and every handoff was another thing to check.
In-browser research, discovery and brainstorming.
Use it for research, discovery, rebuilding context from your tabs, brainstorming and honest critique — and remain the quality gate yourself.
How the 63 was built.
Each metric is scored 1–10 and multiplied by its weight. It scores highest where it's easiest to start and where research goes deepest. Reliability and output quality pull it down, because of the hallucination and the slow responses.
Six months, four phases.
This wasn't run phase by phase. Six months of live use came first; the retrospective sets out what happened against the same four phases every crash test uses.
The claim I started from was Dia's own: on day one it described memory across sessions as a feature, and warned me, “You need to remain the quality gate.” Setup cost almost nothing — one prompt from an open tab, a URL or a brief returned structured, specific work.
Real build work on a hands-on Eleventy site. It couldn't reach my repo or run the build — git and localhost pages were blocked — and images reached it as text descriptions, not pixels. It could advise but not act, and usage ceilings cut some passes short.
Context held for about a week inside a thread; by July Dia said plainly that nothing carried between sessions, so I kept a Notion memory stack. Hallucination and slow responses wore me down. What held up: it rebuilt the full project state from my tabs and mapped 774 commits to a restore point.
Strong in the browser, but short of the output standard I need for build work. Cost-to-value sits in the middle: I got real value from it, but not enough that I'd pay for it.
What the data showed.
Headline numbers from six months of live use. Each traces to the session transcripts and artifacts behind this retrospective — there was no Rig evidence log.
Within a long chat thread. Between separate sessions nothing carried — by July Dia said so plainly.
Across 5 branches and 11 tags, to find a restore point from 19 June.
Master Context, Session Primer and AI Handoff in Notion, maintained purely to hand back the context it dropped.
Where it fought me.
The real-world resistance points that determine whether a tool survives daily use or gets abandoned within a week.
| Onboarding — Cost almost nothing — an open tab, a URL or a pasted brief was enough to start. |
| Memory — Nothing carried between separate sessions; I maintained a Notion stack to hand context back. |
| Build access — Couldn't reach my repo or run the build; git and localhost pages were blocked. |
| Visual work — Images reached it as text descriptions, not pixels, so every visual call came back to my eye. |
| Accuracy — Confident answers I had to check and correct. |
| Speed — Slow responses — the defining frustration across daily use. |
What held. What failed.
- Near-zero setup — structured, specific work from one prompt.
- Research, discovery and everyday brainstorming.
- Rebuilt the full project state from my open tabs (4 Sep).
- Corrected itself when I challenged a detail it had wrong.
- Mapped 774 commits to find a restore point (19 Jun).
- Pushed back on another tool's uncritically positive feedback (9 May).
- Design critique stayed useful without seeing the design.
- Nothing carried between sessions, despite describing memory as a feature.
- No access to my repo or the build; git and localhost were blocked.
- Images reached it as text descriptions, not pixels.
- Read-only, and usage ceilings cut some passes short.
- Hallucination: confident answers I had to check and correct.
- Response delay — the defining frustration across daily use.
You want in-browser research, discovery, brainstorming and honest critique — paired with a coding agent and your own memory layer.
You need hands-on build work, visual QA, or anything that needs it to remember last week.
No. I got real value from it, but not enough to pay for at the standard my work requires.
How this test was run.
| Scope — Dia's built-in AI assistant, used daily as my research, discovery and build partner while I rebuilt crashtestingai.com. |
| Run — Six months of live use, April–October 2026. A pre-Rig retrospective, not a controlled Test Rig run — like CT-001 and CT-002, it predates the Rig. |
| Evidence — Session transcripts and artifacts from that period. Hallucination and response delay are my observations across daily use, not logged session by session. |
| Equipment — Dia retrieved the transcripts, artifacts and quotes, and drafted a self-graded report, which I set aside. Claude prepared the scoring sheet and laid out this report. |
| Verdict — Every score, quote and the verdict are mine. |
Sole tester · CTAI
Dia is a good AI browser. For research, discovery and everyday brainstorming it did a good job. But it hallucinated too often, slow responses were the defining frustration, and it couldn't reach my repo, see a render or take any action — so build work always needed a second tool. Strong in the browser, but short of the output standard I need for build work.