I used Dia for six months as part of my daily work. It became my research, discovery and brainstorming partner while I was working on CTAI.
That meant the test wasn't something I designed in advance. The work came first.
And eventually I started to notice where Dia was genuinely useful, where it became frustrating, and where I kept having to take over. That turned out to be much more interesting than a feature comparison.
Is Dia actually good for everyday AI work?
Yes.
For research, discovery and everyday brainstorming, Dia was useful.
Getting started was almost frictionless. I could point it at an open tab, give it a URL or paste in a brief and get structured, specific work back without much setup. That matters when you're working through ideas. The distance between a question and something useful coming back was small.
On 13 April, my first day with it, I asked Dia to assess its fit for my workflow. It described memory across sessions as a feature and was candid about the role I would still need to play:
“I'll always produce something that looks polished… You need to remain the quality gate.”
That turned out to be an important part of the experience. Because the further I pushed Dia into actual work, the more obvious that quality gate became.
Does Dia remember what happened last week?
This became one of the biggest problems. Within a long chat thread, Dia could hold context well for about a week. But nothing carried between separate sessions.
- April
In April, Dia had described memory across sessions as part of its fit for my workflow.
- July
By July, Dia said plainly that nothing carried between sessions.
The practical consequence was bigger than simply having to repeat myself. I ended up maintaining a Notion stack containing my Master Context, Session Primer and AI Handoff documents purely to give Dia the context it had lost.
The AI was useful. But I was doing work to make the AI useful.
That distinction matters.
What happens when an AI can't access the work?
This was where the ceiling became much more obvious. I was working on a hands-on Eleventy website. Dia could advise me about the work. It couldn't actually get into the project and work on it.
- It couldn't reach my repository.No access
- It couldn't run the build.No access
- Git access was blocked.Blocked
- Localhost pages were security-blocked.Blocked
- Files on my Desktop weren't available to it.Out of reach
So I became the relay. I had to move information between the environment where the work was happening and the environment where Dia could reason about it.
At one point, showing Dia a render meant taking a screenshot, resizing it with sips, and dropping the resulting file into its work folder.
None of that makes Dia a bad browser. But from my point of view, the result was the same:
I had become part of the integration layer.
Can Dia actually see the work it's helping with?
This became particularly obvious when the work turned visual. Dia could discuss design. It could reason about content and structure. But it couldn't see the visual styling in the way I needed it to. Image attachments reached it as text descriptions, not pixels.
As Dia itself put it on 12 May:
“I can read the content and structure… but I can't see the visual styling…”
On a project where visual design mattered, that meant the final judgement still came back to me. I had to look at the thing myself.
Again, the issue wasn't that Dia had no useful design input. It did. The issue was understanding the difference between reasoning about something and actually having access to the thing you're asking it to judge.
Can Dia actually take action?
For the work I was doing, Dia was largely read-only. It could advise. It could analyse. It could help me think through a problem. But it couldn't simply go and make the required change.
As it put it on 7 May:
“I can't set reminders or create calendar events — I only have read access.”
Usage ceilings also cut some passes short mid-edit. That meant build work needed another tool. Eventually, that meant using Cursor and later Claude Code for work inside the repository.
This created a recurring pattern:
- Dia
- Me
- Another tool
- Me
- Dia
Why did Dia become frustrating to use?
Two things came up repeatedly.
The first was hallucination.
Dia would sometimes give me confident answers that I had to check and correct. Not every answer was wrong. But when the work matters, uncertainty has a cost.
The second was speed.
Slow responses became one of the defining frustrations across daily use. A delay when you're casually asking a question is one thing. A delay when you're in the middle of a work loop is different. It breaks momentum.
Those two things — reliability and response time — became much more important over six months than they were during the initial impression.
Where did Dia genuinely save me time?
Despite all of that, Dia was not a failure. Some of the most useful evidence came from the things it did well.
After a holiday break on 4 September, Dia was able to rebuild the full project state from my open tabs. When I challenged a detail it had wrong, it corrected itself.
It mapped 774 commits across five branches and eleven tags to help identify a restore point from 19 June.
On 9 May, it also pushed back on another tool's uncritically positive feedback rather than simply agreeing with it. And even though it couldn't see the visual design directly, parts of its design critique remained useful.
That matters.
How far can an AI tool take you before you have to take over?
This became one of the most useful questions I took from the experience. In the retrospective, I called this the Operator Ceiling.
How far can the tool take the human operator before the human has to take over?
For Dia, the answer was fairly far.
- It could take me through research.
- It could help with discovery.
- It could help me brainstorm.
- It could reconstruct context from open tabs.
- It could critique ideas.
- I had to leave Dia.
- Go to the repository.
- Use another tool.
- Look at the render myself.
- Check the answer.
- Re-establish context.
- Run something.
- Then come back.
That handoff is part of the real-world experience of using an AI tool, whether the product itself considers it part of the product or not.
What happens when one AI tool isn't enough?
This is where the six months changed how I think about AI tools. It is easy to evaluate a tool in isolation. But I don't work in isolation.
I work in workflows.
Dia was useful. Cursor was useful. Claude Code was useful. But usefulness doesn't automatically add up to a good workflow. Sometimes it creates a chain of handoffs. And the more handoffs there are, the more I have to manage the system around the AI.
That creates a hidden cost. The tool might save me time on one part of the job while creating additional work somewhere else.
Would I actually pay for Dia?
After six months of real use:
No.
That doesn't mean I got no value from it. I did. It means the value wasn't enough at the level I needed.
- Onboarding Friction9/10
- UX Clarity5/10
- Core Use Case Depth7/10
- Workflow Integration6/10
- Output Quality6/10
- Reliability / Stability5/10
- Cost-to-Value6/10
- Operator Ceiling7/10
The strongest scores were around getting started and the depth of the core research/discovery use case. Reliability and output quality pulled the overall result down.
And although I got real value from Dia, I didn't get enough value to pay for it at the standard my work required.
What is Dia actually good for?
- In-browser research
- Discovery
- Brainstorming
- Exploring ideas
- Rebuilding context from open tabs
- Honest critique
- Hands-on build work
- Visual QA
- Work that depends on remembering previous sessions
- Repository-level execution
- Getting the final output over the line without another tool or human intervention
That's more useful to me than simply saying Dia is a 63/100 product. It tells me where I'd actually use it.
Does an AI tool have to be bad to fail a crash test?
No.
This is probably the most important conclusion. Dia wasn't bad. It was useful. It just wasn't useful enough for everything I wanted it to do.
A crash test shouldn't be designed to prove that a tool is bad. The useful question is:
Where does it stop being good enough for the person relying on it?
- Sometimes that will be an obvious failure.
- Sometimes it will be a hallucination.
- Sometimes it will be a missing capability.
- Sometimes it will be slow enough to break the workflow.
- Sometimes it will simply be the point where the human has to take over.
That boundary is valuable evidence.
What did six months of using Dia teach me about testing AI?
This experience happened before the CTAI Test Rig existed. It wasn't a controlled experiment. I didn't log every hallucination. I didn't run the same prompts every day. I wasn't following a fixed test script. I was doing my work.
And that is precisely why I don't want to dismiss the experience because it happened before the Rig. It tells me something a controlled test doesn't always capture. It tells me what happened when an AI tool became part of my actual workflow.
But it also taught me that not every test needs to last six months. The value came from the real work, not simply from the length of the calendar.
- A focused testcan expose a breaking point in days.
- A real-work testcan expose a workflow ceiling.
- A longer testcan reveal things that only appear through repeated use.
The duration should follow the question. Not the other way around.
Where does the human still have to take over?
This is why I'm building CTAI around a human verdict. Dia itself told me early on that I needed to remain the quality gate. I agree.
AI can be useful equipment. It can help execute. It can help analyse. It can help find patterns. It can help prepare the next test.
But when the question is:
Would I actually trust this with my work?
the answer still needs to come from the person doing the work. That is the difference between an AI tool producing an answer and an AI tool being good enough to depend on.
What did Dia change about the way I think about Crash Test AI?
The biggest lesson wasn't actually about Dia. It was about the kind of testing I want CTAI to do.
The demo tells you what a tool can do.
Real work tells you what happens when you depend on it.
That's where the interesting evidence appears.
- Does it remember?
- Can it access the work?
- Can it see what you need it to see?
- Can it take the action?
- Does it hallucinate?
- Does it slow you down?
- Where does another tool become necessary?
- Where does the human have to take over?
And ultimately:
Can it produce work I'd actually trust enough to use?
Those are much more interesting questions to me than a feature checklist.
My Dia verdict
After six months of daily use, Dia was genuinely useful for research, discovery and brainstorming. It helped me think, explore, reconstruct context and sometimes challenge assumptions.
But it also hallucinated too often, became frustratingly slow, couldn't access the build environment, couldn't carry context between separate sessions, couldn't see visual work in the way I needed, and couldn't take the actions required to get the work over the line.
For my workflow, that created too many handoffs.
I'd use Dia again for the things it does well. I wouldn't make it the only tool responsible for getting the work over the line.
And that's probably the most useful conclusion a crash test can give you:
Not whether a tool is good or bad. Where it is good enough — and where it isn't.
Test disclosure
- StatusThis is a pre-Rig retrospective (Crash Test 003) based on approximately six months of live use from April to October 2026.
- RunIt was not a controlled Test Rig run. The evidence came from session transcripts and project artifacts from that period.
- LoggingObservations around hallucination and response delay reflect my experience across daily use rather than a session-by-session log.
- EquipmentDia retrieved the transcripts, artifacts and quotes used as evidence. It also drafted a self-graded report, which I set aside. Claude prepared the scoring sheet and laid out the Crash Test report.
- VerdictEvery score, quote and the verdict are mine.
- NextFuture CTAI Crash Tests use the Test Rig to make the investigation more structured, repeatable, evidence-led and traceable.
The crash-test starts where the demo ends