Files
firehose/app/priv/blog/engineering/2026/07-20-improve-diagnostic-in-the-tdd-cycle.md
Firehose Bot b972603143 Add draft blog posts on step size and TDD diagnostics
Both published: false. The 07-20 draft references a screenshot that is
not yet in the repo (roam-export image copying is still in progress).
2026-10-08 09:49:07 +01:00

49 lines
4.3 KiB
Markdown

%{
title: "Improve Diagnostic in the TDD cycle",
author: "Willem van den Ende",
tags: ~w(TDD AI),
description: "When we do TDD in baby steps, we review the diagnostic when we have written the failing test at the start of the cycle. We do our best to make the failure message understandable to our future selves and others. Does this work with Synthetic TDD? If not, what do we do about it?",
published: false
}
---
I am bringing a vibe-TDDed project started a year ago to production. It has 998 tests at the moment (well, 999, I just added one, which led me down a rabbit hole).
When we do TDD in baby steps, we review the diagnostic when we have written the failing test at the start of the cycle. We do our best to make the failure message understandable to our future selves and others. Does this work with Synthetic TDD? If not what do we do about it?,
From hands-off to hand-on, one test at a time
----
I had 998 tests added by a variety of agents. The 999th one, I added by hand. Inspired by Matteo Vaccari, I started refactoring one of the end to end tests towards an internal DSL, with output (steps and dialogs in a workflow) that can work as a conversation starter between me and users, as well as keeping me on track in the myriad of states that go in to seemingly simple workflows.
I quite like the semi-formality that Allium gives me to determine states. I had used it in this project to help me create dynamic forms with Claude Code. I had made various spikes of Dynamic Forms before, but it gets complicated quickly, especially when you want to show some user defined field conditionally based on another user defined fields' value. Allium, a coding agent, tests and iteration got me over the line. Re-reading the Allium spec, we have indeed a small, almost programming language to define and validate forms.
What does failure feel like?
----
One thing that I had not done much in this system is see tests fail and try to understand what the output says. There are a fair number of integration tests, close to the UI. Phoenix Liveview comes with decent support for writing tests, but if the tests' expectation is to find some words in a web page, and the test fails, I got a wall of text and escape characters in return.
The feedback from the failing test I wrote looked like this (copy pasted from the passing test that would fail in a similar non-intuitive way):
![Screenshot of ascii gibberish: escaped html with many 'quot' and css etc. plus a line number of the failure](/static/images/blog/2026/Screenshot 2026-07-20 at 14.06.13.png)
Here I was trying to establish that an error message was present. The smallest step was asserting that success was absent, and then iterating to find a better diagnostic.
Especially for integration tests it matters when and where a failure is triggered.
Matteo Vaccari's talk reminded me of Ward Cunningham's swim system - so besides the cryptic assertion failures I have screenshots in a html page, one step at a time. I am working in small steps, building just what I need. I still do like to have a precise assertion failure, so I get fast feedback right where the test is. I can then refer to the screenshot to see what it looks like.
It tunrs out that getting both the feedback I want in the printed dialog, and the expectation needed a fair amount of refactoring, and I did some learning on phoenix liveview testing.
Backstory - what brought us here
---
WeReview 2, or at least the initial architectural spike for it, was one of those projects that gave me AI whiplash last year. WeReview 0 and 1 have been running since 2016, and 1 is in a kind of local optimum. It just runs in production with close to zero effort, but is not extensible in the directions I want. So a bit over a year ago I spent 5 days with claude code, mostly hands-off to do 80 percent of a rewrite, with a number of quality of life improvements for me, conference organisers, and reviewers. For proposers (people who propose a session) I also want to improve feedback, and that is more or less where I am now, trying to bring WeReview 2 to production.
Further reading
===
My [Enabling a local model to explain images](https://willemvandenende.com/blog/engineering/enabling-a-local-model-to-explain-images-in-pidev) has the sketch by Jon Jagger, naming the "Improve the diagnostic" action as a transition form Red to Red in a TDD State diagram.