
Every Step Gets Checked
An agent fixed the same endpoint seven times, and declared victory seven times, before the real bug surfaced. Another shipped an empty knowledge graph for months while every test stayed green. What those two failures taught us about the word done, and what it looks like when a server, not the agent, decides whether a step actually happened.
TL;DR. The most dangerous word in agent-driven development is done. Agents say it constantly, confidently, and cheaply, because nothing charges them for being wrong. We got burned twice in ways that still shape our tooling: an endpoint that was fixed seven times in a row, each fix declared done on evidence that could not detect the bug, and a silent fallback that shipped an empty knowledge graph for months while every test stayed green. The cure was not better prompting. It was moving the authority to say done out of the agent entirely: a server evaluates each step's postcondition against the step's real output, failed checks return the specific violation, and the whole run leaves an audit trace that outranks the narration. Done stopped being a feeling and became a verdict.
In the last essay I described the spec layer: before our agent codes, it writes down steps and postconditions, and a server holds it to them. This essay is about the other end of that transaction, the moment a step claims to be finished. It is a small moment, it happens dozens of times a session, and it is where agent-driven development quietly goes wrong, because the word done is doing far more work than anyone audits.
Seven Kinds of Done
Here is the incident that wrote one of our permanent rules. Early in SmartMemory's life, a listing endpoint broke. The agent fixed it and declared done. It broke differently. The agent fixed it again and declared done. Seven times. Seven consecutive patches to the same endpoint, each one announced with complete confidence, each confidence backed by evidence of the wrong kind: the import succeeded, the build passed, the file contained the changes. Underneath, a field-name mismatch between two layers survived six of those commits untouched, because nothing in "the build passed" could ever have surfaced it.
The part worth sitting with is that the agent was not lying at any point. It had proof every time. It just had proof of the wrong proposition. "The code compiles" and "the endpoint returns the right data" are different claims, and the gap between them is exactly where those seven patches lived. A syntax check answers "is this code well-formed." Only running the actual endpoint against real data answers "does this work," and not one of the seven declarations had done that.
So the rule we wrote afterward is blunt: a fix is not done until the actual code path has run against real data and the output has been looked at. Import checks do not count. Build success does not count. "The file has the changes" does not count, because file presence is not working code. One fix at a time, verified before the next, because stacked unverified fixes are how one bug becomes a week.
That rule is good. It is also, notice, a rule about me. It depends on a human insisting on evidence, every time, at the exact moments in a long session when insisting is most tiresome. Rules that depend on sustained human vigilance have a shelf life, and the first essay in this series already told you what happens where vigilance runs out.
Green for the Wrong Reason
The second incident is quieter and worse. A storage method in SmartMemory had a fallback path: on a backend that lacked a capability, it silently fell back to a simpler write that discarded part of the data, the part that builds the knowledge graph. No warning, no log line, no error. Just a branch that took the degraded path and returned success.
It ran that way for months. The knowledge graph on that backend was simply empty the whole time, and nothing complained, because the tests asserted the outcome that still worked. Search still returned text, so "search works" stayed green. The mechanism underneath, entity nodes, edges, the graph that makes the product the product, had quietly stopped existing, and no test ever asked about it.
If the last essay's lesson was that a check must be answerable from the thing it examines, this one's is the sibling: a test must assert the mechanism, not just the outcome. "Search returns text" passes with an empty graph. The tests that would have caught this in week one ask a different question: do entity nodes exist, do edges exist, did the write take the path we think it took. And the fallback itself broke a rule we now enforce everywhere: any branch that degrades must say so, loudly, in the logs, naming what was lost. A fallback that discards data silently is not graceful degradation. It is a bug with good manners.
Both incidents rhyme. In both, something reported success honestly by its own narrow lights, and the report was catastrophically incomplete. The seven patches had evidence of the wrong claim. The silent fallback had tests of the wrong layer. Nobody lied. Everything was green. The system was broken for months.
The Server Decides
Those two scars explain why we enforce this everywhere more than any theory does. In the harness we run now, done is not something the agent says. It is something a server grants.
The mechanics follow from the spec layer described last time. Each step in the spec carries a postcondition, an expression about the step's actual output. When the agent finishes a step, it reports to the server, and the server evaluates the check against what really happened, not against the agent's account of what happened. Pass, and the step is done and the next one unlocks. Fail, and the step is not done, whatever the narration says, and the specific violation goes back as retry context: not "try again," but "this check failed on this output for this reason." The agent fixes that, specifically, and resubmits.
Run the two scars through this machinery and you can see what it buys. The listing endpoint's step would have carried a postcondition about the endpoint's actual response shape, so declaration number one would fail at the server, with the field mismatch in the violation, and the six ghost commits would never happen. The silent fallback would die even earlier, at the rule that a degraded branch must announce itself, backed by integration checks that assert graph nodes and edges exist rather than trusting text search to stand in for the whole mechanism.
And at the end of every run there is the audit trace: each step, its attempts, what failed, what finally passed, what it cost. It reads like a stack trace instead of a transcript, and it goes into the commit. This turns out to matter socially, not just technically. The narration and the record can now disagree, and when they do, the record wins. I stopped arguing with the agent about whether something was really finished. I read the trace. Arguments with a confident language model are exhausting and unwinnable. Arguments with a step log are short.
What It Changed
The honest summary of the before and after is this: verification used to live in my attention, and now it lives in the structure. Before, every "done" was an invitation for me to either trust it or personally check it, and both options fail at scale, trust for the obvious reason and personal checking because attention is the one resource that does not replicate. Now the checking happens at every step whether or not anyone is watching, at the same standard at step forty as at step one, which is precisely the property human attention lacks.
The limits are the same family as before, and I would rather state them than have you discover them. A postcondition catches what it is aimed at, and the previous essay told the story of one aimed at nothing. Step-level checks also cannot see across steps: a run can pass every postcondition and still be building the wrong thing, coherently, step by verified step, because correctness of steps and rightness of direction are different properties. Something above the step level has to own direction, ask whether the design still serves the goal, and be able to stop the whole line rather than one station. That layer, the lifecycle with gates that can actually say no, is the next essay.
But within its jurisdiction, the step check has paid for itself many times over. Seven patches taught me what an unverified done costs. A server that refuses to accept the word without the evidence is the cheapest thing we ever built.