
From Goal to Product
Postconditions catch broken steps. They cannot catch a project marching in the wrong direction, one perfectly verified step at a time. The lifecycle layer of our harness: phases as levels of concreteness, gates that can actually say no, competing architects whose first ideas are treated as samples rather than decisions, and a cockpit so project state never lives in terminal scrollback.
TL;DR. The last two essays covered steps: the agent declares its plan in a checkable spec, and a server verifies every step against its postcondition. This one is about the failure class that machinery cannot touch. A run can pass every check and still build the wrong thing, coherently, step by verified step, because correctness of steps and rightness of direction are different properties. Direction is owned by the lifecycle layer: a feature moves through idea, design, blueprint, and implementation, and between phases sit gates that can actually say no. Design is produced by multiple agents under competing mandates, because one agent's first idea is a sample, not a decision. Review is done by a different model than the one that wrote the code, because no phase grades its own homework. And all of it is visible in artifacts, so the state of the project never has to be reconstructed from terminal scrollback.
The origin story in the first essay of this series ended with an architecture: a lifecycle tool rebuilt on top of an enforcement kernel. The second and third essays described the kernel. This one climbs back up to the layer that decides what the kernel should be enforcing in the first place, because there is a failure mode that no postcondition can catch, and it is the expensive one.
Verified Steps in the Wrong Direction
Here is the uncomfortable gap the step machinery leaves open. Every check we have described so far is local. A postcondition knows about its step's output. An audit trace knows what each step did. None of it knows whether the design the steps are executing still serves the goal, or whether the goal itself quietly shifted three sessions ago. A project can be wrong at a level the step checker cannot even express: right code, wrong feature. Correct migrations, wrong schema. A beautifully verified implementation of a design nobody would have approved if anyone had been asked at the right moment.
Humans do not solve this with better unit tests. They solve it with process: someone writes the design down before building starts, someone who is not the author reviews it, and there are points in the life of the work where it is legitimate to say stop. Every one of those is a direction check, not a step check. The lifecycle layer is those direction checks, made structural and made mandatory, for work done by agents.
Phases Are Levels of Concreteness
In our system a feature is never just "in progress." It is in a phase: idea, then design, then blueprint, then implementation. The phases are not paperwork stations. They are levels of concreteness, and the discipline is that work at one level has to finish crystallizing before the next level is allowed to begin. The idea phase produces a stated goal and its boundaries. Design produces the approach: what changes, what does not, what the interfaces look like, what was considered and rejected. Blueprint turns the chosen design into an executable plan, enumerated work with acceptance criteria, which is where "done is defined before the work starts" stops being a slogan and becomes a file. Implementation is the part everyone wants to skip to, and by the time it starts, the interesting decisions are already made, reviewed, and written down.
The journal entry that started all of this, quoted in the origin essay, said the thinking is the work and implementation is the easy part once the thinking is done. The phase structure is that sentence turned into an enforceable shape. An agent cannot start building from a vibe, because building is a phase, and phases have entry conditions.
Between phases sit the gates, and the gates run on the dial described in the origin story: every decision point is set to gate, flag, or skip. Gate means a human decides before anything advances. Flag means the agent proceeds and the human is notified. Skip means the agent proceeds silently. The dial is how the same lifecycle serves a prototype and a production system without changing shape: you do not relax the process, you set the dials. And the escalation rule from the origin still holds, the agent may tighten its own dial when its confidence drops, and may never loosen it. Only the human loosens.
A First Idea Is a Sample
The design phase has a property I have come to think of as the most quietly important decision in the whole system. We do not ask one agent for a design. We dispatch several, in parallel, under competing mandates: one instructed to make the smallest change that could possibly work, one instructed to design it cleanly as if the codebase's future depended on it, one told to find the pragmatic middle. They work independently, and then the proposals are judged against each other and synthesized.
The reason is statistical, not ceremonial. A language model's first design is a sample from a distribution, and treating a sample as a decision is how you end up implementing the model's habits instead of your architecture. Ask once and you get an answer. Ask three ways and you get a space, with disagreements, and the disagreements are the information. When the minimal-change proposal and the clean-architecture proposal agree on something, that something is probably load-bearing. Where they diverge is exactly where a human, or a judge with explicit criteria, should be looking.
The same suspicion of self-agreement drives the review gates. Work does not advance out of implementation because the agent that wrote it is satisfied. It advances when an independent review passes, and the reviewer is deliberately never the model that wrote the work, in practice usually a model from a different provider entirely. Two models from the same family share blind spots, and a blind spot reviewing itself approves itself. The earlier essay in this series called the spec a message from the agent's best self to its worst self. The review gate is the same principle across models: nobody, human or machine, grades their own homework, and the grader should not share the student's eyes.
The State of the Project Is Not in the Scrollback
There is one more job the lifecycle layer does, and it is the one I underestimated most when we started. It remembers.
Agent-driven development has a state problem. The truth about a project, what is in flight, what is decided, what is blocked, what was tried and abandoned, accumulates in the worst possible places: terminal scrollback, chat history, the model's context window, the human's short-term memory. All four evaporate. Come back after a weekend and the reconstruction cost is real, and it gets paid in the most expensive currency there is, attention at the start of a session, exactly when direction is set.
So the lifecycle treats project state as a first-class artifact. Every feature's phase, its gates and their status, its design and blueprint and review verdicts, its completion evidence, all of it lives in files, versioned, generated from what actually happened rather than from what someone remembered to update. The roadmap is not a wish list that drifts from reality. It is a view over the recorded state, and when the state changes, the view changes. Walking away and coming back is cheap, for me and, just as importantly, for the agent, whose next session inherits structured state instead of an archaeology assignment.
Yesterday's essay on the main track argued that the future of AI systems is cognitive infrastructure around models rather than bigger models. This is that argument, applied to the development process itself. Nothing in the lifecycle layer makes the model smarter. All of it makes the work around the model harder to lose, harder to fake, and harder to derail.
What It Costs
The bill, stated plainly, because every layer of this series has one. Process has weight. For a one-file fix, phases and gates are absurd, which is why quick paths exist and the dial exists, and trivial work should sprint past the ceremony. The sharper danger is the one the second essay in this whole project named: a gate you rubber-stamp is worse than no gate, because it manufactures false confidence at the exact moment it stops manufacturing safety. Gates only work if no is a real outcome, taken seriously, acted on. The day you start waving reviews through because the sprint is ending, the lifecycle has become theater, and theater with artifacts is still theater.
And the lifecycle is opinionated. It encodes a claim about how software should come into existence: thinking first, design as an artifact, independent eyes before advancement. If you disagree with the claim, the tool will feel like friction all the way down. I hold the claim because I watched what happened without it, across more than twenty repositories, and the origin essay is the receipts.
The original vision, the one the journal recorded on day one, was "say build me X and the system handles the rest." That sentence sounded like autonomy when we wrote it. What it actually turned out to require is the opposite of hands-off: a structure where every step is checked, every phase is gated, and every decision leaves a trace. The system handles the rest not because the agent is trusted, but because the process is. One more essay in this series, on what happens when the work fans out across many agents and many models, and the budget becomes the constraint that keeps a fleet honest.