
Subagents With a Model Budget
Once the lifecycle fans work out to many agents, a new question appears: which model gets which job? The obvious answer, cheap models for grunt work, is right about the jobs and wrong about the economics. Why per-token pricing inverts on unbounded tasks, how we route by role instead, and why the budget is what keeps a fleet honest.
TL;DR. The last essay in this series. Once you have a lifecycle dispatching work to many agents, you have a fleet, and fleets force a question single sessions never ask: which model does which job? Everyone's instinct is to send grunt work to the cheap model and keep the expensive one for the hard parts. The instinct is right about the jobs and wrong about the economics, because the cost of a model is not its price per token, it is price per token multiplied by the tokens it needs to finish, and the second factor moves more than the first. On open-ended work we have watched a cheaper model reach the same outcome as a stronger one while taking two to three times the steps and tokens, which makes the cheap model the expensive one. The rule that survived contact with our own dispatch data: route by whether the brief bounds the work, not by how hard the task sounds. And put a hard budget on every dispatch, because the budget is what keeps a fleet honest.
This series started with an origin story, then walked up the stack: the spec the agent writes before it codes, the server that checks every step, and the lifecycle with gates that own direction. This last essay is about what happens when that lifecycle stops being one agent in one terminal and becomes what it actually is on a normal day for us: many agents, running in parallel, on different models, from different providers. A fleet. Fleets are wonderful and fleets are how you spend a fortune by accident, and the difference is routing.
The Instinct and the Inversion
Ask anyone how to divide work between a cheap model and an expensive one and you will get the same sensible answer I would have given: send the mechanical work to the cheap model, keep the frontier model for the hard parts. Boilerplate, repetitive edits, scaffolding, test grunt work, all downhill to the small model. Design, diagnosis, judgment, uphill to the big one.
The job classification in that answer is roughly right. The economics hide a trap, and we only saw it clearly once we started reading our own audit traces and calibrating against long-horizon coding benchmarks. The cost of giving a task to a model is not the price on the pricing page. It is price per token times the tokens the model needs to reach an acceptable result, and that second factor is not a constant. It depends violently on how well-specified the work is. On a tight brief the two factors multiply out the way you would hope, and the cheap model is genuinely cheap. On an open-ended task, the pattern we have seen repeatedly is that the mid-tier model gets to the same outcome as the frontier model but takes two to three times the steps and tokens to get there: more re-reading, more retries, more wandering. Equal outcomes at a multiple of the cost, which is a strange way to spell cheaper.
The inversion is easy to say once you have paid for it: per-token pricing rewards models that need fewer tokens, and needing fewer tokens is precisely what capability buys you on hard, underspecified work. On bounded work, capability is wasted and price wins. On unbounded work, price is an illusion and capability wins. The pricing page tells you neither.
Route by the Brief, Not the Task
So the routing rule we actually enforce is not about how hard the task sounds. It is about whether the brief bounds the work.
A bounded brief has three properties: the design is already locked, the files to touch are enumerated, and there is a mechanical pass or fail gate at the end, usually a test suite. Work shaped like that has a known horizon. The agent cannot wander far, because the brief is a corridor, and the gate at the end does not negotiate. This is where the cheaper model earns its keep, and the earlier layers of this series are what make such briefs possible at all. A locked design comes out of the lifecycle's design phase. Enumerated files come out of the blueprint. The mechanical gate is the step postcondition. The harness does not just check the work, it manufactures the boundedness that makes cheap models safe to use.
An unbounded brief is everything else: find the root cause, explore the options, design the approach, figure out why this flakes. If the agent must discover its own path, the horizon is unknown, and an unknown horizon on a per-token meter is a blank check. That work goes to the strongest model available, not as a luxury but as the budget option, per the inversion above.
And one operational tell, the single most useful heuristic we have: watch for iteration without convergence. When a cheap-model dispatch starts circling, retrying, rereading, reformulating, that loop is exactly where the cost blowup lives, and the correct move is to escalate the task to a stronger model, not to feed the loop more retries. Retries are for a step that failed a check once. A model that is orbiting the problem needs a different model, and every retry you grant it is a donation.
Roles, Not Just Tiers
Price and capability are one axis. The other is independence, and it matters most for review.
Our review gates default to a model from a different provider than the one that wrote the work, on purpose. This is not vendor diplomacy. Models from the same family share training lineage, which means they share failure modes, and a reviewer that shares the writer's blind spots approves the writer's mistakes. The previous essays put it as no one grades their own homework. The fleet version is: the grader should not even come from the same school. Heterogeneity is not a compromise in a fleet, it is a feature you select for, because uncorrelated errors are the whole point of a second opinion.
At the other end, the genuinely tiny model has a real job too: transcription-level work, mechanical reformatting, doc and status updates, anything where the instruction is essentially the output. The failure mode to respect there is fidelity on exact values, so those briefs stay literal, and anything that requires the model to hold a large context does not go to the small context window. Matching the model to the role, not just the price to the task, is most of the game.
The Budget Is the Enforcement
All of this would be advisory, a wiki page nobody reads, except for the last piece, which readers of this series will recognize as the kernel doing what the kernel does. Every dispatch carries a hard budget: time and cost, and exceeding it is an exception, not a warning. Retries draw from the same envelope, so a step cannot fail its way into an unbounded bill. A parallel fan-out's total spend is bounded by the flow it belongs to. When the envelope runs out, the next call fails before it starts.
The budget does two jobs at once. The obvious one is financial: no surprise invoices from a loop nobody watched. The subtle one is informational. A budget exceeded is a signal, and it usually means the routing was wrong, a task believed bounded was not, a model believed sufficient was orbiting. The audit traces from every run carry cost and attempts per step, which means the routing table stops being folklore and becomes something you can check against data. Our rules about which tier handles which role came out of exactly that loop: route, measure, get surprised, revise.
Which is also the honest limitation. A routing table is a cache of measurements, and every model release invalidates some of it. The tier that flailed at a problem in spring converges on it by autumn, the price that justified a rule gets cut in half, and a table nobody revalidates quietly becomes a list of expensive superstitions. Ours has been rewritten repeatedly and will be again. The table is disposable. The discipline of measuring is not.
The Whole Argument
This is the end of the series, so let me assemble the five essays into the one thing they were always saying.
The origin story was the diagnosis: agent-built software fails at the boundaries nobody watches, and human vigilance does not scale to twenty repositories. The spec layer makes the agent's intent a checkable artifact instead of a private thought. The step checker makes done a verdict instead of a feeling. The lifecycle owns direction, the failure class steps cannot see. And the fleet layer, this essay, extends the same move to money and models: budgets instead of hopes, measured routing instead of pricing-page instincts, independent reviewers instead of self-agreement.
None of it makes the model smarter. That was never the project. The project was an old piece of engineering wisdom applied to a new kind of worker: build the system so that the property you care about is enforced by structure, not by anyone's sustained attention, including yours, including the model's. An essay before this series started already said it, and after seven months of running the thing, I have not improved on it. Trust the harness, not the agent.