New here? Start this series at Part 1: Most teams reaching for multiple agents do not need them
- Part 1: Most teams reaching for multiple agents do not need them
- Part 2: Write role contracts, not agent personalities
- Part 3: Every multi-agent pattern is an answer to one question: who controls the run
- Part 4: An agent that stops when it feels finished has no stop condition
- Part 5: A wrong answer tells you nothing about which agent was wrong
The invoice arrives first. Then someone opens the trace and finds a worker that called the same search tool sixty-one times with near-identical queries, each time reading the results, deciding they were insufficient, and trying again with two words changed.
Nothing errored and every call succeeded. The system simply had no definition of when to stop, so it stopped when the context window filled.
Permissions belong to the role, not the run
Most multi-agent systems hand every agent the same tool registry, then rely on the prompt to keep each one in its lane. That holds until a worker decides the fastest route to its goal runs through a tool it was never meant to touch.
Scope tools to roles at the point of construction. A summariser gets read access and nothing else. A worker that can write gets exactly the write scopes its contract names. If a role never needs to delete, it should not be able to form the call.
This also gives you something the prompt cannot: an audit answer. When you need to know which agents could have caused a side effect, the allowlist tells you, and you do not have to reason about what a model might have been persuaded to do.
Budgets are per run and they nest
A global rate limit is no substitute for a budget, which is an allocation for one run that gets divided among the agents in it.
Give the run a ceiling in tokens, wall-clock time, money and tool calls, then have the orchestrator subdivide it. Workers receive a share and cannot exceed it. When a worker exhausts its allocation it returns partial results with a flag, rather than quietly asking for more.
That flag matters more than it sounds. Systems without it degrade invisibly: the answer still arrives, it is just built on three of the five sources it was supposed to consult. Parallelism needs the same discipline. Cap concurrent workers, apply backpressure when downstream steps queue, and cancel siblings when the orchestrator already has enough to answer.
Stopping is a decision someone has to write down
Useful termination rules are boring and checkable. Budget exhausted. Done criteria from the role contract satisfied. Maximum iterations reached. The last three attempts produced the same error, or the same tool call with the same arguments. No measurable progress across two cycles, where progress is defined against the output schema and not the agent's own opinion of itself.
Retries need the same rigour. Retry a transient failure with backoff. Do not retry a validation failure, because the input is wrong and running it again just costs more. Escalate to a human when the stop condition fires on something irreversible.
Closing thought
Every one of these rules is code that runs outside the model. Written into a prompt, they are things the system usually does. Written into the orchestrator, they are things the system cannot avoid doing, which is the only version that holds on a bad day.
More on these topics
Deep dive · · 6 min read
Adding a second agent does not add intelligence, it adds a contract
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
Deep dive · · 6 min read
Six ways to wire agents together, and the same three things break every time
The topology gets all the design attention. Ownership, termination and traceability are what decide whether it survives contact with production.
Deep dive · · 6 min read
Most agent controls do not actually control anything
Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.
Discussion