New here? Start this series at Part 1: Most teams reaching for multiple agents do not need them
- Part 1: Most teams reaching for multiple agents do not need them
- Part 2: Write role contracts, not agent personalities
- Part 3: Every multi-agent pattern is an answer to one question: who controls the run
- Part 4: An agent that stops when it feels finished has no stop condition
- Part 5: A wrong answer tells you nothing about which agent was wrong
Open almost any multi-agent repository and the roles are described as characters. A meticulous researcher. A ruthless editor. A pragmatic engineer.
None of that is checkable. You cannot write a test for meticulous, you cannot page someone when ruthless degrades, and when two of these agents disagree there is nothing in the definition that says who wins.
What a role contract actually contains
Four things, and personality is not among them.
Inputs: the exact fields this role receives, and which are required. Outputs: a schema it must produce, not "a summary". Permissions: which tools and which scopes, expressed as a real allowlist. Done criteria: a condition someone else can evaluate without reading the transcript.
Once a role has those four, the prompt becomes short. Most of the prompt bloat in agent systems is compensation for a contract that was never written down.
The handoff is a payload, not a conversation
The most common defect I see is agents passing prose to each other. A researcher hands the writer three paragraphs of findings, the writer picks up the confident sentences and drops the caveats, and the caveats were the part that mattered.
Hand off structured objects. Claims with sources attached. Constraints that were discovered, as well as conclusions reached. Which attempts failed, so the next role does not repeat them. A confidence or coverage marker the receiver can branch on.
If the payload cannot be validated on arrival, the boundary is decorative. Validate it, and reject on the spot rather than letting a malformed handoff propagate three steps downstream where the trace no longer explains anything.
Shared state is a liability before it is a feature
Every team eventually builds a shared scratchpad, and every shared scratchpad eventually fills with stale assumptions that no role owns and no role clears.
Default to private memory per role, with explicit reads. If an agent needs something, it should request it by name rather than browse a common pool. Shared state should hold facts about the run itself: the goal, the budget consumed, the approvals granted. Working notes belong to whoever wrote them.
Conflict needs a rule, not a debate
Two agents will contradict each other, and the resolution cannot be another model call that picks the more fluent answer.
Decide in advance. Sometimes precedence works: the role with direct tool evidence outranks the role that inferred. Sometimes the answer is escalation, meaning the run stops and a human sees both positions with their evidence. What does not work is silent last-writer-wins, which is what you get by default when nobody chose.
Closing thought
If you cannot express a role as inputs, outputs, permissions and done criteria, you have a prompt with a job title rather than a role, and it will behave differently every time the model does.
More on these topics
Deep dive · · 6 min read
Adding a second agent does not add intelligence, it adds a contract
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
Deep dive · · 6 min read
Six ways to wire agents together, and the same three things break every time
The topology gets all the design attention. Ownership, termination and traceability are what decide whether it survives contact with production.
Deep dive · · 6 min read
Most agent controls do not actually control anything
Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.
Discussion