New here? Start this series at Part 1: Most teams reaching for multiple agents do not need them
- Part 1: Most teams reaching for multiple agents do not need them
- Part 2: Write role contracts, not agent personalities
- Part 3: Every multi-agent pattern is an answer to one question: who controls the run
- Part 4: An agent that stops when it feels finished has no stop condition
- Part 5: A wrong answer tells you nothing about which agent was wrong
A user reports that the answer was wrong. You have the answer, the original request, and four agents that each made decisions in between. The output tells you a mistake happened somewhere. It does not tell you where, and rerunning the request produces a different path that succeeds.
Single-agent debugging survives on output inspection. Multi-agent debugging does not, because the interesting decisions are the ones nobody saw.
What the trace has to record
One timeline for the whole run, with every event carrying the run id, the role that emitted it and a parent span, so the causal structure survives when workers execute concurrently.
Record the handoff payloads themselves, both sides. What the sender produced and what the receiver actually consumed after validation, because the delta between those two is where context silently disappears. Record every tool call with its arguments and result, the budget consumed at each step, and the point at which control passed from one role to another.
Then make it replayable. If you cannot reconstruct a run without calling the models again, you cannot investigate a failure you have already stopped reproducing.
Evaluate the path, not just the answer
Output scoring is blind to most coordination failures. These are the ones that reach production:
| Failure mode | What the final answer looks like | What catches it |
|---|---|---|
| Context lost at a handoff | Fluent, quietly missing a constraint | Diff the sent payload against the consumed one |
| Workers duplicating effort | Correct, at four times the cost | Tool call overlap across sibling workers |
| Judge preferring fluency | Confidently wrong | Rubric scored against a labelled disagreement set |
| Plan diverging from reality | Plausible but stale | Check each step's precondition at execution time |
| Silent partial completion | Complete-looking, built on half the sources | Assert coverage against the expected input set |
Score handoffs as first-class units. A handoff eval takes a payload and asks whether it carries everything the next role needs to do its job. That is far cheaper to build than a full trajectory judge, and it catches the majority of real defects.
Recovery is a design decision, not an exception handler
Decide per step what happens when it fails, before you ship. Retry with a different approach. Continue with a partial result and mark the gap. Fall back to a simpler deterministic path. Stop and escalate.
Escalation needs the evidence attached. A human reviewer shown only the disputed output cannot approve responsibly; they need what each role saw, what it decided and why. Put the approval gates on the irreversible steps rather than the uncertain ones. Uncertainty is normal; an irreversible step is the one you cannot undo.
When something does go wrong, review the trajectory rather than the answer. The answer is the last thing that happened, and it is almost never where the problem started.
If a handoff dropped a constraint in your system tomorrow, what would surface it before a user did?
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Deep dive · · 6 min read
Adding a second agent does not add intelligence, it adds a contract
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
Deep dive · · 6 min read
Six ways to wire agents together, and the same three things break every time
The topology gets all the design attention. Ownership, termination and traceability are what decide whether it survives contact with production.
Discussion