A team I worked with replaced a working single-agent research assistant with five agents: a planner, two researchers, a critic and a writer. The new version was slower, cost roughly four times more per run, and when it produced something wrong, nobody could say which agent had introduced the error.

The single-agent version had been fine. Nobody measured it before replacing it.

That is the usual shape of this decision. Multi-agent architecture gets adopted because it reads as advanced on a diagram, before any workflow has demanded it.

What splitting the work actually costs

Every boundary you add is somewhere context can be dropped. One agent carries its whole history in a single window: messy, but complete. The moment work moves to a second agent you have to decide what travels with it, and whatever you leave behind is gone. The receiving agent fills the gap by inference, confidently, and nothing in the trace flags that it did.

The failure surface also grows faster than the agent count. With one agent you debug one trajectory. With four you debug four trajectories plus the handoffs between them, and the expensive bugs live in the handoffs. Latency stacks unless you fan out, and fanning out means somebody has to own synthesis, which is a new step that can fail on its own.

The three reasons that survive scrutiny

Separation has to buy quality, control or throughput. If it buys none of them, it is decoration.

Quality means genuinely different context or tooling. A retrieval agent tuned against your document corpus and a code agent with repository access do different work with different permissions. A critic agent reading the same context as the writer, using the same model, is usually just a second sample with extra steps.

Control means different blast radius. If one part of the job can issue refunds and the rest cannot, splitting it gives you somewhere to attach permissions and an approval gate. That is a control. Telling a single agent to be careful is a preference.

Throughput means real parallelism over independent subtasks. Forty documents summarised concurrently is a good split. Five steps in a fixed line is a workflow, and a workflow with hard-coded edges is cheaper, faster and far easier to test than agents negotiating who goes next.

Decompose the work before you name any agents

Write the steps out as a flat sequence. Mark the ones where the next action genuinely depends on what the previous step returned. Those are the uncertain steps, and they are the only candidates for agency. Everything else is code.

If exactly one step is uncertain, you need one agent inside a workflow rather than a team of them.

Then name the owner of the final output: the single role accountable for whether the answer is correct. Systems that skip this produce results everyone reviews and nobody trusts.

Think about your last agent failure. Could you name the step that caused it, or only the run?

Deep dive · · 6 min read

Most agent controls do not actually control anything

Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.