Ask a team how their agent is kept safe and you will usually get a list. There is a human in the loop. There are guardrails. It retries on failure. There is a spend cap somewhere.

Every item is true, and none answers the question that matters: what does the system do at the moment a boundary is reached?

I have never seen an incident caused by an agent doing something obviously forbidden. They come from the seams: an approval gate a reviewer clicked through in four seconds, a rule that lived only in a prompt, a retry loop that repeated one wrong assumption eleven times, a budget the agent hit halfway through a write.

These six posts sit in the middle of a sixty-day series, and the through-line only surfaced when I read them together: a control counts only if the running system enforces it at the boundary and you can watch it do so, and most controls have never been watched at all.

A checkpoint nobody reads is worse than no checkpoint

"There is a human in the loop" is one of the most reassuring sentences in AI product design and one of the least informative. It confirms a person exists somewhere in the workflow, and says nothing about whether they can change the outcome.

Placement is the real decision, and it runs on two axes: how reversible the action is, and how solid the evidence behind it is. Irreversible plus uncertain is where a human belongs. Reversible plus well-evidenced should be automated, with logs. The diagonal between those corners is where the design work lives, and it is worth mapping deliberately rather than letting each implementer decide it.

Two numbers tell you whether you got it right. If a reviewer approves 99% of what reaches them, the gate is in the wrong place, or it has trained its human to stop reading. And if approving one request takes fifteen minutes of reconstructing what the agent was doing, reviewers stop reconstructing. They approve on instinct, and you find out when something ships with a signature on it.

The second one is a design problem, not a discipline problem: proposed action, evidence, blast radius and a recommendation, on one screen. Reviewers who can decide in thirty seconds decide. Reviewers who have to go digging approve.

A rule in the prompt is a preference, not a boundary

One question separates real controls from documentation: where is it enforced?

"Do not access customer PII" in a system prompt is a request made to a probabilistic system. It can be misread, argued around, or overridden by instructions inside retrieved content. It is a strong default, not a boundary.

Enforced in the system, the same rule looks different. The agent's credentials cannot read the PII store. The tool is not in its registry. The spend cap is applied by the runtime, which rejects the call rather than reasoning about it. Prompts shape behaviour; systems bound it. Sorting your rules into those two buckets is uncomfortable: most teams find more of their safety story in the prompt than they expected.

Then there is the failure direction nobody counts. Everyone tracks escapes, the harm that got through; almost nobody tracks false blocks, the legitimate work the guardrail wrongly refused. That is how a guardrail programme strangles a product silently: users hit a wall, decide the system cannot handle their real work, and go back to doing it by hand. Adoption dies without one incident being logged, and the dashboard stays green throughout.

Which is why a refusal is not a design. A good boundary comes with somewhere to go: a narrower version of the action, a request for the missing authorisation, or an escalation a human can resolve in seconds.

The state worth designing for is "failed halfway"

The dangerous agent state is not failure. It is partial success: three of seven records updated, the notification sent but the database write lost. A clean failure you rerun; a partial one somebody untangles by hand while a customer waits.

Recovery is five distinct mechanisms, and teams routinely implement one and call it recovery. Retry handles transient failures, with backoff and a hard cap. Fallback means a simpler tool, an older source, or a degraded answer that is honest about being degraded. Checkpoints are durable progress markers, so resuming means continuing rather than restarting. Rollback is the compensating action for anything externally visible, and "we will fix it manually" is your rollback story until you write a better one. Escalation means handing to a human with state attached: what was attempted, what succeeded, what is half-done.

The distinction that costs most when missed: retry is for transient failures, not wrong assumptions. An agent that retries a bad plan three times has failed three times more expensively.

None of it counts until you have watched it happen. Kill the API mid-task in staging. Return a malformed response. Time out the third of five tool calls. Whatever it does there is what it will do in production at 2am, and most teams discover on the first attempt that it carried on as though the failed step had succeeded.

Budgets change how an agent thinks, not just what it costs

Every team running agents eventually has the invoice conversation. The only variable is whether it happens calmly, as a design decision, or on a Monday after a weekend in which one stuck agent found no reason to stop.

Limits on time, tool calls, cost, retries and risk are the easy half. What gets skipped is the behaviour at the limit. Hitting a budget should produce an outcome, not an exception: a soft limit triggers a wind-down (summarise progress, return partial results, hand over with state attached), a hard ceiling stops the run. An agent that hits a cap mid-action and dies has manufactured the "failed halfway" mess above, deliberately.

Cost control is the obvious win. Planning is the one that surprises people. An agent that knows it has five tool calls left prioritises and abandons low-value branches, and trace quality changes visibly when the remaining budget is exposed to the agent rather than merely enforced around it. Unbounded agents are not more capable, only less deliberate.

Cap retries separately from total actions. They are the most common way a budget disappears invisibly: each one is individually defensible, and the trace looks like diligence rather than what it is, one failing approach in cosmetic disguise.

Splitting one agent into six is a loan

Plenty of teams that needed one agent built a committee. There are three honest reasons to split, and everything else is enthusiasm with an architecture diagram.

Separation of permissions is the strongest: the agent that reads everything should not be the agent that can write anything, and that is a security property no prompt can provide. Genuine parallelism is the second: thirty documents to analyse independently is real concurrent work; sequential work gains nothing from a crowd. Context isolation is the third: an agent with one job outperforms one juggling five, because "be exhaustive" and "be concise" in the same prompt produce something inconsistently both.

If none of the three applies, the single agent wins. Every extra agent levies a coordination tax paid on every request forever: handoffs that drop context, shared state that drifts, conflicts needing arbitration, and a class of failure that exists only between agents and appears in none of your per-agent metrics. When every component tests clean and the system still fails, the fault is in the seams, and seams are what you bought.

The supervisor is your quality ceiling, not just your router

Most teams land on the supervisor pattern first, for a reason they rarely cite. Not throughput: there is exactly one place to look when something goes wrong, because one agent understood the whole mission and the trace reads like a narrative rather than five simultaneous conversations.

The costs concentrate in the same spot. The supervisor is your latency floor, your single point of failure and, least obviously, your quality ceiling. Workers can only be as good as the task descriptions they receive, so a supervisor that delegates vaguely turns excellent specialists into confused ones. The symptom presents as worker underperformance, so teams tune the worker prompts, which fixes nothing: the defect arrived in the input.

That makes the delegation format the highest-leverage prompt in the system. Also worth watching: whether workers ask follow-up questions or silently guess (both indicate the same defect), and whether the supervisor's tracking state inflates until decomposition degrades. And the supervisor owns the definition of "done", so if its standard slips the whole system inherits the new one.

What this adds up to

Six posts, one argument: a control is only real if the running system enforces it and something specific happens when it fires, so the useful question about any control you have is not whether it exists but what it did the last time it triggered.

None of this is about caution. It is about speed. Teams that give agents the most autonomy almost always have the strictest enforced boundaries: a corridor with real walls is what makes full throttle acceptable. Control is what makes autonomy cheap enough to ship.

Where each control is spelled out

Each is a short standalone post, with the failure modes and what to actually do:

They form the Agent Control and Supervision chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. If you would rather start from the beginning, Start Here lays out all sixty in ten chapters.

Comparison · · 1 min read

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.