"There's a human in the loop" is one of the most reassuring sentences in AI product design, and one of the least informative. It says a person exists somewhere in the workflow. It says nothing about whether that person can actually change the outcome.

The goal is to place the human where they are irreplaceable. Human attention is a scarce resource, and most systems spend it badly.

Placement is a two-axis decision

Impact: can this action be undone in five minutes, or will it be discussed in a postmortem?

Uncertainty: is the agent operating on solid evidence, or on a guess?

High impact plus high uncertainty is where review belongs. Low impact plus high confidence is where automation belongs, with logs. The diagonal between them is where the actual design work lives, and it is worth mapping explicitly rather than letting it emerge from whoever implemented each feature.

Two numbers that tell you if you got it right

Approval rate. If a reviewer approves 99% of what reaches them, the gate is in the wrong place, or worse, it has trained its human to stop reading. Approval fatigue is a real failure mode: a person rubber-stamping 200 requests a day gives you the same risk with extra latency and a name attached for accountability purposes.

Time-to-context. If approving a request requires fifteen minutes of reconstructing what the agent was doing, reviewers will stop reconstructing. They will approve on vibes, and you will not know until something bad ships with a signature on it.

The fix for the second one is design, not discipline: put the proposed action, the evidence behind it, the blast radius, and a recommended decision on one screen. Reviewers who can decide in thirty seconds actually decide. Reviewers who need to go spelunking approve.

Close the loop

Track what reviewers actually do (approve, reject, modify) and feed it back. Rejection patterns are the highest-quality training signal you will ever get about where your agent's judgement is weak, and almost nobody collects them.

If reviewers consistently modify one field before approving, that is a bug report written in behaviour rather than words.

Closing thought

Automation needs accountable checkpoints. But a checkpoint nobody performs attentively is worse than none: same risk, plus false assurance, plus delay.

Fewer, better-placed gates with better context beat comprehensive checkbox theatre every time.

Comparison · · 1 min read

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.