Put two traces side by side (a planning agent and a non-planning agent on the same task) and the difference is immediate.

The non-planner acts, discovers, backtracks, re-fetches data it already had, and meets the critical dependency at step nine, where unwinding costs nine steps of spend. The planner met that dependency at step zero, on paper, for the price of one model call.

Planning is cheap. Wandering is billed per token.

What makes a plan operationally useful

Not eloquence. Three properties:

  • Verifiable steps. Each one can be marked done or not done without interpretation. "Research the market" fails this; "retrieve Q3 revenue for the three named competitors" passes.
  • Explicit dependencies. This is what makes parallel execution safe, and it is what surfaces the blocker before you have spent anything on the branches behind it.
  • Revisability. When step 3 fails, the plan updates from what actually happened. A plan the agent cannot revise is a script, and it will march confidently into a world that stopped matching its assumptions two steps ago.

The second payoff

The plan is documentation.

When you audit a run three weeks later (because a customer complained, or the bill spiked, or someone asks why the system did what it did), the plan tells you what the agent thought it was doing. That intent is half of every investigation, and without it you are reverse-engineering a sequence of tool calls into a hypothesis about motive.

A trace shows what happened, and a plan shows what was supposed to happen. The bug usually lives in the gap between them.

When to skip it

Planning has a cost, and applying it uniformly is its own failure. A single-step lookup does not need a plan; it needs an answer. If the task has one obvious action and no dependencies, a planning step is pure latency.

The trigger worth encoding: plan when actions have cost or are hard to reverse. If a tool call spends money, modifies data, or contacts a customer, spend a few hundred tokens thinking first. If the worst outcome is a slightly slow read, act.

Planning turns goals into work. It also turns an opaque sequence of tool calls into something a human can review before the expensive part starts.

Which planning decision should be visible before your agent is allowed to touch a tool?

Comparison · · 1 min read

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.