New here? Start this series at Part 1: Autonomy is a runtime decision, not a model decision
- Part 1: Autonomy is a runtime decision, not a model decision
- Part 2: Your tool registry is an access control list wearing a different name
- Part 3: Four kinds of agent state, and only one of them is memory
- Part 4: A retried agent job is a second chance to send the same email
- Part 5: The 3am agent run that nobody is watching
The support agent timed out at eight minutes. The job runner retried, as configured. The customer received two refunds and one apology.
Nothing in that sequence was a bug in the usual sense. Every component did what it was told. The system had just never been given semantics for doing the same work twice.
Synchronous is a choice you make once and pay for repeatedly
Agent work is long, bursty and full of calls that can hang. Running it inside the request that triggered it means the browser tab, the load balancer and the process lifetime all become part of your reliability story, and the first deploy during a busy hour kills every in-flight task.
Push it onto a queue with a worker pool and the run gains something it did not have: an identity that outlives any single process. It can be inspected, resumed, cancelled and counted. That is the whole reason to accept the extra machinery.
The idempotency key must come from the work, not the worker
A key generated when a worker picks up a job is a new key on every attempt, which is the same as having none.
Derive it from the intent: this user, this order, this action, this logical request. Then have the side-effecting system enforce it, whether that is a unique constraint, a conditional write, or a provider that accepts an idempotency header. Enforcement in the agent's own memory does not survive the crash you are protecting against.
Design this before the duplicate refund, which means before enabling retries. Retries are trivial to turn on and that is exactly why they get switched on first.
Timeouts, cancellation, and the half-finished job
Every long job needs a maximum runtime, and cancellation that actually reaches the worker rather than just marking a row as cancelled while the process keeps spending money.
Then decide what a partially complete run means. If three of five steps committed, is the job resumable from step four, or must it be compensated and restarted? Both answers are defensible. Not choosing means the answer varies by whichever code path failed.
Backpressure is a cost control
Queue depth is a budget metric before it is a latency metric. Expensive jobs piling up during an incident quietly build a bill that arrives well after the incident is closed.
Give operators depth, oldest message age, failure rate and spend per queue, plus concurrency limits per tenant so one workspace cannot consume the pool. Users need the same visibility at their own scale: queued, running, failed, retrying.
Closing thought
Durable execution is unglamorous work that no demo requires and every production system eventually buys, usually at the price of one duplicated side effect.
Discussion