The feature shipped in March and everyone was pleased. In June, finance asked why the model line had grown eleven times while revenue had not moved, and nobody in the room could say which product surface was responsible.

That conversation is avoidable, but only before launch.

Budget the workflow, not the call

Per-token pricing invites you to reason about one call. Users never experience one call. They experience a workflow that fans out into a retrieval query, three model calls, two tool round trips and a re-ask when validation fails.

So the unit is cost per successful task. Successful is load-bearing: a run that ends in a retry loop still burns tokens, and a metric that counts only completions will understate the real number. Same for latency. The p95 on your model call is not the number the user feels.

Context discipline is the biggest lever most teams have not pulled. Prompts accumulate. An example here, a policy paragraph there, full conversation history because summarising was harder. Nobody deletes anything, because deletion might hurt quality and quality is hard to measure. Six months later, half the spend is re-sending context that changes no answers.

Set the target per workflow

Not every path deserves the same budget. Decide the tier explicitly and write it down:

Workflow classLatency targetCost ceilingDefault approach
Inline assistunder 300 msfractions of a centsmall model, hard cache, no tools
Interactive answer2 to 5 s, streamedcentsmid model, retrieval, capped tool loop
Deep taskseconds to minutestens of centsstrong model, show progress, checkpoint
Batch or offlinehourslowest per unitbatch API, off-peak, retry cheaply

The numbers are yours to set. The point is that somebody chose them before shipping, so a regression becomes a visible breach rather than a drift nobody owns.

The levers, in the order I reach for them

Route first. Most traffic in most products is low risk and does not need your best model. Classify the request, send easy ones down the cheap path, reserve the expensive path for what earns it.

Cache second. Exact-match caching catches more than teams expect and costs almost nothing. Semantic caching helps too, with a real failure mode: a near-miss returning a confidently wrong neighbour. Give it its own eval.

Batch third, wherever latency allows. Then degrade deliberately. Under load, a shorter answer from a smaller model beats a queue timeout, and you should decide which surfaces may degrade before you are inside the incident. Capacity planning is the unglamorous part: provider rate limits, queue depth, behaviour at three times normal load. Test the queue before your users do.

Attribution, or that June meeting

Tag every call with feature, tenant, model and outcome. Report spend along those dimensions weekly, to the team that can actually change it. Cost regressions should reach the same people quality regressions do.

Unit economics you did not choose are still unit economics. You just meet them later, in a meeting you did not schedule.

Explainer · · 1 min read

Why autonomy needs budgets

Why autonomous systems need limits on cost, time, actions, retries, and risk.