Your provider sends a deprecation notice. The model you built on retires in sixty days. You grep the repo and find thirty-one call sites across nine services, each with its own retry logic, its own timeout, and three with prompts pasted inline.
That is the bill for treating the model API as an implementation detail. It arrives all at once.
What the gateway is actually for
A model gateway is more than a wrapper around an SDK. It is the place where five decisions become enforceable: which model serves which workflow, what happens when that model is unavailable, what a call is allowed to cost, what policy and redaction run before the request leaves your network, and how a version change gets rolled out.
It is also the only honest answer to "who called the model, on whose behalf, with which prompt version". Auth and audit belong at a boundary. There is no boundary if every service brings its own.
Routing is a product decision written as config
Not every workflow deserves the same model. Autocomplete and a compliance summary have nothing in common except that both happen to call an LLM.
| Workflow | Model tier | Latency budget | Fallback |
|---|---|---|---|
| Inline suggestions | small, fast | 300 ms | cached template |
| Support drafting | mid | 3 s | smaller model, draft marked unverified |
| Compliance summary | frontier | 30 s | queue for a human, no auto-answer |
The useful column is the last one. Choosing a fallback forces you to say out loud what degraded behaviour is acceptable, and for the compliance row the honest answer is "none, so stop and escalate". A gateway lets you encode that once. Scattered call sites let every team quietly invent their own version of it.
An untested fallback is a guess
Nearly everyone writes the fallback branch. Almost nobody exercises it. Then a provider has a partial outage, requests do not fail cleanly but hang at eleven seconds, and the fallback never fires because it was wired to exceptions rather than to latency.
Force the failure deliberately by killing the primary in staging every week. Watch what a user actually sees, and check whether a degraded answer is labelled as degraded, because a silently worse answer costs more trust than a visible error.
Migration is where it pays for itself
One boundary means you can shadow a new model against live traffic, move one workflow at a time, attribute cost per workflow instead of per invoice, and roll back with a config change rather than a release.
It also makes "which prompt and which model version produced this output" a query instead of an archaeology project. When a customer disputes an answer from three weeks ago, that difference is the entire conversation.
Provider strategy is less about picking the best model than about staying able to change your mind cheaply.
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Discussion