New here? Start this series at Part 1: Logging the answer tells you almost nothing
- Part 1: Logging the answer tells you almost nothing
- Part 2: Your AI feature has unit economics whether you measured them or not
- Part 3: Real users ask questions your test set never imagined
- Part 4: You cannot roll back a prompt you never versioned
- Part 5: Nothing broke, and the system is still getting worse
At eleven at night the answers start going subtly wrong: confident, well formatted, citing documents that do not say what the answer claims.
Three things changed today. A prompt template edit shipped this morning, the vendor rolled a model point release, and an ingestion job failed quietly at six. Which one is it, and which one can you undo right now, on its own, without reverting the other two?
If the answer is "we would have to redeploy", what you have is a deployment that you are calling a rollback.
AI incidents do not fit the service taxonomy
Classic incident categories are about availability. The service is down, latency is up, the error rate spiked. AI systems fail while returning 200s.
The classes worth naming in advance, because you will not invent them under pressure:
- Quality regression. Answers got worse. No error anywhere.
- Retrieval failure. The index is stale, empty, or serving the wrong tenant's documents.
- Tool failure wearing a model costume. A tool returns nothing and the model fills the gap.
- Safety or policy breach. Output that should have been refused, or a leak across a boundary.
- Cost or latency blowout. Usually a retry loop, or a context that quietly doubled.
- Upstream model change. The vendor changed something and your behaviour moved with it.
Each has a different first responder and a different first action. That is the entire reason to write them down.
Independent rollback, or none at all
Every layer that can change behaviour needs its own version and its own reversal path, decoupled from the application release. Prompt templates: versioned, flag-selectable, revertible in seconds. Model choice and parameters: configuration, not code. Retrieval config: index version, chunking, thresholds, reranker on or off. Tools: individually disableable, so one bad integration does not force you to kill the whole feature.
Then the kill switch, and be precise about what it kills. Hiding the UI while the workflow still runs on the backend is not one. Exercise it in production on a quiet afternoon, because an untested kill switch is a belief.
The half that is not engineering
Detection is chronically underinvested. Your availability alerts will not fire here. Alert on quality proxies instead: refusal rate, retrieval hit rate, correction rate, escalation volume, tokens per task.
Support needs a path that is not "file a bug". Give them the failure taxonomy, a trace ID field, and a named person to escalate to. Decide the customer communication rule before you need it: which failure classes get proactive contact, who writes it, who approves it.
Afterwards, a postmortem without evidence is a memory exercise. Traces, versions in play, tenants affected, timeline. Then the step that turns an incident into an asset: every confirmed bad output becomes a regression case, so that failure has to be new next time.
Closing thought
You cannot build rollback during an incident. It is a set of boundaries you drew months earlier, on an ordinary Tuesday when nothing was on fire.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Discussion