- Day 55: Why observability is not optional for AI systems
- Day 56: Why evaluating the model is not enough
- Day 57: Why retrieved content must stay untrusted
- Day 58: Why governance belongs in the architecture
- Day 59: Why production AI is coordinated infrastructure
- Day 60: From model demos to mission-ready AI systems
Every agent team eventually hits the same moment. A user reports a bad outcome from yesterday, and someone asks: can we see what it did?
The answer to that question was decided months earlier, when tracing either was or was not built.
Logging the answer is not tracing
A production trace connects the full causal chain for one request, under one ID, navigable end to end.
| Step | What to capture | What it lets you answer |
|---|---|---|
| Input | request, user context, request ID | Can we reproduce it? |
| Retrieval | queries, filters, source IDs, ranks | Was the evidence there, and did it rank? |
| Model call | prompt, response, tokens, model version | What did it actually see? |
| Tool call | tool, arguments, result, duration | What did it do to the world? |
| Handoff | payload passed between agents | Where did context get lost? |
| Decision | output, confidence, owner | Who is accountable? |
What it converts
Adjectives into locations. "The model hallucinated" becomes "the 2024 policy entered context at step 3 because the freshness filter was not applied." That points to a different owner and a different fix.
Guesses into attribution. The step burning 80% of your tokens becomes visible, and it is rarely the one people predicted.
Apologies into answers. When compliance asks why the system approved something, you replay it rather than reconstructing it from memory and hope.
Capture context without capturing everything
The tension nobody mentions: traces are useful in proportion to their detail, and prompts and retrieved documents routinely contain personal data.
Resolve it deliberately: redact or tokenise sensitive fields, store references rather than payloads where you can, and set retention that satisfies both debugging and privacy. This is a design decision, and making it late usually means making it badly, in one direction or the other.
The retrofit trap
Teams add tracing after the incident that needed it. Which means debugging that incident blind, and instrumenting a system that is now large, live, and load-bearing.
Instrument on day one, while it is small enough that instrumenting is easy.
Closing thought
You cannot operate what you cannot see, and AI failures are particularly good at looking like normal answers. Observability is the control surface for production AI, not a nice-to-have you add once things get serious. Things being serious is when you need it to already exist.
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Explainer · · 2 min read
Real users ask questions your test set never imagined
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Explainer · · 2 min read
Logging the answer tells you almost nothing
A trace that links intent, prompt, retrieval, tools and output is the only thing that makes an AI failure debuggable.
Discussion