A customer tells you the assistant quoted the wrong renewal figure last Tuesday. You open your logs and find the answer text, a timestamp and a user ID. Nothing about which documents were retrieved, which prompt version was live, or whether the pricing tool timed out and the model filled the gap from memory.

You cannot debug that; you can only apologise and guess.

Most AI systems get logged like chat apps: one row per exchange, final answer stored, everything that produced it discarded. That works fine until someone asks why.

Spans, not messages

The unit of AI observability is the span. One user request opens a root trace, and every step underneath gets its own span: the retrieval call, each model call, each tool invocation, the guardrail check, the final render.

Each span carries what it needs to be reconstructed later. Start and end time. Status. The prompt template version and the model version actually served rather than the one written in config. Tokens in and out. For retrieval, the document IDs and scores returned rather than the passages. For tools, arguments redacted or hashed, response code, retry count.

The trace ID then has to escape the system. Put it in the response payload, on the support ticket, on the eval failure record. A complaint that arrives with a trace ID attached is a ten minute investigation. Without one it is a day of guesswork, and the guess is almost always "the model hallucinated" because that explanation is free to reach for.

Half the time it was a tool returning an empty array.

Metrics you keep, content you do not

Four metric families answer most operating questions: tokens per successful task, latency split by span type, error rate by component, and retry count. Retries are the one teams skip and the one that quietly explains both the latency curve and the invoice.

Content logging is where people get hurt. Prompts contain whatever the user typed, retrieved chunks contain whatever is in the corpus, and both land in a store with a long retention policy and broad read access. Decide the redaction boundary while you are designing the span model, before legal asks. A sane default: structured metadata everywhere, raw content only in a short-retention store behind separate access, and a documented list of fields that are never written at all.

Then give support a timeline view over it: a rendered sequence of spans with versions and timings, rather than log search. Support saying "retrieval returned nothing for this query" is a closed ticket. Support saying "the AI is broken" is a week.

Closing thought

An answer is an output. The evidence is the trace, and it only exists if you designed it before you needed it.