What we built

A query-level cache in front of retrieval to cut cost and latency. Cache key: a hash of the normalised question. Hit rate was excellent.

What broke

The policy corpus was re-indexed after a change to parental leave rules. The cache did not know. For eleven days, anyone who asked a question someone else had already asked got the old answer, with a confident citation to a document that no longer said that.

Why nobody noticed

Retrieval recall on the golden set stayed perfect, because the golden set was built from the old corpus. Latency dashboards looked better than ever. There was no metric for "answer freshness".

What we changed

Cache key includes corpus version

Every index build writes a version. The version is part of the cache key. A re-index invalidates everything, which is the correct trade.

We also added a freshness metric: the age of the newest source passage behind each answer, plotted against corpus age.

The lesson

A cache in front of a knowledge base stores answers that were true for one corpus version. Put that version in the key.

Key takeaways

  • Cache keys must include the corpus version as well as the query.
  • Staleness is invisible without a freshness metric on answers.
  • The fix was ten lines, but detecting the problem took eleven days.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.