New here? Start this series at Day 25: Why production knowledge often lives in systems, not PDFs
- Day 25: Why production knowledge often lives in systems, not PDFs
- Day 26: Why evidence is becoming multimodal
- Day 27: Why freshness is part of correctness
- Day 28: Why weak evidence should trigger recovery, not confidence
- Day 29: Why reflection only matters when tied to evidence
- Day 30: Why RAG must be evaluated in parts
A RAG system with stale documents does not fail loudly. It answers beautifully, from last year.
That is what makes staleness the most under-managed failure in production RAG. Retrieval succeeded, citations resolve, and groundedness scores are excellent. Every metric on your dashboard is green, and the answer quotes a policy that was replaced in March.
The failure with no error message
Compare it to the alternatives. A retrieval miss produces a visibly weak answer. A permission leak triggers an incident. An outage pages someone.
Staleness produces a confident, well-sourced, completely correct-looking answer to a question whose truth has changed. Nothing in the system knows. You find out from a customer, from legal, or from the support agent who noticed the bot quoting a discount that ended two quarters ago.
Correctness expires, and for prices, policies, versions, headcount, contractual terms, and anything with a legal dimension, the timestamp is part of the answer.
Run the knowledge base like a newsroom
An archive preserves. A newsroom updates and retires. RAG needs the second posture:
- Ingestion has monitoring and an owner. "The sync job died in April and nobody noticed until August" is a story most RAG teams eventually tell. Alert on documents-ingested dropping to zero, not just on job failures. A job that succeeds while processing nothing is the quieter version.
- Chunks carry effective dates, and retrieval can prefer recent sources for time-sensitive question types.
- Superseded content is expired, not left to compete. When policy v3 lands, v2 must lose. If both sit in the index with equal standing, retrieval will sometimes pick the old one, and it will be right to: v2 may well be the better semantic match.
- Time-sensitive answers show their vintage. "As of the March 2026 policy…" turns a silent failure into a claim the reader can check.
That last one is the cheapest trust mechanism in RAG, and it is routinely skipped because it feels like clutter. But it is what separates an answer a user can validate from one they must simply believe.
Measure the lag
There is a concrete number worth knowing and almost nobody tracks it: how long does it take for a changed document to change your system's answers?
Update one document. Ask the system about it every hour. The gap you measure is the window your users are living in, and it is usually longer than anyone on the team assumed. Index lag plus cache plus embedding queue adds up.
If that number is a day, say so in your product. If it is a month, that is a roadmap item.
Closing thought
Freshness tends to get filed as an infrastructure concern that sits below correctness. For a large class of questions, it is correctness.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion