How to use this
Work top to bottom. The first two items gate everything else: without a golden set and a CI gate, the remaining checks drift within weeks.
Why these twelve
Each one corresponds to a production incident I have seen or caused. The lessons behind them are in the Production RAG track and the stale-cache failure story.
The checks
A golden set of at least 200 questions with known source passages exists and lives in the repo.
Without it every other check is an opinion.
Retrieval recall on the golden set is measured in CI and a drop blocks the merge.
Chunking follows document structure; headings, tables and code blocks are never split.
Hybrid retrieval is on. Dense vectors plus a keyword index.
Product names and error codes do not embed well.
The corpus and the chunking strategy are versioned together.
Every answer carries the source passages it used, visible to the user.
Documents the user is not allowed to read are filtered before retrieval, not after generation.
Retrieval latency is traced per request with OpenTelemetry.
Embedding model version is pinned and recorded with each index build.
A re-index can run without downtime.
The system says "I don't know" when recall is low, and that path is tested.
Someone owns the golden set and reviews additions monthly.
More on these topics
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Discussion