The answer is wrong, so the team edits the prompt. Still wrong, so they try a bigger model. Still wrong, and two weeks are gone.
The relevant document was never retrieved, and nothing downstream could have rescued that. Generation cannot ground itself in text it never received, and no amount of prompt work makes a missing chunk appear.
Retrieval needs its own test set and its own number. Fifty to two hundred real questions, each labelled with the document IDs that should come back, scored on recall at k before a language model is involved at all. That one metric tells you whether a bad answer is a retrieval failure or a generation failure, and teams without it guess wrong most of the time.
Vector search loses exactly the queries enterprises type
Semantic similarity is excellent at paraphrase and poor at identity.
"Error code PS-4471", "the Q3 amendment to MSA-2019-114", part numbers, ticket IDs, surnames: these are tokens that mean one specific thing, and embedding them into a dense vector smears them across a neighbourhood of things that merely look similar. BM25 finds them instantly. Enterprise corpora are full of them.
So run both and fuse the results, usually with reciprocal rank fusion, which needs no score calibration between the two systems. Then rewrite the query before it reaches either index: expand acronyms, resolve pronouns against conversation history, and for ambiguous questions issue two or three variants and union what comes back. A good share of "the model does not understand our domain" complaints are really a query that arrived at the retriever stripped of context the user assumed was obvious.
Embeddings and indexes are versioned dependencies
An embedding model is a dependency that silently rewrites your entire index when it changes.
Version it explicitly, store the model ID on every chunk, and treat migration as a planned operation: build the new index alongside the old one, run the retrieval test set against both, compare recall, then cut over. Re-embedding in place gives you a corpus where half the vectors live in one space and half in another, and similarity scores across that boundary are noise.
Index parameters deserve the same discipline. HNSW ef_search and M, or IVF nlist and nprobe, are latency against recall trade-offs. Do not just inherit the defaults from a tutorial. Changing them changes quality with no error message anywhere. Rebuilds should be rehearsed before they are urgent, because the first time you need one will be mid-incident.
Make retrieval legible while you are at it. A trace should show the rewritten query, the candidates each retriever produced, their raw scores, the fused ranking, and what the permission filters removed. Debugging retrieval without that is archaeology.
Quality problems that feel like model problems are usually recall problems, and recall is something you can measure today.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion