A support assistant tells a customer they qualify for a refund. They do not. The exception sat two paragraphs below the rule the system quoted, and never reached the context.

The first instinct is to call it a hallucination and reach for the prompt. Add "only answer from the provided context". Ship it. The bug returns in a different shape a fortnight later.

It was never a hallucination. Retrieval found a real passage, the model read it correctly, and answered faithfully from evidence stripped of the thing that changed its meaning. No prompt reaches that failure, because it happened before the prompt existed.

Days 19 to 24 sit in that gap, and together they argue something none of them argues alone: once retrieval works at all, the remaining problem becomes choosing the evidence, framing it, and showing which piece produced which claim.

Nearest is not the same as best

Vector search scans millions of chunks in milliseconds, which it can only do because each was embedded once, in advance, knowing nothing about the question it would be asked. What comes back is proximity in a compressed space: a decent way to gather fifty plausible candidates, a poor way to pick the three that should shape an answer.

A reranker does the expensive version, scoring the question and one candidate together to judge how well that passage answers that question. Far more accurate, orders of magnitude slower, impractical across a whole corpus, which is why it belongs behind a cheap wide first pass. Skip it and raw distance decides what the model reads. The model cannot know that passage two was the better match, so it builds a confident answer from whatever arrived, and confidently incorrect output gets misdiagnosed as a model problem.

The caveat is worth as much as the pattern. Reranking only reorders a candidate set. If the right passage is rarely in the top fifty, the problem is recall, upstream. Pull twenty failed queries and check whether the correct passage was in the list at all. That tells you which component you are actually working on.

Citations behave like observability, not polish

Provenance gets filed under presentation, something added once the core experience works. It behaves like logging: the thing you desperately want during an incident and cannot retrofit after one. When an answer causes a problem the questions are always the same three. What did the system claim, what evidence did it use, and was that evidence wrong, stale, out of scope, or simply ignored?

Which means a document title under the answer is not provenance. A useful receipt carries a span rather than a document, because "somewhere in this forty-page PDF" fails the under-a-minute verification test. It carries a source version, because a citation resolving to "the current version" cannot explain an answer given three months ago under different rules. And it carries a retrieval timestamp, because "read the wrong thing" and "read the right thing, which has since changed" are different bugs with different owners.

There is also a metric hiding here that few teams track: citation coverage, the share of material claims backed by a specific span. Uncited claims cluster exactly where the model is improvising connective tissue between real facts, so coverage falling on a class of questions is an early hallucination signal, visible before any user complains.

Match at one size, answer at another

The refund example at the top is a sizing failure, and no chunk size fixes it. Small chunks retrieve well and read badly: the embedding is focused, but a two-sentence fragment pulled out alone has lost the qualifiers that gave it meaning. Large chunks read well and retrieve badly, because the embedding averages six topics into one vague point and precise queries stop matching.

Parent-child retrieval refuses to use one unit for both jobs. Index the small chunks, because that is what makes search precise, but when a child matches, hand the model the parent section it came from, surrounding paragraphs intact. The match happens at street level and the evidence arrives as the whole neighbourhood.

The signal that you need it is specific: your system keeps quoting the rule and missing the exception underneath it. Retrieval was correct. The exception lived in a neighbouring chunk that never restated the rule it was modifying, so it never matched. Anywhere meaning is carried by conditions and carve-outs (policies, contracts, technical documentation, legal and medical content) this stops being an optimisation.

Every extra search adds something that can be quietly wrong

Users do not phrase questions the way documents answer them. Someone types "is this tool actually worth it?" and your knowledge base holds a pricing table, a feature comparison and three case studies, none of which resembles that sentence. Multi-query retrieval attacks the mismatch by generating several sharper questions from the vague one, searching each, then merging.

It works. It also adds a component that rarely appears on the architecture diagram: a rewriter, which is a model making silent judgement calls about what your user meant. When retrieval mysteriously fails on questions that should have worked, the cause is often a rewrite nobody looked at. Too narrow, so the good documents fell outside every variant. Too clever, so it answered a more interesting question than the one asked. Hence the first operating rule, unglamorous: log every generated variant beside the original query.

The cost hides just as well. Four variants, twenty candidates each, all reranked, is one question consuming four retrieval calls and eighty candidate evaluations. That arrives as a slow expensive answer, which teams blame on the model. Cap the variants, route specific questions around expansion, and track which variant produced the evidence that ended up cited.

A bundled question hides the place it broke

"Compare our Q3 churn against our top two competitors and suggest what to fix" is not a question. It is five questions in a trench coat. Embed it and you land at a point in vector space near everything and matching nothing well, so retrieval returns decent material for the loudest part and thin coverage of the rest. The model writes one fluent paragraph across all five topics, grounded on one and improvising four. Nothing in it flags that the competitor numbers came from nowhere.

Decomposition identifies the sub-questions, retrieves for each independently, and merges only once each part has its own support. The retrieval gain is obvious, since each search finally has one target. The operational gain matters more: every sub-answer becomes a checkpoint. When the output is wrong you can point at the step that failed instead of wondering which third of a paragraph to distrust.

It fails in two directions. Decomposing everything trades latency and cost for nothing, because simple questions deserve an answer rather than a five-step plan. Unbounded depth lets one hard query consume an afternoon of compute while the trace shows every step was reasonable. The skill lies in judging when a question is actually a project, rather than in the decomposing itself.

Some questions are not lookups at all

"Which of our clients would be affected if this vendor went under?" is not in any document. It lives in the connections between them, vendor to contracts, contracts to services, services to clients, and similarity search cannot traverse a connection. Vector retrieval answers what text is similar to this. Graph retrieval answers what is connected to this and how, walking a path no similarity score would make, because the destination shares no vocabulary with the question.

So if your hardest questions sound like "what depends on Z", better embeddings will not rescue you. But the cost accounting is brutal and rarely done upfront. Somebody has to build the graph, including the identity resolution that decides Acme Corp and ACME are one node. Somebody has to maintain it, because a stale graph does not degrade gracefully, it answers with total confidence using last quarter's org chart. And a four-hop answer needs its path shown or nobody can check it.

The test is a spreadsheet, not a proof of concept. Take the twenty questions your users most want answered that your system handles worst, and sort them: genuinely relationship-shaped, or document lookups that simply retrieved badly? Mostly the latter, in my experience, and chunking, hybrid search and reranking move those numbers further for a fraction of the cost.

What this adds up to

Six patterns, one argument: each replaces a single opaque retrieval step with a sequence of decisions you can name, inspect and attribute, and that traceability is what makes an answer defensible rather than merely fluent.

Notice what each adds to the system. A reranker deciding which passages matter. A rewriter deciding what the user meant. A decomposer deciding what the real questions are. A graph builder deciding which entities are the same entity. Each is a new place to be wrong, and each makes its call silently unless you log it. Which is why provenance is less the second item on this list than the thing holding the other five up.

It also explains why several of these posts end with the same instruction: check whether you actually need this, on your own queries. Every pattern here buys accuracy with latency, cost and complexity. Applied by default they produce an expensive system that answers easy questions slowly. Applied where the failure evidence points, they are the difference between an answer you hope is right and one you can defend.

The six posts behind this

Each is a short standalone note on my blog, with the specifics and failure modes:

Together they form Evidence-Driven RAG Patterns, the fourth chapter of 60 Days of Production AI Systems. If you would rather take it from the top, Start Here lays out all sixty posts in ten chapters.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.