New here? Start this series at Day 19: Why reranking is where retrieval becomes useful
- Day 19: Why reranking is where retrieval becomes useful
- Day 20: Why every important AI claim needs provenance
- Day 21: Why RAG needs both broad context and precise evidence
- Day 22: Why one user question may need many searches
- Day 23: Why complex AI questions need decomposition
- Day 24: Why some answers live in relationships, not documents
Users do not ask questions the way documents answer them.
Someone types "is this tool actually worth it?" Your knowledge base holds a pricing table, a feature comparison, and three case studies. Not one of them resembles that sentence, semantically or lexically, and a single retrieval pass returns something disappointing.
Multi-query retrieval attacks that mismatch directly: take the one vague question, generate several sharper ones ("pricing tiers," "feature comparison versus alternatives," "customer outcomes"), search each independently, then merge and deduplicate.
The component you just added
The architecture diagram usually skips this part: you now have a rewriter, and the rewriter is a model making silent judgement calls about what your user meant.
When retrieval mysteriously fails on questions that should have worked, the cause is frequently a rewrite nobody ever looked at. Too narrow, so the good documents fell outside all four variants. Too clever, so it answered a more interesting question than the one asked. Or subtly off-intent in a way that is obvious the moment you read it and invisible if you never do.
So the first operating rule is unglamorous: log every generated variant alongside the original query. Rewrites you cannot inspect are failures you cannot diagnose, and they will present as "retrieval is flaky."
Controlling the fan-out
The cost of this pattern hides well. It shows up as a slow, expensive answer, which teams tend to attribute to the model.
Multiply it out: four variants, each retrieving twenty candidates, each reranked. That is one user question consuming four retrieval calls, eighty candidate evaluations, and a merge, versus one call for a simple lookup.
Three controls worth having from the start:
- Cap the variants. Four good rewrites beat ten mediocre ones, and the tail adds cost without adding coverage.
- Route by need. Specific questions with clear terms should skip expansion entirely. Fanning out "what is our office address" is pure waste.
- Attribute the win. Track which variant produced the evidence that ended up cited. If variant four never contributes across a thousand queries, stop generating it.
The test worth running
Take a deliberately broad question and watch the whole fan-out: the variants generated, what each returned, and what actually reached the final answer.
You are checking one thing: did the extra searches find evidence the single search would have missed, or did they find the same documents four times and bill you for the privilege?
Both outcomes are common. Only one justifies the pattern.
Closing thought
One question can genuinely need many searches. It can also just need one good one. Multi-query retrieval is a real fix for vocabulary mismatch and a real way to quadruple your retrieval bill for nothing. Which one you have built is an empirical question.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion