New here? Start this series at Day 25: Why production knowledge often lives in systems, not PDFs
- Day 25: Why production knowledge often lives in systems, not PDFs
- Day 26: Why evidence is becoming multimodal
- Day 27: Why freshness is part of correctness
- Day 28: Why weak evidence should trigger recovery, not confidence
- Day 29: Why reflection only matters when tied to evidence
- Day 30: Why RAG must be evaluated in parts
"Have the model check its own work" sounds circular, and sometimes it is. Whether it helps or just costs money comes down to one thing: what the check is anchored to.
Anchored versus unanchored
Ask a model "is this answer good?" and you get what you would expect: a thoughtful paragraph concluding that the answer is, on balance, quite good. Unanchored reflection is a model grading its own prose using the same intuitions that produced it.
Ask instead: does source 2 support claim 3?, and you have changed the operation entirely. That is verification against an external artifact rather than self-assessment, and it works because verification is a genuinely easier task than generation. The same model that confidently invented a detail can often catch that invention when forced to line the sentence up against the retrieved text.
One is a vibe check; the other is a comparison with a right answer available.
Make it produce a decision, not a paragraph
The most common way this pattern wastes money: the reflection step runs, produces a considered critique, and nothing downstream consumes it. The answer ships unchanged. You have added latency and a token bill for a second opinion nobody acts on.
Wire in consequences explicitly:
- Unsupported claim → removed from the answer, or marked as unverified
- Multiple unsupported claims → trigger another retrieval pass with a reformulated query
- Core claim unsupported → refuse rather than ship
If you cannot name what changes as a result of the reflection verdict, do not add the step yet.
Budget it
Reflection loops have a natural tendency to sprawl. Generate, critique, revise, critique the revision, and so on, each iteration defensible, the total unbounded.
One pass, on answers that matter. Route by stakes: a customer-facing eligibility determination earns a verification pass; an internal search summary probably does not. A reflection loop with no exit condition is latency with good intentions.
Closing thought
Check the answer before sending it, against the evidence, with clear criteria, and with the authority to change what ships.
Reflection tied to sources is one of the highest-value additions to a RAG pipeline. Reflection tied to nothing is the model admiring its own work at your expense.
If your reflection step returned "unsupported" tomorrow, what would actually happen to the answer?
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion