The first eval set is usually forty rows in a spreadsheet, written by one engineer in an afternoon, drawn from queries they happened to remember. It is genuinely useful for about six weeks. Then it starts lying, quietly, in the direction of whatever the team already believes.
What actually belongs in the set
Representative traffic is only the floor. If the set mirrors production volume, it will be ninety per cent easy cases and the average will drown every failure worth finding.
Build it as three deliberate slices and score them separately:
- Common cases, enough to notice a broad regression
- Long tail: odd formats, multi-part questions, the languages nobody planned for
- Adversarial: injection attempts, questions with no answer in the corpus, requests that should be refused
The third slice is the one teams skip, because adding it makes the dashboard look worse. That is precisely its job.
Labels are only as stable as the guide
Two reviewers will label the same borderline answer differently, and neither of them is wrong. They are applying different definitions because nobody wrote one down.
So write the reviewer guide before labelling, with worked examples of a pass, a fail and the ambiguous case that sits between them. Measure agreement between reviewers on a shared subset. Persistent disagreement is a specification bug rather than a people problem, so fix the guide, then relabel.
Version the data and the rules together
The biggest time sink is a comparison across versions. The labelling guidance is tightened in March, the set is relabelled, and in April someone compares the new score to February's as though nothing changed. The model may not have moved at all.
Treat the dataset and its guide as one versioned artefact. Every example carries provenance: where it came from, who labelled it, under which version of the rubric. When a comparison spans a version boundary, say so on the chart instead of hoping nobody notices.
The set should also grow on purpose. Production failures that reached a user get triaged back in as new cases, with the corrected output attached. That way the set reflects your users rather than your assumptions.
Keep it out of the tuning loop
An eval set that has been used for prompt tuning is a training set wearing a costume. Every round where someone read the failures and edited the prompt to fix them has fitted that prompt to that data, and the score has stopped being evidence.
Hold back a slice nobody optimises against and nobody browses. Rotate it rarely and deliberately. It is the only number you can quote outside the team without a caveat.
All of this needs one owner: the guide, the versions, the intake of new failures. Shared custody means nobody notices when it goes stale.
When did your eval set last gain an example that came from a real user complaint?
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Explainer · · 2 min read
Real users ask questions your test set never imagined
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Discussion