Most AI features go through a moment where the demo lands perfectly. Someone types a well-formed question, the answer comes back articulate and correct, the room agrees it is basically done, and the conversation moves to timelines.
Then it ships and the numbers are quietly bad. Nothing throws an error. The model did not get worse between the demo and the release. Adoption just fails to happen, and nobody can point at the reason.
I spent days seven through twelve of a sixty-day series inside that gap, and reading them back together, the pattern is uncomfortable: almost every property that made the demo impressive was a constraint the team had not added yet.
Free text demos better than a schema. A blank chat box demos better than a form with four fields. Answering from training memory demos better than refusing because nothing in your sources supports the claim. The demo is not a smaller version of the product. On several axes it is the opposite of one. Here is what has to go back in.
Fluency is the only thing you get for free
Every other failure mode in software announces itself. Exceptions, timeouts, red tests, a 500 in the logs. A confident wrong answer renders beautifully, in the same tone as a correct one, because fluency and accuracy come out of the same mechanism and only one of them is optimised for at generation time.
The dangerous outputs are never the strange ones; strange ones get caught. The expensive ones are a real policy with one wrong number, or a perfectly good answer to a question nobody asked.
And they do not arrive as bug reports. They get absorbed, either by a user who acts on them or by one who notices, says nothing and goes back to doing the job manually. Trust drains long before a ticket is opened.
So the constraint is not a better model. It is a decision about which claims must carry evidence before a human sees them, and a designed path for what the system says when it does not know. Left undesigned, the default behaviour of a model handed an unanswerable question is to answer it anyway.
Prose is a lovely answer and a terrible interface
The moment a model's output feeds another system rather than a person, free text stops being a feature and becomes a parsing problem you maintain forever.
The useful reframe is that structure is not a formatting preference, it is a set of checks you can run before anything irreversible happens. A decision field limited to a closed set means an unexpected value gets caught rather than executed. A field for missing information gives the model a legitimate way to say the evidence was thin instead of inventing something. A field for evidence identifiers makes a claim traceable, and asking for a field pulls harder towards grounding than any instruction in the prompt.
Two things get skipped almost universally. The first is order: teams write the prompt, inspect the output, then build a schema around whatever the model emitted, so the contract ends up shaped by the model rather than by what downstream code needs. The second is the failure branch. Validation is half a contract; the other half is what happens when it fails. Retry with the error attached, fall back to a narrower call, route to a human, fail cleanly. Any of those beats the default, which is partially valid output flowing into code that assumed otherwise.
The demo only ever visits one state
Users never experience your benchmark scores. They experience the wait, the blank box before they know what to type, what happens when it breaks, and whether they can tell a grounded answer from a guess.
A model can be excellent on the benchmarks and terrible at every one of those.
Every one is a state the demo never visits, because the demo is one pass down the happy path run by someone who knows the system. Eight seconds of nothing feels broken. A user facing an empty box does not know what the feature can do, so they ask something trivial and conclude it is trivial. A wrong answer with no route to retry or escalate is a dead end, and people leave dead ends.
The measurement follows. Answer quality is a component metric. Whether the user finished the job is the product metric, and the two diverge constantly: a correct answer that took ninety seconds, or that nobody could verify, completed nothing. Once a team tracks completion, the priority list reorders itself, and the model is rarely at the top.
Conversation is an interface choice, not a starting position
Almost every prototype begins with a chat box, because it is the fastest thing to build and the most impressive thing to show. That is a demo optimisation mistaken for an architecture.
Chatbot, workflow and agent are not a sophistication ladder. They are three answers to one question: how much is this system allowed to do on its own? Two mismatches recur. A chatbot bolted onto a job that needed a workflow makes a user who knows exactly what they want describe it in prose, guesses at the parameters and produces something adjacent, where a four-field form would have been faster and auditable. An agent given autonomy the task never needed takes steps that were fixed and knowable and decides them at runtime, so they vary, cost more and resist tracing.
The asymmetry makes this urgent. Starting narrow and widening is easy. Starting wide and narrowing is a product regression, because users have learned they can ask anything, and every unhandled request now reads as failure rather than out of scope.
Compare the patterns on what you will live with, not how they present: cost per interaction, latency, debuggability, who owns it when it misbehaves. Then write one sentence saying what the system may do unsupervised. If that sentence is hard to write, the choice has not been made.
Retrieval is a decision about authority
RAG usually gets described as letting the model look things up. Accurate, and not very useful: it names the mechanism and skips the decision. The decision is what your system is allowed to answer from.
Left alone, a model answers from training, everything it absorbed, frozen at a cutoff, with no way to know which parts apply to you. Ask about your refund policy and it will confidently produce a refund policy, a plausible average of every one on the public internet. Retrieval does not make the model smarter. It stops the model guessing about things it never knew, a narrower claim than the marketing makes and a much more useful one.
In exchange, answer quality stops being mostly a model property and becomes mostly a retrieval property. A weak model with excellent evidence gives you something decent. A strong model with the wrong three paragraphs gives you an articulate, entirely wrong answer that looks more trustworthy than the first, because fluency scales with model quality while correctness scales with what you fed it. Most escalations filed as hallucination dissect into the model faithfully summarising the wrong pages.
Nobody demos the ingestion pipeline
Everyone wants to tune the retriever. Nobody wants to clean the documents. The order of operations is unforgiving: retrieval quality cannot exceed ingestion quality, and no embedding model or reranker un-breaks content that entered the index broken.
What makes this the most expensive item on the list is that ingestion problems do not fail at ingestion time. The job runs, the count goes up, the dashboard is green. They surface weeks later at query time wearing someone else's uniform: a retrieval miss, a hallucination complaint, an unauthorised disclosure. The team debugs the retriever and rewrites prompts, while the actual defect is that page four of a PDF became a column of single characters in March.
So the gate needs inspection, not a turnstile. Validate extraction against its output, not its exit code. Strip boilerplate, because navigation menus and cookie banners embed just as happily as your content. Capture metadata at entry, because you cannot filter at query time on a field you never stored. And write an exclusion policy, since deciding what stays out is half the job and gets none of the attention.
The cheapest diagnostic available is to open your index and read twenty random entries. Not the config, the stored text. If a human winces at what is in there, the model is already answering from it, and it never winces.
What this adds up to
Six posts, one argument: a demo is judged on its best output, a product is judged on what it does when nothing goes to plan, and the work between those two is almost entirely the work of adding constraints back.
Evidence rules constrain what reaches a user unverified. Schemas constrain what leaves the model. Designed failure states constrain how bad the bad path gets. Pattern choice constrains what the system may do unsupervised. Retrieval constrains what it may answer from. Ingestion constrains what counts as a source at all.
None of those make for a better demo. Together they are the difference between a feature people try once and one they build their work around.
Where the detail lives
Each of these is a short standalone post on my blog, with the failure modes and the tests you can run this week:
- Why confident AI answers still need evidence
- Why production AI needs structured outputs
- Why a capable model is still not a product
- When to use a chatbot, workflow, or agent
- Why RAG is an evidence design problem, not a buzzword
- Why retrieval quality starts before search
They form the second chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. To take it from the top, Start Here lays out all sixty posts in ten chapters.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion