Ask a team what they would change to make their AI feature more reliable, and most will answer with a model name. A newer one, a bigger one, the one that just topped a benchmark.
Then you look at the actual failures and almost none of them are model failures.
The system fed the model the wrong evidence. Or too much of it. Or the prompt was ambiguous in a way nobody noticed until a user phrased something slightly differently. Or the output was free text where the next step needed structure, so a downstream parser guessed and guessed wrong.
I spent the first six days of a sixty-day series on this, and the through-line surprised me more than it should have: the model is the smallest replaceable part of a production AI system, and almost every reliability problem lives in the layer around it.
Here is what that layer consists of.
The model is one component, not the product
Swapping models changes the ceiling of what your system can do. It does not change whether your system knows which document to look at, what it is allowed to do, or what happens when it is wrong.
Those are architecture questions, and a better model answers none of them. This is why "upgrade the model" so often produces a system that fails in exactly the same shape, slightly more fluently. The failure was never in the generation step.
The reframe that helps is to write down what behaviour the model is genuinely responsible for, and treat everything else as your problem. In most production systems that list is shorter than people expect. The model turns evidence into prose, or a request into a structured decision. Deciding which evidence, validating the decision, enforcing what happens next: none of that is the model's job, and none of it improves when the model does.
There is a practical test. If you replaced today's model with one from eighteen months ago, which of your current bugs would still exist? For most teams the honest answer is nearly all of them, which tells you where the engineering is.
Fluency is not reasoning, and the difference is operational
An LLM generates one token at a time, each conditioned on everything before it. There is no plan being executed and no capacity to revise a sentence already produced.
This sounds like trivia until you notice what it rules out. A confident, well-formed answer carries no signal about whether the underlying claim is true, because fluency is a property of the generation process rather than evidence of correctness. The same mechanism that produces a correct answer produces a fabricated one, with identical polish.
Two things follow directly. You cannot use "it reads well" as a quality gate, which is what most manual review actually does. And you cannot rely on the model to notice mid-answer that it has drifted, because noticing would require exactly the backwards look the architecture does not perform.
Both of those have to be built around the model. A verification step that checks claims against retrieved sources. A schema that constrains what a valid answer even looks like. Neither is exotic, and neither arrives by choosing a better model.
Tokenisation is a product constraint wearing a technical costume
Text becomes tokens before the model sees it. Tokens are what you pay for, what fills the context window, and what determines how long a response takes.
The operational sting is that token counts have nothing to do with how important the text is. A rambling retrieved document and a critical safety instruction cost the same per token and compete for the same space. Teams meet this when a feature that behaved in testing becomes three times more expensive against real traffic, because real documents are longer and messier than the tidy samples someone picked for the demo.
It also lands unevenly. Text in some languages tokenises far less efficiently than English, so the same feature can cost noticeably more per request for part of your user base, and hit context limits sooner. That is a product decision that arrived through an implementation detail, which is the pattern worth watching for.
More context frequently makes things worse
The instinct with a large context window is to fill it. Put everything in, let the model work out what matters.
It does not behave that way. Relevant information competes with irrelevant information for the model's attention, and adding weak evidence alongside strong evidence reliably dilutes the answer. A system that retrieves twenty documents and passes all twenty will often perform worse than one that retrieves twenty, ranks them, and passes three.
The failure is quiet, which is what makes it durable. Nothing errors. The answer still arrives, still fluent, now subtly shaped by the four irrelevant documents that happened to be semantically nearby. You only catch it by comparing against a version that passed less.
Treat context as a budget you spend deliberately rather than a container you fill. The question at each step is not "could this be relevant" but "is this more useful than the thing it is crowding out".
A vague prompt becomes a vague system
Prompts stop being text the moment real users depend on them. They become operating instructions, and they inherit every ambiguity left in them.
"Be helpful and concise" reads fine to a human because we silently fill in the intent. A model fills it in differently on Tuesday than on Monday, and differently again for a user who phrased their question in an unusual way. The variation was always present. Testing did not reveal it because the people testing knew what they meant, and unconsciously phrased their inputs to suit.
The useful exercise takes ten minutes. Hand your prompt to a colleague and ask them what a technically compliant but unhelpful response would look like. Not what it should produce. What it could produce while honouring every word. That reading is available to the model too, and it is the one that turns up once real inputs stop resembling your test cases.
Sampling settings are product decisions
Temperature and the related sampling parameters get treated as a knob for writing style. They are really a decision about how much variation your users should experience.
A support assistant that answers the same policy question two different ways has a consistency bug, not a personality. A brainstorming tool that returns identical output every time is broken in the opposite direction. Same parameter, opposite correct settings, and the deciding factor is what the product is for, not what the model prefers.
The version that causes real trouble is a single global setting across a system that does several jobs. One value covering both the summarisation path and the creative-drafting path guarantees one of them is wrong. These belong per use case, chosen deliberately, and written down next to the reason.
The part you actually control
Six ideas, one argument: reliability in AI systems is an engineering property of the surrounding system, and you can improve it enormously without touching the model at all.
That is good news, because the surrounding system is the part you own. You cannot make the model better. You can absolutely decide what evidence reaches it, what shape its output has to take, how much variation is acceptable, and what happens when it is wrong.
It also explains a pattern that otherwise looks strange from the outside: teams shipping genuinely reliable AI features on unremarkable models, while better-resourced teams with frontier access ship things that impress in a demo and frustrate in week two. The difference is rarely access. It is that one team treated the model as a component and engineered the rest, and the other treated the model as the product and kept waiting for the next release to fix things that were never the model's fault.
Read the detail
Each of these is a short standalone post on my blog, with the specifics, the failure modes, and what to do about them:
- Why model choice is rarely the first production AI problem
- Why LLMs feel intelligent but still generate one step at a time
- Why tokenization quietly affects cost, limits, and reliability
- Why more context can make an AI system worse
- Why vague prompts become vague systems
- Why creativity settings are product decisions
They are the opening chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. If you would rather start from the beginning, Start Here lays out all sixty in ten chapters.
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Architecture pattern · · 1 min read
Designing a Multi-Agent System for Real-World Use
How to design and run a multi-agent system in production, using the supervisor pattern, with its trade-offs and where human approval belongs.
Discussion