A team spent six weeks building an agent. It read a refund request, checked the order, applied the refund policy, issued the refund, and emailed the customer. Five steps, always the same five, in the same order, governed by rules the finance team wrote down in 2019.
They gave it a planner and a tool loop anyway. Now the same request sometimes takes three tool calls and sometimes eleven, and a failure means reading a trace instead of a stack trace.
The autonomy they paid to build behaves like a bug.
Four patterns answer four different questions
Chatbot fits when the user wants a conversation and success is a useful reply. State is the transcript. The risk is that it answers confidently from nothing.
Workflow fits when you already know the steps. The model does the parts software is bad at (classification, extraction, drafting) at fixed points you control. Control flow stays in your code, which means it stays testable.
Agent fits when the path genuinely cannot be known in advance. The model chooses the tools and the order. You buy flexibility and you pay in determinism, cost variance and debugging time.
RAG fits when answers must be grounded in a corpus that changes faster than any release cycle. It is a grounding strategy, and that is why it composes with the other three.
The common error is treating agent as the advanced tier of the other three. It is a trade, and most teams make it without noticing one was on offer.
The question that decides it
Can you draw the steps on a whiteboard?
If you can, do not pay a model to rediscover them at runtime on every request. Encode them. Autonomy is worth its cost only where the inputs are messy and the branching is too real to enumerate in advance.
The second question is about blast radius. A pattern that can only produce text is a different risk class from one that can move money or delete records. Autonomy and irreversibility are a bad combination, and the fix is usually a narrower pattern rather than a stronger prompt.
Combining without overbuilding
Real systems mix. A workflow with a RAG step is the most common production shape I see and the most underrated: deterministic control flow, grounded content at one node. A chatbot that hands off to a workflow the moment intent is clear beats a chatbot that tries to do everything conversationally.
The discipline is to add each pattern for a stated reason. Looking more serious with all four on the diagram does not count. Every pattern you add brings its own failure mode, its own eval set, and its own on-call story.
Choose the smallest pattern that solves the user's actual problem, then let evidence push you up.
Look at the AI feature you shipped most recently: if you had to defend the pattern choice to someone paying the inference bill, what would you say?
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Architecture pattern · · 1 min read
Designing a Multi-Agent System for Real-World Use
How to design and run a multi-agent system in production, using the supervisor pattern, with its trade-offs and where human approval belongs.
Discussion