Most AI architecture reviews ask one question: does it do the thing? Everyone nods, it ships, and the next four questions arrive as incidents.

What the diagram should show

Four boundaries, which are the four days behind this one.

A gateway every model call passes through, so routing, cost and audit have one home. A stated pattern for each feature (chatbot, workflow, agent or RAG), chosen deliberately. A control plane holding policy, prompt versions, memory rules and eval configuration, separate from the request path that consumes them. And explicit human and tenant boundaries, with approvals carrying evidence and events carrying identity.

If a reviewer can point at each of those on the page and name its owner, the review can move on. If any of them is implied, that is where the first incident will start.

Every dependency needs a written degraded mode

Ask what happens when the model provider is slow, the vector index is stale, a tool API is down, or a tenant sends ten times its usual volume. Answer with what the user sees, in words you would put in the interface.

Degraded is not the same as broken. A grounded answer with a "sources may be out of date" banner is a good degraded mode. A confident answer built on an index that stopped updating on Tuesday is a bad one, and the difference is entirely in whether someone designed it. Write the degraded mode next to each arrow on the diagram, and the gaps become embarrassingly obvious.

Rollout and rollback are design decisions

A prompt change is a production change. It deserves a canary, a metric, and a reversal that does not require a release train.

That means versioning prompts, models and retrieval configuration as artefacts you can pin. It means shadow traffic before a cutover so you compare on real inputs instead of a demo set. And it means a rollback that is one action, because the moment you need it you will be reading a dashboard at an awkward hour.

The readiness review

Five questions, and a name attached to each answer.

What is the cost and latency budget for this workflow, and what enforces it? What quality signal would tell us this regressed, and who watches it? Which dependency, failing, produces the worst user experience, and what is the degraded mode? How is a bad rollout reversed, and how fast? Who is paged, and what can support see without an engineer?

Capacity belongs here too. Rate limits are shared across your product, so one enthusiastic new feature can starve an old one. Budget them per workflow before adoption teaches you the hard way.

Closing thought

Feature completeness is the easiest thing to review and the least predictive of how a system behaves under real load. The rest of the list takes an hour, and it is the hour that decides whether launch week is exciting or quiet.

Quiet is the goal.

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.