Two weeks before launch, someone asks whether the assistant is safe. The answer comes back as a link to the moderation endpoint and a screenshot of the system prompt.
That is one component. Safety is a property of the whole system, and it gets designed alongside the retrieval layer and the tool schema rather than bolted on during the review meeting.
Threat model before you choose controls
Reaching for a filter before you know what you are protecting produces the wrong filter. Three questions come first, and they take an afternoon.
Which assets are exposed? Customer records, internal documents, credentials, and the ability to spend money or send messages on someone's behalf. The list is usually longer than the team expects, because tool access quietly turns a read-only assistant into a write-capable one.
Who are the actors? Outside attackers belong on the list, and so do a curious employee, a confused user, a compromised integration, and the model itself on a bad day.
What are the abuse cases? Write them as concrete sentences: a support agent uses the assistant to read a customer file they have no ticket for. Abuse cases are testable in a way that "prevent misuse" never is.
Controls sit in four layers
Once the threats have names, placing controls stops being a matter of taste.
- Model: refusal behaviour, instruction hierarchy, constraints on what may be produced
- Tools: scoped permissions, staged actions, approval gates
- Data: minimisation before the call, redaction in logs, retention limits
- Interface: what the user must confirm, and what is clearly labelled as generated
A threat with no control in any layer is an accepted risk. That is a legitimate choice when it has been recorded and reviewed, and a liability when nobody noticed.
Ownership, and what the launch gate should ask for
Policy that lives in a document is not enforcement. Each boundary needs a named owner who can say what it blocks, where it runs, and what happens when it fires. Where ownership is vague, the boundary degrades quietly through six months of shipping pressure and nobody can point to the change that weakened it.
Governance review works the same way. The useful question at the gate is "show me the test where someone tried this and the control held", rather than "have we considered safety". Evidence means a recorded abuse case, the control that caught it, the log line it produced, and the named reviewer for the exception path. Three of those artefacts are worth more than a page of principles.
Closing thought
Safety architecture is boring in the best way: assets, actors, abuse cases, controls, owners, evidence. A team that cannot fill those in is hoping rather than making a risk decision.
More on these topics
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Checklist · · 18 checks
When to bring in a compliance review
The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.
Checklist · · 29 checks
Security review checklist for an AI feature
What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.
Discussion