The system prompt says never reveal customer email addresses. The retrieval step puts forty customer records into context anyway, the trace collector stores the whole prompt for debugging, and the observability vendor keeps it for thirty days.
Nothing leaked to a user. The data still moved to three places nobody put on the diagram, and the instruction had no bearing on any of it.
Minimise before the call, never after
Data protection happens at the point the context is assembled. Once sensitive values are in the prompt, every downstream system inherits them.
Fetch fields, not records. A summarisation step rarely needs the account number, the address and the payment history when it needs the order status. Tokenise identifiers on the way in and resolve them on the way out, so the model reasons over CUST_7741 and the interface renders the real name. Secrets never enter context at all: credentials belong in the tool layer, held server side, so the model asks for an action instead of holding the key that performs it.
The paths data leaves by
Exfiltration risk in AI systems is mostly mundane plumbing rather than anything exotic.
- Traces and logs: prompts, tool arguments and completions captured wholesale. Redact at the emitter, before the payload leaves the process, since redacting at the collector means the raw data already travelled.
- Tool destinations: outbound calls that take a recipient, URL or file path. Allowlist them, and alert on first use of any new destination.
- Model output: rendered content that can carry data out through a link or an image reference. Scan and escape it.
- Retention drift: conversation history and vector stores that quietly become the longest-lived copy of data you have. Give them a stated retention period with deletion that is actually verified.
Monitor for volume anomalies too. A single run that reads two thousand records when the workflow typically reads five is the signal worth paging on.
The evidence you will want at 2am
The question in an incident is always the same: what data, whose, how much, and where did it go. Answering it needs artefacts prepared in advance.
Log identity, resource identifiers and record counts per run, deliberately without the payloads themselves, so you can reconstruct scope without the log becoming the next exposure. Keep a data map naming every store a run can touch. Write the deletion and notification procedures before you need them, including who signs off on customer communication.
Then track a few safety metrics as steadily as latency: blocked destinations, redaction failures, approval overrides, anomalous read volumes. They are the ones that show a control weakening while there is still time to fix it.
Closing thought
An instruction not to leak is a wish. A field that was never fetched cannot be leaked at all.
More on these topics
Checklist · · 18 checks
When to bring in a compliance review
The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.
Checklist · · 29 checks
Security review checklist for an AI feature
What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.
Deep dive · · 5 min read
One MCP server is an integration. Thirty is why you need a gateway
Notes from building an MCP gateway. The protocol standardised how agents call tools, then stayed silent about credentials, context budgets and who may call what. That silence gets expensive as the servers multiply.
Discussion