An internal agent was asked to tidy up a test project. It did exactly that. It also deleted the shared fixtures directory three other projects imported, because the directory sat inside its working path and nothing said it was off limits.

There was no jailbreak, no injection and no bug, just a correctly executed action with a blast radius nobody had measured.

Narrow the tool before you narrow the prompt

Excessive agency is usually an interface problem. A generic run_sql or write_file tool inherits every capability of the system underneath it, and then a paragraph of prompt is asked to hold the line. That paragraph is guidance that happens to work most of the time.

Workflow-shaped tools beat general ones. Ship refund_order(order_id, amount) with a server-side cap rather than execute_transaction. The narrow signature is easier to log, easier to rate limit, and easier to argue about at review time, because its worst case can be written down in a sentence.

For anything with real consequences, split the action in two. The agent produces a diff, a preview, a staged record. A separate call commits it. Staging turns an irreversible operation into a reviewable artefact, and it makes the approval question concrete: a human sees what will change rather than a description of intent.

Attach approval to properties of the request, evaluated in code. Amount over a threshold, records touched above a count, an external recipient, a first-time destination. Never attach it to the model's own assessment of how important the action is.

Containment gets built before you need it

Sandbox the execution environment as if the agent were untrusted code, because functionally it is. Explicit workspace root, no ambient credentials in the environment, an egress allowlist, and bounds on CPU and wall clock. That fixtures directory would have survived a workspace root one level tighter.

Rate limit per user, per workflow and per tool over a rolling window, then alert on the shape of the traffic rather than the totals. Repeated near-identical calls, a sudden widening of argument ranges, or a burst against a rarely used tool are the signals that arrive before the damage.

Two things stay ready. A kill switch per tool that operations can flip without a deployment, and that gets exercised on a schedule so nobody discovers it is broken during an incident. And a documented reverse for every write path. Where an action genuinely cannot be undone, that is precisely the action that requires a human signature.

What is the most damaging single action your agent can take right now without a person seeing it first, and how long would undoing it take?

Checklist · · 18 checks

When to bring in a compliance review

The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.

Checklist · · 29 checks

Security review checklist for an AI feature

What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.