Skip to content

Five-day tracks

5 Days of AI Safety Engineering

Safety as architecture rather than moderation: threat models, agent identity and scope, untrusted input, blast radius and data exfiltration.
10 min in total
  1. 01

    Part 1 · Explainer · 2 min read

    Safety is an architecture decision, not a moderation setting

    Assets, actors, abuse cases, controls, owners, evidence. Six things a team can actually write down before launch.

  2. 02

    Part 2 · Explainer · 2 min read

    Your agent should not be a superuser

    Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.

  3. 03

    Part 3 · Explainer · 2 min read

    Untrusted text does not get to give orders

    Injection defence is layered validation and clear authority, not a better-worded system prompt.

  4. 04

    Part 4 · Explainer · 2 min read

    Blast radius is a design parameter

    Narrow tools, staged writes, sandboxes and a tested kill switch. Containment is built before it is needed.

  5. 05

    Part 5 · Explainer · 2 min read

    You cannot instruct a model into data protection

    Minimise what enters context, watch the paths data can leave by, and keep the evidence an incident will demand.