New here? Start this series at Part 1: Autonomy is a runtime decision, not a model decision
- Part 1: Autonomy is a runtime decision, not a model decision
- Part 2: Your tool registry is an access control list wearing a different name
- Part 3: Four kinds of agent state, and only one of them is memory
- Part 4: A retried agent job is a second chance to send the same email
- Part 5: The 3am agent run that nobody is watching
The first state store is always one table called sessions with a JSON blob in it. It works, right up until someone asks how long a customer's uploaded invoice is retained, and the honest answer is "the same as everything else, because it is all in the blob".
Agent systems store at least four distinct things. They are routinely treated as one because they arrive through the same code path.
What you are actually storing
| State type | Typical lifetime | Safe to reuse on resume? | Deletion driven by |
|---|---|---|---|
| Conversation state | The session | Yes, it is the transcript | User or session expiry |
| Task state | The run | Only the committed steps | Run completion plus a debug window |
| Durable memory | Indefinite | Yes, but re-verify facts | Explicit user action or policy |
| Artifacts | Longer than the run | Read yes, rewrite no | Workspace lifecycle plus retention rules |
Once the columns differ, the argument for a single ungoverned store disappears. A transcript that lives as long as a session and a customer document that must be purgeable within thirty days do not belong in the same bucket just because the same request created them.
Resume is where the mixing hurts
The interesting question is what a resumed run is allowed to trust.
A run that dies mid-flight and restarts will happily reload its cached plan, its earlier tool results and its intermediate reasoning. Some of that is still true. The plan probably is. A tool result captured ninety seconds before a timeout probably is not, because the world moved while the job was dead.
So mark state as durable or ephemeral at write time rather than deciding at read time. Committed steps and their idempotency keys survive a resume. Cached reads and half-finished intermediates get discarded and re-fetched. Getting this wrong is quiet: the run completes, the output looks plausible, and it was computed against a snapshot nobody realises is stale.
Artifacts need provenance, not just a bucket
Files, plans, generated reports and exported data outlive the run that produced them, which means someone will eventually hold one and need to know where it came from.
Store every artifact with the run ID that created it, the identity the run acted under, the workspace it belongs to, and the inputs it derived from. That is four fields, and they turn an unexplained file into something you can audit, attribute, expire or delete on request.
An artifact you cannot trace back to a run is a liability you are storing on purpose. The hard part is deciding, per state type, how long it lives and who may reach it, and that has to happen before the traffic does.
More on these topics
Deep dive · · 6 min read
Adding a second agent does not add intelligence, it adds a contract
Six posts on multi-agent systems, and the failures were never inside an agent. They were between two of them.
Explainer · · 1 min read
Why shared memory becomes a permission problem
Why shared memory becomes an access-control problem in multi-agent systems.
Explainer · · 2 min read
Why agent memory is not one database
Why useful agent memory has to be separated, scoped, and governed.
Discussion