Agent State Persistence: Checkpointing and Resumable Workflows
Agent state persistence explained: checkpoints, resumable runs, replay safety, versioned state schemas, and how to resume without repeating side effects.
A run that dies at step seven of ten should be a resume event, not a tragedy. The difference is state persistence: what the system saved, what it can reconstruct, and what it refuses to repeat.
The engineering is old. Distributed systems have checkpointed long-running workflows for decades, and the patterns are proven. What agents add is sharpness: state schemas change often, actions have side effects, and the decision component is nondeterministic, so the classic questions (what to save, when to replay, how to migrate) become load-bearing.
This article covers the state model, what a checkpoint must contain, the resume contract, replay safety, versioning, and the side-effect discipline that makes all of it safe.
Five Kinds of State
“State” gets used for five different things, and conflating them is the root of most resume bugs:
| Kind | What it is | Lifetime | Persistence duty |
|---|---|---|---|
| Conversation history | Messages across turns | Session | Compact and summarize, do not checkpoint raw |
| Agent working state | Task, trajectory, flags, budgets | One run | Checkpoint every step |
| Workflow state | Position in the larger pipeline | Workflow | Checkpoint at boundaries in hybrid systems |
| Checkpoint | Serialized snapshot of working state | Until superseded | Atomic, versioned, complete enough to resume |
| Durable business state | Real-world outcome records | Forever | Belongs in systems of record, never only in checkpoints |
The bottom row prevents a category of incident: the refund exists in the agent’s checkpoint but not in the payments system, or worse, exists in the payments system but only in a checkpoint nobody can find. Business outcomes live in their systems of record; checkpoints are run-scoped recovery artifacts, not ledgers. The memory layer, what persists across sessions for recall, is a separate discipline from all five: agent memory.
What a Checkpoint Must Capture
A checkpoint is resumable only if it answers six questions. Omit any one, and resume is a guess:
- Schema version. Which state shape this checkpoint speaks. Without it, migration is archaeology.
- Step position. Which iteration completed last, so resume skips what is done.
- Structured working state. The task, findings, flags, and plan status, as typed data, not prose buried in prompts.
- Budgets consumed. Steps, tokens, and wall-clock already spent. A resumed run without this history gets fresh limits, and the budget system quietly doubles.
- Pending side effects. Actions attempted with unknown outcomes, each with its idempotency key, so resume checks before acting.
- Run metadata. Model and tool versions, so the resume decision knows whether the world changed.
Checkpointing Cadence and Cost
The cadence rule is simple: checkpoint after every loop iteration, before the next model call. Any coarser cadence converts a crash into lost steps, and lost steps into repeated side effects. The write cost is trivial relative to a model call, and the storage discipline is ordinary: keep the latest checkpoint plus a bounded history window for debugging, prune the rest, and put retention rules on the store because checkpoints contain run content.
The Resume Contract
Resume is a contract, and writing it down is what separates resumable systems from systems that merely store blobs:
- Resume continues from the last completed step. Completed steps are never re-executed.
- Resume re-checks budgets with consumed amounts. The run’s remaining allowance is computed from the checkpoint, not restarted.
- Resume re-validates and re-authorizes. Tool permissions are checked at resume time, because roles and policies may have changed while the run was parked.
- Resume surfaces partial completion honestly. A run that resumes carries its history; the user and the trace see what was done before the pause.
- Pending effects are checked before action. Anything in flight with an unknown outcome is resolved through its idempotency key before the loop continues.
Human approvals are the most common resume trigger, and the pattern is worth naming: an approval pause is persisted state waiting for input, not a blocked thread holding memory. Durable-execution runtimes make this a first-class property, LangGraph interrupts persist the graph and resume it with the decision, as covered in LangGraph for agents.
Replay Safety: Do Not Assume It
Resume and replay are different operations with different safety requirements, and confusing them is the classic persistence accident.
- Resume continues from a checkpoint without re-executing completed work. Safe when the contract above holds.
- Replay re-executes from the start or from an earlier checkpoint: for debugging, for time-travel inspection, or because someone typed “run it again.” Replay re-runs side effects.
Replay against production side effects is safe only when the side effects were designed for it: idempotency keys that deduplicate, read-only tools, or a sandboxed execution mode. Absent those, a replayed run re-sends the email and re-issues the refund with fresh enthusiasm. The rule is absolute: never replay against production writes, and treat any replay capability as read-only by default.
There is also a subtler problem the nondeterministic component introduces: a replayed agent may choose different steps than the original run, because the model is not a deterministic function. Replay is therefore for inspection and analysis, not for “fixing” a run by redoing it, and the retry rules from reliability engineering apply unchanged: a timeout means unknown outcome, and unknown means check, not repeat.
Versioning and Migrations
Agent state schemas change more often than ordinary workflow schemas, because prompts, tools, and findings structures evolve weekly. The versioning discipline that keeps old checkpoints usable:
- Version tag on every checkpoint. Written at save time, checked at load time. The single most valuable line of data in the whole subsystem.
- Lazy migration on load. Old checkpoints migrate to the current schema when resumed, with migration code reviewed like any other code path.
- Graceful rejection. Checkpoints too old to migrate are rejected with an explicit “restart required” rather than silently misinterpreted.
- A dual-read window during major changes. New code reads both schema versions during rollout, so in-flight runs finish on the shape they started with.
The frameworks document the same discipline: LangGraph covers persistence, checkpointers, graph migrations, and backward compatibility explicitly because checkpointed state is the interface between deploys, and durable-execution platforms such as Temporal are built entirely around the guarantee that code resumes exactly where it left off, across crashes and version changes: LangGraph documentation.
A Checkpoint Sketch
The shape below is illustrative; production systems use a checkpointer from a runtime or a database with transactional writes. The fields are the six-item contract made concrete:
import json
from dataclasses import dataclass, field
from typing import Any, Optional
@dataclass
class Checkpoint:
run_id: str
schema_version: int
step_index: int
working_state: dict[str, Any]
tokens_spent: int
steps_spent: int
pending_effects: list[str] = field(default_factory=list) # idempotency keys
model_version: str = ""
def to_json(self) -> str:
return json.dumps(self.__dict__)
@classmethod
def from_json(cls, raw: str) -> "Checkpoint":
return cls(**json.loads(raw))
def save_checkpoint(store, checkpoint: Checkpoint) -> None:
"""Illustrative store interface: put(run_id, step, payload)."""
store.put(checkpoint.run_id, checkpoint.step_index, checkpoint.to_json())
def load_latest(store, run_id: str) -> Optional[Checkpoint]:
raw = store.latest(run_id)
return Checkpoint.from_json(raw) if raw else None
The Store, the Retention, and the Migration Playbook
Three implementation decisions turn the model into a subsystem:
Choosing the store
A transactional database beats a blob store for the checkpoint itself: per-run keys with step indexes, atomic writes, and cheap latest-checkpoint reads are ordinary database operations. Blob storage still earns its place for oversized artifacts referenced by the checkpoint row, but the row wants transactional semantics, because a half-written checkpoint is a corrupted resume.
Retention with a story
Keep the latest checkpoint plus a bounded history per run for debugging, and expire the rest on a schedule aligned with compliance rules. Completed runs can usually expire earlier than in-flight ones, and a restricted dataset’s checkpoints may need deletion that cascades through telemetry as well. Write the retention rules before anyone asks for them under audit.
The migration playbook
When the state schema changes, the sequence is mechanical: add the new field with a default so old code ignores it; bump the schema version; write and test the migration function against old-checkpoint fixtures; run a dual-read window where resumed runs migrate lazily on load; then close the window and reject the unmigratable explicitly. The sandboxed replay from earlier is how the migration gets tested without touching production side effects.
Resume across deploys
Workers drain on deploy, and runs resume on the new version carrying checkpoints written by the old one. That is exactly why schema versions and model versions live inside the checkpoint: the resumed run can tell what it was built with, and the operator can decide whether to let it finish on the old shape, migrate it, or restart it honestly. Version mixing inside resumed runs is one of those bugs that only appears in production, because only production deploys mid-run.
How Real Systems Do This
- Checkpointing is a runtime feature, not a weekend project. Teams use LangGraph checkpointers or durable-execution platforms, and spend their engineering time on the schema and the side-effect rules instead of the write path.
- Approvals park runs as persisted state. The run resumes with the human’s decision injected, minutes or days later, without a blocked thread in sight, integrated with the deployment architecture.
- Idempotency keys are on every write, always. Resume, redelivery, and retry share one mechanism, reviewed as part of the tool contract rather than bolted on later.
- Old checkpoints migrate lazily with the version tag. A schema change ships with migration code and a dual-read window, and “restart required” is an explicit, handled outcome rather than a crash.
- Partial completion is a first-class status. Budget-exhausted and crash-resumed runs report exactly what was done, so users and downstream systems never mistake resumption for completion.
Decision Framework
For any agent going to production:
- Which runs are worth resuming? Long runs with side effects and approval pauses: yes. Sub-second lookups: restart them, do not resume them.
- What does a checkpoint contain? The six items: schema version, step position, working state, budgets consumed, pending effects, run metadata.
- What is the cadence? Per loop iteration. Coarser cadence is a policy to convert crashes into repeated effects.
- Are side effects replay-safe? Every write keyed, or replay disabled. If neither is true, the system is not allowed to replay, and the UI must not offer it.
- How is the schema versioned and migrated? Tag, lazy migration, dual-read window, explicit rejection for the unmigratable.
- Where do checkpoints live, and for how long? Transactional store, retention rules, and access control, because checkpoints contain run content.
- What does the user see on resume? The run reports its prior progress. Silent resumption is indistinguishable from amnesia.
When NOT to Use This
- Single-call and short read-only tasks. Restarting is simpler, cheaper, and safer than resuming. Persistence machinery here is overhead with no failure mode to prevent.
- Agents whose re-runs are cheaper than resumes. If the whole task costs three seconds and zero side effects, crash recovery is “run it again,” and that is a fine design.
- Schemas still changing weekly. Version the schema once the shape stabilizes; before that, checkpointing every revision multiplies migration surface for little value.
- Side effects that cannot be keyed. If the downstream system accepts no idempotency mechanism, resume must verify-then-act with human escalation, and pretending the writes are safe to repeat is how duplicate incidents are scheduled.
Common Mistakes
- Checkpointing the transcript, not the working state. The consequence: resumes that reread history but lack the structured facts and budgets, and behave like a new run with old paper.
- No schema version on checkpoints. What follows: the first schema change silently corrupts every in-flight run.
- Budgets outside the checkpoint. The price: every resumed run gets fresh limits, and the cost system doubles exactly when reliability is being tested.
- Resume without re-authorization. In practice: a parked run resumes under policies that changed while it slept, including the ones that would now forbid it.
- Replay against production writes. The result: duplicated side effects, delivered by the recovery system with the best of intentions.
- Unbounded checkpoint retention. The damage: a growing store of run content with no deletion story, which is a privacy project nobody scheduled.
Key Takeaways
- Five kinds of state, five duties: history compacts, working state checkpoints, workflow state persists at boundaries, checkpoints version and expire, business state lives in systems of record.
- A resumable checkpoint contains schema version, step position, structured working state, consumed budgets, pending effects with keys, and run metadata.
- Resume continues without re-execution. Replay re-executes and is unsafe against production writes unless every side effect is idempotent by design.
- Version every checkpoint from day one; migrate lazily with a dual-read window; reject the unmigratable explicitly.
- Idempotency keys are the mechanism that makes resume, retry, and redelivery one shared safety system rather than three failure modes.
- Use the machinery that exists: runtime checkpointers or durable-execution platforms. Own the schema, the migrations, and the side-effect rules regardless.
FAQ
What is agent state persistence?
Saving enough structured run state, after every loop iteration, that a crashed, paused, or timed-out run can continue from the last completed step instead of restarting. It covers the state schema, checkpoint contents, resume and replay semantics, and the versioning that keeps old checkpoints usable as the schema evolves.
What is the difference between resume and replay?
Resume continues from the last checkpoint without re-executing completed steps, re-checking budgets and permissions as it goes. Replay re-executes from the start or an earlier checkpoint, which repeats side effects, and is only safe against production systems when every write is idempotent by design.
How often should an agent checkpoint?
After every loop iteration, before the next model call. Coarser cadence converts crashes into lost work, and lost work into repeated side effects. The write cost is trivial next to a model call, and the storage discipline is ordinary retention management.
How do you version agent state schemas?
Tag every checkpoint with its schema version at save time, check the tag at load time, migrate old checkpoints lazily through reviewed migration code, run a dual-read window during major changes, and reject the unmigratable with an explicit restart rather than misinterpreting them.
What makes a checkpoint resumable?
Six things: schema version, step position, structured working state, budgets already consumed, pending side effects with their idempotency keys, and run metadata such as model and tool versions. Omit any one and resume becomes a guess with a UI.
Conclusion
State persistence is what converts agent failures from tragedies into pause events. The machinery is old and proven, the frameworks ship it, and the hard parts are agent-specific: a schema that evolves weekly, actions that touch the real world, and a component that cannot be relied upon to retrace its own steps.
The disciplines that make it work are unglamorous: typed state, six-field checkpoints, per-step cadence, version tags written before they are needed, and idempotency keys on every write so that resume, retry, and redelivery are one safety system rather than three incident reports.
Persist what the loop needs to continue, and key what the world needs to forgive.
Last updated on 2 October 2026
