Agent Memory: Short-Term, Long-Term, and Episodic

Agent memory for AI agents: short-term, long-term, and episodic stores, with write policies, retrieval, freshness, and how memory differs from context.

Executive Summary: Agent memory is mainly a retrieval problem: the right memory often exists but never surfaces when the task needs it. This post separates context, run state, and persisted memory, covers write, retrieval, and expiry policies, and recommends designing retrieval scoring on relevance, recency, and importance before choosing storage.

Give an agent a transcript and it will repeat yesterday. Give it a memory system and it can act like it knows you: it remembers that you prefer morning calls, that the refund already happened, and that this bug is the same bug from Tuesday.

The word “memory” hides three engineering problems that teams routinely conflate: what the model sees now, what the current run carries, and what the system persists beyond runs. Teams that treat the transcript as memory ship agents that get slower and worse as conversations get longer.

This article covers the memory model for AI agents: the taxonomy of short-term, long-term, and episodic memory, the two policies that decide whether memory helps at all (write and retrieval), freshness, and an implementation sketch you can adapt.

One claim anchors everything: memory is useful only when the system can retrieve the right state at the right time. Storage is the easy part. Retrieval is the product.

Memory vs Context vs State

Three layers get called “memory” in casual conversation. They behave differently, fail differently, and belong to different subsystems:

Layer Lifetime Example What goes wrong when confused
Working context One model call The assembled prompt this iteration Bloat: everything gets stuffed into the window
Run state One agent run Trajectory, step count, flags Unresumable runs; see the anatomy of an AI agent
Short-term memory One session Recent turns, intermediate results Grows until quality and cost degrade together
Long-term memory Across sessions Preferences, facts, episodes Stale or poisoned entries surface with confidence

The context window is not memory. It is a budget. What enters it should be selected, not accumulated, and the selection discipline is covered in context window management. The run state and its checkpoints are an execution concern, covered in agent state persistence. Memory is the cross-session layer, and it is this article’s subject.

The Memory Taxonomy

Borrowing the vocabulary from cognitive science, filtered through what is actually buildable:

Type What it holds Example Typical store
Working context What the model sees this call Assembled prompt with selected memories None, rebuilt per call
Short-term Current session working set Last N turns, task scratchpad Session store
Semantic (long-term) Facts and preferences “User prefers email summaries” Key-value or document store
Episodic (long-term) Events with time and provenance “Refund of $42 issued on Oct 2” Append-only event log
Procedural How tasks get done Learned workflows, resolved-incident summaries Skill or playbook library

The distinction that saves the most debugging time is semantic versus episodic. A fact has one current value and should be superseded when contradicted. An episode is an event that happened and should never be rewritten, only reinterpreted. “The user’s plan is Premium” is a fact. “The user downgraded from Enterprise on March 3” is an episode. Storing facts as episodes, or episodes as facts, produces agents that cite expired states as current truth.

Short-Term Memory: The Working Set

Short-term memory is the session’s working set: recent turns, intermediate results, and the task scratchpad. It lives outside the prompt and feeds the prompt. The design question is compression: as the session grows, what gets summarized, what gets pruned, and what stays verbatim.

The failure mode is silent. Nothing crashes when short-term memory grows past usefulness. The agent just starts ignoring earlier instructions, misreading earlier conclusions, and costing more per turn. Summarizing older turns at fixed intervals, while keeping recent turns verbatim, is the standard mitigation, and it belongs to the context budgeting discipline rather than to memory proper.

Long-Term Memory: What Survives the Session

Long-term memory is what the system persists and retrieves across sessions. It is not a transcript archive. It is a curated set of records with structure, provenance, and permissions. Three properties separate a memory system from a log:

  • Structured records. Each memory has a type, a subject, an importance, a timestamp, and a source. Free-text blobs are unfilterable and unrankable.
  • Provenance. Every record knows where it came from: a user message, a tool result, a model summary. Provenance is what makes a memory auditable and correctable.
  • Access scope. Memories are per-user or per-tenant. A memory system without scope boundaries is a data leak with extra steps.

Episodic Memory and the Memory Stream

Episodic memory records what happened, when, and in what context. It is the layer that lets an agent say “we already tried that on Tuesday” instead of repeating the investigation.

The reference design is the Generative Agents paper (Park et al., 2023), which gave LLM agents a memory stream: a growing log of natural-language observations, retrieved by scoring candidates on relevance, recency, and importance, then periodically synthesized into higher-level reflections. The paper’s ablations are worth reading, because removing any scoring component visibly degrades behavior: Generative Agents: Interactive Simulacra of Human Behavior.

A second useful reference is MemGPT (Packer et al., 2023), which treats the context window the way an operating system treats RAM: a small fast tier, managed by paging data in and out of a larger slow tier. It is the cleanest articulation of the idea that the context window is a cache, not the memory itself: MemGPT: Towards LLMs as Operating Systems.

Write Policies: What Deserves Persisting

Writing everything is the default mistake. A memory system that persists every message becomes a noise store where retrieval cannot find signal. Write policies need three components:

  • Selection. Persist user-stated preferences, durable task outcomes, corrections, and surprising events. Do not persist transient chatter, restatements, or routine tool results.
  • Validation. A model-proposed memory is a hallucination candidate. Extract facts through a schema, tie them to a source, and reject anything untraceable. A poisoned memory poisons every future run that retrieves it.
  • Deduplication and contradiction. Near-duplicate facts merge. Contradictions resolve explicitly: the new fact supersedes the old, and the superseded fact expires or becomes an episode of the change.

The write path is also where privacy lives. If the agent stores user data, users need a way to see it and delete it. A memory system without user-facing controls is a compliance problem on a timer.

Retrieval Policies: The Right State at the Right Time

Retrieval is where memory systems succeed or fail, and it has more moving parts than storage:

  1. Query generation. The current task must become one or more retrieval queries, generated by the model or derived by the orchestrator from task entities. This is the same machinery as RAG query generation, covered in RAG for agents.
  2. Scoring. Rank candidates by relevance to the query, recency, and importance. Pure similarity search ignores that yesterday matters more than last year for status questions, and less for preferences.
  3. Top-k selection. Return a handful of memories, never the whole store. Retrieval without a limit is context pollution with extra steps.
  4. Assembly. Retrieved memories enter the prompt labeled as memories, with timestamps, so the model can weigh them rather than treat them as ground truth.

Freshness: How Memories Expire

Stale memory is worse than no memory, because it arrives with the confidence of a system that “remembers.” Freshness needs rules per type:

  • Facts with TTLs. Plans, prices, and statuses change. Expire them on a schedule or on the events that invalidate them.
  • Event-driven invalidation. A plan-change event should retire the old plan fact immediately, not at the next TTL sweep.
  • Episodes never expire, but their weight decays. What happened stays true forever; how much it should influence behavior decays through recency scoring.

The audit for any memory system: take ten questions a returning user would ask, and check whether retrieval surfaces the current truth, the expired truth, or nothing. If expired truth wins, freshness is broken and the agent will confidently misinform.

An Implementation Sketch

The sketch below is illustrative and deliberately small: it shows the record shape, the write policy, and the retrieval scoring in one place. A production version replaces the list with a database, the lexical overlap with embeddings, and adds tenant scoping to every operation.

import time
from dataclasses import dataclass
from typing import Optional

@dataclass(frozen=True)
class MemoryRecord:
    content: str
    kind: str                    # "fact", "episode", "preference"
    subject: str                 # user or entity id, the access scope
    importance: float             # 0.0 to 1.0, set by the write policy
    created_at: float             # unix timestamp
    source: str                  # provenance: message or tool call id
    expires_at: Optional[float] = None

class MemoryStore:
    """Minimal store. Production: database, embeddings, tenant checks."""

    def __init__(self) -> None:
        self._records: list[MemoryRecord] = []

    def write(self, record: MemoryRecord) -> None:
        # write policy: drop near-identical duplicates for the same subject
        for existing in self._records:
            same = (existing.subject == record.subject
                    and existing.kind == record.kind
                    and existing.content == record.content)
            if same:
                return
        self._records.append(record)

    def retrieve(self, subject: str, query: str, k: int = 5) -> list[MemoryRecord]:
        now = time.time()
        scored = []
        for record in self._records:
            if record.subject != subject:
                continue
            if record.expires_at is not None and record.expires_at < now:
                continue  # freshness: expired facts never surface
            relevance = _overlap(query, record.content)
            recency = 1.0 / (1.0 + (now - record.created_at) / 3600.0)
            scored.append((relevance + 0.5 * recency + record.importance, record))
        scored.sort(key=lambda pair: pair[0], reverse=True)
        return [record for _, record in scored[:k]]

def _overlap(query: str, content: str) -> float:
    """Cheap lexical relevance. Production: use embeddings."""
    query_terms = set(query.lower().split())
    content_terms = set(content.lower().split())
    if not query_terms:
        return 0.0
    return len(query_terms & content_terms) / len(query_terms)

Notice what the sketch does not do. It never lets the model write directly: writes arrive through a policy that stamps provenance and importance. It never returns everything: the k limit is part of retrieval, not an afterthought. And it never treats the store as shared: subject scoping is checked on both write and read.

How Real Systems Do This

  • Localization at scale is the taxonomy in one example. In my content localization agent, market-specific adaptation must stay consistent with brand guidance across thousands of pages: brand rules are the durable facts that persist across page runs, and each page adaptation is an episode. The separation earns its keep every time the pipeline runs.
  • Support and customer agents pair a small preference store with an append-only episode log. Each turn retrieves a few records per type; episodes answer “did we already handle this,” facts answer “what is true now.”
  • Coding and operations agents persist resolved-incident summaries as episodic memories, then retrieve them when similar errors appear, turning past fixes into a searchable playbook.
  • Consumer assistant products expose the memory itself: users can list, edit, and delete what the assistant remembers. The write path runs through a strict schema, and deletion is honored everywhere memories fan out.
  • Long-document workloads use tiered context management in the MemGPT style: the window holds the active working set, and a controller pages summaries and evidence in and out.
  • Mature teams measure retrieval. They keep a set of test questions with known expected memories and alert when the right memory stops surfacing, because retrieval quality decays silently as stores grow.

Decision Framework

Answer these in order before writing any memory code:

  1. What must the agent recall next session? If the honest answer is nothing, stop. Short-term compression plus run state is enough, and you have saved yourself a subsystem.
  2. Is it a fact or an episode? Facts get superseded when contradicted. Episodes get appended forever. Choosing wrong produces agents that cite expired states as current truth.
  3. Who validates writes? Model-proposed memories need a schema, a source, and a rejection path. Unvalidated writes are how hallucinations become permanent.
  4. How does retrieval score? Relevance plus recency plus importance, with a hard top-k. If the answer is “similarity search only,” status questions will surface last year’s expired truth.
  5. When do memories expire? Every fact type needs a TTL or an invalidation event. “Forever” is not a freshness policy.
  6. Can users see and delete stored memories? If not, this is a product and compliance decision you are making by accident.
  7. How do you know retrieval works? A fixed set of test questions with known expected memories is the minimum. Unmeasured retrieval decays silently as the store grows.

When NOT to Use This

  • One-shot tasks. If each run starts from scratch, long-term memory adds write cost and privacy surface for zero recall value.
  • Deterministic workflows. Fixed pipelines need run state and checkpoints, not memory. The persistence machinery for that is agent state persistence.
  • Data you are not allowed to retain. If retention rules prohibit storing the content, a clever memory layer does not neutralize the policy. It creates evidence.
  • Retrieval you cannot measure. Without test questions and expected surfaces, you will ship a memory store whose quality nobody can see degrading, and users will feel it before your dashboards do.

Common Mistakes

  • Treating the transcript as memory. The damage: cost grows with every turn while the useful recall stays flat, and the agent rereads its own noise forever.
  • Writing everything. The fallout: a store where retrieval cannot find signal, which behaves like no memory at a higher price.
  • No write validation. The cost: one hallucinated “fact” persisted once and confidently repeated in every future session.
  • No freshness rules. What you get: the agent tells a customer their old plan, their old price, and their old address, with total confidence.
  • Retrieval without limits. Where it lands: context pollution on every turn, and the model weighing a hundred stale records instead of three current ones.
  • No user controls. The consequence: a trust and compliance incident the first time a user discovers what the assistant “knows” about them.

Key Takeaways

  • Memory is a retrieval problem, not a storage problem. Design scoring before schema before engine.
  • Context is a per-call budget, state is run data, memory is curated cross-session records. Three layers, three disciplines.
  • Facts supersede when contradicted; episodes append forever. Choosing the wrong type produces confident misinformation.
  • Write policies matter more than storage engines: selection, validation with provenance, deduplication, contradiction handling.
  • Score retrieval on relevance, recency, and importance, and cap it with top-k. The memory stream design from the Generative Agents paper remains the reference.
  • Stale memory is worse than no memory. Every fact needs a TTL or an invalidation event.
  • User-visible memory with deletion is table stakes, not a nice-to-have.

FAQ

What is agent memory?

Agent memory is the structured, persisted state an AI agent retrieves across sessions: facts, preferences, and events with timestamps, provenance, and access scope. It is not the context window, and it is not the raw transcript. It is a curated store queried selectively when a task needs it.

What is the difference between short-term and long-term memory in AI agents?

Short-term memory covers the current session: recent turns and intermediate results, compressed as the session grows. Long-term memory persists across sessions as validated facts and episodes with freshness rules. Short-term degrades gracefully into summaries; long-term degrades dangerously into stale truth if freshness is ignored.

How do AI agents remember across sessions?

Through an explicit write path and read path. On the write side, the system extracts validated facts and events with provenance and importance. On the read side, each task generates queries, the store scores candidates by relevance, recency, and importance, and a handful of records enter the prompt labeled as memories.

Is RAG the same as agent memory?

No. RAG is retrieval infrastructure over a corpus, usually external knowledge. Memory is persisted state about users, entities, and past runs. They share the retrieval machinery, which is why query generation and scoring look identical, but they answer different questions: RAG answers “what does the knowledge base say,” memory answers “what do we know about this user and this task.”

How do you stop an agent from remembering wrong things?

Validate writes against a schema with provenance before persisting, score and cap retrieval, expire facts on TTL or invalidation events, and let users correct or delete stored memories. Then measure: keep test questions with known expected memories and alert when the right record stops surfacing.

Conclusion

Agent memory is a small database with an unusually hard retrieval problem attached. The storage is ordinary engineering. The difficulty is deciding what deserves persisting, surfacing the right records under the right task, and retiring the ones that stopped being true.

The systems that work treat memory as a product surface: visible to users, measured by tests, and bounded by freshness. The systems that fail treat it as an append-only transcript and discover that everything a transcript gives, it eventually takes back through cost, noise, and confident error.

A memory system is working when the agent finds the three facts that matter and never mentions the ten thousand that do not.

Last updated on 4 October 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *