Context Window Management: Summarization, Compression, Pruning
Context window management for AI agents: budgeting, summarization, compression, and pruning, so agent runs keep the right state in the model's working set.
Agent teams treat the context window as free storage until the day it is not. Then the agent forgets its instructions, misreads its own earlier findings, and costs more per step than the first ten steps combined. Nothing crashed. The window just filled with the wrong things.
The context window is a budget, and long agent runs are budget discipline problems. Select what enters the window instead of accumulating it. What stays should earn its place. What leaves should leave a receipt.
This article covers context window management for agents: what belongs in the window, how to budget it, and the three levers (summarization, compression, and pruning) that keep long runs coherent and affordable. It is the per-call sibling of agent memory: memory decides what the system knows, context management decides what the model sees.
The Context Window Is a Budget, Not a Feature
A larger window feels like a solution, but it changes the economics, not the discipline. Bigger windows cost more per call, and the cost compounds because the whole context is reprocessed on every iteration of the loop. An agent that carries 80,000 tokens of trajectory through ten model calls pays for 800,000 tokens of input, regardless of what the task needed. Why every token is reprocessed comes down to how models consume input, explained in how LLMs work: tokens, context, and embeddings.
Two failure modes hide behind a full window. The obvious one is the provider error, the run that dies at step seven because observation four was a 40-kilobyte JSON dump. The subtle one is quality decay: attention spreads across everything present, so irrelevant context does not sit quietly. It dilutes. Agents get worse as their windows get fuller, and the decline is gradual enough to blame on the model.
The research framing worth internalizing is MemGPT’s operating-system analogy: the context window is fast memory, and everything else is slow memory, with explicit paging between them (Packer et al., 2023). You can read the paper as engineering advice: MemGPT: Towards LLMs as Operating Systems. The window is a cache of the run’s working set, and caches need eviction policies.
What Belongs in the Window
Ask of every token: does the next decision need this? The answers sort into a budget table that should exist before the run starts:
| Section | Contains | Typical share | Policy |
|---|---|---|---|
| System instructions | Role, rules, tool guidance | 10 percent | Fixed, never summarized |
| Task | The user request and constraints | 5 percent | Verbatim, never summarized |
| Working state | Structured facts: plan, findings, flags | 15 percent | Explicitly versioned |
| Evidence | Retrieved passages, tool results | 30 percent | Bounded at the source, citeable |
| Trajectory | Recent turns and observations | 40 percent | Recent verbatim, old summarized |
The shares are a starting point, not a law. The principle: never compress the task and instructions, because losing them loses the run. The trajectory is the compressible majority, because old steps are scaffolding once their outcomes live in working state.
Summarization: Controlled Lossy Compression
Summarization is the primary lever for the trajectory section. Old turns become a summary; recent turns stay verbatim. Two design choices decide whether it helps or quietly destroys the run.
First, what the summary must keep: constraints, decisions already made, their outcomes, and open questions. What it may drop: the transcript of how it got there. A summary that omits a constraint produces an agent that violates it, confidently, on step nine. Second, the shape: prefer structured summaries over prose. A fixed schema (task state, decisions taken, constraints in force, open questions) is stable across repeated compactions. Prose summaries of prose summaries lose fidelity every generation, like photocopies of photocopies.
def compact_history(messages: list[dict], keep_recent: int,
summarize) -> list[dict]:
"""Summarize old turns, keep recent turns verbatim. Illustrative."""
if len(messages) <= keep_recent:
return messages
old, recent = messages[:-keep_recent], messages[-keep_recent:]
summary = summarize(old) # constraints, decisions, open questions
return [{"role": "system",
"content": f"Progress so far: {summary}"}] + recent
Trigger compaction on thresholds, not vibes: a turn count or a token watermark. Log compaction as a trace event. A run whose context silently changed between step four and step five is a debugging mystery until the trace says otherwise.
Pruning and State Promotion
Summarization shrinks history. Pruning removes it, and the companion move is promotion: when a turn produces a durable fact, write it to structured state and drop the turn that produced it.
The distinction matters for long runs. A step that queried the order status has produced one durable fact: the order shipped two days ago. That fact belongs in the working state section as a field, and the 3,000 tokens of query, reasoning, and raw API response that produced it belong nowhere. Plans deserve the same treatment: a plan is an object with steps and status, not a paragraph in the transcript, which is the point made in planning and reasoning patterns.
Promotion has a second benefit that outlives the run: anything promoted to state can be persisted, resumed, and tested. Anything left in history disappears the moment history compacts, and no checkpoint can capture it usefully.
Compression and Output Discipline
The cheapest compression happens before ingestion. Bound every observation at the source: project tool results to needed fields, cap retrieval passages at assembly, and truncate with markers rather than silently. This is the discipline described for tool results and retrieval assembly in RAG for agents, and it is the highest-leverage context decision, because a payload that never enters the window never costs anything again.
For what must enter, compress structurally: shorter field names in serialized state are trivial but real, and deduplicated repeated content, such as the same document retrieved twice, beats reprocessing it twice. Provider-side prompt caching can reduce the cost of stable prefixes, but it reduces the bill, not the window’s attention budget. Treat it as a cost lever, not a discipline substitute.
Handling Overflow Gracefully
The provider should never be the one to discover overflow. Track estimated tokens per iteration before the call, and walk a degradation ladder when the budget is threatened:
- Prune completed-step content whose outcomes already live in state.
- Summarize the next-oldest chunk of trajectory.
- Drop the lowest-value evidence section, marked as dropped.
- Escalate or split the task rather than proceed with a hollowed-out context.
The last step is the one teams skip, and it is the one that matters. Silent truncation produces an agent that reasons over gaps it cannot see, which manifests as confidently wrong actions, not as context errors. A run that cannot fit its needed context should say so and stop or split, the same escalation honesty required anywhere else in agent engineering.
The Cost Connection
Context discipline is cost discipline. Input tokens are reprocessed on every iteration, so window size multiplies by step count: a 30,000-token context carried through eight model calls pays for 240,000 input tokens before any output is counted. The levers rank as follows:
- Smaller context beats cheaper models: it reduces cost on every call and usually improves quality by removing noise.
- Prompt caching helps where prefixes are stable and the provider supports it. It discounts the bill, not the attention budget.
- Model routing sends compaction and routine steps to cheaper models, reserving the strong model for decisions. The full treatment is in cost control: token budgets, caching, and model routing.
Teams that treat context management as a cost project discover it is also the quality project. The same tokens that inflate the bill dilute the attention, and removing them fixes both at once.
How Real Systems Do This
- Support agents keep a structured case state object and compact conversation history on a turn threshold. Nobody summarizes the instructions or the task; the case state carries the durable facts.
- Coding agents window the repository instead of loading it: file views, diffs, and focused excerpts, with findings promoted to state. The context is a working set of code, not a transcript of a filesystem.
- Research agents prune search transcripts once findings are recorded. The old queries and result pages leave the window; the synthesized findings, with sources, stay.
- Long-running systems log compaction events in the trace, so a quality difference between step four and step five has a visible cause instead of a suspected one.
- Practitioner guidance agrees on leanness. Anthropic’s engineering guidance on building effective agents recommends the simplest solution that works and warns that added complexity must be justified, which is the same rule applied to tokens.
Decision Framework
For any agent that runs more than a few steps:
- What does the next decision need? Select on this basis. Everything else is a candidate for pruning or summary.
- What are the section budgets? Write them down: instructions, task, state, evidence, trajectory, with shares.
- What is never compressed? Instructions and task. Loss there is loss of the run.
- When does compaction trigger? A token or turn threshold, logged in the trace.
- Where do durable facts go? Promoted to structured state, never left in history.
- What happens at overflow? The ladder: prune, summarize, drop with markers, escalate or split. Never silent truncation.
- Is the window measured? Tokens per iteration and coherence failures, tracked per task, tell you when the budget is wrong before the users do.
When NOT to Use This
- Short runs. A three-step agent with bounded observations needs a budget awareness and nothing more. Compaction machinery is overhead with no return.
- Single-call tasks. There is no trajectory to manage. Assemble the context, call once, done.
- No state schema to promote into. Promotion needs a destination. Define the state contract before building compaction on top of prose.
- No measurement. Without tokens-per-iteration and coherence metrics, context decay is invisible until users report that the agent “forgot” something it was never given room to remember.
Common Mistakes
- Everything verbatim forever. The result: cost that scales with history and quality that decays with it, blamed alternately on the model and the bill.
- Prose summaries without structure. The damage: fidelity loss on every compaction, until the agent’s own history is a rumor.
- Summarizing the instructions or the task. The fallout: a technically running agent that has quietly lost its purpose.
- Silent truncation. The cost: reasoning over invisible gaps, which surfaces as confident nonsense rather than an error.
- No overflow ladder. What you get: the provider error is the first warning, and the run dies holding state that cannot be resumed.
- Promotion deferred. Where it lands: durable facts live in compacted-away history, and resumability plus auditability die with them.
Key Takeaways
- The context window is a per-call budget that multiplies by step count. Fill it by selection, not accumulation.
- Budget by section: instructions and task are never compressed; working state is explicit; evidence is bounded at the source; the trajectory is the compressible majority.
- Summarize old turns with structured summaries that preserve constraints, decisions, and open questions. Trigger on thresholds and log the event.
- Prune and promote: durable facts belong in structured state, not in history. What is promoted can be persisted, resumed, and tested.
- The cheapest compression is refusing oversized payloads at the source, in tool results and retrieval assembly.
- Handle overflow with a ladder: prune, summarize, drop with markers, then escalate or split. Never let the provider discover your overflow.
- Context discipline is cost discipline: smaller contexts usually beat cheaper models on both cost and quality.
FAQ
What is context window management for AI agents?
The discipline of deciding what the model sees on each call of a run: budgeting the window by section, summarizing old trajectory, pruning completed steps, promoting durable facts to structured state, and bounding evidence at the source, so long runs stay coherent, resumable, and affordable.
How do you summarize agent history?
With structured summaries on a fixed threshold: keep recent turns verbatim, compress older turns into a fixed schema that preserves constraints, decisions made, outcomes, and open questions. Avoid prose summaries of prose, because each compaction generation loses fidelity.
What is the difference between pruning and compression?
Pruning removes content from the window entirely, usually completed steps whose outcomes already live in structured state. Compression shrinks content that stays, through summarization or bounded re-assembly. Prefer pruning when the information has a durable home, and compression when it does not.
How do you prevent context overflow in agents?
Bound observations at the source, budget sections with written shares, trigger compaction on thresholds before the window fills, track estimated tokens before each call, and keep a degradation ladder ending in escalation or task splitting. Overflow should be a handled condition, never a provider error.
Does a bigger context window solve the problem?
No. A larger window raises the cost of every iteration and does not fix selection. Attention still spreads across everything present, so irrelevant context degrades decisions, and the bill scales with the fill. Discipline is the solution at every window size.
Conclusion
Context management is the least visible and most compounding discipline in agent engineering. Every other subsystem (tools, retrieval, memory, planning) deposits its output into the window, and the window reprocesses all of it on every step. A team that manages it well pays for exactly what each decision needs. A team that does not pays for everything, forever, with interest in both cost and quality.
The mechanics are simple enough to state in one paragraph: budget by section, keep instructions and the task verbatim, summarize the old trajectory into structured summaries, prune what state already holds, bound everything at ingestion, and escalate rather than truncate. The discipline is applying them before the first long run teaches them to you.
Fill the window by selection, not accumulation. What the model no longer needs should be state, summary, or gone.
Last updated on 1 October 2026
