Context Compaction
Context compaction is the practice of shrinking an AI agent's accumulated context — conversation turns, tool calls, tool results and reasoning — by summarising or removing older material so that the agent can keep working within its model's context window. Anthropic's engineering guidance defines it narrowly as "taking a conversation nearing the context window limit, summarizing its contents, and reinitiating a new context window with the summary"; in practice the term covers a family of techniques for deciding what a long-running agent should still be looking at.
Compaction exists for two reasons. One is capacity: an agent that reads files and calls tools for hours produces more tokens than any window holds. The other is quality and cost: Anthropic's documentation states that "response quality degrades as a conversation grows", and every token left in context is paid for again on each turn. Compaction is one tool within context engineering, and it is lossy by construction.
Approaches
| Approach | What happens | Main risk |
|---|---|---|
| Summarisation | Older turns are replaced by a model-written summary; recent turns may be kept verbatim | The summary omits a detail a later step needs |
| Tool-result clearing | Old tool outputs are replaced by placeholders while the record of the call remains | The agent re-runs tools to recover what was cleared |
| Sub-agent delegation | A sub-agent works in its own window and returns a condensed result to the parent | The parent never sees evidence behind the result |
| External memory | The agent writes notes or files outside the window and reads them back on demand | Notes go stale or are never consulted |
| Recursive decomposition | The input is held outside the prompt and processed by recursive sub-calls over pieces of it | More calls, and dependence on the model's decomposition |
Anthropic's September 2025 engineering post calls tool-result clearing "one of the safest lightest touch forms of compaction", since stale tool output can usually be regenerated. Sub-agents compress by architecture: the same post describes sub-agents that use tens of thousands of tokens exploring and return a distilled summary of roughly 1,000 to 2,000 tokens, the mechanism behind the orchestrator-worker pattern. External memory, which the post calls structured note-taking, moves state into files that survive any number of compactions (see agentic memory). Recursive language models go furthest, treating the long input as data to be examined programmatically; Zhang, Kraska and Khattab (arXiv 2512.24601, revised May 2026) report handling inputs up to two orders of magnitude beyond the model's window this way.
Documented Vendor Features
As of October 2026 Anthropic and OpenAI both expose compaction as an API feature rather than leaving it entirely to the agent harness. Anthropic labels its features beta.
Anthropic. The Claude API documents three related mechanisms. Compaction at a token threshold has the server summarise earlier turns when input reaches a trigger — 150,000 input tokens by default, with a 50,000 minimum — and return a compaction block that replaces what came before; the default summarisation prompt can be replaced. Compaction on demand lets the application decide when to request a summary, can keep recent turns word for word, and can run in the background. Context editing clears content by rule instead of summarising it: one strategy clears the oldest tool results once input passes a threshold (100,000 tokens by default, keeping the three most recent tool uses), another clears old thinking blocks. The summarisation pass is billed as its own iteration, and clearing tool results invalidates the cached prompt prefix.
OpenAI. The Responses API documents server-side compaction, enabled by setting a compact_threshold under context_management, and a standalone /responses/compact endpoint that takes a full window and returns a compacted one. The compaction item it returns is encrypted and, in the documentation's words, "not intended to be human-interpretable"; it is passed back unmodified on the next request.
The designs differ in a way that matters for debugging: Anthropic's summarisation prompt is replaceable, while OpenAI's compaction item is opaque.
Failure Modes
Lossy summaries. A summary keeps what the summariser judged important at the time. Constraints stated once, the reason an approach was rejected, or the exact text of an error are typical casualties, and the agent gets no signal that something is missing. Under Anthropic's threshold compaction on its newest models, thinking from before the compaction block is not carried forward, so "the summary is all the model has of that earlier work".
Behavioural instability. In a preliminary empirical study, Min et al. (arXiv 2608.06503, August 2026) report that recurrent compression "can weaken the influence of recent interactions, increasing blocked actions, repeated exploration, and instability across runs". The results are initial and come from one benchmark, AppWorld.
Regressions from context-management changes. Anthropic's April 2026 postmortem on Claude Code quality describes a change meant to clear older thinking once in sessions idle for more than an hour; a bug cleared it on every subsequent turn instead, making the agent appear "forgetful and repetitive" and causing repeated cache misses. The change had passed code review and tests: context-management code degraded an agent with no change to the model.
Cache and cost interactions. Rewriting the start of the context discards the prompt cache built on it, so a compaction can make the next turns more expensive even though they are shorter (see AI model pricing). With cached input priced at a small fraction of fresh input, keeping a long, stable history can be cheaper than compacting it.
Practical Guidance
The defensible practices follow from the failure modes. Prefer reversible reductions, such as clearing regenerable tool output, to irreversible ones; keep recent turns verbatim; and write plans, decisions and constraints to external memory before compacting. Treat the compaction prompt, thresholds and clearing rules as part of the system under test: changes belong behind the same agent evals as a model upgrade, measured across repeated runs. Larger windows postpone compaction without removing the need for it.
Further Reading
- Effective context engineering for AI agents — Anthropic, September 2025
- Compaction overview — Anthropic documentation, accessed October 2026
- Compaction at a token threshold — Anthropic documentation, accessed October 2026
- Context editing — Anthropic documentation, accessed October 2026
- Compaction guide — OpenAI documentation, accessed October 2026
- Postmortem on Claude Code quality issues — Anthropic, April 2026
- Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability — Min et al., arXiv, August 2026
- Recursive Language Models — Zhang, Kraska and Khattab, arXiv, December 2025 (v3 May 2026)