Recursive Language Models vs Context Windows
ComparisonThere are two ways to give a model a very long input. One is to rely on the context window: put everything in the prompt and let attention do the work. The other is a Recursive Language Model (RLM): keep the input outside the prompt, in a code environment, and let the model write programs that examine it and call the model again on fragments. Long context is a property of the model; recursive decomposition is an inference strategy wrapped around it.
The evidence points in both directions. The RLM authors report large gains over direct prompting, including on inputs that fit in the window. A later independent preprint reports the opposite for short inputs: when the context already fits, recursion often makes results worse than simply calling the base model. With 1M-token windows now standard on several frontier models, the practical question is less whether an input fits and more whether the task is dense enough to justify the overhead.
Feature Comparison
| Dimension | Recursive Language Models | Long context window |
|---|---|---|
| What it is | An inference paradigm: the prompt lives in a REPL as a variable and the model decomposes it programmatically and recursively (Zhang, Kraska and Khattab, arXiv v3 May 2026) | A model capacity: all text the model can reference in a single request, including its own output |
| Maximum input | Tested to 6M–11M tokens; reported to handle inputs up to two orders of magnitude beyond the window | Up to 1M tokens on current Claude models listed in Anthropic's documentation, with older models at 200K (as of October 2026); the RLM paper used GPT-5 at 272K |
| Behaviour as input grows | Reported to degrade more slowly than the base model as length and task complexity rise | Accuracy and recall degrade as token count grows ("context rot" in Anthropic's documentation; Chroma measured it across 18 models) |
| Position sensitivity | Not publicly documented | Performance is highest when relevant information is at the beginning or end and drops in the middle (Liu et al., 2023) |
| Inputs that fit in the window: authors' result | OOLONG at 131K tokens with GPT-5: 56.0 | Same task, base GPT-5: 44.0 |
| Inputs that fit in the window: independent result | "Often degrade performance relative to the base model" within the window (Alizadeh et al., March 2026) | The base model is the stronger option at shorter lengths in that study |
| Is recursion the active ingredient? | Disputed: a self-reflective program search without sub-calls matched or surpassed RLMs (Alizadeh et al.); the RLM authors report sub-calls help most on information-dense inputs | Not applicable |
| Cost per query | OOLONG, GPT-5: $0.43 average with high variance ($0.85 standard deviation) | OOLONG, GPT-5: $0.14 average; Anthropic bills 1M-context requests at standard pricing |
| Latency | Multiple model calls, synchronous in the reference implementation; no standard benchmark | Longer inputs have higher time to first token (Google's Gemini documentation); a single call |
| Repeated queries over the same input | Not publicly documented | Context caching reduces the cost of reusing the same tokens (Gemini and Claude documentation) |
| Infrastructure | Code sandbox plus an orchestration library | None beyond the model API |
| Model requirements | Depends on the model writing correct code against the environment | Any model with a sufficiently large window |
Detailed Analysis
The case for decomposition
Large windows do not mean uniform attention. Liu et al. (2023) showed that models use information best at the start or end of a long context and markedly worse in the middle. Chroma's July 2025 study of 18 models found that "model performance varies significantly as input length changes, even on simple tasks". Anthropic's own documentation states that "more context isn't automatically better" and names the effect context rot.
The RLM paper builds on that weakness. Its headline figure shows GPT-5 degrading as input length and task complexity grow while the RLM holds up, and its abstract claims that RLMs outperform vanilla frontier models "even for shorter prompts". On OOLONG, a 131K-token task that fits comfortably inside GPT-5's 272K window, the authors' table reports 44.0 for the base model and 56.0 for the RLM. On OOLONG-Pairs, where the work grows quadratically with input size, the base model scores 0.1 and the RLM 58.0. For inputs that do not fit at all, such as BrowseComp-Plus at 6M–11M tokens, the base model cannot run and the comparison is with other scaffolds.
The critique: recursion can hurt when the input fits
In March 2026, Alizadeh, Shojaee, Cho and Farajtabar (Apple) published the most direct challenge. Their abstract states: "for context lengths within the model's window, RLMs with recursion often degrade performance relative to the base model". Sweeping context length on OOLONG and LongBench-v2 with GPT-5 and Qwen3-Coder, they find the RLM "noticeably more sensitive to context length" and conclude that recursive decomposition "can introduce unnecessary overhead when the context is already manageable".
The same paper questions what produces the gains at long lengths. In its main table, with GPT-5 as backbone, the RLM without sub-calls scores higher than the full RLM on LongBench CodeQA (65.2 vs 59.5) and BrowseComp-Plus (89.7 vs 86.0), and lower on OOLONG (50.5 vs 53.0). Their alternative, which selects among candidate programs using self-consistency, reasoning length and verbalized confidence, improved on the RLM by up to 22% under the same time budget. The RLM authors' ablation is partly consistent with this: they credit the REPL for handling long inputs and sub-calls for information-dense ones, and note one case where the no-sub-call variant won.
The two papers therefore agree on more than their framing suggests. Offloading the input to a programmable environment helps at long lengths. Whether recursive self-calls add to that depends on the task, and at short lengths the evidence conflicts. The critique is a single preprint that has not been peer reviewed, and its reproduction of the RLM may differ in configuration from the original.
Cost and simplicity
A single long-context call is the simplest system there is: no sandbox, no orchestration, predictable billing, and caching when the same document is queried repeatedly. Google's guidance for 1M-token prompts is plain: put the query at the end and cache reused context. The RLM authors report costs comparable to baselines at the median but more expensive on average because of outlier trajectories; on OOLONG the reported averages are $0.43 for the RLM against $0.14 for the base call. Where the input exceeds the window, the alternative is not one call but compaction, retrieval or chunking, each with its own losses.
Best For
Input fits in the window and the question is a lookup
Long context windowOne call is cheaper and simpler, and the independent critique finds recursion often degrades results at these lengths.
Input exceeds the window by a wide margin
Recursive Language ModelsThe base model cannot ingest it; the RLM was tested at 6M–11M tokens.
Dense aggregation or pairwise comparison across the whole input
Recursive Language ModelsThe RLM paper reports its largest margins here (OOLONG-Pairs: 58.0 vs 0.1 for GPT-5).
Many questions against the same document
Long context windowContext caching lowers the cost of reuse; equivalent reuse for RLM trajectories is not publicly documented.
Semantically intensive reading, such as dialogue or narrative understanding
Long context windowAlizadeh et al. report RLMs are less effective where heuristic program search is insufficient and broad contextual understanding is needed.
Mid-length inputs, roughly 100K tokens up to the window limit
Both / dependsThe RLM authors report gains at 131K; the critique reports frequent losses below that. Test both on the actual task.
Strict latency budgets
Long context windowA single call avoids sequential sub-calls, although very long prompts raise time to first token.
The Bottom Line
If the input fits and the task is retrieval-like, use the context window directly. That is the cheaper option, and the one independent study of the question found recursion often made such cases worse. If the input does not fit, or the task requires touching most of a long input, a programmatic approach has clear support: both the RLM paper and its critic find that moving the context into a code environment beats direct prompting at long lengths.
What remains unsettled, as of October 2026, is the contribution of recursion itself. The proposing paper and the critique disagree on short inputs and on whether sub-calls or program selection drive the gains. Anyone evaluating an RLM should include two baselines: the plain long-context call, and the same REPL loop with sub-calls disabled.
Further Reading
- Recursive Language Models (Zhang, Kraska, Khattab; v3 May 2026) – arXiv
- Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context (Alizadeh et al., March 2026) – arXiv
- Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023) – arXiv
- Context Rot: How Increasing Input Tokens Impacts LLM Performance (July 2025) – Chroma
- Context windows – Claude API documentation
- Long context – Gemini API documentation