Recursive Language Models vs Retrieval-Augmented Generation
ComparisonRecursive Language Models (RLMs) and retrieval-augmented generation (RAG) are two answers to the same problem: a model has to work with more text than it can usefully read in one pass. RAG selects first and reads second: a retriever picks a handful of passages from an index and the model generates from those. An RLM does not select in advance. The whole input is loaded as a variable in a code environment, and the model writes programs that inspect it, split it, and call the model again on the pieces.
The short answer: RAG is the cheaper and far more established choice when the answer lives in a few passages of a large, changing corpus. RLMs are aimed at tasks where most of the input matters, such as aggregating or cross-referencing across thousands of entries, and where a top-k retrieval step would discard needed evidence. The public evidence for RLMs is promising but, as of October 2026, rests on a small number of papers and has been challenged.
Feature Comparison
| Dimension | Recursive Language Models | Retrieval-Augmented Generation |
|---|---|---|
| Core mechanism | The prompt is treated as part of an external environment; the model writes code to examine and decompose it and recursively calls itself on snippets (Zhang, Kraska and Khattab) | A retriever fetches passages from an index and a generator conditions on them; the original formulation combined parametric memory with a dense vector index of Wikipedia (Lewis et al.) |
| First published | arXiv preprint, 31 December 2025; third version 11 May 2026 | arXiv preprint, 22 May 2020 |
| Selection step | Decided at inference time by model-written programs (search, slicing, sub-calls) | Decided by the retriever before generation |
| Preprocessing | None required: the input is loaded into a REPL variable | A corpus must be chunked and indexed, and the index kept current |
| Input scale tested | 6M–11M tokens per task on BrowseComp-Plus (1K documents); the authors report handling inputs up to two orders of magnitude beyond the context window | Bounded by the index rather than the context window; the 2020 paper indexed Wikipedia |
| Information-dense aggregation | OOLONG, GPT-5: 56.0 (authors' table) | OOLONG, GPT-5, CodeAct agent with BM25 retrieval: 38.0 (same table; not a tuned dense-retrieval pipeline) |
| Head-to-head evidence | BrowseComp-Plus (1K), GPT-5: 91.3% | Same benchmark, CodeAct with BM25: 51.0%. One paper, one retrieval configuration |
| Cost profile | Average $0.99 per BrowseComp-Plus query with GPT-5 (standard deviation $1.22); the authors warn of "exploding sub-call costs" in outlier runs | "Significantly lower cost" than feeding full long contexts (Li et al., 2024); per-query figures depend on the pipeline |
| Runtime requirements | A code sandbox; the reference library's default local REPL runs Python exec() and is described as unsuitable for production | A retriever and index (for example a vector database); no code execution |
| Latency | Sub-calls are synchronous in the reference implementation, which the authors list as a limitation; no standard latency benchmark | Not publicly documented as a general figure; depends on retriever and generator |
| Independent challenge | An Apple preprint (March 2026) reports that program search without recursion matches or beats RLMs | Long-context models given the full text outperform RAG on average when resourced sufficiently (Li et al., 2024) |
| Maturity | Open-source MIT-licensed library; research-stage | Six years of deployment; a standard production pattern |
Detailed Analysis
Selecting evidence before reading versus while reading
RAG, introduced by Lewis et al. in May 2020, commits to its evidence before the generator sees anything. That is its efficiency and its weakness. If the retriever surfaces the right passages, the model reads a few thousand tokens instead of millions. If the question needs evidence the retriever did not rank highly, or needs nearly all of the corpus, the generator never sees what it needs. Related techniques such as semantic search over embeddings and GraphRAG improve the selection step but keep the same shape.
The RLM paper (Zhang, Kraska and Khattab; v3, May 2026) moves selection inside the loop. The model can grep the input, sample it, partition it and hand each partition to a fresh model call, then combine the returned values in code. Retrieval becomes one strategy the model may choose rather than a fixed stage, which is why the two approaches are better seen as overlapping than as opposites: an RLM can run a keyword search, and an agentic RAG system can issue several queries.
What the benchmarks show, and their limits
On BrowseComp-Plus with 1,000 documents per task (6M–11M tokens), the RLM authors report 91.3% for an RLM built on GPT-5 against 51.0% for a CodeAct agent equipped with BM25 retrieval. On OOLONG, a task that requires processing most of a 131K-token input, their table gives 56.0 for the RLM and 38.0 for the retrieval agent. Both gaps favour the RLM, and the OOLONG gap fits the mechanism: aggregation over every entry is the case top-k retrieval handles worst.
Three caveats apply. The retrieval baseline is a single lexical (BM25) configuration inside a coding agent, not a tuned dense-retrieval and reranking pipeline, so the result does not establish that RLMs beat RAG in general. The comparison comes from the paper proposing the method. And a March 2026 preprint from Apple researchers (Alizadeh et al.) argues that "recursion itself is not the primary driver of performance": a self-reflective program search without sub-calls matched or exceeded RLMs in their tests, and RLMs were weaker on semantically intensive tasks. No study opened for this page compares an RLM with a modern production RAG stack.
Cost, latency and operations
RAG's costs are front-loaded: building and refreshing an index, then a cheap query path. Li et al. (July 2024) found that long-context models beat RAG on average quality but that RAG's "significantly lower cost remains a distinct advantage". RLM costs are back-loaded and variable. The authors report an average of $0.99 per BrowseComp-Plus query with a standard deviation larger than the mean, describe median runs as cheap and outlier trajectories as expensive, and name runaway sub-call costs as an open problem.
Operationally, RAG needs retrieval infrastructure such as vector databases; an RLM needs a sandbox that executes model-written code. The open-source rlm library supports local, Docker and several cloud sandboxes, and warns that its default local environment should not be used in production.
Trainability
RAG components are trained separately or jointly, and the practice is well understood. For RLMs the first results are recent: the paper's RLM-Qwen3-8B, post-trained on about 1,000 filtered trajectories, is reported to outperform its base model by 28.3% on average. This is a small-scale result from the method's authors.
Best For
Question answering over a large, changing knowledge base
Retrieval-Augmented GenerationThe answer usually sits in a few passages, the index can be refreshed without retraining, and per-query cost stays low.
Aggregation across every record in a long input
Recursive Language ModelsCounting, classifying or pairing entries needs most of the input; the RLM paper's OOLONG results favour programmatic decomposition over retrieval.
Inputs of millions of tokens with no index built
Recursive Language ModelsAn RLM loads the raw input as a variable and was tested at 6M–11M tokens without preprocessing.
Predictable per-query cost and latency
Retrieval-Augmented GenerationA fixed retrieve-then-generate path is easier to budget than trajectories whose cost variance the RLM authors themselves flag.
Environments where model-written code cannot run
Retrieval-Augmented GenerationRLMs require a code sandbox; RAG does not execute generated code.
Multi-hop research over a fixed document set
Both / dependsThe published comparison favours the RLM, but only against one BM25 configuration; a tuned retrieval pipeline has not been tested head to head.
Citations and source attribution
Retrieval-Augmented GenerationRetrieved passages give a direct audit trail; attribution practices for RLM trajectories are not publicly documented.
The Bottom Line
RAG remains the default for knowledge-base question answering: it is mature, inexpensive at query time and easy to audit. RLMs address a different failure, where the task needs the whole input and a retriever would throw away evidence. On the two benchmarks where the RLM authors include a retrieval baseline, the RLM wins by a wide margin, but that baseline is narrow and the comparison has not been independently reproduced against a strong RAG system.
A reasonable reading as of October 2026: use RAG when a small fraction of a corpus answers the question, consider an RLM-style loop when the work is aggregation or cross-referencing over a very long input, and treat retrieval as a tool the recursive loop can call rather than a competitor to it. Budget controls on sub-calls and a hardened sandbox are prerequisites, not refinements.
Further Reading
- Recursive Language Models (Zhang, Kraska, Khattab; v3 May 2026) – arXiv
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., May 2020) – arXiv
- Recursive Language Models Meet Uncertainty: Self-Reflective Program Search for Long Context (Alizadeh et al., March 2026) – arXiv
- Retrieval Augmented Generation or Long-Context LLMs? (Li et al., July 2024) – arXiv
- rlm: inference library for Recursive Language Models – GitHub