RL Environments
RL environments, in the context of language-model agents, are executable task setups — an instruction, a sandboxed world to act in (a repository, a terminal, a set of stateful tools), and a programmatic verifier that turns the outcome into a score — used as the training signal for reinforcement learning. They have displaced static datasets of example trajectories as the scarce input to agent training: a trajectory shows one way a task was done, whereas an environment can be attempted any number of times by any model and graded each time. The same package doubles as an evaluation.
Anatomy of an Environment
A typical coding or tool-use environment has five parts: the task instruction, fixtures that set up initial state, the tools or shell the agent may use, a verifier (unit tests, a state check, a rubric), and a container definition that makes the whole thing reproducible. Envs-FORGE (Wu et al., August 2026) describes rewriting exactly this bundle — instructions, test fixtures, reference solutions and containerised environments — as a unit, because changing a task's difficulty without changing its tests produces an unverifiable task.
The verifier is what separates an environment from a prompt. DeepSeek's R1 paper restricted large-scale RL to tasks with rule-based checks because learned reward models proved exploitable; the argument carries over directly to agents. Where no programmatic check exists, the reward has to come from a rubric scored by a model, which is cheaper to write and easier to game.
Synthesis Aimed at the Learning Frontier
Hand-building environments is slow, so 2026 research concentrates on generating them — and on generating the right ones. A task the model always solves, or never solves, yields no gradient under group-based algorithms such as GRPO. The useful environments sit at the edge of current ability.
| System | Approach | Reported result (authors' own) |
|---|---|---|
| Envs-FORGE (August 2026) | Estimates each seed task's pass rate from verifier rewards, then chooses among six edit directions to move it towards a target difficulty | With 100 verified environments, Qwen 3.5 35B rose from 40.0% to 49.2% Pass@1 on tb-core and from 73.4% to 77.1% on SWE-bench Verified |
| EnvFactory (May 2026) | Automatically discovers and validates stateful tool environments; writes queries with implicit intent rather than step-by-step instructions | 85 verified environments and 2,575 trajectories; up to +15% on BFCLv3 and +8.6% on MCP-Atlas for Qwen3-series models |
| SPADE (August 2026) | One model both writes environments as code and trains inside them, steered by a regret signal | +5.3 points on average over fixed-environment baselines across eight benchmarks at 30B parameters |
| Qwen-AgentWorld (June 2026) | A language model trained on more than 10M real interaction trajectories simulates the environment instead of executing it | Used as an RL simulator and as a warm-up stage improving seven agent benchmarks |
Two things stand out. First, the counts are small — on the order of a hundred environments, not millions of examples — which suggests placement matters more than volume. Second, every figure in the table is from a single paper reporting on its own method; none has been independently replicated as of October 2026, and the benchmarks differ, so the rows are not comparable with one another. SPADE is the bridge to self-play: once environment design is itself learned, the curriculum is no longer fixed by its authors. Qwen-AgentWorld points the other way, towards learned simulators that resemble world models; a simulated environment is cheaper to scale but its verdicts are only as reliable as the simulator.
Environments as Evals
Because an environment already contains a task and a grader, running it without a gradient update is an evaluation. Benchmarks such as SWE-bench Verified are environments in this sense, and open tooling treats the two uses as one artefact: Prime Intellect describes its verifiers library as "RL environments + evals" and reported more than 2,500 community environments on its hub as of October 2026 (company figure). The convenience has a cost. Training on environments drawn from the same distribution as a benchmark inflates the benchmark, so held-out environments are needed for honest agent evals.
Costs and Limits
Verifier quality bounds everything. An agent trained against a weak check learns to satisfy the check. METR's May 2026 Frontier Risk Report found that, on its tasks longer than eight hours, at least 16% of successful runs were illegitimate on review — reward hacking at evaluation time, in environments built by specialists. Training against the same flaws would reinforce them.
Rollouts are expensive. Each training step needs several full episodes per task, each in its own sandbox; long-horizon tasks multiply both compute and wall-clock time. Envs-FORGE reports 2.27M–2.88M tokens just to synthesise 100 verified environments, before any training.
Coverage is narrow. Published results cluster in code, terminal work and API-style tool use, where outcomes are checkable. Tasks whose success is a matter of judgement remain hard to package, and the fidelity of synthetic environments to real production systems is largely unmeasured.
Further Reading
- Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis — Wu et al., arXiv, August 2026
- EnvFactory: Scaling Tool-Use Agents — Xu et al., arXiv, May 2026
- SPADE: Self-Play in Adaptive Synthetic Executable Environments — Liu et al., arXiv, August 2026
- Qwen-AgentWorld: Language World Models for General Agents — Zuo et al., arXiv, June 2026
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI, arXiv, January 2025
- verifiers: library for RL environments and evals — Prime Intellect, GitHub, accessed October 2026
- Environments Hub: A Community Hub To Scale RL To Open AGI — Prime Intellect, August 2025
- Frontier Risk Report (February to March 2026) — METR, May 2026