# Orchestrator–Worker Pattern

> The orchestrator–worker pattern is a multi-agent design in which a lead agent delegates separable sub-tasks to workers with isolated contexts.

Source: https://metavert.io/orchestrator-worker-pattern  
Published: 2026-10-07  
Updated: 2026-10-07

The **orchestrator–worker pattern** is a [multi-agent](https://metavert.io/multi-agent-systems) architecture in which a lead agent decomposes a task, delegates separable sub-tasks to worker agents that each run in their own isolated context, and then synthesizes what they return. Anthropic's December 2024 guide to agent design describes it as a workflow in which "a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results", and distinguishes it from plain parallelization by its flexibility: the sub-tasks are not fixed in advance but decided by the orchestrator from the input.

### How It Works

The orchestrator holds the user's goal, the plan and the running synthesis. Each worker receives a brief, works through its own sequence of tool calls in a fresh [context window](https://metavert.io/context-windows), and hands back a condensed result rather than its full transcript. Workers typically do not talk to each other; coordination flows through the lead. The hierarchy is usually kept shallow, because every extra level adds another lossy summary between the evidence and the decision.

Frameworks expose the pattern in different ways. The [OpenAI Agents SDK](https://metavert.io/openai-agents-sdk) lets one agent call another as a tool, keeping control with the caller, or hand the conversation off entirely. [Microsoft Agent Framework](https://metavert.io/microsoft-agent-framework) ships sequential, concurrent, handoff, group-chat and Magentic orchestrations as prebuilt workflows. The pattern itself is independent of any framework.

### When It Helps

Two properties account for nearly all of the real gains.

**Parallel breadth.** When a task splits into independent lines of inquiry, such as researching twenty companies or searching many repositories, workers can pursue them at once, and the total amount of material examined can exceed what one context could hold. Anthropic's June 2025 account of its Research feature is the best-known report: a lead Claude Opus 4 with Claude Sonnet 4 subagents "outperformed single-agent Claude Opus 4 by 90.2% on our internal research eval". That figure is vendor-reported, comes from one internal evaluation of breadth-first research queries, and was achieved by spending far more compute, as the same article makes clear.

**Context isolation.** A worker can read fifty files or pages, discard the dead ends, and return a few hundred words. The orchestrator's context stays small and on-topic, which matters because model performance degrades as a context fills with irrelevant material. Seen this way, the pattern is a form of [context engineering](https://metavert.io/context-engineering) as much as a division of labour.

### When a Single Agent Is Better

The cost is tokens and coordination. Anthropic's article reports that agents typically use about 4 times more tokens than chat interactions and multi-agent systems about 15 times more, and it says plainly that the approach fits poorly where agents must share the same context or where sub-tasks have many dependencies, noting that most coding tasks have fewer truly parallelizable parts than research.

Controlled comparisons sharpen the point. Tran and Kiela (arXiv 2604.02460, April 2026) held the reasoning-token budget equal and compared single-agent systems with multiple multi-agent architectures across three model families (Qwen3, DeepSeek-R1-Distill-Llama and Gemini 2.5). They found that single agents "consistently match or outperform" multi-agent systems on multi-hop reasoning when reasoning tokens are held constant, and argue from the Data Processing Inequality that a single agent with perfect context utilization is the more information-efficient design. Their analysis also predicts where the pattern earns its keep: multi-agent systems become competitive when a single agent's effective use of its context is degraded, or when more compute is spent. The study covers multi-hop reasoning only, so it should not be read as a verdict on long tool-using tasks.

A second paper points the same way. Jwalapuram et al. (arXiv 2606.13003, June 2026) report that automatically designed multi-agent systems "consistently underperform" chain-of-thought with self-consistency "despite being up to 10x more expensive", and that expert-designed architectures beat automatically generated ones. Taken together, the evidence suggests that much of the reported multi-agent advantage is additional compute, and that the honest baseline for any fan-out is a single agent given the same budget. The [single-agent vs multi-agent comparison](https://metavert.io/compare/single-agent-vs-multi-agent-systems) covers this in more depth.

| Task property | Favours |
| --- | --- |
| Many independent sub-questions; breadth exceeds one context | Orchestrator–worker |
| Sub-tasks generate large volumes of disposable intermediate reading | Orchestrator–worker |
| Tightly coupled steps; every part needs the same evolving state | Single agent |
| Fixed token budget; sequential multi-hop reasoning | Single agent |
| Wall-clock time matters more than cost | Orchestrator–worker |

### Handoff Design

Most failures of the pattern are failures of the brief. A worker knows only what the orchestrator writes down, and the orchestrator sees only what the worker sends back. Anthropic's write-up lists what each delegation needs: an objective, an output format, guidance on the tools and sources to use, and clear task boundaries.

The return path deserves equal care. Results should be structured, cite the evidence they rest on, and separate what was verified from what was inferred, because a worker's mistake arrives in the same confident register as its findings. An orchestrator that passes its own hypothesis to a worker tends to get that hypothesis back confirmed, so independent checks should receive the raw material rather than the conclusion. Each handoff is one more step at which an error can enter unobserved, which makes this an [agent reliability](https://metavert.io/agent-reliability) concern.

### Budgets

Fan-out multiplies spend, so it needs explicit limits: how many workers, how many tool calls each, and when to stop. Anthropic describes embedding effort-scaling rules in the lead agent's prompt, with simple fact-finding assigned one agent making 3 to 10 tool calls, direct comparisons 2 to 4 subagents making 10 to 15 calls each, and complex research more than 10 subagents with clearly divided responsibilities. The general rule is to scale worker count to the breadth of the task and to test any orchestrated system against a single agent given the same total token budget.

## Related Topics

- [Multi-Agent Systems](https://metavert.io/multi-agent-systems) — The broader family this pattern belongs to
- [Agent Orchestration](https://metavert.io/agent-orchestration) — Hierarchical, pipeline and swarm coordination
- [Single-Agent vs Multi-Agent Systems](https://metavert.io/compare/single-agent-vs-multi-agent-systems) — The equal-budget comparison in detail
- [Context Engineering](https://metavert.io/context-engineering) — Isolation as a way to keep contexts clean
- [Context Compaction](https://metavert.io/context-compaction) — The single-agent alternative for long tasks
- [Agent Reliability](https://metavert.io/agent-reliability) — Each handoff is another point of failure
- [OpenAI Agents SDK](https://metavert.io/openai-agents-sdk) — Agents as tools and handoffs
- [Microsoft Agent Framework](https://metavert.io/microsoft-agent-framework) — Prebuilt orchestration workflows
- [LangGraph](https://metavert.io/langgraph) — Graph-based supervisor patterns

## Further Reading

- [Building Effective Agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, December 2024
- [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) — Anthropic, June 2025
- [Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets](https://arxiv.org/abs/2604.02460) — Tran and Kiela, arXiv, April 2026
- [The Illusion of Multi-Agent Advantage](https://arxiv.org/abs/2606.13003) — Jwalapuram et al., arXiv, June 2026
- [OpenAI Agents SDK documentation](https://openai.github.io/openai-agents-python/) — OpenAI, 2026
- [Agent Framework workflow capabilities and orchestrations](https://learn.microsoft.com/en-us/agent-framework/workflows/) — Microsoft Learn, July 2026
