The Orchestrator-Worker Pattern, Explained
RedHub AI Editorialupdated October 4, 20266 min read

Jump to a section7
The orchestrator-worker pattern puts one agent in charge. That lead agent, the orchestrator, plans a task, hands pieces of it to specialist worker agents, reads what comes back and decides whether to delegate again, until the task is done. It's the right shape when you can't know the sub-tasks in advance: a research request that might need three lookups or thirty, a coding change that might touch one file or a dozen. It's also the riskiest pattern to run in production, because the orchestrator decides as it goes. Without limits it cannot override, it can loop, over-delegate or burn through budget.
TL;DR: An orchestrator agent plans and delegates to worker agents, then assembles their output. Use it for open-ended tasks whose sub-jobs aren't known upfront. Give each worker a narrow scope, a clear brief and only the tools it needs. The pattern fails through runaway behavior, so enforce a delegation-round limit, a worker budget, a token and time ceiling, and a flagged partial-result path. The orchestrator plans, and deterministic code enforces the limits.
How it works
Anthropic's December 2024 post Building effective agents (opens in a new tab) defines it in one line: "a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results." In practice the flow has three moving parts:
- The orchestrator receives the goal and breaks it into sub-tasks. Its job is planning, delegating and integrating, not doing the detailed work itself.
- The workers are specialist agents with narrow scopes and only the tools their role needs. A search worker searches; a code worker edits code. Narrow workers are testable, and a mistake stays contained.
- The integration step assembles worker outputs into the final result. Often that's the orchestrator itself, sometimes a dedicated merge agent. It reconciles, de-duplicates and decides when the answer is complete.
What separates it from a plain parallel fan-out is who decides the split. Anthropic puts it this way: the subtasks "aren't pre-defined, but determined by the orchestrator based on the specific input." You'll also see the pattern called a supervisor or manager pattern. The shape is the same.
Why it goes wrong
Anthropic's June 2025 write-up on its multi-agent research system (opens in a new tab), which uses this pattern, lists what its early versions did: spawning 50 subagents for simple queries, "scouring the web endlessly for nonexistent sources," and distracting each other with excessive updates. Each failure traces back to the same root. The orchestrator decides dynamically, and dynamic decisions without limits drift.
- Delegation loops. A worker fails or returns something ambiguous, the orchestrator re-delegates the same task, it fails again, and the system spins until something else times out.
- Worker sprawl. Faced with a big task, an eager orchestrator spawns far more workers than the job needs. Each one adds tokens, latency and another chance to contradict the rest.
- Vague briefs. The same write-up found that short instructions like "research the semiconductor shortage" left subagents misreading the task or running the same searches as each other. Without a clear brief, Anthropic wrote, "agents duplicate work, leave gaps, or fail to find necessary information."
- Silent truncation. The usual band-aid, a single hard iteration cap, stops the runaway but often cuts the work off mid-task and returns an incomplete answer with no signal that it was cut.
The guardrails that make it safe
None of this means avoiding the pattern. It means wrapping it in limits the orchestrator can't argue with:
- A delegation-round budget. Cap how many planning-and-delegation cycles the orchestrator gets. When the cap is hit, stop and return the best answer so far.
- A worker budget. Cap total workers spawned per run, so a big task can't become a hundred agents.
- A token and time ceiling. Hard cost and latency limits enforced in code, not requested politely in the prompt.
- A structured brief for every worker. An objective, an output format, the tools and sources to use, and clear task boundaries. Those are the four things Anthropic's write-up says each subagent needs.
- A flagged partial-result path. When any limit trips, return the assembled partial answer with a clear flag that it's incomplete, never a silent cut-off that looks finished.
- Scoped worker tools. Each worker gets only its role's tools, so a confused worker can't act outside its lane.
A runaway orchestrator also needs a way to be stopped from outside the run. Our guide to the AI agent kill switch covers what that takes.
Prompt rules or code limits?
Anthropic did not fix its sprawl problem with hard caps alone. It wrote scaling rules into the orchestrator's prompt: simple fact-finding gets one agent with 3 to 10 tool calls, direct comparisons might need 2 to 4 subagents with 10 to 15 calls each, and complex research might use more than 10 subagents with clearly divided responsibilities. That's guidance the model reads, not a limit code enforces.
The two approaches do different jobs, and you likely need both. Prompt rules shape the typical run, so the orchestrator sizes its effort sensibly most of the time. Code limits catch the run where the model ignores its own rules. A prompt rule alone fails on exactly the input you didn't anticipate, and a code cap alone can cut off a legitimately large task. Where to set the caps depends on your task mix, and the only way to find the line is to log how many rounds and workers real runs use.
When to choose it over the simpler shapes
Reach for orchestrator-worker only when the sub-tasks genuinely can't be known in advance. If you can lay out the steps ahead of time, a sequential or parallel pattern is simpler, cheaper and easier to reason about. The hierarchical pattern buys flexibility, and you pay for it in complexity and risk. The single-agent starting point still comes first: a well-instrumented single agent often handles open-ended tasks fine.
Ship the orchestrator-worker pattern with its guardrails
Orchestrator-worker is the first of the five patterns in the Agent Orchestration Cookbook, written out in full code for the Claude Agent SDK and LangGraph, with failure notes on where it breaks and a control-flow eval that checks no worker result is dropped or double-counted.
Get the Agent Orchestration Cookbook — $79Pairs well with
The AI Spend Runaway & Billing-Safeguard Gate ($49) checks whether a runaway AI bill would actually be stopped and returns SAFEGUARDED, EXPOSED or RUNAWAY RISK, with a hard spend cutoff as the deciding safeguard. The Agent Side-Effect & Blast-Radius Checkpoint ($89) grades whether an agent action is safe to run unattended and returns RUN UNATTENDED, RUN WITH APPROVAL or DO NOT AUTOMATE. The Agent Reliability Harness ($149) evaluates agent runs at the trajectory level, including step efficiency and cost, and returns SHIP, HOLD or FIX.
More in this guide
What is the orchestrator-worker pattern?
A hierarchical orchestration where a lead orchestrator agent plans a task, delegates pieces to scoped worker agents, reads their results and integrates them, delegating more if needed. It's also called the supervisor or manager pattern.
When should I use it?
When the sub-tasks aren't known until the work starts, such as open-ended research or coding changes of unknown scope. If you can plan the steps in advance, a sequential or parallel pattern is cheaper and safer, and a single agent comes before any of them.
Why is it considered risky?
The orchestrator decides as it goes, so it can loop on a failing sub-task, spawn far more workers than needed or run up cost. Anthropic reported that early versions of its research system spawned 50 subagents for simple queries.
Should the orchestrator do the work too?
No. Keep the orchestrator on planning and integrating, and let scoped workers do the detailed work with only their role's tools. Merging planner and worker loses the isolation and testability that make the pattern worth using.
How do I stop it from looping forever?
Enforce limits in code: cap delegation rounds, cap total workers, and set a hard token and time ceiling. When a limit trips, return the partial answer with a clear incomplete flag. Prompt rules on effort help with typical runs but don't replace code limits.


The gate this post refers to, drawn from the tool’s own logic. See the tool.