Why Multi-Agent Systems Fail (and How to Catch It)
RedHub AI Editorialupdated October 4, 20266 min read

Jump to a section7
When researchers annotated more than 1,600 traces from 7 multi-agent frameworks, the multi-agent failure modes they found fell into three groups: system design issues, inter-agent misalignment and task verification. In practice that means most failures sit in the glue between agents rather than inside one model. An agent loops on an edge case the demo never hit. A hand-off passes malformed data the next agent trusts. A reflection step doubles cost with no gain in quality. An orchestrator spawns runaway workers. You catch them the way you catch other software failures: an eval harness with golden cases, failure notes per agent, deterministic gates at every hand-off, and a cost and latency budget someone actually watches.
TL;DR: The common failure modes are loops, bad hand-offs, cost blow-ups, cascading errors, runaway orchestrators and silent truncation, and most of them live in the glue between agents. Catch them with a golden-case eval harness, schema-checked hand-offs, hard step and token budgets, and per-agent failure notes. Gates catch the failures you can name, but not every failure has a gate, so verification of the final result still matters.
What the research found
The study is Why Do Multi-Agent LLM Systems Fail? (opens in a new tab), first posted to arXiv in March 2025 and revised since. Its authors built a taxonomy they call MAST from an analysis of 150 traces, then released a dataset of more than 1,600 annotated traces collected across 7 popular multi-agent frameworks. The taxonomy names 14 failure modes in three categories: system design issues, inter-agent misalignment and task verification.
The paper opens with a blunt observation: despite the enthusiasm, the performance gains of multi-agent systems on popular benchmarks "are often minimal." Multi-agent designs can still pay off, but a working demo is weak evidence that the extra agents are earning their cost.
The failure modes that actually bite
| Failure mode | What happens | How to catch it |
|---|---|---|
| Infinite or long loops | An agent retries the same failing step forever, or two agents ping-pong | A hard step budget and loop detection, plus a golden case built from the input that first caused it |
| Broken hand-offs | One agent passes malformed or partial data; the next trusts it and produces a confident wrong answer | Schema-validate every hand-off and fail loudly at the broken seam |
| Cost blow-ups | A reflection or delegation step multiplies tokens with no measurable quality gain | A per-pattern token and latency budget; A/B the step and cut it if it doesn't earn its cost |
| Cascading errors | A small early mistake compounds through a pipeline into a large final one | Validate intermediate outputs, not just the final result |
| Runaway orchestrators | A hierarchical planner over-delegates or loops, spawning far too many workers | Delegation-round and worker budgets enforced in code |
| Silent truncation | A hard cap stops a runaway but returns an incomplete answer with no signal | A partial-result path with an explicit incomplete flag |
Nearly every row is a glue problem. The agents do roughly what they were told. The system around them failed to constrain, validate or budget them. Cascading errors are the hardest to see, because each step looks reasonable on its own. Anthropic's June 2025 write-up on its multi-agent research system (opens in a new tab) puts it plainly: in agentic systems, "minor changes cascade into large behavioral changes."
The demo-to-production wall
Take a composite case, assembled to illustrate the pattern rather than drawn from one team. An engineer builds a multi-step agent in four days, and it works in dev. In production it loops on an edge case in ticket 23. Someone adds max_iterations=10 as a band-aid, and ticket 47 hits the cap and fails silently. Someone adds a reflection step that doubles token cost without moving quality. Someone reads a blog post about supervisor patterns and rewrites half the system. Six weeks later there's still no eval harness, no failure-mode notes and no cost dashboard, and the pilot gets shelved.
Every fix in that story treated a symptom. None of them answered what the system does when one agent returns garbage, and none left behind a test that would catch the same failure next time.
How to catch failures before your users do
- Build a golden-case eval harness. Collect real inputs, especially the ones that broke something, with expected outcomes and a scoring rule. Run it on every prompt or pattern change, so a fix to step 1 that breaks step 3 fails a test instead of a customer.
- Schema-check every hand-off. Data passed between agents should be validated structured objects. A failed check retries or escalates at that seam instead of poisoning everything downstream.
- Write failure-mode notes per agent. For each agent: where it breaks, why, and what the system does about it (retry, degrade or escalate). This turns a middle-of-the-night outage into a case someone already answered.
- Watch a cost and latency budget. Track tokens, p50 and p95 latency, and retry rate per pattern. A step that doubles cost for no quality gain should show up in a dashboard, not a bill.
- Add deterministic gates. Confidence thresholds, human-handoff triggers and hard step budgets are plain code around the probabilistic agents.
Single agents fail quietly too: happy-path passes, drift and wrong tool calls. Our guide to AI agent reliability covers what to monitor for each one, and every agent you add is one more to watch.
Where gates stop helping
The list above catches failures you can name in advance: a malformed object, a loop past its budget, a cost spike. The MAST authors are more cautious about their own findings. Their abstract says the failures they identified "require more sophisticated solutions" than a quick fix, and one of their three categories is task verification itself, the step where a system checks whether the job was done.
A schema check confirms the shape of an answer. It can't confirm the answer is right. For that you need verification of the final result against expected outcomes, which is the job of the golden cases, and sometimes a human reviewer on high-stakes outputs. We don't know of a gate that removes the need for that last step, and we'd be suspicious of anyone selling one.
Start with patterns that ship with their failure notes
For each of its five patterns, the Agent Orchestration Cookbook writes down five ways it breaks, how to catch each and how to mitigate it, next to a per-pattern eval script that checks the control flow against a mock model.
Get the Agent Orchestration Cookbook — $79Pairs well with
The Agent Reliability Harness ($149) evaluates recorded agent runs at the trajectory level (tool choice, argument validity, step efficiency, cost and policy) and returns a SHIP, HOLD or FIX verdict you can gate CI on. The Prompt Regression Lab ($89) snapshots a baseline, diffs every prompt change and fails CI on any regression. If a step does retrieval, the RAG Retrieval Grader ($89) scores that step with Recall@k, Precision@k, MRR and nDCG and names every miss.
More in this guide
Why do multi-agent systems fail in production?
Mostly at the seams between agents: loops on untested edge cases, malformed hand-offs the next agent trusts, cost-doubling steps with no quality gain, cascading errors and runaway orchestrators. These come from missing scaffolding around the agents more often than from the models.
Is there a published taxonomy of multi-agent failure modes?
Yes. The 2025 paper "Why Do Multi-Agent LLM Systems Fail?" introduced MAST, which names 14 failure modes in three categories: system design issues, inter-agent misalignment and task verification. Its dataset covers more than 1,600 annotated traces from 7 frameworks.
Why does it work in the demo but break live?
The demo exercises the happy path, and production is every path the demo skipped. An agent loops on an input the demo never sent, or a hand-off fails a check nobody wrote. Teams that get past this build eval harnesses, failure notes and cost budgets before shipping.
How do I catch a looping agent?
Enforce a hard step budget and loop detection in code, and turn the input that first caused the loop into a golden test case so a future change can't bring it back. Pair any cap with a flagged partial result so a cut-off never looks finished.
What's the single highest-value thing to add?
A golden-case eval harness. Once you can run realistic inputs with expected outcomes on every change, most regressions surface before release, and you can measure whether a pattern change helped. Schema-checked hand-offs come a close second.


The gate this post refers to, drawn from the tool’s own logic. See the tool.