Multi-Agent Architectures: Supervisor, Swarm, and Hierarchical
Supervisor, swarm, and hierarchical multi-agent architecture compared on coordination, shared state, latency, cost, and how failures spread between agents.
A second agent is never free. It adds a communication channel, a state to keep consistent, and a new path for failures to travel. Multi-agent architecture is the discipline of paying that price only when the problem pays you back.
Three patterns cover nearly every production multi-agent system: supervisor, where one agent routes work to specialized workers; swarm, where peers hand tasks directly among themselves; and hierarchical, where managers delegate to managers who delegate to workers.
This article is the pattern catalog: the mechanics of each pattern, how agents communicate and share state, and how failures propagate through each topology. Whether multiple agents are worth running at all is the economic question answered separately in when multi-agent beats single-agent.
When Two Agents Are Actually Two Agents
The first design decision is not which pattern. It is whether a boundary exists at all. A legitimate agent boundary means at least one of these differs across the two agents:
- Prompts. Genuinely different instructions, personas, or expertise, not cosmetic role names.
- Tools. Different capabilities, especially different permission tiers: one reads, one writes.
- Context isolation. One agent’s working context would poison the other’s: full document sets, noisy logs, long histories.
If two “agents” share all three, they are one agent with extra steps, and the pattern you choose is irrelevant because the boundary is decorative. Decomposition itself, how to split a task into subtasks worth delegating, is a planning skill covered in planning and reasoning patterns.
Supervisor: One Router, Many Workers
The supervisor pattern puts one coordinating agent above a set of specialized workers. The supervisor decomposes the task, routes subtasks, collects results, and decides when the work is done. Workers never talk to each other; all traffic flows through the coordinator.
+------------+
Task -->| Supervisor |
+-----+------+
+---------+---------+
| | |
v v v
+--------+ +--------+ +--------+
|Worker A| |Worker B| |Worker C|
|(search)| |(write) | |(verify)|
+--------+ +--------+ +--------+
| | |
+---------+---------+
v
results, then supervisor decides:
done, retry, or re-route
What you gain: a single place for policy. Routing decisions are auditable in one trajectory, workers can carry different permission tiers, and one worker’s failure is contained by a supervisor that sees it as a structured result. Anthropic’s production pattern catalog describes the same shape as orchestrator-workers (a central LLM that decomposes and delegates dynamically) in building effective agents, and it is the natural fit when subtask types are known but their order and mix are not.
What you pay: the supervisor holds a growing picture of all results, so its context is the system’s bottleneck. A supervisor that must understand every worker’s output deeply will eventually run out of window, and its routing errors propagate everywhere.
def supervisor_loop(task: str, workers: dict, model, max_steps: int = 8) -> str:
"""Illustrative supervisor: route subtasks, collect results."""
results: list[str] = []
for _ in range(max_steps):
decision = model.plan_next(task, results, workers=list(workers))
if decision.done:
return decision.answer
worker = workers[decision.worker]
outcome = worker.run(decision.subtask) # each worker has its own budget
results.append(f"{decision.worker}: {outcome}")
return "stopped: supervisor budget exhausted"
The model client and worker interfaces are illustrative. The structural points are not: the supervisor owns the global budget, workers own their local budgets, and worker failure returns to the supervisor as data, never as an exception that kills the run.
Swarm: Peers Handing Off Directly
The swarm pattern removes the coordinator. Each agent declares what it handles, and any agent can hand the task to another when the conversation or task moves outside its scope. Control travels with the handoff: the receiving agent continues from the transferred context.
+--------+ handoff +--------+
| Agent | ---------------->| Agent |
| A | | B |
+--------+ +---+----+
^ | handoff
| +--------+ v
+--------| Agent | <------+
control | C |
+--------+
What you gain: no bottleneck and natural specialization. Routing decisions are made locally by the agent closest to the context, which suits customer-facing flows that bounce between narrow specialists: billing, technical, retention.
What you pay: no global view. Nothing in the system inherently knows where the task is or where it has been, so handoff loops (two agents politely returning the task to each other forever) are the pattern’s signature failure. Termination rules, maximum handoff counts, and an explicit escalation target are not optional. Tracing is also harder: the run’s narrative is scattered across the handoff chain.
Hierarchical: Layers of Delegation
The hierarchical pattern stacks supervisors: a top-level manager decomposes into domain-level managers, who decompose into workers, whose results aggregate back upward, summarized at each layer.
Depth pays when the task graph is genuinely layered: a research task across five domains, each domain splitting into subdomains, each subdomain into searches and reads. One flat supervisor cannot hold that shape; its context would overflow long before the leaf work finishes.
Depth costs in two currencies. Latency stacks with every layer, since each delegation is a model call plus a wait. And summarization at each boundary is lossy: constraints dropped in a mid-level summary reappear as contradictions at the leaves. The pattern’s failure mode is quiet divergence, where two branches work from subtly different understandings of the same instruction.
Communication and Shared State
Multi-agent systems move information two ways, and the choice decides your consistency semantics.
- Message passing. Agents exchange explicit messages or handoffs. Provenance is clear, causality is traceable, and nothing changes silently. Cost: every transfer is a deliberate act with its own latency.
- Shared state. Agents read and write a common store, the blackboard model. Fast and flexible, but it imports distributed-systems problems: clobbered writes, stale reads, and no inherent answer to “who changed this and when.”
Either way, three rules keep multi-agent state sane. Prefer structured artifacts over conversation histories: a typed, versioned result beats a transcript fragment, and it survives retrieval into any agent’s context. Give every artifact exactly one owner: single-writer ownership eliminates the clobbering class of bugs. And repeat hard constraints in the artifacts themselves, because constraints that live only in a parent agent’s prompt do not survive summarization at the next layer.
How Failures Propagate
In a single agent, a failure is a bad step. In a multi-agent system, a failure is misinformation that other agents consume and compound. The containment toolkit is the difference between a bad run and a poisoned one:
| Failure | Most at risk | Symptom | Containment |
|---|---|---|---|
| Worker error or hallucination | All patterns | Poisoned downstream conclusions | Structured failure to the coordinator, verification on critical outputs |
| Handoff ping-pong | Swarm | Task bounces between two agents | Maximum handoff count, termination rules, human escalation target |
| Mis-delegation | Supervisor | Wrong specialist, wasted budget | Routing evaluation suite, supervisor re-routes on poor results |
| Summarization loss | Hierarchical | Leaf work contradicts root constraints | Constraints carried in artifacts, verification at layer boundaries |
| Budget exhaustion | All patterns | Run stalls mid-task | Per-agent budgets inside a global run budget |
| Duplicated side effects | Swarm, hierarchical | Two agents execute the same write | Idempotency keys, single-writer ownership |
| Deadlock | Hierarchical | Layers wait on each other | Timeouts on every delegation, no synchronous mutual dependency |
Pattern Comparison
| Dimension | Supervisor | Swarm | Hierarchical |
|---|---|---|---|
| Coordination | Central router | Peer handoffs | Layered delegation |
| Global view | One place: the router | Nowhere inherent | Per layer, lossy upward |
| State model | Router-held results, worker artifacts | Context travels with handoffs | Summaries between layers |
| Cost profile | Router call per delegation | Call per handoff | Multiplied calls per layer |
| Latency | Sequential per subtask | Fast for simple handoffs | Stacks with depth |
| Failure signature | Router misjudgment, contained worker errors | Ping-pong loops, lost task | Divergent branches, quiet contradictions |
| Best fit | Known subtask types, permission tiers | Narrow specialists, routing-like flows | Large, layered task graphs |
One note on implementation: these are topologies, not products. LangGraph builds supervisors naturally (a coordinator graph whose nodes are worker subgraphs), as covered in LangGraph for agents and the official LangGraph documentation. CrewAI ships hierarchical processes for crews, and Microsoft Agent Framework, the successor to AutoGen, supports handoff and group collaboration workflows. The pattern you choose should survive the framework you use, per choosing a framework.
How Real Systems Do This
- Customer-facing systems run swarm-style specialists: billing, technical, retention, each narrow, with maximum-handoff rules and a human escalation target rather than infinite politeness between agents.
- Engineering and analysis platforms run supervisors: a coordinator with named workers, per-worker tool sets and permission tiers, and results returned as artifacts rather than chatter.
- Large research workloads go hierarchical: domain managers fan out to subdomain workers, and citation-carrying artifacts aggregate upward, so summaries remain verifiable.
- Every serious deployment shares two habits: a run-scoped trace that spans all agents, and per-agent budgets nested inside a global run budget. Systems without either are unfalsifiable when they fail.
Decision Framework
Work through in order, and stop at the first decisive answer:
- Do subtasks genuinely differ in prompts, tools, or permissions? If no, you have one agent and a diagram. Return to the single loop.
- Is there a natural central planner that understands the whole task? If yes, supervisor. The routing lives in one auditable place.
- Is the work routing-like, bouncing between narrow specialists? If yes, swarm, with termination rules and an escalation target from day one.
- Is the task graph too large for one router’s context? If yes, hierarchical, with constraints carried inside artifacts.
- What travels between agents? Structured artifacts with single-writer ownership. Histories and summaries are fallbacks, not defaults.
- How are failures contained? Per-agent budgets, structured failure reports, verification at every boundary where output becomes another agent’s input.
- Can you trace the whole run across agents? If not, build that first. Multi-agent systems without run-scoped tracing are undebuggable by design.
When NOT to Use This
- Decorative boundaries. Same prompts, same tools, same permissions means one agent with extra communication overhead and a more expensive failure analysis.
- Org-chart cosplay. Agents modeled on the company’s team structure inherit the meetings without the judgment. Model the task’s shape, not the reporting line.
- No cross-agent tracing yet. Ship the observability before the second agent, or the first incident becomes an attribution argument.
- Tight latency budgets. Every delegation is a model call plus a wait. Interactive paths that cannot absorb that should stay single-agent.
- Unowned shared state. If no one can be named as the single writer of each artifact, the clobbering bugs are already scheduled.
Common Mistakes
- Agents as the org chart. What follows: coordination tax paid for zero capability gained, and a system diagram that impresses exactly the people who never debug it.
- Full-history forwarding. The price: token costs that scale quadratically with agent count, and cross-talk where one agent’s noise becomes another’s instruction.
- No per-agent budgets. In practice: one looping worker silently consumes the entire run budget while the coordinator waits forever.
- Handoffs without termination rules. The result: two agents returning the task to each other until a timeout stops the politeness contest.
- Shared mutable state without ownership. The damage: overwritten artifacts, stale reads, and failures that reproduce only in production.
- Tracing deferred to “later.” The fallout: incidents where “which agent did that” has no answer, and the fix is faith-based.
Key Takeaways
- An agent boundary is justified only by a difference in prompts, tools, or permissions, or by context isolation. Everything else is decoration.
- Supervisor gives you central policy and contained failures; its context is the bottleneck. Swarm gives you local routing and no global view; termination rules are mandatory. Hierarchical fits large layered graphs; latency and lossy summaries stack with depth.
- Move structured artifacts, not conversation histories, and give every artifact a single writer.
- Failures in multi-agent systems are misinformation that compounds. Contain them with per-agent budgets, structured failure reports, and verification at boundaries.
- Run-scoped tracing across every agent is the entry ticket. Multi-agent without traceability is unfalsifiable when it fails.
- The pattern is a topology, not a product: it should survive your framework choice.
FAQ
What is a supervisor multi-agent architecture?
One coordinating agent decomposes a task, routes subtasks to specialized workers, collects their results, and decides when the run is complete. Workers never talk to each other. The supervisor is the single place for routing policy, permission tiers, and audit, and its context is the system’s scalability limit.
What is the difference between a swarm and a supervisor?
A supervisor centralizes routing: all work flows through one coordinator, giving a global view and contained failures. A swarm distributes routing: peers hand the task directly among themselves based on what each handles, which removes the bottleneck but loses the global view and introduces handoff-loop risk.
When should you use a hierarchical multi-agent system?
When the task graph is too large and layered for one router: multiple domains, each splitting into subdomains and leaf tasks. Depth buys context headroom at the cost of stacked latency and lossy summaries between layers, so constraints must travel inside artifacts, not only in prompts.
How do agents in a multi-agent system share state?
Through message passing, explicit handoffs with clear provenance, or a shared store, which is fast but needs ownership discipline. In both cases, prefer typed, versioned artifacts with a single writer over conversation histories or mutable blobs, and repeat hard constraints inside the artifacts themselves.
How do you stop agents from looping between each other?
Enforce termination mechanically: a maximum handoff or delegation count, per-agent step budgets inside a global run budget, timeouts on every delegation, and an explicit escalation target. Never rely on prompts to end a loop; rely on the orchestrator to count and cut.
Conclusion
A multi-agent system is a distributed system whose components happen to be nondeterministic. The patterns (supervisor, swarm, hierarchical) are coordination topologies for it, and each has a signature strength and a signature failure. Choosing well means matching the topology to the task’s actual shape, then paying the unglamorous costs: artifacts with owners, budgets at every level, and a trace that spans the whole run.
The architectural honesty that keeps these systems maintainable is the boundary test. Boundaries exist to separate prompts, tools, and permissions, and every boundary that fails the test should be removed, because each one is a communication channel, a consistency problem, and a failure path you are maintaining for nothing.
Every agent boundary must pay rent: different prompts, different tools, or different permissions. Otherwise it is one agent wearing three name badges.
Last updated on 2 October 2026
