Skip to content

Does a multi-agent AI system beat a single agent?

Unity Horizon6 min read

Not at a matched token budget. The 90.2% result was measured against an agent spending a quarter of the tokens. Checked 9 October 2026.

Not at a matched token budget. The number everyone quotes, a multi-agent system beating a single agent by 90.2%, came out of a comparison where the winning side was allowed to spend about four times the tokens. Three controlled studies published in 2026 held the budget equal, and the result turns over.

The mechanism sits in the same post as the 90.2%. Anthropic found that "token usage by itself explains 80% of the variance" in performance on the BrowseComp evaluation. Budget is most of what these comparisons measure, and hardly any of them hold it still.

What did the 90.2% multi-agent result measure?

Anthropic's engineering post of 13 June 2025 reports that "a multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on our internal research eval". Read further down the same post and it prices itself. "In our data, agents typically use about 4x more tokens than chat interactions", and "multi-agent systems use about 15x more tokens than chats". Divide the second by the first. The architecture that won was spending roughly 3.75 times the tokens of the architecture that lost.

The post is candid about it. These systems "require tasks where the value of the task is high enough to pay for the increased performance", and "in practice, these architectures burn through tokens fast". Neither sentence survives the trip into a slide. Both models in the comparison are also gone now. Opus 4 and Sonnet 4 sit on Anthropic's pricing page as retired rows served on partner clouds, checked 9 October 2026.

The subagents did not out-think the single agent. They out-spent it.

What happens when every architecture gets the same token budget?

Three teams held it still during 2026, on different models and different benchmarks.

Study Setup Result
Xu et al., 18 January 2026, arXiv 2601.12307 Seven benchmarks across coding, maths, QA, domain reasoning and tool use. One agent in multi-turn conversation against homogeneous workflows sharing the same base model The single agent matches the workflows and picks up efficiency from KV cache reuse that the workflow discards
Tran and Kiela, 2 April 2026, arXiv 2604.02460 FRAMES and MuSiQue filtered to 4-hop questions. Qwen3-30B-A3B, DeepSeek-R1-Distill-Llama-70B and Gemini 2.5, with reasoning tokens held constant Single agent matches or beats multi-agent at every budget from 500 thinking tokens up. Averaged at a 1,000-token budget, 0.42 against 0.38
Fu et al., 4 June 2026, arXiv 2606.05670 Ten reasoning, coding and tool-use benchmarks on GPT-4.1, under one shared execution protocol controlling the benchmark loader, tool access, answer format and usage accounting Five of six multi-agent systems trail the matched single-agent baseline by 2.56 to 11.29 accuracy points and sit at higher-cost points. The sixth falls inside one-run confidence

Tran and Kiela also report the single agent using fewer thinking tokens than the system it beat, 800 against 889 on Qwen3 with FRAMES at a 1,000-token budget. Then they name the artefact that pushes the other way. Visible thought text from Gemini 2.5 plateaus below the requested budget, so a multi-agent system making several calls surfaces more reasoning for the same nominal number. The budget control was leaking in favour of the fan-out.

Multi-agent did win in their data below 100 thinking tokens. The authors discount that band, because at that size neither side produces reasoning worth scoring.

Why does fanning out spend so much more?

A homogeneous multi-agent system is one model talking to itself through a protocol. Xu et al. put the waste precisely: agents in these frameworks differ only in prompts, tools and their position in the workflow, and the same work run as one multi-turn conversation reaches the workflow's performance with an efficiency advantage the authors put down to KV cache reuse.

The bill arrives twice. Once in money, once in capacity. Every subagent starts on a prefix no cache has seen. On the Claude API only uncached input counts against input tokens per minute; tokens written to the cache count at full value and tokens read from it do not count at all. We worked that through in How many AI agents can you run at once before you hit a rate limit?, where a cached 40-turn session draws 109,500 input tokens from the per-minute limit instead of 2,430,000.

Estimated, not measured: a warm turn in that session draws about 2,500 tokens against the limit, while five subagents each carrying a 12,000-token system prompt and tool catalogue draw 60,000 in the minute they launch. Your numbers will differ. Fan-out turns cache reads, which cost nothing against the limit, into cache writes, which cost full value. A team that spent a month lifting one agent's cache hit rate hands the gain back on the afternoon somebody adds the second agent.

What does a multi-agent run cost at list prices?

Take the anchor we have used since the piece on automating a government portal that has no API: a research run consuming 500,000 input and 50,000 output tokens on Claude Opus 5, at $5 and $25 per million, costs $3.75. Hold the work fixed and multiply the tokens by 3.75.

Same research job Input tokens Output tokens Cost
One agent, Opus 5 500,000 50,000 $3.75
Fan-out, Opus 5 throughout 1,875,000 187,500 $14.06
Fan-out, Opus 5 lead carrying 500,000 of the input, Sonnet 5 subagents carrying the other 1,375,000 1,875,000 187,500 $7.88

Arithmetic over published list prices, checked by Unity Horizon on 9 October 2026. Sonnet 5 is $2 and $10 per million, so moving the subagent share down a model class recovers about six dollars of the ten.

Ten dollars, then. On a task worth running at all that is about ten minutes of a EUR 60 an hour engineer, which makes money the weakest argument against the second agent. The accuracy claim nobody controlled is the strong one. Capacity is where the rest of the bill lands.

Where does running several agents in parallel still win?

One figure in the Anthropic post is not a token comparison, and it survives all three studies. The changes described there "cut research time by up to 90% for complex queries", with the lead agent spinning up "3-5 subagents in parallel rather than serially". Wall-clock time is what parallel work buys. No budget buys it for a single agent. The same post draws the far edge of it: domains needing shared context with many dependencies between agents "are not a good fit for multi-agent systems today", and "most coding tasks involve fewer truly parallelizable tasks than research".

A second agent buys minutes. It does not buy accuracy.

Two kinds of team read that sentence differently. A team whose task is a breadth-first sweep over sources that do not depend on each other, with a deadline measured in minutes and an answer worth real money, collects the wall clock and pays the token bill without feeling it. A team whose agent drives one browser through one stateful session, where step four needs what step three saw on screen, pays the bill and collects nothing. We built OFFSET, which files declarations directly to the ΑΑΔΕ portal. A filing queue is the second kind. One session, one login, work that runs as a line rather than a fan.

The orchestration diagram is the cheapest slide in the deck and the most expensive line on the bill.

What should you ask before adding a second agent?

Every study above measures benchmark work. None of them puts a computer-use agent in front of a real browser and a real rate limit, which is where most of our own work sits. We have not published a matched-budget comparison of our own either. If you have run one on a tool-heavy production agent, with the token counts recorded next to the accuracy, that number is worth more than this article.

So the question the 90.2% skipped is the one to put on Monday's list. How many tokens does your single agent spend on the task today, and did the architecture being sold to you get the same number? Settling that comparison before the build is the kind of work we scope.

Common questions

Does a multi-agent AI system beat a single agent?

Not when both get the same token budget. The 90.2% figure in circulation came from a multi-agent system spending about 15 times the tokens of a chat, measured against a single agent spending about 4 times. Three controlled studies published in 2026 held the budget equal and found the single agent matching or beating the multi-agent architectures.

How many more tokens does a multi-agent system use?

Anthropic reports agents using about 4 times more tokens than chat interactions and multi-agent systems about 15 times more, which puts a fan-out at roughly four times the token spend of one agent on the same work. On Claude Opus 5 list prices of $5 and $25 per million tokens, a research run costing $3.75 as a single agent costs about $14.06 fanned out, or about $7.88 with Claude Sonnet 5 subagents under an Opus 5 lead.

When is a multi-agent system worth building?

When the work is genuinely breadth-first and wall-clock time is the binding constraint. Anthropic reports cutting research time by up to 90% on complex queries by running 3 to 5 subagents in parallel, and says in the same post that domains needing shared context, with many dependencies between agents, are not a good fit.

Why do multi-agent benchmark comparisons favour multi-agent systems?

Because most of them do not hold compute constant. Anthropic found that token usage by itself explains 80% of the performance variance on the BrowseComp evaluation, so a comparison that lets one architecture spend more tokens has measured the budget instead of the architecture.

Ready when you are.

Tell us what slows your business down. We'll show you what intelligent software can do about it, usually within a week.

Book a callhello@unity-horizon.com