Cognition and Anthropic Don’t Actually Disagree About Multi-Agent Systems
Two posts from two frontier labs, one day apart, read as opposite advice: don’t build multi-agent systems; multi-agent systems cut our research time by 90%. They’re both right — they’re describing different shapes of task. The difference is a decision rule, not a preference.
- Subject
- When Are Multi-Agent Systems Worth It?
- Published
- 18 AUG 2026
- Reading time
- 9 min
- Film
- Watch · 4:24
In this post · 5 sections
On June 12, 2025, Cognition — the company behind the AI coding agent Devin — published a post titled “Don’t Build Multi-Agent Systems.” The next day, Anthropic published a detailed writeup of the multi-agent architecture behind its own research product, reporting that it beat a single-agent baseline by 90.2% on their internal evaluation. Read next to each other, as they were by a lot of people that week, they look like a direct contradiction from two labs that both know what they’re talking about.
They’re not disagreeing. They’re describing two different shapes of task, and the shape of the task is the entire decision.
Cognition’s case
Cognition’s argument is specific and mechanistic, not a general suspicion of agents working together: “actions carry implicit decisions, and conflicting decisions carry bad results.” Split a task across agents that can’t see each other’s full context, and each one is making small judgment calls — what a variable should be named, what a function should assume about its caller — that the others never see and can’t reconcile with. Their post walks through a case where two subagents, working from the same instruction without shared context, each built a plausible-looking piece of a game and produced two versions that quietly contradicted each other. Their conclusion: “share context, and share full agent traces, not just individual messages.”
Cognition AI, “Don’t Build Multi-Agent Systems,” Jun 12, 2025. cognition.ai/blog/dont-build-multi-agents
Anthropic’s case
Anthropic’s architecture is a lead agent that plans, then spawns subagents to work in parallel, each with its own isolated context window, recombining their outputs at the end. For their research product specifically, that beat single-agent Claude Opus 4 by 90.2% on their internal research eval. Running three to five subagents concurrently rather than serially, and letting each call several tools at once, cut research time by up to 90% on complex queries against their own earlier sequential version. The speed comes from the concurrency. The isolated context windows are what make the concurrency safe — each subagent independently chasing a different piece of a genuinely parallelizable question instead of stepping on the others’ judgment.
Anthropic doesn’t hide the cost. Their own numbers: agents typically use about 4× the tokens of a single chat call, and multi-agent systems about 15×. That is not a rounding error — it is the actual price of the architecture, stated plainly by the team that built it, precisely because the payoff was worth it for this specific task shape.
Anthropic, “How we built our multi-agent research system,” Jun 2025. anthropic.com/engineering/multi-agent-research-system
The variable that actually predicts the winner
Not team preference, not which lab’s blog post you read most recently. One question: does the task decompose into genuinely independent subtasks, or does it need one evolving line of judgment?
- Genuinely parallel — research breadth across independent sources, extraction across a batch of unrelated documents, evaluating the same input against several unrelated criteria — is where isolated context is a feature. Nothing any one subagent decides constrains what another one is allowed to decide, so there’s nothing to conflict over.
- One evolving thread — most coding, most customer conversations, anything where step four’s correct behavior depends on a judgment call made at step one — is where splitting the context is where the conflicts Cognition describes come from. The work was never actually parallel; it was serial work that got forced into parallel agents anyway.
Multiply your expected query volume by the multiplier before you commit to either shape. A 15× token cost that finishes a task in a tenth of the wall-clock time is an easy call at research scale. The same 15× on a task that didn’t actually need splitting is pure waste, paid on every call, forever.
What we constrain to keep it debuggable
We’ve built the shape Anthropic describes for real: an on-prem multi-agent harness, run on Celery and Redis, for data that could not leave the building. The architectural discipline that keeps a system like that debuggable has nothing to do with the model and everything to do with the plumbing around it — queues with explicit task boundaries, so a subagent’s scope is a hard edge, not a convention; idempotent retries, so a task that fails halfway through doesn’t leave a duplicate side effect behind when it’s retried; and a merge step that is itself a defined, inspectable operation, not an implicit “the last agent to finish wins.”
The test to run on your own roadmap this week
Before adopting a multi-agent pattern for anything on your current roadmap, write down, specifically, what the subagents would each own, and check whether any of them needs to know a decision another one is about to make. If the answer is yes for any pair, you have one task wearing a multi-agent costume — give it one thread and a bigger context window instead. If every subagent’s scope is genuinely blind to what the others are doing, and the task is big enough that the 4–15× token cost is a rounding error against the time saved, that’s the shape multi-agent architecture was actually built for.
Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.