Cognition and Anthropic Don’t Actually Disagree About Multi-Agent Systems

Two posts from two frontier labs, one day apart, read as opposite advice: don’t build multi-agent systems; multi-agent systems cut our research time by 90%. They’re both right — they’re describing different shapes of task. The difference is a decision rule, not a preference.

Subject
When Are Multi-Agent Systems Worth It?
Published
18 AUG 2026
Reading time
9 min
In this post · 5 sections
  1. 01Cognition’s case
  2. 02Anthropic’s case
  3. 03The variable that actually predicts the winner
  4. 04What we constrain to keep it debuggable
  5. 05The test to run on your own roadmap this week
The same argument in 4:24. Our graphics, AI narration.

On June 12, 2025, Cognition — the company behind the AI coding agent Devin — published a post titled “Don’t Build Multi-Agent Systems.” The next day, Anthropic published a detailed writeup of the multi-agent architecture behind its own research product, reporting that it beat a single-agent baseline by 90.2% on their internal evaluation. Read next to each other, as they were by a lot of people that week, they look like a direct contradiction from two labs that both know what they’re talking about.

They’re not disagreeing. They’re describing two different shapes of task, and the shape of the task is the entire decision.

Cognition’s case

Cognition’s argument is specific and mechanistic, not a general suspicion of agents working together: “actions carry implicit decisions, and conflicting decisions carry bad results.” Split a task across agents that can’t see each other’s full context, and each one is making small judgment calls — what a variable should be named, what a function should assume about its caller — that the others never see and can’t reconcile with. Their post walks through a case where two subagents, working from the same instruction without shared context, each built a plausible-looking piece of a game and produced two versions that quietly contradicted each other. Their conclusion: “share context, and share full agent traces, not just individual messages.”

Cognition AI, “Don’t Build Multi-Agent Systems,” Jun 12, 2025. cognition.ai/blog/dont-build-multi-agents

Anthropic’s case

Anthropic’s architecture is a lead agent that plans, then spawns subagents to work in parallel, each with its own isolated context window, recombining their outputs at the end. For their research product specifically, that beat single-agent Claude Opus 4 by 90.2% on their internal research eval. Running three to five subagents concurrently rather than serially, and letting each call several tools at once, cut research time by up to 90% on complex queries against their own earlier sequential version. The speed comes from the concurrency. The isolated context windows are what make the concurrency safe — each subagent independently chasing a different piece of a genuinely parallelizable question instead of stepping on the others’ judgment.

Anthropic doesn’t hide the cost. Their own numbers: agents typically use about 4× the tokens of a single chat call, and multi-agent systems about 15×. That is not a rounding error — it is the actual price of the architecture, stated plainly by the team that built it, precisely because the payoff was worth it for this specific task shape.

Anthropic, “How we built our multi-agent research system,” Jun 2025. anthropic.com/engineering/multi-agent-research-system

A single accumulating context thread compared with a lead agent fanning out to isolated subagents that converge at a merge point, alongside the real token-cost multiplier of each shapeAGENTSONE THREAD OR MANYTASK SHAPE DECIDESDWG Nº 02TOKENS SCALE ON SPLITSINGLE THREADone accumulating thread, no split to reconcileRESULTOne outputright for one coherent line of judgmentMULTI-AGENTLEADLead agentAGENT AAGENT BAGENT Cisolated contextMERGE✓ CLEAN MERGE✕ CONFLICTTOKEN COST, RELATIVE TO ONE CHAT CALLCHAT~1×SINGLE AGENT~4×MULTI-AGENT~15×Shared context, one voiceSplit work, verify mergesPrice it before you build
A single thread has nothing to reconcile, because it was never split — the right shape for one coherent line of judgment. A multi-agent fan-out can outrun it on genuinely parallel work, but every isolated context it spawns is a place a merge can go wrong, and every subagent is tokens spent whether or not the split was justified. The lead and subagent marks alternate Claude and Codex on purpose — both actually run this shape, and no single vendor owns the pattern.

The variable that actually predicts the winner

Not team preference, not which lab’s blog post you read most recently. One question: does the task decompose into genuinely independent subtasks, or does it need one evolving line of judgment?

  • Genuinely parallel — research breadth across independent sources, extraction across a batch of unrelated documents, evaluating the same input against several unrelated criteria — is where isolated context is a feature. Nothing any one subagent decides constrains what another one is allowed to decide, so there’s nothing to conflict over.
  • One evolving thread — most coding, most customer conversations, anything where step four’s correct behavior depends on a judgment call made at step one — is where splitting the context is where the conflicts Cognition describes come from. The work was never actually parallel; it was serial work that got forced into parallel agents anyway.

Multiply your expected query volume by the multiplier before you commit to either shape. A 15× token cost that finishes a task in a tenth of the wall-clock time is an easy call at research scale. The same 15× on a task that didn’t actually need splitting is pure waste, paid on every call, forever.

What we constrain to keep it debuggable

We’ve built the shape Anthropic describes for real: an on-prem multi-agent harness, run on Celery and Redis, for data that could not leave the building. The architectural discipline that keeps a system like that debuggable has nothing to do with the model and everything to do with the plumbing around it — queues with explicit task boundaries, so a subagent’s scope is a hard edge, not a convention; idempotent retries, so a task that fails halfway through doesn’t leave a duplicate side effect behind when it’s retried; and a merge step that is itself a defined, inspectable operation, not an implicit “the last agent to finish wins.”

The failure mode Cognition names doesn’t disappear at the infrastructure layer. Queues and retries make a multi-agent system operable. They don’t make conflicting subagent decisions reconcile themselves — that still has to be a designed step, with a real answer for what happens when two independent agents disagree, not an assumption that they won’t.

The test to run on your own roadmap this week

Before adopting a multi-agent pattern for anything on your current roadmap, write down, specifically, what the subagents would each own, and check whether any of them needs to know a decision another one is about to make. If the answer is yes for any pair, you have one task wearing a multi-agent costume — give it one thread and a bigger context window instead. If every subagent’s scope is genuinely blind to what the others are doing, and the task is big enough that the 4–15× token cost is a rounding error against the time saved, that’s the shape multi-agent architecture was actually built for.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.

Book an AI Architecture Review