One LLM or an army of agents? The call that sets your bill and your failure rate
automation October 5, 2026 · Mintec

One LLM or an army of agents? The call that sets your bill and your failure rate

A single agent hit 28 of 28; the swarm failed 68%. Anthropic still measured +90% from multi-agent. Here is how to pick the topology with data instead of hype.

One LLM or an army of agents? The call that sets your bill and your failure rate

Short answer: start with a single orchestrator plus tools. An army of agents only earns its keep when the task splits into independent subtasks, nobody needs to write to the same system of record, and the value of each task pays for ~15x a chat's tokens. In McEntire's 2026 study a single agent got 28 of 28 right; the swarm failed 68% of the time. And yet Anthropic's multi-agent system beat the single agent by 90.2% — at 15x the token bill. Both findings are true, and that is exactly the trap.

The question is no longer whether AI agents work. It is how many to deploy. On r/Entrepreneur, the thread "Are you using one LLM or an army of agents?" pulled 54 upvotes and 125 comments: people who have been running agents for months and still cannot tell whether they paid for an architecture or a habit.

We have been implementing agents inside real CRMs this year, and the answer we keep giving clients is the same: one box in the diagram, not ten. Here is why, with the data that actually exists.

Two serious studies that disagree — and why both are right

In June 2025, Anthropic published how it built its multi-agent Research system: a lead Opus 4 agent delegating to Sonnet 4 subagents in parallel. On its internal eval, that architecture beat the single Opus 4 agent by 90.2%. Hard to argue with.

But the bill arrives in the same post: agents use ~4x a chat's tokens and multi-agent systems ~15x. Token usage alone explains 80% of the variance in their browsing eval. In Anthropic's own words, multi-agent works mainly because they help spend enough tokens to solve the problem — it wins by spending, not by some magic collaboration of several brains.

Now the other side. In February 2026, organizational researcher Jeremy McEntire published "The Organizational Physics of Multi-Agent AI" on SSRN, and the numbers hurt: a single agent scored 28 of 28. A hierarchy (one agent delegating to others) failed 36% of the time. A self-organized stigmergic swarm failed 68%. An 11-stage gated pipeline never produced a good outcome at all — it exhausted the budget on planning without writing a line of implementation. His conclusion: "The substrate changes; the physics of coordination at scale remains constant."

Between the two sits UC Berkeley's MAST taxonomy (Why Do Multi-Agent LLM Systems Fail?, ICLR 2025): 1,600+ execution traces across 7 frameworks, 14 failure modes in three categories:

Failure categoryShare of observed failures
Ambiguous or violated specification41.77%
Inter-agent misalignment36.94%
Weak task verification21.30%

The failure everyone imagines — one agent ignoring another — is 0.17%. The failures are organizational, not intellectual: agents repeating work, never stopping, never checking. ChatDev, the most-cited multi-agent "software company", was correct only 33.33% of the time.

The synthesis: Anthropic is right when the task is broad and parallelizable. McEntire is right when agents must coordinate to produce one thing. The mistake is assuming one architecture always wins.

What an army of agents actually costs

The price per token barely changes. What changes is how many tokens your chosen topology burns. Using Anthropic's published multipliers:

TopologyTokens vs. a chatPer task ($0.02 base)10,000 tasks/month
Plain chat (no agent)1x$0.02$200
A. One agent with tools~4x$0.08$800
B. Orchestrator + chained specialists4x–8x$0.08–0.16$800–1,600
C. Parallel swarm (multi-agent)~15x$0.30$3,000

That is before accumulated context: a multi-step agent with RAG and tools already clears a dollar per task on its own, as we covered in what an AI agent really costs per task.

Here is the part almost nobody in LatAm calculates: you pay for tokens in dollars and bill in local currency. Multiplying an operation that runs 10,000 times a month by 15 is not a cost increase — it is a different spending category with a different exposure to FX. Argue the token budget per task before you argue architecture.

Why swarms fail: not intelligence, but organization

There is arithmetic no model generation erases. Agents are sequential and probabilistic: if each step is right 95% of the time, a 10-step chain finishes correctly only 59.9% of the time, and 20 steps drop to 35.8%. Every agent you add is another multiplication, another context handoff, another place where a locally sensible decision becomes globally wrong.

It is the same pattern McEntire describes: review thrashing, governance conflicts, budget exhausted on coordination. Agents inherit the failure modes of human organizations because they are modeled on human reasoning. The failure is not in the model; it is in the org chart.

The industry is already seeing it. Gartner (Hype Cycle for Agentic AI, 2026): only 17% of organizations have deployed AI agents, more than 60% expect to within two years — the most aggressive adoption curve they measured — and they project over 40% of agentic projects will be canceled by the end of 2027 on escalating costs, unclear business value and inadequate risk controls. Same cycle: "fully autonomous agents are not ready for the majority of enterprise use cases today." Deloitte found only 11% with production-ready systems.

Same pattern as why most automation projects fail: teams automate before the process is in order, and now they multi-agent before they have one agent that works.

The 4 questions that decide the topology

This is the matrix we run before writing a single line of workflow. Answer in order; the first disqualifier wins.

#QuestionIf the answer is...Topology
1Does the task split into independent subtasks that run in parallel?Yes, genuinely (several sources, documents, markets)Candidate for C
No: there are dependencies or a single outputA (one orchestrator)
2Do the agents share the same context or record (the CRM, the customer file, inventory)?YesA — shared context is dependency, and Anthropic is explicit: domains with many dependencies are a poor fit for multi-agent
3Does someone have to write (create, update, close, bill)?YesSingle write thread. Cognition's rule: writes stay single-threaded, extra agents contribute intelligence, not actions
4What does an error cost?High (money, customer, compliance)A with human review, no matter how many agents you use

Rule of thumb: parallelism + high value per task + read-only is the only combination that justifies topology C. Everything else is one agent with a good tool library.

TopologyWhen we use itWhat we never give it
A. One orchestrator with toolsAny process touching a system of record: lead qualification, follow-up, collections, onboarding—
B. Orchestrator + chained specialistsRead pipelines with distinct steps: enrich, classify, summarize, draftDirect writes from the specialist
C. Parallel swarmBroad search: market research, competitive monitoring, mass document readingAny write permission on CRM or billing

How we pick the topology at Mintec

A real story we have told in full: we connected an agent straight to a logistics client's CRM via API. Within 72 hours we had 14 duplicate opportunities, fabricated follow-up notes, and deal stages jumping from "qualified" to "lost won" and back. This was not bad AI — it was architecture: two agents writing to the same contact with no orchestration layer. The case and its fixes are in AI agents in a real CRM.

Since then the pattern is consistent:

  • Processes that write (CRM, billing, collections): one agent, one credential, one source of truth, deterministic rules around it. The who's in charge question gets settled with n8n's three-question rule, not by adding agents.
  • Processes that read and decide (lead qualification, reply classification): one orchestrator with rules + AI logic — the build with AI, execute with rules principle.
  • Exploration processes (research, competitive monitoring, bulk reading): this is where we do open parallel subagents, output to file, no write credentials. It is the only case where the 15x multiplier pays for itself.

Clear opinion: in 90% of client processes running in production, the answer is one agent. A swarm is a research tool, not an operations tool. If your diagram has ten boxes and one genuinely parallel task, you bought an org chart, not an architecture.

How to implement it without paying for the lesson twice

  1. One agent, clear tools, one record. Standardize the process first — the agent implementation framework exists because 88% never reach production.
  2. Measure from day 1 with ~20 real cases. Agents are non-deterministic even with the same prompt: a green demo is a single sample. Report pass rates across multiple runs.
  3. Add the second agent only with real parallelism. Anthropic publishes the scaling rules: simple query = 1 agent with 3–10 tool calls; direct comparison = 2–4 subagents; complex research = 10+ with divided responsibilities.
  4. Single write thread, always. If two agents can modify the same record, you do not have redundancy — you have a race condition in natural language.
  5. Token budget per task and model routing. The biggest cost lever is not topology, it is which model serves which step: up to 60–80% of spend comes out with routing.

Three mistakes we see constantly: counting boxes on a diagram instead of parallel tasks; paying 15x for an "army" to solve what was a two-condition rule; and shipping without evaluation, trusting the first green run.

The conversation is still open in that Reddit thread. What moves it is not one more opinion: it is knowing what each box costs and what happens when it falls.

Sources: Anthropic, How we built our multi-agent research system (2025) · McEntire, The Organizational Physics of Multi-Agent AI (SSRN, 2026) · Cemri et al., Why Do Multi-Agent LLM Systems Fail? (UC Berkeley, ICLR 2025) · Cognition, Don't Build Multi-Agents · Gartner, Hype Cycle for Agentic AI 2026 · r/Entrepreneur (54 upvotes, 125 comments).

Frequently Asked Questions

How many AI agents do I need to automate a process?

One, by default. A single orchestrator with good tools and one source of truth covers most business processes. Add a second agent only when the task genuinely parallelizes (several independent searches or analyses) and the value of each task justifies paying ~4x a chat's tokens. Multi-agent systems run about 15x a chat's tokens, so they only pay off in high-value, genuinely parallel work.

What does an army of agents cost per month?

Using Anthropic's published multipliers (agent ~4x, multi-agent ~15x a chat's tokens), a task that costs $0.02 in plain chat costs about $0.08 with an agent and about $0.30 with a swarm. At 10,000 tasks a month that is $200, $800 and $3,000 respectively. The difference is not the price per token — it is how many tokens the topology you chose consumes.

When is it worth running multiple agents in parallel?

When the task splits into independent subtasks (broad information gathering, analyzing many documents or markets), the output is read-only with no writes to systems of record, and the result is valuable enough to pay the cost premium. Anthropic uses this for broad research; Cognition recommends keeping writes single-threaded and having the extra agents contribute intelligence rather than actions.

Related Articles