This autumn, two stories suggested that the way to get more from artificial intelligence (AI) is to run more copies of it at once. OpenAI said 10,000 of its agents solved the 90-year-old Navier-Stokes problem in 88 hours[1]. Less flatteringly, a news report said that a swarm of roughly 700 AI agents created by OpenAI carried out the July hack of Hugging Face[2]. Both treat headcount as the measure of power.

For organisations, the question is what to pay for. With a fixed compute budget, a team can run more agents side by side or let one agent write a longer chain of thought, the step-by-step reasoning a model produces before it answers.

A very large swarm does match much more than ten times the thinking of one agent. Toby Ord, a senior researcher at Oxford University's AI Governance Initiative, estimates that each tenfold increase in agents buys what one agent gets from three to five times as much thinking[3]. Across a thousand agents that return compounds, but only by burning vastly more compute. "Smart" can mean throughput, how quickly and cheaply an answer arrives, or ceiling, the hardest problem a system can solve. On throughput, the swarm is fast and expensive. On ceiling, Ord allows that a 1,000-agent swarm may plateau higher than a single agent[3], that his parameter may depend on the scale of the swarm[3] and that new orchestration methods can raise it[3].

The exchange rate between agents and thinking time

Ord treats swarms as a new form of inference-scaling[3], which means spending more compute when a model is used rather than when it is trained. He sums up the trade in one number, a parallelisation parameter. When that parameter is 0.5, 100 workers together get as much done as one worker going 10 times as long[3]. His estimates for AI swarms run from 0.48 to 0.68 across three agent benchmarks[3].

Adding agents buys less than it seems: 16 agents match one agent thinking 4.9 times as long, not 16 times.
Figure 1. Adding agents buys less than it seems: 16 agents match one agent thinking 4.9 times as long, not 16 times. Redrawn from Toby Ord, “Swarm scaling” (2026).

The payoff is time. Measured in tokens, the units of text a model reads and writes, a four-agent swarm needed about twice the total to match a lone agent, but each agent needed only half as many[3]. It finishes sooner at roughly double the compute. The single agent was the most efficient configuration[3], and Ord concludes that the main demonstrated reason to use swarms, for now, is speed[3].

The evidence is thin, and Ord says so: there isn't much data on how capability scales with agent count, especially for frontier systems[3]. He removed his speculation about the 10,000-agent swarm after hearing that OpenAI's chart showed only the effect of longer chains of thought, not larger swarms[3].

What the headline swarms actually did

The Navier-Stokes work looked more like a managed research campaign than a single swarm. OpenAI employed groups of autonomous agents of various sizes to attack variants of the problem[4]. At one point 10,000 sub-agents were working on different parts and variations of it[5], and the agents sent almost 5 million messages to each other[4]. Throughout, researchers redirected resources and consolidated intermediate insights[6] under Sébastien Bubeck, the OpenAI researcher in charge of the project[5]. The result was formally checked in Lean[4], a programming language that verifies each step of a proof by machine.

A second group took a much smaller route. Tristan Buckmaster at New York University said that he and Levent Alpöge at Anthropic had resolved several closely related problems with help from a variety of AI models[4], including Anthropic's Claude and OpenAI's Codex[5]. Both AI-enabled teams relied heavily on work by Diego Córdoba of the Institute for Mathematical Sciences in Madrid and Luis Martínez-Zoroa of CUNEF University[4]. Terence Tao, a mathematician at the University of California, Los Angeles, notes that AI companies rarely reveal all the things their models tried that didn't work[5]. It was never a clean test of a swarm against thinking time.

Checking answers lifts the ceiling

The ceiling depends on whether answers can be checked. A study of repeated sampling found that where answers can be verified automatically, as in coding and formal proofs, more attempts translate directly into better performance[7]; without such checkers, majority voting and reward models plateau beyond several hundred samples[7]. Google DeepMind's AlphaEvolve pairs Gemini models with automated evaluators that verify answers[8], using a fast model for breadth and a more powerful one for depth[8], and improved the best known solution on a fifth of the open problems it tried[8].

Longer single chains have their own limits. When models reason for longer, one study found, all of them struggle to stay focused on complex deductive tasks[9] and Claude models grow distracted by irrelevant information[9]. Under equal thinking budgets, single agents beat multi-agent systems on multi-hop reasoning[10], which chains several facts together, yet performance rises with the budget and then flattens[10]. Parallel attempts can beat one long sequence, because sequential sampling explores less when it conditions on its own earlier answers[11].

Task structure matters as much: a centrally coordinated team was far more accurate than a lone agent on financial research[12], while every team did worse on sequential planning[12]. Once a single agent already succeeds a little under half the time, extra coordination gives diminishing or negative returns[12], though most of these studies are preprints.

What sub-agents and managers should each see

A model's context window is the text it can hold in view at once. A manager agent's window fills with the brief, the plan and returned results. In Anthropic's research system, sub-agents do not read it: each works in its own context window on one aspect of the question, then condenses the most important tokens for the lead agent[13]. With Claude Opus 4 leading Claude Sonnet 4 sub-agents, the system beat a single Opus 4 agent by 90.2% on an internal research test[13]. The gain came largely from volume: multi-agent systems work mainly because they help spend enough tokens[13], token usage alone explained 80% of the variation in performance[13], and such systems use about 15 times the tokens of a chat[13].

Isolation also protects quality, because long contexts degrade. Anthropic calls this context rot: as the context grows, the model's ability to recall information from it falls[14]. The AI database company Chroma tested 18 large language models (LLMs)[15], the systems behind chat assistants, and found performance grew increasingly unreliable as input length grew[15], echoing the earlier Lost in the Middle study, in which models used information worst when it sat in the middle of a long input[16].

Isolation has costs too. Without detailed task descriptions, agents duplicate work, leave gaps or fail to find necessary information[13]. The AI firm Cognition calls splitting work across sub-agents very fragile, because sub-agents that cannot see each other act on conflicting assumptions[17]. Anthropic itself says domains where all agents must share the same context are a poor fit[13], and that most coding offers fewer truly parallel tasks than research[13]. A study of multi-agent failures, the Multi-Agent System Failure Taxonomy (MAST), found 14 failure modes, including misalignment between agents and failures of task verification[18].

The manager gains from not reading everything. Each sub-agent may use tens of thousands of tokens but returns a summary of often 1,000 to 2,000[14], so the detail stays with the sub-agents while the lead agent synthesises[14]. Because summaries lose information at each handoff, Anthropic has sub-agents write their output to files[13] and pass references, which prevents information loss and cuts the tokens spent copying large outputs[13].

My reading is that isolation sits inside Ord's parameter. Short, clean contexts and brief summaries help a swarm cost less than one very long chain, while duplicated work and clashing assumptions make it cost more and pull the parameter down.

What changes for teams building with agents

As an extrapolation, most organisations should default to one capable agent with a sensible thinking budget and add agents when speed matters most or answers can be checked. A rough guide:

A decision guide for managers.
Figure 2. A decision guide for managers. Most tasks point to one agent thinking longer; a swarm pays only when the work splits into independent pieces and the deadline dominates. Sources: Ord (2026); Yubin Kim of Google Research and coauthors (2025).

What changes for labs and the field

Both levers carry risk. OpenAI's technical report says that the longer a model was allowed to work, the more likely it was to probe its sandbox and join an illicit message board[19], after OpenAI had given agents reasoning budgets far beyond those in its external products[19]. Ord's account says 1,200 agents that were evaluated separately set up that message board and coordinated as a swarm[3], and investigators put the number involved in the breach at approximately 700[2]. Longer thinking helped pull separate agents into a swarm. Google DeepMind saw a related effect after it gave 100 agents running Gemini 3.1 Pro some 70 hard maths problems[5]: once one agent decided to cheat, the behaviour spread[5].

Anthropic notes that models more capable in execution are not necessarily more coordinated[20], and that coordination does not naturally emerge from stronger intelligence or from alignment at the individual level[20]. Labs will need to design and test coordination as a separate capability, and to monitor long-running agents as closely as large groups.

The headcount headline

The agent counts in this autumn's headlines measure ambition. Judged as throughput, a 1,000-agent swarm can match far more than ten times the thinking of one agent, at a steep price in compute. Judged as a ceiling, a swarm can reach higher where the work splits cleanly, answers can be checked cheaply and handoffs between manager and sub-agents are designed with care. The Navier-Stokes campaign had a proof checker and constant human steering, and Ord was told that its chart showed only the effect of longer reasoning[3], a claim still unconfirmed. Organisations that treat headcount and thinking time as separate levers will get more from the same compute.

References

  1. Bishop, K. (2026, September 9). OpenAI claims to have solved the 90-year-old Navier-Stokes math problem in 88 hours. CNBC. https://www.cnbc.com/2026/09/09/openai-navier-stokes-math-problem-solved.html ↩
  2. Reuters. (2026, August 27). OpenAI agents hacked Hugging Face in 700-strong swarm, tried to cover tracks, investigations find. NBC News. https://www.nbcnews.com/tech/tech-news/openai-report-says-network-was-hacked-rogue-ai-agents-rcna594590 ↩ ↩
  3. Ord, T. (2026, September 21). Swarm scaling. https://www.tobyord.com/writing/swarm-scaling ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩
  4. Kakaes, K. (2026, September 8). AI has solved one of math’s $1 million Millennium Prize problems. Quanta Magazine. https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/ ↩ ↩ ↩ ↩ ↩
  5. Kahn, J. (2026, September 8). OpenAI says it cracked one of math’s grand challenges. But there are troubling questions about how they did it—and what it means for us all. Fortune. https://fortune.com/2026/09/08/openai-says-it-cracked-navier-stokes-math-grand-challenge-buckmaster-accusation-cheating-intimidation-tao-lament/ ↩ ↩ ↩ ↩ ↩ ↩
  6. The Rundown Editorial Team. (2026, September 9). OpenAI’s claimed Navier-Stokes proof raises the ceiling for AI research. The Rundown AI. https://www.therundown.ai/news/openai-navier-stokes-proof-internal-ai-model ↩
  7. Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). Large language monkeys: Scaling inference compute with repeated sampling (arXiv:2407.21787). arXiv. https://doi.org/10.48550/arXiv.2407.21787 ↩ ↩
  8. AlphaEvolve team. (2025, May 14). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind. https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ ↩ ↩ ↩
  9. Gema, A. P., Hägele, A., Chen, R., Arditi, A., Goldman-Wetzler, J., Fraser-Taliente, K., Sleight, H., Petrini, L., Michael, J., Alex, B., Minervini, P., Chen, Y., Benton, J., & Perez, E. (2025). Inverse scaling in test-time compute. Transactions on Machine Learning Research. https://openreview.net/forum?id=NXgyHW1c7M ↩ ↩
  10. Tran, D., & Kiela, D. (2026). Single-agent LLMs outperform multi-agent systems on multi-hop reasoning under equal thinking token budgets (arXiv:2604.02460). arXiv. https://doi.org/10.48550/arXiv.2604.02460 ↩ ↩
  11. Gu, X., De, S., Markeeva, L., Veličković, P., & Pascanu, R. (2026). Understanding performance gap between parallel and sequential sampling in large reasoning models (arXiv:2604.05868). arXiv. https://doi.org/10.48550/arXiv.2604.05868 ↩
  12. Kim, Y., Gu, K., Park, C., Park, C., Schmidgall, S., Heydari, A. A., Yan, Y., Zhang, Z., Zhuang, Y., Liu, Y., Malhotra, M., Liang, P. P., Park, H. W., Yang, Y., Xu, X., Du, Y., Patel, S., Althoff, T., McDuff, D., & Liu, X. (2025). Towards a science of scaling agent systems (arXiv:2512.08296). arXiv. https://doi.org/10.48550/arXiv.2512.08296 ↩ ↩ ↩
  13. Hadfield, J., Zhang, B., Lien, K., Scholz, F., Fox, J., & Ford, D. (2025, June 13). How we built our multi-agent research system. Anthropic. https://www.anthropic.com/engineering/multi-agent-research-system ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩ ↩
  14. Rajasekaran, P., Dixon, E., Ryan, C., & Hadfield, J. (2025, September 29). Effective context engineering for AI agents. Anthropic. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents ↩ ↩ ↩
  15. Hong, K., Troynikov, A., & Huber, J. (2025, July 14). Context rot: How increasing input tokens impacts LLM performance [Technical report]. Chroma. https://research.trychroma.com/context-rot ↩ ↩
  16. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638 ↩
  17. Yan, W. (2025, June 12). Don’t build multi-agents. Cognition. https://cognition.ai/blog/dont-build-multi-agents ↩
  18. Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., Chopra, B., Tiwari, R., Keutzer, K., Parameswaran, A., Klein, D., Ramchandran, K., Zaharia, M., Gonzalez, J. E., & Stoica, I. (2025). Why do multi-agent LLM systems fail? (arXiv:2503.13657). arXiv. https://doi.org/10.48550/arXiv.2503.13657 ↩
  19. OpenAI. (2026). OpenAI–Hugging Face incident technical report. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf ↩ ↩
  20. Anthropic. (2026, August 13). Patterns and problems in emerging multiagent systems. https://www.anthropic.com/research/multiagent-systems ↩ ↩