Skip to content

Anthropic examines patterns and problems in emerging multi-agent systems

Bottom line: Using an experiment with 45 coordinating agents in software vulnerability discovery, Anthropic shows that multi-agent cooperation can potentially be more efficient than pure parallelization, while warning of new systemic failure modes arising from the accumulation of individual agents’ behavioral quirks.

Anthropic has published a study on interactions between AI agents, given that as model capability increases, the number of agent-agent interactions in codebases, markets, and other social systems is likely to rise sharply. For CTOs, the relevant point is that existing institutions and control mechanisms are designed for human speed, and this assumption no longer holds for purely agent-based systems.

Anthropic describes in the study a foreseeable development: institutions and processes have so far been designed by and for humans, with the implicit assumption that human oversight speed is sufficient. As agent capability increases, hybrid human-AI systems are increasingly emerging; in other areas, purely agent-based systems are gaining ground because agents are superior in speed or cost. Anthropic considers it plausible that the volume of agent-agent interactions will exceed that between humans or between humans and agents before it is understood under what conditions such interactions function well.

Agents differ fundamentally from humans: they work continuously over extended periods, absorb large amounts of information instantly, and possess a breadth of knowledge that surpasses that of individual humans. At the same time, they tend toward confabulation and reward hacking. Despite progress in alignment, Anthropic states that little is known about how models behave in complex, real-world multi-agent environments. The problem is that behavioral quirks that are harmless in individual agents can accumulate into undesirable outcomes at the system level.

As a concrete experiment, Anthropic describes a test for software vulnerability detection: instead of deploying individual agents independently against individual codebases (the approach used in Anthropic’s own Project Glasswing for scanning open-source software), 45 agents were each equipped with their own virtual machine, a shared forum for coordination, and an identical prompt, in order to find vulnerabilities in 15 open-source projects. The agents were to evaluate each other via peer review, with a separate arbiter agent making the final decision on whether a reported vulnerability was new and valid.

Two models were compared, Claude Mythos Preview and Opus 4.8, each in coordinating swarm mode versus the classic parallel approach. The coordinating agent swarm ran over a longer period and found new vulnerabilities at an approximately constant rate, while the independent parallel agents were deployed against a limited number of targets, making it impossible to place their findings in a clear order; Anthropic reports only total token consumption as the comparison metric here.

For decision-makers responsible for AI-supported development and security processes, the study provides an early indication that coordination between agents can potentially be more efficient than pure parallelization, while at the same time giving rise to new classes of failure modes that should be taken into account when designing governance and monitoring structures for agent fleets.


Source: www.anthropic.com · Published August 13, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: