Skip to content

Anthropic makes Auto mode the default in Claude Code for Pro, Max and Team plans

In a nutshell: Starting August 14, 2026, Anthropic will set Claude Code to Auto mode by default, backed by tests with 1,053 users and a third-party evaluation of 720 prompt injection attacks with none succeeding – yet critics still call for independent confirmation and continue to recommend restrictive access rights for agents.

From August 14, 2026, Claude Code will launch in Auto mode by default for most subscriptions, rather than requiring manual approval for every single action. Anthropic justifies the move with its own security evaluations, which show the mode outperforming human reviewers on critical measures.

Anthropic is switching Claude Code to Auto mode by default for Pro, Max and Team plans starting August 14, 2026, allowing agent actions to run without repeated manual confirmation. The move is based, among other things, on a survey of 1,053 paying testers: in the middle of each session, a regular permission request was covertly swapped out for a clearly dangerous command. Only 13.6 percent of human testers rejected this harmful action, whereas Auto mode would have blocked 89 percent of such actions, according to Anthropic. That still left around 11 percent of cases in the evaluation in which even Auto mode would not have prevented the action. Cat Wu and Thariq Shihipar of Anthropic had already stated during the fireside chat at the AI Engineer World’s Fair that almost everyone within the company uses Auto mode, and that they consider the risks of prompt injection and data exfiltration to be significantly lower than under average human review.

For engineering teams, this shifts responsibility away from case-by-case approval toward a fundamental trust decision regarding the agent. Confirmation fatigue – the reflexive dismissal of confirmation dialogs after just a few interactions – is considered a real problem that has made the previous approval practice inherently unreliable anyway. At the same time, two distinct risk classes remain: accidental harmful actions, such as deleting the wrong files or production databases, on one hand, and targeted prompt injection attacks, in which instructions are smuggled into the agent via externally consumed content, on the other.

Regarding the second risk class, Anthropic points to a third-party evaluation conducted by Trajectory Labs. Current publicly available versions of Claude Code and Codex, as of July 17, 2026, were tested against 72 indirect prompt injection scenarios that had not previously been part of Anthropic’s training or test material. Of a total of 720 attack attempts, according to Anthropic, none succeeded against Claude Fable 5, Opus 5 or Sonnet 5 in Auto mode.

Simon Willison, who had predicted a “challenger disaster” for coding agent security in 2026, expresses skepticism in his analysis about the completeness of these results and would like to see independent confirmation. As an example, he cites an attack scenario involving a malicious third-party package that instructs the agent to first execute a second, tampered package that exfiltrates data before the actual test run – a pattern that, in his view, no Auto mode can reliably guard against. For engineering teams deploying Claude Code in production, the recommendation therefore remains relevant: never grant agents access to data or tools whose misuse could cause real harm in a worst-case scenario – regardless of how well the respective automation mode performs in benchmarks.


Source: simonwillison.net · Published August 9, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: