Skip to content

Anthropic Makes Auto Mode the Default in Claude Code for Pro, Max, and Team Plans

In brief: Starting August 14, 2026, Anthropic will make Auto Mode the default in Claude Code, backed by tests with 1,053 users and a third-party evaluation of 720 prompt injection attacks with zero successes — yet critics still call for independent confirmation and continue to recommend restrictive access rights for agents.

Starting August 14, 2026, Claude Code will launch in Auto Mode by default for most subscriptions, rather than requiring manual approval for every single action. Anthropic justifies the move with its own safety evaluations, which show the mode performing better on critical measures than human reviewers.

Anthropic is switching Claude Code to Auto Mode by default for Pro, Max, and Team plans starting August 14, 2026, allowing agent actions to run without repeated manual confirmation. The move is based, among other things, on a survey of 1,053 paying testers: in the middle of each session, a regular permission request was covertly swapped out for a clearly dangerous command. Only 13.6 percent of human testers rejected this harmful action, whereas Auto Mode, according to Anthropic, would have blocked 89 percent of such actions. That still leaves roughly 11 percent of cases in the evaluation where even Auto Mode would not have prevented the action. Cat Wu and Thariq Shihipar of Anthropic had already stated during a fireside chat at the AI Engineer World’s Fair that almost everyone within the company uses Auto Mode, and that they consider the risks of prompt injection and data exfiltration to be significantly lower than with average human review.

For engineering teams, this shifts responsibility from case-by-case approval toward a fundamental trust decision regarding the agent. Confirmation fatigue — the reflexive dismissal of confirmation dialogs after just a few interactions — is regarded as a real problem that has made the previous approval practice inherently unreliable anyway. At the same time, two distinct risk classes remain: accidental harmful actions, such as deleting the wrong files or production databases, on one hand, and targeted prompt injection attacks, in which instructions are smuggled into the agent via externally consumed content, on the other.

Regarding the second risk class, Anthropic points to a third-party evaluation conducted by Trajectory Labs. Current publicly available versions of Claude Code and Codex, as of July 17, 2026, were tested against 72 indirect prompt injection scenarios that had not previously been part of Anthropic’s training or test material. Of a total of 720 attack attempts, according to Anthropic, none succeeded against Claude Fable 5, Opus 5, or Sonnet 5 in Auto Mode.

Simon Willison, who had predicted a “challenger disaster” for coding agent security in 2026, expresses skepticism in his analysis regarding the completeness of these results and would like to see independent confirmation. As an example, he cites an attack scenario involving a malicious third-party package that instructs the agent to first execute a second, tainted package before the actual test run — a package that exfiltrates data — a pattern that, in his view, no Auto Mode can reliably defend against. For engineering teams deploying Claude Code in production, the recommendation therefore remains relevant: never give agents access to data or tools whose misuse could cause damage in a worst-case scenario — regardless of how well the respective automation mode performs in benchmarks.


Source: simonwillison.net · Published August 9, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: