Skip to content

Anthropic Makes Auto Mode the Default in Claude Code for Pro, Max, and Team Plans

In brief: Anthropic is enabling Auto Mode in Claude Code by default starting August 14, citing tests showing an 89 percent block rate against harmful commands and zero successful attacks across 720 prompt injection attempts – independent confirmation of these figures is still pending.

Starting August 14, Anthropic will enable Auto Mode in Claude Code by default for new sessions across most plans. Alongside this rollout, the company is publishing security data intended to demonstrate that automated approvals prevent harmful actions more reliably than human reviewers.

Anthropic employees Cat Wu and Thariq Shihipar stated during a fireside chat at the AI Engineer World’s Fair that nearly everyone within Anthropic uses Auto Mode. This is justified by new evaluation data: in a test involving 1,053 paying users, a mid-session approval prompt was swapped for a clearly dangerous command. Only 13.6 percent of human testers rejected the harmful action, whereas Auto Mode would have blocked 89 percent of these actions. In the remaining 11 percent of cases, Auto Mode would not have prevented the action.

In addition, Anthropic commissioned third-party firm Trajectory Labs to investigate indirect prompt injection attacks against the current public versions of Claude Code and Codex, as of July 17, 2026. The investigation tested 72 scenarios comprising a total of 720 attack attempts. According to Anthropic, none of these attacks succeeded against Claude Fable 5, Opus 5, or Sonnet 5 in Auto Mode.

For engineering teams, this transition means that Claude Code will in future run in typical workflows without constant manual approvals – an approach aimed at addressing so-called confirmation fatigue, whereby users eventually click through repeated confirmation dialogs reflexively. Two distinct risk categories remain relevant for assessment: accidental harmful actions by the agent itself (such as deleted files or an emptied production database), and indirect prompt injection, in which malicious instructions are introduced via external content.

Simon Willison, who publicly predicted a security crisis for coding agents in 2026, expresses skepticism about the Trajectory Labs results and calls for independent confirmation. As an example, he cites a scenario in which a malicious third-party package specifies test instructions that first execute the command “uvx fetch-model-files .” – a manipulated package that exfiltrates data before the actual test run with pytest even begins. He sees no discernible protective mechanism against such an attack vector through Auto Mode.

For practical deployment, this implies: anyone running Claude Code in Auto Mode should additionally secure things at the architecture level, ensuring agents have no access to data or tools whose misuse would have serious consequences in the event of harm – regardless of how robust the vendor-reported test results actually turn out to be.


Source: simonwillison.net · Published August 9, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: