The Study Behind the Switch
Anthropic's case for flipping the default rests on a controlled study of 1,053 paid professional testers: human reviewers caught only 13.6% of dangerous commands, while auto mode's classifier caught 89% [1]. Head-to-head, auto mode blocked 800 commands that humans had approved, while humans blocked only 6 commands auto mode had allowed [1]. The underlying failure mode isn't laziness so much as fatigue: Anthropic's own production data shows users approve 93% of permission prompts regardless of actual risk [2]. That matters most for long-running autonomous sessions, where manual review gets leaned on hardest exactly as fatigue sets in - the moment oversight is weakest is the moment agents are doing the most unsupervised work. Anthropic frames this less as a UX upgrade and more as a correction: constant approval prompts were producing an illusion of oversight rather than real oversight.


