The 89 vs 13.6 Split That Made the Case

Anthropic didn't flip Auto Mode to default on a hunch. In a controlled study of 1,053 paid testers, human reviewers caught just 13.6% of dangerous commands under manual approval, a rate that fell to roughly 5% once a session had accumulated 50 prompts - the classic approval-fatigue curve. The separate classifier that now screens every risky tool call caught 89% of the same commands [1].
That gap sits next to a habit Anthropic's own telemetry surfaces: users approve 97% of permission prompts regardless of content, which makes the old safeguard closer to a formality than a real check [1]. An independent developer analysis made a blunter version of the same point months earlier, arguing the human-in-the-loop had effectively vanished before Auto Mode existed because manual prompts were routinely clicked through without scrutiny [2].
Anthropic frames the flip in outcome terms too: the share of sessions containing an unrequested harmful action fell from 6.3% under manual approval to 2.4% under Auto Mode, and Team and Enterprise customers running Auto Mode ship about 25% more pull requests than those who don't [1]. Trade press coverage corroborated the timeline and the headline classifier numbers ahead of the August 14 switch [3].


