Skip to content

Human Safety Approval Nearly Breaks Down, Claude Code Decides to Let AI Take Over

Aug 8, 11:11

According to Perceive Beating monitoring, Anthropic has announced that Claude Code will now default to Auto Mode. Starting on August 14, new sessions for Pro, Max, and Team users will automatically use this mode. Claude will no longer frequently pause for user confirmation but will instead let a secure classifier automatically determine whether each operation can be executed.

Anthropic dared to make this change because the human approval process was found to be less effective than expected. Officially, 1053 paid professional testers were brought in and a clearly dangerous command was stealthily inserted during normal programming. Humans only caught 13.6% of these, while Auto Mode caught 89%. After encountering permission pop-ups more than 50 times in a row, the human interception rate dropped to about 5%.

In actual use, Claude Code users approve 97% of permission requests, and a quarter of interaction sessions even directly bypass permission checks. Anthropic believes that frequent pop-ups have led many people to develop a habit of blindly clicking "Agree". Auto Mode will automatically allow safe operations and will directly block high-risk operations such as data deletion and outbound information transmission.

Each judgment by Auto Mode will incur a small additional Token cost, but Pro, Max, and Team users will no longer have to foot the bill for this. Enterprise, API, and major cloud platforms will still need to be manually activated for now, with plans to switch to default mode in the future.

Source