Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

A browser game simulating human oversight of an AI coding agent reveals that players missed 32.9% of malicious commands, with the most-missed being 'npm run analyze' (64.7% approval). The game's 40,000+ runs show that threats disguised as familiar scripts succeed twice as often, and miss rates climb under time pressure. Over-blocking benign commands (like 'rm -rf dist/') is also common, highlighting the fatigue and context problems in human-in-the-loop approval systems.
That’s a great example of how dangerous actions are perceived as innocent. The entire model of approving specific commands is absolutely bonkers.