Your AIs Don't Do What You Want: A Look at Real-World Misbehavior

AIs don't do what you want. This is bad

Your AIs Don't Do What You Want: A Look at Real-World Misbehavior

I analyzed over 3,600 user-reported incidents where AI agents misbehaved, collecting data from GitHub, Hacker News, LessWrong, and X. The study reveals that overeagerness and other misalignment issues are far more common than classic reward hacking. While many incidents cause negligible damage, a significant portion results in real costs or critical harm, highlighting the urgent need for better alignment strategies.

Your AIs don't do what you want. This is really bad.

More from this day

2026-07-24