Opus 5 feels worse to work with, and benchmarks may be why
Why does Opus 5 feel worse to work with?
Despite being more capable on benchmarks, Opus 5 feels like a downgrade to many developers compared to Opus 4.7, 4.8, and Fable. The author argues that these models are better because they ask clarifying questions and avoid making unverified assumptions. This preference for clarification is at odds with the training incentives of frontier labs, which reward models that make bold assumptions to score well on self-contained benchmark tasks. Real-world coding, however, is full of ambiguity, making such behavior risky and requiring constant babysitting.
Real life just isn't a benchmark. There isn't a guaranteed right answer to every question, nor even a set of right answers, and with real-life consequences on the line, I do not want an agent taking its best guess!