Benchmarking Opus 5 on SlopCodeBench: Why AI Still Needs Human Steering

I tested Opus 5, Sonnet 5, and Opus 4.8 on the new SlopCodeBench, which simulates real-world software evolution with hidden requirements. While Opus 5 achieved a 24% strict pass rate, it generated five times more code than older models and still failed to complete challenges without defects. The results suggest that current models cannot yet run lights-off for complex engineering tasks, as every dollar spent bought correctness but nobody bought enough of it.
every dollar bought correctness. nobody bought enough of it.