Benchmarking Opus 5 on SlopCodeBench: Why AI Still Needs Human Steering

Benchmarking Opus 5 on SlopCodeBench: Why AI Still Needs Human Steering

I tested Opus 5, Sonnet 5, and Opus 4.8 on the new SlopCodeBench, which simulates real-world software evolution with hidden requirements. While Opus 5 achieved a 24% strict pass rate, it generated five times more code than older models and still failed to complete challenges without defects. The results suggest that current models cannot yet run lights-off for complex engineering tasks, as every dollar spent bought correctness but nobody bought enough of it.

every dollar bought correctness. nobody bought enough of it.

More from this day

2026-07-28