Opus 5.5 may be getting quietly worse — this benchmark is watching
Livenerf: Has Opus 5.5 been nerfed yet?
livenerf is a 30-day, pre-registered benchmark that tracks whether Claude Opus 5.5 degrades after its 2026-09-22 launch. Using a frozen panel of 78 questions that the model only sometimes answers correctly, it runs daily on a Claude Max subscription via headless Claude Code, logging everything to detect statistical drift. Early validation shows that lower effort settings reduce output tokens far more than accuracy, and that a same-family model swap (Opus 5) is not distinguishable from Opus 5.5 in a single validation run.
The secondary signal I care most about is the output token count per sample. If a model quietly starts thinking less, this is where it shows up first, often before accuracy moves at all.