Claude Opus 5.5: The Real Cost of Intelligence

Claude Opus 5.5 Intelligence, Performance and Price Analysis

Claude Opus 5.5: The Real Cost of Intelligence

Artificial Analysis breaks down Claude Opus 5.5 across its Intelligence Index v4.3.2, which bundles ten evaluations including AA-Briefcase, GDPval-AA, Terminal-Bench 4.0, and Humanity's Last Exam. The analysis maps intelligence against cost per task, output token use, cache-hit pricing, and context window, revealing where the model sits on the Pareto frontier of performance versus price.

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
  1. simonw

    This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

    I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

    I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

    Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

  2. breckenedge

    Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.

  3. hglaser

    Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.

    Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

  4. mckirk

    Why do I get the feeling this 'pacing' will be one of "yeah alright guys, let's pace ourselves while I'm ahead".

  5. linuxrebe1

    Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off on a tangent rather than dealing with the problem. In that respect, 4.8 was a lot more stable.

More from this day

2026-09-22