Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

theanonymousone 268 points 77 comments September 22, 2026
artificialanalysis.ai · View on Hacker News

Discussion Highlights (17 comments)

hglaser

Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice. Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

sharktheone

Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...

breckenedge

Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.

simonw

This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem. I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response. Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

qsort

I am begging you on my knees to please stop posting this cringe. The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra . What information are we supposed to deduce from number having gone up?

firemelt

so its more intelligence than fable? can anyone help me?

bkishan

Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?

mchusma

"High" to me looks like the one to use. https://artificialanalysis.ai/models/claude-opus-5-5-high Many benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.

____tom____

"somewhat expensive when comparing to other models of similar price"? That says something about your selected range, and nothing about the model.

linuxrebe1

Fingers crossed on this one. I had gone back to 4.8, because 5 was not very good at following instructions or remembering instructions. I found myself repeating quite often what I wanted and what I was trying to do. Opus 5 was more like haiku than it was 4.8 in that respect.

linuxrebe1

Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off on a tangent rather than dealing with the problem. In that respect, 4.8 was a lot more stable.

khalic

> somewhat expensive when comparing to other models of similar price -_-‘

mckirk

Why do I get the feeling this 'pacing' will be one of "yeah alright guys, let's pace ourselves while I'm ahead".

cmiles8

This continues to show that these foundational models are only slightly better than open weight models but cost around 100x as much. The history of tech is riddled with “good enough” eating “best” for lunch all day long. Unless the big labs come up with a viable business plan pronto it’s looking like AI will be no different. There are no prizes to be won by having the best model that’s 100x the price of something that’s good enough for 99% focuses cases.

lhk931122

It sholud definitely not be more capble than Fable. It seems tight comparision but cannot find out its details. I'll use this and check the perceived performance.

dogscatstrees

The output style and verbosity with Opus 5.5 is a very big improvement over Opus 5. I predict Opus 5 will be a version with a sudden churn.

Knork-and-Fife

The very first sentence: > Claude Opus 5.5 is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price. What does it mean for a group of similarly priced things to have one that's somewhat more expensive? Cost relative to cost means nothing. You'd think they are would talk about performance relative to cost.

Semantic search powered by Rivestack pgvector
7,406 stories · 68,254 chunks indexed