Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
aarondong
220 points
130 comments
July 24, 2026
Related Discussions
Found 5 related stories in 851.7ms across 14,736 title embeddings via pgvector HNSW
- Claude Opus 4.7 Intelligence, Performance and Price Analysis Topfi · 33 pts · April 18, 2026 · 62% similar
- GLM-5.2 is the new leading open weights model on Artificial Analysis himata4113 · 831 pts · June 17, 2026 · 59% similar
- Meta AI chief says their coming LLM has caught up with OpenAI's flagship model maxloh · 13 pts · July 03, 2026 · 54% similar
- Claude Sonnet 5 – benchmark results lucamark · 39 pts · June 30, 2026 · 53% similar
- OpenCode – Open source AI coding agent rbanffy · 607 pts · March 20, 2026 · 52% similar
Discussion Highlights (17 comments)
aarondong
Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
claude-ai
On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb). Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
firasd
Very interesting that one of the components is "AA-Omniscience Index" AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max) I've thought for a while that Gemini 3.x has 'big model smell'
sggyamg
It's new, normal.
andy99
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
chmod775
The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
hoppp
I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.
zormino
I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.
theplumber
Why do I find it dummer/even more superficial than opus 4.8? It just continued a session and I had to stop it because it become obviously “lost”
nu11ptr
I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
didibus
What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max. That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
kristopolous
I posted this before but I have a really simple shell tool to keep up with these charts over at https://github.com/day50-dev/aa-eval-email This also works $ curl day50.dev/art-analysis.sh | bash Artificial analysis knows about my tool and I'm working with them on getting their API improved.
zuzululu
i used for several hours now and my verdict is that its no better or worse than sol its surprisingly bad at UI which is unexpected its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)
XCSme
Twice the cost for 4% more intelligence, is it worth it?
zkmon
I think a more useful metric would be intelligence per dollar spent.
anigbrowl
Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song. It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.
NamlchakKhandro
This company sniffs it's own farts too much tbh