Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
aarondong
220 points
130 comments
July 24, 2026
Related Discussions
Found 5 related stories in 64.5ms across 5,804 title embeddings via pgvector HNSW
- Artificial Analysis Intelligence Index v4.2 nojs · 94 pts · September 05, 2026 · 60% similar
- GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index wertyk · 21 pts · September 03, 2026 · 56% similar
- GLM-5.3 Artificial Analysis Benchmarks apitman · 114 pts · August 18, 2026 · 56% similar
- Benchmarking Opus 5 on SlopCodeBench dhorthy · 216 pts · July 27, 2026 · 55% similar
- Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index wertyk · 324 pts · August 12, 2026 · 55% similar
Discussion Highlights (17 comments)
aarondong
Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
claude-ai
On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb). Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
firasd
Very interesting that one of the components is "AA-Omniscience Index" AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max) I've thought for a while that Gemini 3.x has 'big model smell'
sggyamg
It's new, normal.
andy99
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
chmod775
The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
hoppp
I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.
zormino
I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.
theplumber
Why do I find it dummer/even more superficial than opus 4.8? It just continued a session and I had to stop it because it become obviously “lost”
nu11ptr
I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
didibus
What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max. That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
kristopolous
I posted this before but I have a really simple shell tool to keep up with these charts over at https://github.com/day50-dev/aa-eval-email This also works $ curl day50.dev/art-analysis.sh | bash Artificial analysis knows about my tool and I'm working with them on getting their API improved.
zuzululu
i used for several hours now and my verdict is that its no better or worse than sol its surprisingly bad at UI which is unexpected its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)
XCSme
Twice the cost for 4% more intelligence, is it worth it?
zkmon
I think a more useful metric would be intelligence per dollar spent.
anigbrowl
Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song. It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.
NamlchakKhandro
This company sniffs it's own farts too much tbh