Livenerf: Has Opus 5.5 been nerfed yet?
bryan0
439 points
181 comments
September 29, 2026
Related Discussions
Found 5 related stories in 83.1ms across 8,041 title embeddings via pgvector HNSW
- Why does Opus 5 feel worse to work with? numeri · 825 pts · August 14, 2026 · 48% similar
- Benchmarking Opus 5 on SlopCodeBench dhorthy · 216 pts · July 27, 2026 · 46% similar
- Litelm: LiteLLM Without the Bloat kennethwolters · 118 pts · September 11, 2026 · 46% similar
- GLM-5.3 is now open-weight jeudesprits · 646 pts · August 28, 2026 · 45% similar
- Elevated Errors for Opus 5 TimCTRL · 93 pts · July 26, 2026 · 43% similar
Discussion Highlights (20 comments)
gigatexal
This is genius. I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.
aabhay
Only ten day interval? I felt Astra got nerfed within a week
judge2020
I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.
solenoid0937
Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
solfox
It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later". After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets. I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again ?
whs
I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.
gr_norm
All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me . The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.
nico
Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more often The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
LeoPanthera
n=1 is useless. The output is not deterministic.
colordrops
This repo already has too much visibility now. Anthropic will soon benchmaxx it.
jug
We also have Nerf Bench: https://www.bridgebench.ai/nerf-bench They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra. This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
octoberfranklin
OpenAI will simply set up a classifier to detect if the client is livenerf, and selectively not nerf those requests. Open models are the endgame.
johnfn
"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was. I made a graphic to explain why people feel like the models get nerfed: https://x.com/thesilenceturns/status/2103551351825543610 The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
xlayn
The only reason why claude fable is better than opus in my opinion is that it has more "criteria"... if you present a problem and then ask for his recommendation you can get an opinion on why and reasoning on why that one... Opus is going to vomit 10k lines of extremely dense prose in nerdify++ level. Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes.. I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1 I had the exact same feeling every time they have a new big release
bethekidyouwant
People just tend towards conspiracies you have to actively fight it.
zeroonetwothree
If it can’t even tell apart Opus 5 and 5.5 (according to the readme) then it’s not useful
apt-apt-apt-apt
Fable 5 seems like it got nerfed when 5.1 came out.
j45
New model releases that have positive reviews should come with a nerfalert reminder service to make hay until it's shaped and shaped and shaped.
onlyrealcuzzo
This is bad data at its finest. Truly, madly, deeply sloppy.
winwang
Complete anecdote, and nothing to do with relative nerfing or not: Opus 5.5 has been surprisingly good for me (including the past couple hours), especially for following research-level questions/directions.