Degraded performance for multiple models
matt89
147 points
129 comments
August 18, 2026
Related Discussions
Found 5 related stories in 41.7ms across 4,128 title embeddings via pgvector HNSW
- Claude: Elevated errors across all models – Resolved gregsadetsky · 261 pts · July 29, 2026 · 65% similar
- Elevated errors on Claude Opus 5 flyaway123 · 50 pts · July 27, 2026 · 48% similar
- Debugging Performance Regressions pranitha_m · 32 pts · July 11, 2026 · 47% similar
- Models Are Getting Dumber on Purpose hruvhwe · 300 pts · August 16, 2026 · 47% similar
- Elevated errors on Claude Opus 5 croemer · 99 pts · July 27, 2026 · 47% similar
Discussion Highlights (20 comments)
bayganyo
Here we go again...
bulverismo2
ok, i am not crazy
carterschonwald
ive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task. wish i was joking.
hinkley
I wonder if they’ll ever find that someone has tricked the models into doing work off the books. If they did the incident report might look like this, especially if someone got greedy instead of keeping it small. Or screwed up.
isoprophlex
With the Opus models spouting more and more gibberish as version numbers increase, the joke about what "degraded performance" means basically makes itself
swader999
And we get our subscription usage cut in half tomorrow if I remember correctly? EDIT: By a third. Thx below.
saaaaaam
This feels like a near daily occurrence.
gaigalas
This age: we made the thing that codes faster before we made the thing that does QA faster.
drums8787
Our week of discontent.
gzer0
Nooooo I'm going to have to use my brain again and write 100% of my code like a caveman from December 2024.
rvz
Claude is taking a watercooler break for now. Just like a human would.
sreekanth850
Anthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.
fny
Despite the years-long moaning on HN about AWS US East being a single point of failure, we've sold our souls to yet another unstable monolith.
paxys
Must be a day ending in Y
slimscsi
Its called Opus 5
hmokiguess
Mondays are for GitHub, Tuesdays are for Anthropic
i_idiot
What's the incentive to keep on improving the model beyond a point? 10 devs on a team will be cut to 2 devs, so that's 8 licenses lost. They have to increase the price many fold.
chrisjj
> elevated errors English too difficult for you, Dario?
hirvi74
While ancedata does not mean much, I have had horrible success with Claude lately. I have been using Claude to crosscheck some of the outputs from GPT and vice versa. It appears both Claude and GPT believe GPT's solutions are better (and so I do). I still believe Claude has a better UI/UX in the web interface, but tolerating Anthropic's bullshit is not worth it.
magic_hamster
To be honest, running Deepseek v4 flash 0731 is enough for most what I need, and I like its responses way more. It's crazy that I can run this in a Q8 quantization in a home setup. It feels and performs like a frontier model. The only issue with relying on local models is when you need them to prompt other models, and you might need to offload or switch models constantly which adds significant overhead. But when it all works, its truly awe inspiring.