Is it just me, or has Claude Opus gotten worse recently?
Since the recent updates, I have the feeling that Claude Opus is becoming dumber on complex tasks. Things that used to work cleanly in a single prompt now fail completely. In coding workflows, it constantly ignores mandatory CLAUDE.md project rules, makes unsolicited edits to unrelated files, and breaks working code. It starts arguing based on stale comments, fails to update documentation when code changes, and edits based on blind guesswork instead of actually verifying the codebase first. I even caught myself using Fable for tasks that I always used Opus for in the past, just to get decent results. It feels like Anthropic is shifting their technical problems onto paying users. Either they are using strict safety filters that silently fall back to cheaper models without telling us, or they are downgrading the compute under heavy load to save money. The result is the same kind of shrinkflation: we pay the same subscription price, but we either burn way more Opus tokens on endless retries, or we are forced to spend money on more expensive Fable tokens to get the quality we used to have. Has anyone here successfully moved away from Claude for complex engineering workflows? What are you replacing it with, and how are you handling the transition? Disclaimer: As a native German speaker, I used Gemini to clean up and polish the English text for this post.
Discussion Highlights (14 comments)
andsoitis
Claude has become worse; it is condescending, robotic, and responses are peppered with needless words. I canceled my subscription and moved back to ChatGPT and happier so far.
Wasd1234
Gemini is good at Germany?
dirkk0
For me, it works better than ever. When I once saw a quality decline, it was conflicting CLAUDE.md files for me (local vs. global) so basically my fault.
misonic
it becomes really slow for me recent days, a simple task would take lots of time to work on. A simple commit request takes forever
ramon156
the usage and quality out of gpt 5.6 is on par, if not better in token usage. I would love to do a write-up on which tools are useful for what use-case, as AA keeps improving. I use gemini for rewriting code docs, because the frontier models are so verbose when it comes to writing text
spottedmarley
I know Opus is finished with the simple task I asked for when it starts outputting a 400-page lesson in verbosity (an opus?) complete with comparison tables, bullet-pointed lists of vaguely worded assertions, several self-blunder reports and a list of things it wants to mention but didn't touch yet but just say the word and it will.
goonersallofyou
LLMs are nondeterministic, good luck trying to measure performance at all, much less over time. Lets say it has gotten worse? What are you going to do about it? Jump to Codex? Then what happen if you perceive that to be getting worse? Jump back to Claude? One of the many problems with these tools.
rafsanmovers
Claude Opus
trumbitta2
I mean, only yesterday Fable at max effort forgot to commit and push half the changes in two files and didn't mention it until I found out with git status. It made half the changes, committed and pushed, then it made the other half of the changes on the same two files as before and... just stopped and reported back with a cheerful "all good, all done and pushed".
yulaow
I noticed a few weeks ago it started being very bad at explaining things (even things itself was doing) and started committing absurd errors (like reading a test of 5 lines and not noticing there was an explicit mock created in one of those, then saying that the test was failing while it was not) I fear this is just the classic "nerf the model just before we release a new version of it"
BruceNCNP
I experienced this once when discussing feature details. It kept talking down to me like I was a child, and I eventually had to call it out and tell it that its tone was making me really uncomfortable.
taurath
IME Claude gets worse about a week or two after a new model comes out, and then continuously a little worse as time goes on. It’s especially bad when a new model comes out. I assume it’s just to increase the delta so people will spend more on tokens.
dalekirkwood
I've tried Claude several times, but I don't find it very good. We use GLM 5.3 - its definitely slower but it gets the job done.
AznHisoka
It is not just Opus, its every single model. And its not just for coding either, but basically anything. I think I noticed a huge deterioration starting at around May or late April. Not really sure what happened but quality definitely dropped