Fable 5 – Median thinking declined in August
espeed
378 points
267 comments
September 21, 2026
Related Discussions
Found 5 related stories in 73.5ms across 7,302 title embeddings via pgvector HNSW
- Fable 5 dropped from subscriptions before the promised July 19 date Ryan_Li · 21 pts · July 17, 2026 · 59% similar
- Fable extended until 19 July protoman3000 · 85 pts · July 12, 2026 · 57% similar
- Fable 5 finds major bugs in 10 year old open source game networking libraries gafferongames · 14 pts · July 18, 2026 · 54% similar
- Working on Economics with Fable 5 Wilsoniumite · 68 pts · September 07, 2026 · 52% similar
- Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher u1hcw9nx · 716 pts · September 13, 2026 · 51% similar
Discussion Highlights (20 comments)
alexjplant
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable). I wonder what their official explanation for this behavior is.
CamperBob2
How do you measure thinking tokens? They don't send those back to the client.
r2-129
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too. Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference. Buy decent coffee instead of your $200 subscription and sidestep all the scams.
mlmonkey
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
cloudking
How do you create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.
llmslave
I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good
theplumber
It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.
jotato
Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it. For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host" I had to tell it to ssh into the server and run journlctl to check it Anecdotal, I know, but they all seem to be less capable with time. _edit_ I use the same reasoning level of `medium`
bpodgursky
The smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open. (ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)
matheusmoreira
Anthropic is straight up scamming its users at this point.
espeed
The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems ( https://news.ycombinator.com/item?id=48742153 ), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.
Waterluvian
I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening. Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now." The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue. Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
underlipton
Gemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too. I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this. EDIT: And, frankly, I can't wait. I'm tired of the sketchy and dishonest way these companies are behaving.
vb-8448
They want transparency from everyone else but not for them ... you don't say.
tamimio
This is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.
saejox
This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too. Tests their intelligence, not their diligence. Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.
bix6
So in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?
rcr-anti
I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.
jesse_dot_id
The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed. AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
varispeed
I stopped using Fable long time ago. It's worse than Sonnet. Opus is not much better. This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam. If I pay for Fable, I should get full, not nerfed model at honest pricing. Regulators should investigate them. OpenAI is no different. Astra has basically the same problem.