Mercury 2.5 LLM hits 770 tokens per second
Retro_Dev
85 points
51 comments
September 23, 2026
Related Discussions
Found 5 related stories in 74.8ms across 7,510 title embeddings via pgvector HNSW
- Mercury 2.5 Topfi · 163 pts · September 08, 2026 · 72% similar
- Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU neomindryan · 266 pts · July 15, 2026 · 52% similar
- LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacB Alephinitesimal · 15 pts · August 21, 2026 · 48% similar
- Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) mikeayles · 12 pts · August 10, 2026 · 48% similar
- Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1 cdnsteve · 60 pts · September 10, 2026 · 47% similar
Discussion Highlights (13 comments)
rvz
The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
walrus01
Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
bearjaws
If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking. I've used it on a few for fun projects and its decent but the speed is crazy to watch.
nylonstrung
I honestly think the diffusion LLM approach is a dead end It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models Still unclear for what, if any use cases this is pareto frontier
low_tech_punk
it's stupid fast!
sharktheone
this feels like "we got the same benches as gpt-oss-120b but are also potentially slower while saying it is great"
entrope
> Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price. Well priced when compared to other models of similar price, eh? Are we allowed to call this slop, even if the output is not directly from an LLM?
nextaccountic
At some point the bottleneck becomes tool calling.. and as such, it's preferably if the model is co-hosted (in the same datacenter, at least) with your code repository and all other reference/context it needs (full documentation for most ecosystems, maybe even a copy of common crawl to minimize web fetch usage, etc)
whalesalad
I used this a few days ago and thought something must be wrong with how fast it was responding. "Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price." this is so funny. So when you have a stupid model that is fast - what do you use it for?
freakynit
I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max. Even GPT-OSS-20B performs way better than this in my own attempts to use it. I really really wanted to use this because it offers incredible speeds and pricing combinations. But nop.. I still am not using it.. not even for basic tasks.
seduerr
Cerebras is fast…?
the_arun
Chat Jimmy clocks at 17K tokens per sec burning LLM into the Chip - https://chatjimmy.ai/ - Source: https://theashishmaurya.medium.com/taalas-the-startup-that-p...
jjcm
I'm still sad that we haven't seen a new Taalas style chip a la https://chatjimmy.ai/ . Smaller models are good enough now to make that insane burst of tokens so useful.