Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

snehesht 696 points 320 comments October 04, 2026
github.com · View on Hacker News

Discussion Highlights (20 comments)

snehesht

I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here. https://huggingface.co/Qwen/Qwen3.8-Flash-Next

esafak

Has anyone calculated the effective intelligence of these quantized models? I think publishing benchmarks with quantized models should become standard practice.

quietFalcon

Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?

gdevenyi

I had this working with the FreeToken inference engine a month ago when they launched. https://github.com/FlashML-org/FreeToken

deadbunny

> Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository. And I thought piping to bash was bad

prettyblocks

I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).

hypfer

Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller? The Readme doesn't say, but it's all AI generated, so..

panny

I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.

0xbadcafebee

Lol, sure, if you quant it to hell (Q2) it'll go real fast... They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers. It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

Tepix

Q2 quantization. Not interested.

kamranjon

Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well. https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

ryan_glass

Anyone know how it compares to GLM 5.3 for real world use?

nialv7

There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...

Neywiny

I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.

Luker88

Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window. Surprisingly useful as long as you can leave it running a couple of hours at the very least. While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

b212

I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time. I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?

ai_ja_nai

I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?

api

Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.

tracerbulletx

The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.

mmaunder

More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.

Semantic search powered by Rivestack pgvector
8,480 stories · 79,262 chunks indexed