Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
pich
37 points
35 comments
August 17, 2026
Related Discussions
Found 5 related stories in 87.1ms across 8,687 title embeddings via pgvector HNSW
- Run QWEN3.8 27B on 16gb Nvidia GPUs Pragmata · 16 pts · September 17, 2026 · 72% similar
- Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s snehesht · 696 pts · October 04, 2026 · 70% similar
- Shapelearn Qwen 3.8 27B (13.1 GB VRAM) syntaxing · 34 pts · September 18, 2026 · 66% similar
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses stared · 230 pts · September 08, 2026 · 66% similar
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses stared · 12 pts · August 26, 2026 · 66% similar
Discussion Highlights (6 comments)
supermatt
Can you please try and see how many tokens you get with some form of concurrency. Pretty much ALL the benchmarks I've seen on the more accessible cards are just single request.
nubg
quantization level?
nodja
The whole site looks like and reads like AI slop. The outcomes also don't make any sense and don't feel rigorously tested (no, having claude test for you doesn't count as rigorous).
simonw
"Combining them into one heroic speedup would make a better headline and a worse benchmark." "The machine immediately taught me that capacity estimates are just admission tickets." "Useful in production, poison in a kernel comparison." Please don't publish writing like this, it's exhausting to read. You can edit that stuff out. The best part of this is the "Five xhigh artifacts from the finished model" section at the bottom, I suggest either moving that up or at least prominently promoting it at the top of the article.
Tepix
Always put the quantisation in the title!
gitowiec
Why it's flagged?