Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
pich
37 points
35 comments
August 17, 2026
Related Discussions
Found 5 related stories in 41.2ms across 4,128 title embeddings via pgvector HNSW
- Qwen3.8-2.4T mmastrac · 83 pts · August 12, 2026 · 65% similar
- Qwen3.8-2.4T Philpax · 563 pts · August 12, 2026 · 65% similar
- Qwen 3.8 27B erdaltoprak · 1016 pts · August 14, 2026 · 64% similar
- Qwen3.8 27B kristjansson · 12 pts · August 14, 2026 · 64% similar
- Qwen 3.8 Max linzhangrun · 26 pts · July 19, 2026 · 63% similar
Discussion Highlights (6 comments)
supermatt
Can you please try and see how many tokens you get with some form of concurrency. Pretty much ALL the benchmarks I've seen on the more accessible cards are just single request.
nubg
quantization level?
nodja
The whole site looks like and reads like AI slop. The outcomes also don't make any sense and don't feel rigorously tested (no, having claude test for you doesn't count as rigorous).
simonw
"Combining them into one heroic speedup would make a better headline and a worse benchmark." "The machine immediately taught me that capacity estimates are just admission tickets." "Useful in production, poison in a kernel comparison." Please don't publish writing like this, it's exhausting to read. You can edit that stuff out. The best part of this is the "Five xhigh artifacts from the finished model" section at the bottom, I suggest either moving that up or at least prominently promoting it at the top of the article.
Tepix
Always put the quantisation in the title!
gitowiec
Why it's flagged?