Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU

pich 37 points 35 comments August 17, 2026
piszczek.pl · View on Hacker News

Discussion Highlights (6 comments)

supermatt

Can you please try and see how many tokens you get with some form of concurrency. Pretty much ALL the benchmarks I've seen on the more accessible cards are just single request.

nubg

quantization level?

nodja

The whole site looks like and reads like AI slop. The outcomes also don't make any sense and don't feel rigorously tested (no, having claude test for you doesn't count as rigorous).

simonw

"Combining them into one heroic speedup would make a better headline and a worse benchmark." "The machine immediately taught me that capacity estimates are just admission tickets." "Useful in production, poison in a kernel comparison." Please don't publish writing like this, it's exhausting to read. You can edit that stuff out. The best part of this is the "Five xhigh artifacts from the finished model" section at the bottom, I suggest either moving that up or at least prominently promoting it at the top of the article.

Tepix

Always put the quantisation in the title!

gitowiec

Why it's flagged?

Semantic search powered by Rivestack pgvector
4,128 stories · 37,281 chunks indexed