Run QWEN3.8 27B on 16gb Nvidia GPUs
Pragmata
16 points
4 comments
September 17, 2026
Related Discussions
Found 5 related stories in 68.2ms across 7,105 title embeddings via pgvector HNSW
- Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU pich · 37 pts · August 17, 2026 · 72% similar
- Qwen 3.8 Max linzhangrun · 26 pts · July 19, 2026 · 71% similar
- Qwen3.8-Max Bluestein · 25 pts · September 02, 2026 · 68% similar
- Qwen3.8-2.4T mmastrac · 83 pts · August 12, 2026 · 68% similar
- Qwen3.8-2.4T Philpax · 563 pts · August 12, 2026 · 67% similar
Discussion Highlights (1 comments)
Pragmata
I've been using it for the past few days, and it runs really well! I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second. Genuinely very usable, and fully local!