Run QWEN3.8 27B on 16gb Nvidia GPUs

Pragmata 16 points 4 comments September 17, 2026
github.com · View on Hacker News

Discussion Highlights (1 comments)

Pragmata

I've been using it for the past few days, and it runs really well! I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second. Genuinely very usable, and fully local!

Semantic search powered by Rivestack pgvector
7,105 stories · 65,136 chunks indexed