I'm not paying $20 for ChatGPT or Claude because a free local LLM does
hsnewman
43 points
17 comments
October 07, 2026
Related Discussions
Found 5 related stories in 157.3ms across 8,795 title embeddings via pgvector HNSW
- Students prefer Gemini over ChatGPT and Claude for AI essays in blind tests pasharayan · 20 pts · August 26, 2026 · 54% similar
- Fewer Americans Pay to Use LLMs Than Still Pay to Play World of Warcraft SLHamlet · 26 pts · August 27, 2026 · 50% similar
- Understanding ChatGPT Work gmays · 114 pts · August 31, 2026 · 50% similar
- Study: Claude, ChatGPT Offer Different Shopping Prices Based on Wealth sbulaev · 95 pts · October 07, 2026 · 49% similar
- ChatGPT ad targeting is garbage hermitcrab · 36 pts · September 02, 2026 · 48% similar
Discussion Highlights (11 comments)
kittikitti
This is a great article! Thank you for sharing. I also recommend opencode that has built in features for local model hosting.
braggerxyz
20$/month vs. 1200$ gpu. That's a lot of months you can pay the subscription
WarmWash
"I spent $1100 on a graphics card so I can save $20/mo running a mid/low intelligence model at 33tk/s with 16k context"
mrandish
Thanks for writing and sharing. Since I also have a 4070 Ti Super, I'm always interested in hearing people's experiences with local models that fit 'middle-ground' hardware. I kind of feel left out because most posts I see seem to either be about clever ways of making older, lower-end cards usable, leveraging mega-CPU RAM (128GB) or how aweseome high-end GPUs like 4090/5090 can be.
kristianp
> llama.cpp serves it over localhost at speeds that stop being a complaint after a few minutes of use. "That stop being a complaint"? That's a strange construction. I wonder if qwen 4 will make these smaller cards more viable by allowing the ngram storage to be hosted on CPU RAM.
reddit_clone
Ok. What can I do with an M4 MBP, with 48G memory?
smcleod
If you were only paying $20 for LLMs previously then you were getting very low usage / little done.
mongrelion
I have gone down this route but I'm afraid that 16K of context is nothing for agentic coding tasks. Even if you use a smarter model directly from an inference provider to plan all the work and then use your own 16GB GPU to execute the plan, managing those 16k tokens becomes unusable at some point even with a minimalistic harness like pi. For other agentic tasks it's nice, and more than just the budget, it's the data sovereignty that you gain (imagine analysing tax forms, contracts, other legal documents, etc.) in my opinion.
wafflemaker
It's two different things to be honest. Local model can be a local assistant, take care of your todos, calendar, private stuff in general. Things you shouldn't share with OpenAI or other Big Tech. Big Tech models for anything else.
MisterMunchkin
I use the free versions on mobile, and openrouter on desktop. It’s way cheaper than paying a subscription, and I can always use the latest models.
gentile
OP could be doing a lot better with a MoE model (uses some RAM as well). I'm doing Qwen 3.6 35BA3B-NVFP4 (a larger ~20GB model) in 8GB VRAM and 20 GB RAM (I also get more than 16k context, and more t/s).