Show HN: Shoehorn – Quantize any model down to run on your machine

rhgraysonii 40 points 7 comments August 18, 2026
notactuallytreyanastasio.github.io · View on Hacker News

Working on Mac, Linux, and Windows now. I include a simple GUI to find new models and get things built and set up. It is working quite well across a few models for me. The GitHub README and DESIGN.md files go into detail of the how/why and it's working remarkably well so far. https://github.com/notactuallytreyanastasio/shoehorn

Discussion Highlights (3 comments)

mbuchel-hn

does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?

hmokiguess

Reminds me of https://github.com/AlexsJones/llmfit

jaylane

tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

Semantic search powered by Rivestack pgvector
4,128 stories · 37,281 chunks indexed