Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
JonSchneider
344 points
113 comments
September 17, 2026
Related Discussions
Found 5 related stories in 69.7ms across 7,105 title embeddings via pgvector HNSW
- Lossless model compression experiment: GLM-5.2 in 25% less memory hambandit · 16 pts · July 20, 2026 · 62% similar
- Bonsai 27B: A 27B-Class model that runs on a phone xenova · 522 pts · July 14, 2026 · 61% similar
- Show HN: UL-SMF – Open-source linear-complexity ~300x KV-cache compression liventruth · 11 pts · August 17, 2026 · 50% similar
- DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression mfiguiere · 94 pts · September 17, 2026 · 49% similar
- Show HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios) nonadhocproblem · 140 pts · July 15, 2026 · 49% similar
Discussion Highlights (20 comments)
kamranjon
Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!
abraxas
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
Aurornis
These are small enough that you can run them entirely in the browser https://huggingface.co/spaces/webml-community/ternary-bonsai... Remember to clear the downloaded weights afterward. Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.
z2
I'd love to see a Bonsai model start with a 100B+ parameter model and get that down to <30 GB. But maybe at that point we call it Topiary?
JonSchneider
I'm hoping they release an 8B v2 based on the Qwen 3.8 series in the near future - that would give us a really powerful model that could be run directly on users phones.
simonw
If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-... This should work: cd /tmp # Get the Prism macOS runtime curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz tar -xzf bonsai-runtime.tar.gz # Get the ~5.95 GB GGUF model: curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf # Run the server, I used port 8331 ./llama-prism-b10685-7dffb15/llama-server \ -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \ --port 8331 -ngl 99 -fa on -c 32768 Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this: uvx llm openai endpoint http://127.0.0.1:8331/v1 \ --model bonsai-2-27b --responses hi That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".
miffy900
I really wish people would stop saying N times smaller than something when making a comparison; that makes no sense - it's 1/9th (11.11%) the size. You don't get a smaller quantity by multiplying by a number greater than 1.0. You could instead reverse the subjects being compared - "the original model is 9x bigger than this new smaller, efficient model" or some such. That makes sense. I keep seeing this being used when people talk about efficiency or performance gains and it's just very unintuitive language.
2001zhaozhao
I think if they made this for Qwen3.8-Next it could fit in a single 5090?
danbrooks
Nice! Does anyone know how this compares to the Unsloth quantizations of this model? https://unsloth.ai/docs/models/qwen3.8#run-qwen3.8-guide
adrian17
> Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better? https://news.ycombinator.com/item?id=49611128
logicallee
(In case anyone remembers the compression post from yesterday[1], I checked and this one doesn't qualify for further compression - it's not zero-biased at all.) [1] https://news.ycombinator.com/item?id=49732931
Havoc
Cautiously optimistic. The V1 was noticeably weak on world knowledge but here the 3.8 base model is geared more towards reasoning than world knowledge anyway so might not matter as much
jedbrooke
Running at about 7-8 tok/s (~60 tok/s prefill) on a Mac Mini M2 16GB. So far feels smarter than Bonsai 1 27B, it’s slightly larger than the Q1_0 quant. Super exciting stuff :)
flutetornado
GPT Astra did some benchmarking on the DGX Spark. Speed: 34.38 tokens/sec for generation. Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens. Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth.
circularfoyers
I wonder how their talks with Apple went. Having this run on the TPU opposed to just the GPU, which drains a significant amount of battery life by comparison, is what I'm really interested in.
cmrdporcupine
What I'd love to see is this done for DS4.1 Flash. That would bring it down to the point where it can fit in 128GB on things like the Spark or Strix Halo.
hedora
RAM requirements? My current rule of thumb is “a byte per parameter”, but I doubt this runs in 1/9th that (~ 3GiB). Also, perf speedup?
g023
They need to make a Big Bonsai, something at the enterprise levels that can compete with DSV4 Flash etc.
blactuary
What is never totally clear with a lot of these releases is the scope of what it's good at. Models that can run with good speed on affordable consumer hardware for coding only is the dream. I am never going to use this for writing, images, or "general knowledge". Coding only
redox99
I tried their WebGPU version and it immediately started looping. Yeah "near lossless" my ass. Plus the reasoning that it looped on was clearly wrong and unlike the non quantized 27B