Sub-1-Bit LLM Compression via Latent Factorization
brainless
81 points
23 comments
October 08, 2026
Related Discussions
Found 5 related stories in 102.1ms across 8,906 title embeddings via pgvector HNSW
- Breaking the 1.58-bit Barrier for Ternary LLMs matt_d · 177 pts · September 16, 2026 · 61% similar
- Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint JonSchneider · 344 pts · September 17, 2026 · 53% similar
- Show HN: UL-SMF – Open-source linear-complexity ~300x KV-cache compression liventruth · 11 pts · August 17, 2026 · 53% similar
- Lossless model compression experiment: GLM-5.2 in 25% less memory hambandit · 16 pts · July 20, 2026 · 52% similar
- The smallest edge AI device for local LLMs alex-moon · 21 pts · September 07, 2026 · 50% similar
Discussion Highlights (8 comments)
badatnames
Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means
nico
Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?
big-chungus4
Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory
nbutton762
Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!
augment_me
Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget. I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance
bArray
Has anybody tested this? Are there any available computed models to test?
cpldcpu
I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities². It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties. ²As to why, I have seen few explanations. But the empirical evidence is there.
hgoel
I understand that the interest in these extreme quants stems from wanting to maximize the capability the average user can get from a local LLM in this era of ludicrously expensive memory, but I wonder if this is maybe targeting the wrong axis? We've been seeing various optimizations towards streaming, that have been much more impactful in the local AI space, e.g. MoE models where the less busy experts are offloaded to slow RAM or even pruned entirely, engram tables that can be read from NVMe instead of sitting around in RAM etc. Maybe the trick with these extreme quants would be to increase the total parameter count while quanting individual weights, such that maybe the active parameter count comes down, or streaming weights from RAM or disk becomes more efficient, or cache behavior improves? Say, replacing a single 4bit/weight matmul with 3 1bit/weight operations that produce a much closer result than a single 1bit/weight matmul would.