Show HN: UL-SMF – Open-source linear-complexity ~300x KV-cache compression
liventruth
11 points
1 comment
August 17, 2026
Related Discussions
Found 5 related stories in 92.0ms across 8,795 title embeddings via pgvector HNSW
- DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression mfiguiere · 94 pts · September 17, 2026 · 62% similar
- Show HN: Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload arnav__1 · 21 pts · July 26, 2026 · 59% similar
- Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) mikeayles · 12 pts · August 10, 2026 · 57% similar
- Cache-to-Cache: Direct Semantic Communication Between LLMs (2025) rochansinha · 78 pts · September 18, 2026 · 57% similar
- Show HN: We Beat MLPerf: Modern Storage for KV Offload and LLM Training arnav__1 · 35 pts · September 05, 2026 · 54% similar
Discussion Highlights (1 comments)
colingauvin
This should be trivially demonstrable if it actually works. Cosine similarity is not a good metric, however, for attention compression.