Show HN: UL-SMF – Open-source linear-complexity ~300x KV-cache compression
liventruth
11 points
1 comment
August 17, 2026
Related Discussions
Found 5 related stories in 44.4ms across 4,128 title embeddings via pgvector HNSW
- Show HN: Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload arnav__1 · 21 pts · July 26, 2026 · 59% similar
- Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) mikeayles · 12 pts · August 10, 2026 · 57% similar
- Show HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios) nonadhocproblem · 140 pts · July 15, 2026 · 54% similar
- Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5 pcbmaker20 · 58 pts · August 13, 2026 · 54% similar
- Show HN: Reame – a CPU inference server that gets faster as it runs targetbridge · 48 pts · July 11, 2026 · 51% similar
Discussion Highlights (1 comments)
colingauvin
This should be trivially demonstrable if it actually works. Cosine similarity is not a good metric, however, for attention compression.