14× faster embeddings: how we rebuilt the ONNX path in Manticore
snikolaev
13 points
0 comments
July 03, 2026
Related Discussions
Found 5 related stories in 144.2ms across 13,723 title embeddings via pgvector HNSW
- Executing programs inside transformers with exponentially faster inference u1hcw9nx · 17 pts · March 12, 2026 · 54% similar
- Manticore Search 27.1.5: Auth, sharding, conversational and faster vector search snikolaev · 34 pts · June 22, 2026 · 54% similar
- Making LLM Training Faster with Unsloth and NVIDIA segmenta · 114 pts · May 07, 2026 · 52% similar
- Real-time LLM Inference on Standard GPUs: 3k tokens/s per request NicoConstant · 202 pts · May 29, 2026 · 49% similar
- Accelerating Gemma 4: faster inference with multi-token prediction drafters amrrs · 521 pts · May 05, 2026 · 48% similar