Benchmarking retrieval for agents on messy real-world company knowledge
emil_sorensen
25 points
3 comments
October 02, 2026
Related Discussions
Found 5 related stories in 85.9ms across 8,345 title embeddings via pgvector HNSW
- Kimi K3: second only to Fable 5 on AA-Briefcase wertyk · 28 pts · July 22, 2026 · 58% similar
- Terminal-Bench-Science: Evaluating AI agents on scientific research workflows matt_d · 61 pts · August 28, 2026 · 58% similar
- Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases theanonymousone · 193 pts · September 12, 2026 · 57% similar
- Show HN: Benchmark your eng team's AI agent maturity in 5 minutes adamgold7 · 13 pts · July 14, 2026 · 55% similar
- Auto-autoresearch: self-improving agents on Karpathy's NanoChat benchmark hyperparticle · 11 pts · September 15, 2026 · 52% similar
Discussion Highlights (2 comments)
emil_sorensen
OP/founder of kapa (YCS23) here. Happy to answer any questions :)
75ziadilrodilax
enzinho bravo