Kimi K3: second only to Fable 5 on AA-Briefcase
wertyk
28 points
2 comments
July 22, 2026
Related Discussions
Found 5 related stories in 281.1ms across 14,369 title embeddings via pgvector HNSW
- Kimi K3 Intelligence, Performance and Price Analysis theanonymousone · 51 pts · July 16, 2026 · 74% similar
- Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA piotrgrabowski · 478 pts · July 21, 2026 · 73% similar
- Kimi K3, and what we can still learn from the pelican benchmark droidjj · 305 pts · July 17, 2026 · 70% similar
- Kimi K3 may be an important inflection point for AI jger15 · 16 pts · July 17, 2026 · 69% similar
- Kimi K3: Open Frontier Intelligence mfiguiere · 30 pts · July 16, 2026 · 66% similar
Discussion Highlights (2 comments)
eugene3306
What harness do they use for testing? Back in 2025 it was common to test models in a different harnesses. I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc. Why did it come out of fashion ? EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
threatripper
Can we assume that the test is still private when it was run on many cloud providers?