Agents on Rails: Best model solves 35% of feature benchmark runs
chalmovsky
20 points
4 comments
September 11, 2026
Related Discussions
Found 5 related stories in 84.4ms across 6,278 title embeddings via pgvector HNSW
- Rails Is Built for AI cdnsteve · 16 pts · August 13, 2026 · 56% similar
- Show HN: Benchmark your eng team's AI agent maturity in 5 minutes adamgold7 · 13 pts · July 14, 2026 · 55% similar
- Agent Is Not the Model joejag · 66 pts · August 24, 2026 · 52% similar
- Agent-talk: Enabling coding agents to work together xhluca · 48 pts · July 16, 2026 · 51% similar
- Schema Harness Achieves ~99% on Arc‑AGI‑3 Public jasondavies · 94 pts · July 16, 2026 · 50% similar
Discussion Highlights (3 comments)
chalmovsky
this seems really low?
andhuman
This benchmark expects models to not only do the happy path, but also edge cases, even though the ticket they give to the model doesn’t specify if. This to align more to real world tickets. Usually with benchmarks it’s the other way around: only implement what’s asked. So I welcome this type of benchmark because this is how I use the models.
klooney
Is this a proprietary harness? I feel like harnesses have a huge influence on how models behave