Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
cdnsteve
60 points
27 comments
September 10, 2026
Related Discussions
Found 5 related stories in 64.9ms across 6,164 title embeddings via pgvector HNSW
- Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra seelos · 393 pts · September 10, 2026 · 69% similar
- Qwen3.8 27B scores 52 on Artificial Analysis anana_ · 329 pts · August 17, 2026 · 53% similar
- GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index wertyk · 21 pts · September 03, 2026 · 50% similar
- German AI consortium releases Soofi S, an open 30B model yogthos · 12 pts · July 15, 2026 · 49% similar
- German AI consortium releases Soofi S, an open 30B model that tops benchmarks amai · 129 pts · July 16, 2026 · 49% similar
Discussion Highlights (6 comments)
varispeed
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
forgot-my-pw
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso. Still quite impressive though.
walrus01
Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system.
ChrisArchitect
Related: Cognition launches new SWE-2 model https://news.ycombinator.com/item?id=49645443
captainregex
my personal experience with swe has been…suboptimal. I am not sure how much I buy these benchmarks and it has a very “just blurt it out even if it’s probably not right” style but hey it’s free.
samusiam
Which is pretty much a useless (i.e., saturated, contaminated) benchmark now.