Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
seelos
393 points
165 comments
September 10, 2026
Related Discussions
Found 5 related stories in 74.7ms across 6,164 title embeddings via pgvector HNSW
- Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1 cdnsteve · 60 pts · September 10, 2026 · 69% similar
- Cerebras CS-4 sunils34 · 177 pts · August 19, 2026 · 55% similar
- Advancing the price-performance frontier with GPT‑5.6 tedsanders · 541 pts · July 30, 2026 · 55% similar
- GPT-6 Astra kibae · 1582 pts · September 03, 2026 · 54% similar
- Cognition (Devin) raises $2B at $48B valuation meco · 21 pts · September 08, 2026 · 53% similar
Discussion Highlights (20 comments)
mydreamof
Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
_doctor_love
SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.
Tsarp
"SWE-2 is post-trained from Kimi K3"
scronkfinkle
Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
monkeydust
As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft . There is an English translation somewhere.
postalcoder
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
TheJCDenton
> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
nullbio
Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.
ltsSmitty
Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at
llmslave
At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!). I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around. Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job
eyeris
Wonder if this was the model that drove factoring the rsa-260 The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.
bobtheborg
SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it. Looking forward to 2 -- maybe it'll be usable
bluelightning2k
I like Cognition as a company and hope they succeed. Seemingly excellent engineering org. I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.
pkilgore
Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
gruez
Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.
m3kw9
I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed.
CyLith
I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.
wqash71
The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all.
Take8435
Post made by account 2 days ago.
alansaber
Fair enough that they did a "propoganda and censorship" eval but not sure why i'd care about that in my highly juiced SWE kimi FT.