Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

seelos 393 points 165 comments September 10, 2026
cognition.com · View on Hacker News

Discussion Highlights (20 comments)

mydreamof

Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?

_doctor_love

SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.

Tsarp

"SWE-2 is post-trained from Kimi K3"

scronkfinkle

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

monkeydust

As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft . There is an English translation somewhere.

postalcoder

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

TheJCDenton

> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

nullbio

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

ltsSmitty

Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at

llmslave

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!). I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around. Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

eyeris

Wonder if this was the model that drove factoring the rsa-260 The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.

bobtheborg

SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it. Looking forward to 2 -- maybe it'll be usable

bluelightning2k

I like Cognition as a company and hope they succeed. Seemingly excellent engineering org. I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.

pkilgore

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

gruez

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

m3kw9

I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed.

CyLith

I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.

wqash71

The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all.

Take8435

Post made by account 2 days ago.

alansaber

Fair enough that they did a "propoganda and censorship" eval but not sure why i'd care about that in my highly juiced SWE kimi FT.

Semantic search powered by Rivestack pgvector
6,164 stories · 56,060 chunks indexed