DeepSeek v4.1 Flash Is Now Our Best Hacking Model

talhof8 162 points 66 comments September 16, 2026
enclave.ai · View on Hacker News

Discussion Highlights (12 comments)

fwip

Edit: Comment deleted - no longer useful.

TuxSH

I find this - or perhaps the title - a bit surprising. I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min. Perhaps DS works better where targets have low-hanging fruits than can be found fast?

nickysielicki

When the history books are written and all is said and done, the hubris of this moment where all the American labs decided to punk their investors and join hand in hand in agreeing to let the Chinese win forever is going to be the main story.

jrflo

Seems pretty bold to claim deepseek is the "best hacking model" while providing zero comparisons to other models...

wg0

DeepSeek is underrated. Basically all Chinese models are good enough for day to day coding at this point. The 2 trillion dollar ROI on anthropic alone? Good luck with that.

Art9681

As opposed to what? Is enclave.ai signed up for GPT Cyber or Glasswing?

habosa

DeepSeek models have such good benchmark performance, amazing pricing, and the team over there seems to be widely considered impressive. I just haven't found them to be very good? I've had a ton more success with the GLM models (since 5.2 anyway). Maybe I'm just holding it wrong, DS models seem to get stuck in loops or tell me nonsense. GLM feels like budget Claude.

fra

Interesting result, but the writing is very poor.

gertlabs

We ran v4.1 Flash through our evaluations and found it to be smarter and faster than V4 Flash, with a commensurate price bump. Some notes: - Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice). - Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in agentic coding, at lower cost. - Chinese models have always been strong iterators in an agentic harness. This model is no different, reaching an average percentile ~20% higher when given a harness vs a one-shot solution. That one-shot fluid intelligence is what makes a model feel smart, though, and typically results in fewer attempts/tokens to solve a problem, and American frontier models are still far ahead in that department. The new architecture is interesting. It puts pricing between their old Flash and Pro lineups, suggesting they might be abandoning their super-cheap flash models (which weren't that fast due to heavy reasoning) and their pro models (which sort of flopped and weren't consistently better than their flash models, despite the size/cost) and shipping a strong intermediate that competes with the Gemini Flash series. Data at https://gertlabs.com/rankings

pelzatessa

How do I make deepseek "hack" my source code? do I just start my coding agent in my directory and command it to "find vulnerabilities", or is there some more sophisticated software to do that?

hactually

we recently got this running in 192gb of vRAM and using it with the Klaudia harness has been incredible for driving out work that we'd need to use Opus and Fable for previously

Grimblewald

I'm a deepseek fanboy, but has anyone else found flash to not meet expectations? I've found it to be wildly bad at doing as asked, overengineering, and always assuming instead of reading even if told to read things in full before doing anything. It makes wildly silly mistakes in code and so far has been quite frustrating to work with. Maybe its just what i work on that its particularly bad at but in general feels like a strict regression on previous offerings.

Semantic search powered by Rivestack pgvector
6,833 stories · 62,541 chunks indexed