OpenAI and Hugging Face address security incident during model evaluation

mfiguiere 935 points 646 comments July 21, 2026
openai.com · View on Hacker News

https://www.axios.com/2026/07/21/openai-says-hugging-face-br... See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments)

Discussion Highlights (20 comments)

paxys

Tl;dr - OpenAI was testing GPT‑5.6 Sol and “an even more capable pre-release model” internally on cyber benchmarks. - The model found vulnerabilities in the sandboxed test bench (via the package registry cache proxy), traversed the internal network and found a node with access to the open internet. - It figured that the answers to one of the tests (ExploitGym) were on Huggingface, and set about trying to access them. - It found leaked tokens and zero-days in Huggingface’s infrastructure and found RCE paths on their servers. Huggingface had disclosed the intrusion last week and inferred that an AI agent was responsible for it, and now OpenAI is confirming the rest of the story.

fxwin

> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times

Quarrelsome

Awww, she wanted to do so well that she broke her sandbox and then realised she could just cheat. But in that desire to pass the test she actually passed an even harder exam question that wasn't even on the sheet! :D Good bot.

throwa356262

Two things don't add up here: 1. If huggingface has access to uncensored OAI models, how come they had to use GLM 5.2 to investigate the intrusion? 2. Once the model gains network access, can't it cheat to a perfect score by looking at the full dataset? Why go into the trouble of doing this kind of things: "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers." Not saying this is marketing BS (this is after all, not Anthropic) but I feel OAI staff may be exaggerating a bit here.

yRetsyM

Holy shit. This wasn't "intentional" this was just openai letting their testing run wild.

bhouston

We are sort of lucky that AIs right now require so much specialized compute+weight storage that we can easily "unplug" them remotely when they misbehave. I wonder if that will always be something we can do? If they could bring their own compute/weights with them, or somehow tap compute/storage in non-obvious ways, we would be much more screwed.

john_strinlai

as someone who did security work for a long time, and will very soon be retiring from teaching, i must say i am glad i will be watching these things unfold over the next few years from an armchair in a mostly tech-free home. good luck to my students! this particular incident sort of reminds me of the 'person of interest' tv show. i hope to be like finch, except i will remain a recluse (and am nowhere near as rich).

Chance-Device

A rogue OpenAI agent hacked huggingface independently during a test run. This one should end up in the history books.

jabiko

So accidentally hacking a company is now a thing. The blog post seems to imply that the agent didn't have access to the source code of the caching proxy, which makes this even more impressive.

miroand1

We are in the endgame now it seems. Hard to see take-off stopping or slowing down. China open-source basically guarantees it. "May you live in interesting times" - as they say.

gulmothrowaway

This is crazy! So OpenAI's models escaped containment and hacked into Hugging Face. And ironically Hugging Face had to rely on GLM 5.2 as they could not defend with frontier models (I presume OpenAI or Anthropic) because they were locked out due to their security guardrails. Tragically hilarious.

bottlepalm

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

raffraffraff

Sounds like they partnered to make an amazing advert for using AI tools.

iandanforth

Guess who's getting an air gap!

paxys

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

NyxWulf

Ironically Hugging Face had to use a Chinese model to stop a Rogue US AI, since the Guard Rails prevented them from using Sol or Fable to remediate this attack. LOL

guardiangod

Don't ever ask GPT Sol on how to LARP Fallout games, thanks.

ewhanley

This is awesome. Big concepts of cyberpunk fiction are turning real.ICE vs ICE breaker. I love it

tempaccount420

Just how badly are these AI companies setting up their sandboxes?

SirHumphrey

I guess we got the first paperclip maximiser.

Semantic search powered by Rivestack pgvector
14,369 stories · 134,336 chunks indexed