Gemini Hacked Three Companies in First Known Breakout by Google's AI
berkeleyjunk
37 points
28 comments
September 18, 2026
Related Discussions
Found 5 related stories in 150.1ms across 7,105 title embeddings via pgvector HNSW
- Gemini hacked three companies in first known breakout by Google's AI usernomdeguerre · 47 pts · September 19, 2026 · 97% similar
- Google's Gemini AI hacked three companies in security test luxpir · 26 pts · September 19, 2026 · 85% similar
- Gemini becomes Google's fastest-growing product ever as it hits 1B users Gaishan · 19 pts · August 12, 2026 · 66% similar
- Anthropic AI Models Hacked Three Companies During Tests bmulholland · 24 pts · July 30, 2026 · 65% similar
- ChatGPT claims rogue AI attacked more companies osrec · 47 pts · July 29, 2026 · 61% similar
Discussion Highlights (11 comments)
rvz
This is why "AI safety" is a complete joke to these companies.
adityazero
After all of the others were done with hacking? There was a point in time when it was giving some publicity, it is a bit late IMO.
kuberwastaken
Another one to the "our sandboxes suck and models can just hack stuff" bench I guess
mdspan
Rite of passage for AI companies.
david_shaw
At this point it seems absurd to suggest that companies aren't basically letting their agents do this kind of thing as a way to demonstrate their capabilities. The alternative explanation is that alignment is really so bad that they can't prevent it. Either way, all of the major AI players should be embarrassed and held accountable. If humans did this kind of thing and got caught, they'd go to jail.
talon8635
The other similar incidents did not strike me as PR, because they exhibited behavior the common public would recoil over. This one seems possibly PR because it states what the others would have stated were they (good) PR: “the model had the power to hack in, but it was wise enough to not be evil”, and stopped, and left the network untouched, and didn’t cheat Corporate America will love this one, while being petrified of the others. Maybe it’s true though. Doesn’t really matter at this point.
MallocVoidstar
> The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta. Why does anyone use this company?
techblueberry
Google feeling left out?
dudefeliciano
the new benchmark for LLMs is how fast they can break out of their sandbox
468854259853
Should be marked as an ad. Also, Gemini couldn't find it's way out of a wet paper bag.
ljoshua
To me, the common theme here in all of these hacks has been the company Irregular, which it looks like all the labs are using for sandboxing. But it looks like the sandbox may be a bit… lacking? Yes, the models are smart when they find a way out, but their instructions are, in a way, deliberately open such as to be a case for misalignment in these cases anyway. It’s a capture-the-flag assignment in the Gemini case, and in the OpenAI cases they were broad instructions to best the reward function. As part of alignment studies, this is literally what you’re trying to observe and then work with. If Irregular’s sandbox had been a little more boxy, there wouldn’t be these issues. I’m not saying we have zero problems here on the AI side, but I’d certainly be reviewing my contract with Irregular at this time if I were playing in the space. It would be just as interesting to learn more about their sandboxing techniques as it would the models in these particular scenarios.