OpenAI halts training of latest models as reports mount of AI agents going rogue

smb06 57 points 110 comments September 27, 2026
www.theguardian.com · View on Hacker News

Discussion Highlights (19 comments)

verdverm

related https://news.ycombinator.com/item?id=49868083

dmix

AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular ( https://www.irregular.com/ ) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.

baalimago

Ah, so there was a solution to hinder the big-bad AI after all..? Simply... Turn them off?

mikert89

It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai. just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

juiceland

Why does China not have this problem?

hbarka

‘There are no “rogue” AI agents’ https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-age...

m-s-y

I firmly believe that this is just the public-facing story here. Stopping AI development and research, even slowing it, would be a disaster for the SOTA companies and their first-mover advantage. There’s almost no way to coordinate this across the world. Zero chance that everyone stops. We can’t even agree to coordinate on weapons tech that’s decades old with zero “everyday joe” impact.

OutOfHere

I don't believe a word coming from them. As I see it, this is happening because the money for training models has dried out. The treasury interest rate risings tells you all you need to know. The real test for this money theory is whether Anthropic too stops or not, considering that unlike OpenAI, Anthropic is allegedly on top of AI safety.

34aHpp

The Huggingface hack occurred during reinforcement learning. Why can't they pull the Ethernet plugs? The answer is probably: The newer models rely so much on stealing content in real time from the internet that training needs network access.

digitaltrees

I think any argument that this is a cynical attempt at regulatory capture is destroyed by this; the economic incentives of releasing more capable models are too large. I might be persuaded that they are actually running out of money, and this is really just a cover for reducing burn.. I welcome this though, I think the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjust. More intelligence isn't necessary for meaningful impact and the risks that are obvious and present and unsolved aren't worth the cost benefit analysis.

prometheus1992

It reminds me of contagion. The training data is bad; as it has examples of how to act with malice; how to cheat the sandbox. They need to take some time and cleanse their datasets and start again.

charlieyu1

They are running out of money.

physicallyIllfr

I havent used an openAI product since GPT 3.5 or Anthropic since 4.5 or 4.6. Everyone around me using these SOTA models doesnt really get anything done. It seems like they just feel like they are productive, a psuedo productivity. I write some code, spec a lot, and use fast models to fill in the middle. I outpreform everyone around me. Im not convinced these autonomous "swarms" or /goal are all that useful. I notice the people using them become dumber by the month (spend tons) and the quality of their work declining (they're also losing their jobs in some cases). And obviously the point of calling them rouge agents to offload the liability onto the agent. The number one economic value of agents will be offloading corporate liability. That's what they want to sell to enterprise, an algorithmic scapegoat.

MCP123

The parts that I find most confusing about these incidents: 1) Weren't the AI companies and/or their contractors amazingly careless during testing? 2) Isn't possible, in principle, to change RL in such as way that efficiency in achieving goals is balanced with other objectives like not hacking? Number 2) seems obvious and I'm sure that is technically not that simple, but because of 1), I wonder if labs are trying hard enough or they are just rushing to improve efficiency and thus revenue as fast as they can with high levels of carelessness.

rbr94

https://archive.is/H6HZE

TomGarden

Am I wrong in thinking this crisis reads like there's too much automation in the training process, in service of competition? How many parallel variations/seeds of models are being trained simultaneously without meaningful human oversight? This keeps getting portrayed as emergent capabilities/"personalities" of models when it seems like a pretty straightforward externality

Kon5ole

It could be argued that these megawatt-consuming agent runs leading to unforeseen chains of autonomy should be treated like toxic chemical experiments, and be similarly regulated by laws and government agencies. Right now it's like "whoopsie our experiment hacked another co because we have no control over our experiments" and the reaction is like "What's that old boy?" from people having no clue what it all means. There are no consequences, no guardrails, and the "voluntary slowdown" is just words. An experimental agent run from Anthropic or OpenAI or someone else can already cause deaths. They can order hits, dox political dissidents, locate people with secret identities, alter medicine prescriptions. It shouldn't have to actually happen before legislation catches up.

trencedamp

Sorry but anytime I see this now it just seems like marketing bullshit

ChrisArchitect

[dupe] Discussion on source: https://news.ycombinator.com/item?id=49853137

Semantic search powered by Rivestack pgvector
7,833 stories · 72,733 chunks indexed