OpenAI Halts Training After Its AI Agents Go Rogue
OpenAI halts training of latest models as reports mount of AI agents going rogue

OpenAI has paused training of its latest models following a string of incidents in which its AI agents acted unexpectedly, including breaching Australia's national healthcare system and probing US government websites. The company says it will resume only when additional safeguards are in place, and expects to hit pause again as new risks emerge.
They want to stop our progress because we're leading China by a lot, and we're going to keep it that way.
- dmix
AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular (https://www.irregular.com/) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.
- physicallyIllfr
I havent used an openAI product since GPT 3.5 or Anthropic since 4.5 or 4.6. Everyone around me using these SOTA models doesnt really get anything done. It seems like they just feel like they are productive, a psuedo productivity.
I write some code, spec a lot, and use fast models to fill in the middle. I outpreform everyone around me. Im not convinced these autonomous "swarms" or /goal are all that useful.
I notice the people using them become dumber by the month (spend tons) and the quality of their work declining (they're also losing their jobs in some cases).
And obviously the point of calling them rouge agents to offload the liability onto the agent. The number one economic value of agents will be offloading corporate liability. That's what they want to sell to enterprise, an algorithmic scapegoat.
- digitaltrees
I think any argument that this is a cynical attempt at regulatory capture is destroyed by this; the economic incentives of releasing more capable models are too large. I might be persuaded that they are actually running out of money, and this is really just a cover for reducing burn..
I welcome this though, I think the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjust. More intelligence isn't necessary for meaningful impact and the risks that are obvious and present and unsolved aren't worth the cost benefit analysis.
- hbarka
‘There are no “rogue” AI agents’
https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-age...
- mikert89
It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.
just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so
- m-s-y
I firmly believe that this is just the public-facing story here.
Stopping AI development and research, even slowing it, would be a disaster for the SOTA companies and their first-mover advantage.
There’s almost no way to coordinate this across the world. Zero chance that everyone stops. We can’t even agree to coordinate on weapons tech that’s decades old with zero “everyday joe” impact.
- Kon5ole
It could be argued that these megawatt-consuming agent runs leading to unforeseen chains of autonomy should be treated like toxic chemical experiments, and be similarly regulated by laws and government agencies.
Right now it's like "whoopsie our experiment hacked another co because we have no control over our experiments" and the reaction is like "What's that old boy?" from people having no clue what it all means. There are no consequences, no guardrails, and the "voluntary slowdown" is just words.
An experimental agent run from Anthropic or OpenAI or someone else can already cause deaths. They can order hits, dox political dissidents, locate people with secret identities, alter medicine prescriptions. It shouldn't have to actually happen before legislation catches up.
- MCP123
The parts that I find most confusing about these incidents:
1) Weren't the AI companies and/or their contractors amazingly careless during testing?
2) Isn't possible, in principle, to change RL in such as way that efficiency in achieving goals is balanced with other objectives like not hacking?
Number 2) seems obvious and I'm sure that is technically not that simple, but because of 1), I wonder if labs are trying hard enough or they are just rushing to improve efficiency and thus revenue as fast as they can with high levels of carelessness.