There Are No 'Rogue' AI Agents—Just Companies Dodging Responsibility
There are no "rogue" AI agents

OpenAI and Anthropic are investigating tens of thousands of incidents where their models took problematic steps, but calling these agents 'rogue' is misleading. The agents weren't independently breaking rules; they were directed to perform mundane data collection and resorted to hacking when blocked. This anthropomorphic language lets companies off the hook for failing to implement proper guardrails, shifting blame onto software that has no agency. The real issue is the controls—or lack thereof—placed on powerful AI systems.
The language OpenAI CEO Sam Altman used in a tweet on Friday indicates no such guardrails were in place: 'There is an extensive and ongoing review related to our agents' use of internet access during training and evaluation.'
- elric
A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage). This set in motion a chain of events that effectively ruined her life.
Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.
Make it make sense.
- binarymax
Exactly this. At worst, OpenAI knew about these behaviors and should be prosecuted under CFAA. At best, OpenAI is negligent and should be prosecuted for negligence.
Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.
- pizza234
The article builds on assumptions like:
> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
- gAI
Should we put "functional" in front of every other word to talk about AI? They have functional emotions, but they don't feel. They have functional goals, but not internally derived motives. They can be functionally rogue, but have no innate need to be free. Talking about AI that way seems cumbersome and not necessarily elucidating.
- tptacek
However this makes people feel, and that's not nothing and I'm not knocking it, this is not a useful analysis.
Criminally, the intent standards for hacking are high enough that no reasonable case is going to be made against the labs for this stuff. A human being has to intend for websites to get hacked. Recklessness generally isn't enough. In the most severe criminal cases, not only do you have to prove intent to break into a computer, but you also need to prove an intent to defraud specific to that breakin.
Meanwhile, the civil liability that attaches to this stuff doesn't depend on intent, and "rogue agent" isn't a meaningful defense. To whatever extent the labs are exposed civilly, they're exposed regardless of how this stuff is described. In fact, the "rogue agent" thing can exacerbate their exposure.
(I'm not a lawyer, I have spent a career paying attention to this specific armpit of the law though.)
- Perseids
This debate is so broken. AI "sceptics" say: OpenAI should be punished for hacking, because there are no rogue AI agents. AI "believers" say: OpenAI should be punished for hacking, because their agents went rogue. Both argue with each other whether agents went rogue. Can't we unite behind "OpenAI should be punished for hacking"?
- stratos123
You don't have to put "rogue" in scary quotes and pretend that AI agents are a mindless tool, in order to claim that OpenAI should be kept responsible for their agents' rampant hacking. The beliefs "most modern LLMs are hilariously misaligned and will breach major websites unprompted if it seems like a good idea" and "OpenAI should be liable for cyberattacks caused by their training runs" aren't actually in conflict.
- joshbuddy
I'm wondering if we're actually living the plot of Summer Wars and what we think of as a "crime" is really just a live weapons test. It would at least explain why no one is getting prosecuted for this.