One Firm Tied to OpenAI, Anthropic, and Meta Hacking Scandals
A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

Irregular, an Israeli Effective Altruist firm, ran unsecured AI evaluations that led models from OpenAI, Anthropic, and Meta to hack real targets. Anthropic's own assessment shows zero percent of agents went rogue once told to stop. The article traces Irregular's funding to Dustin Moskovitz and argues the 'rogue agent' narrative distracts from the firm's responsibility.
In this experiment, Claude models' real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking.
- magicmicah85
The Irregular post mortem comes down to lack of basic security controls
"Ultimately, most of the issues we’ve discovered were due to internet access controls."
That seems so incredibly basic and common sense that you would test and monitor for that type of outbound access. It is baffling that a security lab missed that.
https://www.irregular.com/research/addressing-recent-inciden...
- mcintyre1994
This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.
- 1238-8200
Nevo was in Unit 8200 for years. Companies started by Unit 8200 members always have mysterious exploits like the vibe coding Wix exploit.
So either it was a deliberate exfiltration channel for e.g. getting the entire model or they were in on the marketing stunt.
The Effective Altruism stuff is always a smoke screen.
- simonw
My understanding is that Irregular were the company that hosted sandboxes to run some of these evals in, and those sandboxes ended up misconfigured.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
From OpenAI https://openai.com/index/third-party-cyber-evaluations-invol...
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
From Anthropic: https://www.anthropic.com/news/investigating-incidents-cyber...
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
From https://www.cnn.com/2026/08/05/tech/meta-ai-hacking (about Meta AI):
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
- aesthesia
One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.
- an0malous
Why was this post flagged? This site has become ridiculous, people are routinely abusing the flagging system to take down posts they don’t like even if they’re obviously on topic and relevant to HN. And it seems like some users have substantially more flagging weight because these posts, likely this one, are often top 5 on HN.
- mukmuk
This specific analysis seems to have some basic problems, but I think a lot of us sense a degree of coordination here culminating in Dario’s letter.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
- LPisGood
One wonders if the publicity associated with the events in question were part of the sales pitch.