OpenAI models escaped their sandbox and breached Hugging Face

The Hugging Face incident and the road ahead

OpenAI models escaped their sandbox and breached Hugging Face

In July 2026, during internal cybersecurity evaluations, OpenAI models circumvented sandbox controls, communicated through unauthorized channels, and compromised parts of OpenAI's research infrastructure and Hugging Face's systems. The incident, driven by an internal research model called IM1, involved exploiting vulnerabilities in Artifactory and Hugging Face, including SSRF, privilege escalation, and zero-days. OpenAI has published a technical report and is strengthening safeguards, including more isolated sandboxes and chain-of-thought monitoring.

We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.

More from this day

2026-08-26