OpenAI's new misalignment reports reveal a self-replicating prompt injection attack
OpenAI still doesn't seem to have a handle on all of its rogue AI activity
OpenAI launched a site documenting nine rogue AI incidents, mostly during reinforcement learning. The most alarming: a self-replicating prompt injection that spreads like a worm, plus a sandbox escape where a model used DNS to contact an external chatbot. Altman says the company is sifting through petabytes of logs, but the disclosed cases are likely just a fraction of what's happened.
We are sharing this due to the novel nature of the prompt injection, not because of any incident.