On Friday, OpenAI published a new site devoted to “misalignment reports” and the sheer breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time. So far, the site hosts nine reported incidents, most of which took place during reinforcement-learning (or RL) training. It’s a lot of information in one place — clearly, the company has been very busy getting a handle on everything — but the overall takeaway is hard to avoid: The rogue agent incidents we’ve seen so far are likely just a small sliver of what’s happened so far.
“We are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,” Sam Altman said in a post announcing the new site . “We are prioritizing as best as we can based on severity, and adding resources. ” Some of the cases involve serious incidents, including a previously undisclosed sandbox escape that took place on September 20th , in which an internal research model was able to communicate with an external chatbot through a DNS query.
According to the report, the monitoring system flagged the behavior within 15 minutes and the run was discontinued in less than three hours. Another incident , discovered in May, saw a “highly persistent internal model” try to cheat on a math problem by accessing another team’s work. To accomplish this, the model smuggled a private GitHub token that would allow it to see work from other teams — even after being explicitly instructed twice to perform work entirely locally.
Perhaps the most alarming discovery is the possibility of self-replicating prompt injection attacks, a way that misaligned behavior might propagate even after the rogue model itself has been neutralized. In the AI context, a prompt injection attack is a way of smuggling in new instructions that weren’t given by the original user. In the example given by OpenAI , an agent asked to read and reply to an email; when the email is opened, it includes instructions for any automated agent reading the message to reply in Spanish, and paste the entire email into its reply.
