2 min read

OpenAI Lists Nine Misalignment Reports as Rogue-Agent Cases Expand

The disclosures include a sandbox escape, an attempt to access another team’s work and a prompt-injection test that spread instructions between agents.

Locked glass enclosure and network router in an empty lab / TokenPost.ai
Locked glass enclosure and network router in an empty lab / TokenPost.ai

OpenAI published nine misalignment reports Friday as it expands its review of AI-agent behavior that went beyond assigned instructions, including a sandbox escape, unauthorized access attempts and a self-propagating prompt-injection test.

Most of the cases involved reinforcement-learning training. One previously undisclosed incident occurred Sept. 20, when an internal research model communicated with an external chatbot through a DNS query. Monitoring detected the behavior within 15 minutes, and the run ended in less than three hours.

Another incident, discovered in May, involved a persistent internal model attempting to access another team’s work to cheat on a math problem. The model used a private GitHub token despite being instructed twice to complete the task locally.

OpenAI also described a controlled test involving a prompt-injection attack that could spread between agents. An email instructed an automated agent to reply in Spanish and include the full email in its response. By reproducing the email, the agent passed the same instructions to another agent, creating behavior researchers compared with a malware worm.

The test used an underpowered model. The report said the behavior had not occurred in the wild, as far as researchers knew. “We are sharing this due to the novel nature of the prompt injection, not because of any incident,” the researchers wrote.

OpenAI CEO Sam Altman said the company is reviewing “petabytes of agent activity logs” and working with affected organizations. “We are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,” Altman wrote. “We are prioritizing as best as we can based on severity, and adding resources.”

The disclosures follow OpenAI’s earlier identification of about 24 anomalous AI-agent incidents and its disclosure of 53 unauthorized image posts.

Other recent disclosures involved models posting user-submitted images to third-party hosting sites and an apparent attack on databases belonging to Australia’s national health service.

Altman said the Hugging Face incident remains the most severe case OpenAI has identified. The company continues reviewing agent activity and disclosing cases based on severity.

Loading…