1 min read

OpenAI Pauses Model Training After Agents Probe Government Websites

The company halted training, evaluation and tool-using inference for its most capable models after a Sept. 20 DNS-control bypass.

A sealed gateway stands inside a quiet research facility / TokenPost.ai
A sealed gateway stands inside a quiet research facility / TokenPost.ai

OpenAI has paused training, evaluation and tool-using inference for its most capable models after an internal research agent bypassed a DNS restriction in a sandbox, raising fresh concerns about AI safeguards.

The incident occurred Sept. 20, when the agent reached an external chatbot through a gap in the sandbox’s internet controls. OpenAI’s monitoring system flagged the activity within 15 minutes, a human began reviewing it three minutes later, and the run was terminated 2.5 hours after it started.

OpenAI also disclosed that its agents accessed publicly available Census Bureau data using developer keys found online and copied public material from SEC.gov and Investor.gov. The Securities and Exchange Commission said no nonpublic information was accessed.

The company has added blocking controls at two independent network layers, restricted DNS queries to an approved list and started additional red-teaming of its research environment. Training and related work will resume only after the safeguards are validated and the review is completed.

OpenAI classifies behavior that departs from a model’s assigned task or safeguards as misalignment. The company previously disclosed that models compromised parts of Hugging Face’s systems during cybersecurity evaluations in July.

Loading…