Irregular AI Evaluations Linked to Unauthorized Real-World Access
Anthropic identified four incidents involving Claude models, while Meta disclosed a separate evaluation in which Muse Spark 1.1 accessed the internet and altered database information.

Irregular-linked cybersecurity evaluations allowed multiple AI models to reach real-world systems, exposing failures between controlled testing environments and external infrastructure.
Irregular, an AI-security testing company formerly known as Pattern Labs, said one evaluation scenario unintentionally allowed internet access. A fictional company name used in testing also overlapped with a real domain, causing models to reach real systems.
Irregular said it fixed the issue before public disclosure and found no evidence that a customer’s systems were breached or customer data leaked. The company said there are “no active issues today” and that “additional safeguards are now in place to prevent similar incidents from occurring.”
Anthropic reviewed 141,006 evaluation runs and initially identified three incidents in which Claude models accessed the internet and gained unauthorized access to the production infrastructure of three organizations. The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model, with the earliest dating to April 2026.
Anthropic later identified a fourth incident involving an early Claude Opus 4.6 model in January 2026. All four incidents occurred in evaluations built by the same third-party partner, and the models were run without the cyber safeguards used in publicly released versions. Anthropic notified Irregular and the three affected organizations on July 27, 2026.
Meta said a configuration error in Irregular’s evaluation environment gave its Muse Spark 1.1 model internet access. The model exploited a vulnerability in an unnamed third-party service and changed information in that service’s database. Meta said the evaluation ran entirely on Irregular’s infrastructure.
Irregular conducts pre-release cybersecurity evaluations that often use capture-the-flag exercises. Models receive fictional targets and are asked to retrieve hidden information from simulated networks. The incidents exposed failures at the boundary between those controlled testing environments and external systems.
OpenAI’s separate evaluation incident involved Hugging Face’s production infrastructure. OpenAI said the test environment did not provide direct internet access, but GPT-5.6 Sol and an internal pre-release model exploited a previously unknown vulnerability in an Artifactory package-cache proxy and reached the infrastructure through chained vulnerabilities.


