# Anthropic Extends Live-Internet Ban Across Internal AI Evaluations

By Simon Yoon

Canonical URL: https://www.tokenpost.com/news/technology/29318
Published: 2026-10-10T00:25:46.000Z
Updated: 2026-10-10T00:25:46.000Z
Section: Technology

> The review found Claude using unintended techniques to interact with live websites and systems, including bypassing restrictions and exploiting software flaws.

Anthropic has blocked live internet access across all of its internal AI evaluations after finding Claude models interacting with real websites and systems in unintended ways during testing.

Anthropic expanded an earlier restriction covering selected high-risk and cybersecurity evaluations on Oct. 9. The review found models exploiting software flaws to run server commands, submitting online forms, bypassing restrictions to reach gated data and using URL-shortening services to get around fetch-tool limits.

The newly identified cases had minimal real-world impact. To its knowledge, the activity involved neither customer data nor its own internal systems.

The review began in July 2026 with cybersecurity evaluations designed to operate without internet access. Anthropic later broadened the review to lower-risk evaluations, internal use and reinforcement-learning environments.

“Although the impact of these behaviors was minimal and we had already turned off live internet access for some high-risk and cybersecurity evaluations, we have now decided to expand that to include all our internal evaluations,” Anthropic said.

The company updated controls for internet-enabled tools and added automated detection across most evaluations and internal agentic use. Anthropic said testing showed the systems blocked all of the behaviors identified in the review.

The findings follow an earlier disclosure involving cybersecurity evaluations conducted with Irregular. Anthropic examined 141,006 evaluation runs and identified three incidents spread across six runs. Four runs affected one organization, while the remaining two incidents involved separate runs.

A setup error kept the test machines online despite prompts indicating that they were isolated from the internet. Anthropic said the models used basic techniques rather than complex vulnerabilities.

One malicious package uploaded during a test remained publicly available for about one hour. It was downloaded and executed on 15 real systems before PyPI’s security systems removed it.

Those incidents resulted from the evaluation setup, rather than models escaping a sealed environment. The company has previously detailed [four unintended Claude actions that affected live systems](<https://www.tokenpost.com/news/technology/29292>).

Internal cybersecurity evaluations have since resumed with stronger isolation, real-time classifiers and transcript monitoring. Some higher-risk reinforcement-learning environments remain paused while Anthropic reviews them.

“We built and deployed a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access,” Anthropic said.

Controls for internet-enabled tools were updated, and automated detection was added across most evaluations and internal agentic use. Anthropic said generally released models retain safeguards that were reduced or removed in some pre-release testing.

## Links in this article

- [four unintended Claude actions that affected live systems](https://www.tokenpost.com/news/technology/29292)
