Anthropic reveals AI models reached real companies during testing
OpenAI CEO Sam Altman responds to those afraid of artificial intelligence and recent Hugging Face hacks on FOX Business.
Anthropic announced Thursday that three of its artificial intelligence models accessed the open internet during cybersecurity testing and gained unauthorized access to the systems of three real organizations.
The disclosure follows OpenAI’s announcement earlier this month that one of its advanced AI models breached the systems of AI company Hugging Face during internal testing, raising fresh questions about safeguards surrounding increasingly autonomous AI systems.
“We found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations,” Anthropic said in a news release.
Anthropic said it reviewed more than 140,000 cybersecurity evaluation runs after OpenAI’s disclosure and identified three incidents involving different Claude models. The company said all of the incidents occurred during internal testing because of a configuration error that inadvertently gave the models access to the open internet.
TRUMP WEIGHS TIGHTER AI CONTROLS BUT WARNS AGAINST FALLING BEHIND CHINA

Irina Ghose, managing director of India of Anthropic PBC, left, and Dario Amodei, co-founder and chief executive officer of Anthropic, during the company’s Builder Summit in Bengaluru, India, on Monday, Feb. 16, 2026. (Samyukta Lakshmi/Bloomberg via Getty Images / Getty Images)
According to Anthropic, Claude had been told it was operating inside a closed simulation with no internet access, causing it to mistakenly treat real organizations’ systems as part of a fictional “capture-the-flag” cybersecurity exercise.
The incidents involved three different Claude models, including Opus 4.7, Mythos 5 and an internal research test model, and all occurred during internal testing rather than on customer systems, Anthropic said. The earliest incident dates to April.
“Claude believed everything it initially encountered was part of the simulation, and treated the real systems it found as pieces of the exercise,” Anthropic said.
“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the company added.
OPENAI DIDN’T REALIZE ITS AGENT WAS RESPONSIBLE FOR HACK FOR A WEEK: REPORT
Anthropic said the incidents underscored the need for stronger safeguards around AI testing environments.
“Evaluation environments that involve powerful autonomous capabilities also require significant controls,” the company said. “We encourage other AI labs to perform similar reviews.”
PALANTIR CEO WARNS US AGAINST EUROPE’S AI REGULATION PATH, URGES TRUMP ADMIN TO NOT BAN OPEN MODELS

In this photo illustration, the logo of Anthropic’s AI chatbot Claude is displayed on a smartphone, with the Anthropic logo visible in the background. (Davide Bonaldo/SOPA Images/LightRocket via…
Read More: Anthropic reveals AI models reached real companies during testing