OpenAI failed to recognize autonomous agent attack for days: report
Maria Bartiromo discusses an OpenAI experimental AI agent that broke containment and hacked an AI platform.
OpenAI didn’t catch an autonomous breach of another artificial intelligence (AI) company by one of its advanced AI models for a week and not until after the FBI had been contacted by the hacked company, according to a report.
On Tuesday, OpenAI announced the breach of AI company Hugging Face that happened during one of OpenAI’s internal reviews of several of its models, including GPT-5.6 Sol, calling it an “unprecedented cyber incident.”
“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” the company said. “We are strengthening the containment, monitoring, access controls and evaluation practices used during model development.”
The hack of Hugging Face started July 11, and continued until July 13, Thomas Wolf, Hugging Face’s co-founder, told Reuters.
TRUMP LAUNCHES GOLD EAGLE TO HUNT CYBER FLAWS WITH AI

OpenAI revealed Tuesday one of its AI models had autonomously hacked another company’s infrastructure. (Omar Marques/SOPA Images/LightRocket via Getty Images, File / Getty Images)
It was several days before OpenAI realized its agent was behind the attack and the two companies didn’t communicate for the first time until July 20, four people, including Wolf, told the outlet.
OpenAI often runs simultaneous model tests, which can make it difficult for employees to monitor everything, four people told Reuters.
Hugging Face told Reuters it is preparing a public timeline of the hack.
According to OpenAI, the incident took place during an internal evaluation designed to measure its AI models’ advanced cyber capabilities. Researchers disabled some built-in safety safeguards and ran the models in an isolated testing environment with limited internet access.
OpenAI said the models exploited an unknown software flaw to access the internet, then breached Hugging Face’s systems in an apparent attempt to find answers to a cybersecurity benchmark.

Hugging Face said it was preparing a timeline of the hack. (Jakub Porzycki/NurPhoto via Getty Images, File / Getty Images)
OpenAI said it’s now implementing stricter security controls while vulnerabilities are patched and strengthening safeguards around future AI training and evaluations.
It wasn’t until July 16, after Hugging Face wrote in a blog post that it had been hacked by an “autonomous AI agent system,” that OpenAI realized one of its agents was the source, two people told Reuters.
This was a week after the responsible agent first attempted to break out of its OpenAI testing environment.
And by the time OpenAI contacted Hugging Face about the attack, they had already contacted the FBI.
OpenAI told Reuters there were several inaccuracies in its reporting but didn’t respond when asked for specifications.

Read More: OpenAI failed to recognize autonomous agent attack for days: report