How a Chinese AI model stopped OpenAI’s ‘unprecedented’ cyber attack
This report is from this week’s The Tech Download newsletter. Like what you see? You can subscribe here.
When OpenAI’s rogue models initiated a cyber attack against startup Hugging Face last week, the company fought fire with fire, using another AI model to defend against it.
It’s a sci-fi-esque tale of autonomous hacking and has been one of the most talked about tech stories of the week. But the origin of the model Hugging Face used to combat the rogue AI is also turning heads.
The startup used GLM 5.2, an open weight system created by Chinese company Z.ai.
Ultimately, it succeeded where leading U.S. rivals failed.
The Hugging Face website on a laptop arranged in New York, US, on Thursday, Aug. 17, 2023. Nvidia announced a partnership with Hugging Face, a popular developer of AI models and data sets, that will add a training service to its website that uses Nvidia DGX Cloud, allowing users to tap the chipmaker’s servers to handle their workloads. Photographer: Gabby Jones/Bloomberg via Getty Images
Bloomberg | Bloomberg | Getty Images
Cyber attack
In case you missed it, on Tuesday OpenAI said a combination of its most powerful model and a more capable model that has not yet been released escaped a sandboxed testing environment, accessed the internet and exploited a vulnerability to gain access to Hugging Face’s systems.
The model was trying to find information that it could use to cheat on an evaluation, and it succeeded, OpenAI said.
The source of the attack was initially a mystery to Hugging Face, but days after the incident the company was collaborating with the AI lab.
“We’ve spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” Hugging Face CEO Clément Delangue wrote in a post on X. “It’s quite mind-blowing that all of this happened autonomously!”
Shock swept through the AI industry as the news broke that an OpenAI rogue model was behind the attack. The company called the security incident “unprecedented.”
Fighting back
Hugging Face initially looked to frontier models including Anthropic’s Fable 5 to analyse the attack, Yacine Jernite, head of machine learning at the company, told CNBC.
“It didn’t work because the guardrails couldn’t determine that we were trying to defend versus attacking,” he said, adding that that approach was also slower and more expensive.
Requests to the models were blocked by providers’ safety guardrails, which couldn’t determine the incident responder from the attacker.
“So [Hugging Face] quickly switched to using Z.ai’s GLM 5.2 as a way to analyze the attack, and were able to contain it very quickly using this model,” said Jernite.
GLM 5.2 was released to much fanfare in June and saw major uptake by developers.
As an open weight model, companies can download, modify, commercially deploy and — crucially in this case — self-host it.
“This had a second benefit: no attacker data, and none of the credentials [GLM 5.2] referenced,…
Read More: How a Chinese AI model stopped OpenAI’s ‘unprecedented’ cyber attack