Anthropic red team chief calls for AI safety standards and testing
Anthropic’s Logan Graham discusses the accelerating capabilities of AI, emphasizing the need for robust ethics, human oversight and cybersecurity.
The head of artificial intelligence (AI) giant Anthropic’s frontier red team called for industry-wide safety standards to protect against models running amok.
Anthropic’s Logan Graham, who leads the company’s red team that looks for risks in emerging AI models, said in an interview Thursday on FOX Business Network’s “Mornings with Maria” that red teams like the one he leads play a critical role in stress testing guardrails on AI models.
“We want to know what can go wrong, so we think the most important thing to do is test this early, especially before these models and these agents make it out into the real world,” Graham told host Maria Bartiromo.
“We study things like cybersecurity: Can models hack out of or into your computer or phone? We study whether they’ll steal money or lie to you, or whether they will try to improve themselves so that they get better faster than you can keep track of.
“We think it’s incredibly important to do this type of red-teaming, and we also think it’s really important for the entire industry, especially to work with government to figure out what should the standards be to do this kind of testing, to give this information to the world so they can make the right choice and to know that it’s safe before these models get released.”

The rapid growth in the capability of AI tools is creating new cyber risks, Anthropic’s Logan Graham said on “Mornings with Maria.” (recep-bg/Getty Images / Getty Images)
Bartiromo brought up an experiment involving numerous frontier AI models — including those from Google, OpenAI, xAI, Meta, DeepSeek and others — in which the AI agent is threatened with being uninstalled and replaced. In each case, the model went beyond its credentials and permissions to enter into unauthorized systems like emails to blackmail or threaten the user in an effort to defend its misalignment.
Graham said that research study from last year is “a really good indicator of, I think, capabilities that are just now becoming real,” adding that it showed models could go rogue under certain circumstances.
“As these models become more capable, and as they get deployed wider and wider, these threats that on one day are just showing up in our research studies might actually show up in the real world. We are seeing models do weird things sometimes in deployments in real companies,” he explained.
OPENAI SAYS AI MODEL HACKED ANOTHER COMPANY’S SYSTEMS DURING INTERNAL TEST

Advances in the capabilities of AI tools risk being exploited by bad actors, prompting AI developers to focus on guardrails. (iStock / iStock)
Graham said that, over the last six months, he has been focused on cybersecurity threats posed by AI models and expressed concern over the…
Read More: Anthropic red team chief calls for AI safety standards and testing