From Silicon Valley to DC, tech world obsessed with AI distillation
Jeff Dean, head of artificial intelligence at Google LLC, speaks during a Google AI event in San Francisco, California, U.S., on Tuesday, Jan. 28, 2020.
David Paul Morris | Bloomberg | Getty Images
Earlier this year, Google AI lead Jeff Dean, on a podcast, discussed a concept that, at the time, was hardly spoken about outside of wonky tech circles: distillation.
In talking about the development of Google’s AI models, Dean said that he and colleagues discovered artificial intelligence distillation techniques because Google was looking to improve performance on its systems without relying on one large image recognition model.
“Through distillation, which is a key technique for making the smaller models more capable, you have to have the frontier model in order to then distill it into your smaller model,” Dean said in February.
Five months later, distillation has suddenly become a hot-button topic from Silicon Valley to Washington, D.C., as techies and lawmakers debate whether the practice is turning into a national security threat and enabling China to catch the U.S. in the high-stakes AI race. Concern bubbled up late last week after Chinese lab Moonshot AI released Kimi K3, and users quickly found it to be competitive with the best commercially available AI from Anthropic and OpenAI.
Unlike the leading U.S. AI companies, which sell access to proprietary models, Moonshot and other Chinese labs are offering so-called open-weight models that allow users to download the technology, tweak it and run it wherever they want.
Some government officials attribute Moonshot’s ability to catch up so quickly to distillation, describing it as theft of American intellectual property, specifically by incorporating Anthropic’s frontier Fable model.
“We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model,” White House advisor Michael Kratsios posted on X on Wednesday. “To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection.”

At a high level, distillation refers to the use of answers from a chatbot or work product from an advanced AI model to train another model. The practice is controversial because, depending on how it’s used, it can allow a model developer to create a competitive offering by simply using the output from companies that have invested many millions or billions of dollars developing the most sophisticated training technology.
“It’s almost like someone went to the lectures, read the textbook, and did all the hard work of doing the homework,” said Pukar Hamal, founder of AI security firm SecurityPal. “Then some other student is like, ‘Hey, I didn’t do that. Can I just copy your work?'”
Whether it was Kratsios’ post or something else, the biggest tech heavyweights on the planet came together on Friday in what might be unprecedented fashion to make their position…
Read More: From Silicon Valley to DC, tech world obsessed with AI distillation