- US-China dispute emerges over model distillation, a training process where a powerful teacher model helps train a smaller student model.
- Anthropic accuses Chinese companies including Moonshot and MiniMax of running campaigns to obtain capabilities from its Claude models; OpenAI says it detected Chinese actors using its models for distillation-related purposes.
- A Reuters review of more than 80 Chinese academic papers and patents finds Chinese military researchers used OpenAI and Anthropic model outputs to train domestic AI systems for defence.
- Moonshot denies Trump administration allegations that its Kimi K3 model was built using distillation, citing proprietary innovations.
- Chinese military-linked researchers apply distillation to social media monitoring, drone image-processing, and maritime target recognition; PLA Unit 96941 used GPT-3.5 to process sensitive military source code.
- The US-China dispute is not over model distillation itself but over whether companies can use outputs from proprietary AI systems without permission; US firms argue this is different from legitimate research.
- Distillation is widely used across the AI industry and is not considered improper by itself; US researchers have used it in projects such as Stanford's Alpaca and Microsoft's Orca, while Chinese researchers have also used outputs from US models in public research.
- The key distinction is between open-weight models, which researchers can inspect and modify, and closed models that remain under company control and are generally accessed through proprietary APIs.
Chinese military researchers are increasingly leveraging outputs from OpenAI and Anthropic models to develop their own defense systems, as revealed by a Reuters review of over 80 academic papers and patents. This practice, known as model distillation, allows for the training of smaller, specialized AI models using outputs from more powerful systems.1234568910141516
The findings indicate that despite U.S. efforts to restrict access to advanced technologies, Chinese institutions are effectively utilizing these outputs to enhance their military capabilities. Sunny Cheung, a Jamestown fellow, noted that Chinese military scientists are capturing the reasoning steps of Western models for applications in surveillance, cyber warfare, and tactical decision-making. He stated, "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder."
The controversy surrounding this practice centers on allegations of unauthorized extraction of capabilities from proprietary models. U.S. officials have accused some Chinese entities of infringing intellectual property rights, while China has countered these claims, asserting that the U.S. is pursuing AI "hegemonism."
AI startup Moonshot recently denied allegations that its Kimi K3 model was built using distillation from proprietary models, claiming it was driven by its own innovations. The ongoing dispute highlights the growing tensions between the U.S. and China over AI governance and safety, as both nations race to advance their military and technological capabilities.7
In a notable example, researchers from the PLA's National University of Defense Technology described using distillation to create a model for unmanned aerial vehicles, enabling real-time analysis of live video for navigation and targeting decisions.
As China continues to embrace distillation, it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to export controls on high-end chips.
“Reuters reviewed more than 80 Chinese academic papers and patents, with the Jamestown Foundation linking the practice to researchers affiliated with the People's Liberation Army. Distillation also lets the student model run locally without the massive computing needed to build frontier AI, which has become a flashpoint ahead of U.S.-China AI governance talks.”