Is China's AI Catching Up Through "Distillation"? Industry Pushes Back on U.S. Characterization
nashnova research
The US government and several AI labs blame China's rapid AI progress on 'distillation' — training on the outputs of superior models — and call it theft. But industry voices, including a co-author of the foundational Transformer paper, are pushing back: Chinese models now beat top US models on some benchmarks, and you can't copy your way past the leader.
What is 'distillation,' and why is it suddenly the flashpoint?
Distillation — using a stronger AI model's outputs to train a smaller or newer model — is a standard technique in machine learning.
The US government and labs like Anthropic accuse Chinese AI companies of distilling American models at industrial scale and frame it as intellectual-property theft.
This means → Washington is redefining a routine technical method as a property-rights violation — a key narrative shift in the US-China AI rivalry.
What exactly is the US alleging?
Anthropic's head of threat intelligence, Jacob Klein, says there is "an entire illicit ecosystem" trying to access Claude and other models.
Anthropic's report this month named Alibaba, Moonshot, and DeepSeek, accusing them of training their own models through "illegal distillation."
The US Cybersecurity and Infrastructure Security Agency (CISA) went further, calling China's extraction of proprietary capabilities a "systematic" effort and labeling industrial-scale distillation as "central, not supplementary" to China's AI strategy.
Why is the industry pushing back?
Cohere CEO Aidan Gomez says Chinese AI models have reached "world-class" level and the US lead is "closing fast." Gomez co-authored the 2017 paper *Attention Is All You Need*, which laid the foundation for modern AI.
His core argument: "You cannot copy or distill your way past a competitor — you can only close the gap." Chinese models already outperform the best US models on some benchmarks. This means → distillation alone cannot explain the full extent of China's progress; independent innovation must be part of the picture.
Former White House senior AI policy adviser Sriram Krishnan adds a technical point: ChatGPT and Claude themselves were trained on internet content. "Distillation has always been central to how computer science works," he says. In plain terms = every large model "distills" existing human knowledge; singling out China for the same practice is logically shaky.
What is really at stake in this debate?
The core dispute is not technical — it is about policy narrative. The claim that "distillation is the main driver of China's progress" is a key pillar supporting US AI export controls on China.
This means → if the industry increasingly concludes that China has built independent innovation capacity, the logic behind "restricting exports will contain Chinese AI" starts to erode.
China's Ministry of Commerce has rebutted the "industrial-scale distillation" allegation, though the full text of its response has not yet been disclosed. This reflects a significant information gap between the two sides on the factual record.
市场有风险,内容仅供研究参考,不构成投资建议。
