China's AI Industry Questions Kimi K3 Benchmark Fraud and Distillation Controversy
N.R. Finch
After Moonshot AI released Kimi K3, detailed allegations of benchmark gaming and unauthorized distillation emerged from within China's own developer community — a dispute that now implicates the broader Chinese LLM industry, not just one company.
How strong is K3's scorecard, really?
Kimi K3 scored 1679 on the Arena front-end coding leaderboard, above Anthropic Fable 5's 1631 and GPT-5.6 Sol's 1618.
Moonshot AI positions K3 as a 2.8-trillion-parameter open-weight model, but concedes its overall performance still trails the strongest closed-source systems.
This means → K3's pitch is "approaching the U.S. frontier," not surpassing it. A coding-leaderboard lead does not equal across-the-board superiority.
Developer tests cited by 36Kr praised K3's UI generation but flagged slower speeds on some comparable workloads.
What exactly do the allegations claim?
A Chinese developer using the GitHub handle "RainPPR" published a detailed but unverified document describing the current Chinese AI industry as an era of "mass distillation."
The document alleges that Zhipu AI's GLM team extracted Fernet-encrypted "reasoning data blocks" from Anthropic's system in April–May 2026. In plain terms = they allegedly pulled out the hidden chain-of-thought that the model keeps private, then used it to train their own systems.
It further alleges Moonshot AI narrowed the gap with Fable through contaminating evaluation data, identifying assessment prompts from system logs, and routing some benchmark requests to Fable's original model.
The document claims Moonshot disbanded its reinforcement-learning team, shifted to supervised fine-tuning, and used "architectural innovation" as cover for reliance on distilled outputs.
Do these allegations hold up?
The claims are serious but lack independent audit evidence. Neither Zhipu AI nor Moonshot AI has responded publicly.
Some circulating screenshots showing "Kimi identifying itself as Claude" have been flagged as potentially fabricated, casting doubt on the "identity contamination" narrative.
This means → the situation is "detailed allegations, no hard proof" — the charges can neither be confirmed nor dismissed.
Is distillation itself wrong? Where is the line?
Distillation — training a smaller model on a larger model's outputs — is a standard practice in AI development and is not inherently improper.
The real questions are two: whether the data was obtained by bypassing access restrictions, and whether benchmarks and marketing accurately disclosed reliance on another party's system.
RainPPR's document also names Tencent, Alibaba, MiniMax, and DeepSeek as having pursued larger-scale distillation projects. This reflects a dispute that has expanded from one company to the entire Chinese LLM industry.
Can the premium pricing hold?
K3's output price is ¥100 per million tokens, roughly 270% higher than Kimi K2.6 and above several major Chinese competitors.
Input pricing sits at ¥20 per million tokens (cache miss) or ¥2 (cache hit).
Moonshot AI's logic: a frontier Chinese model, even with open weights, deserves a fair commercial return. This means → Moonshot is betting the market will pay a premium for "near-U.S.-level" capability.
But if the performance edge cannot be validated beyond coding — across reasoning, reliability, tool use, and latency — the premium strategy's ability to sustain user scale will be the key test of K3's commercial viability.
Content is for reference only, not financial advice.