Moonshot AI Seeks to Purchase More Nvidia Blackwell Chips to Train Next-Gen Models

Alina Collins
Published todayAbout 10 min read

Moonshot AI is seeking more Nvidia Blackwell chips to train Kimi K4, its next-generation model — significantly larger than the just-released 2.8-trillion-parameter open-source K3. This means China's top AI labs are growing more dependent on American chips, not less.

01

How was K3 trained?

Kimi K3 launched on July 16 with 2.8 trillion parameters, making it the world's largest open-source model to date.
Training relied on Nvidia chips, including the latest Blackwell series — Nvidia's newest-generation AI training chips, a major leap beyond the prior H100 line.
Senior White House official Michael Kratsios said K3 was trained in a Thai data center. But a person familiar with the matter said part of the training actually took place inside China.
This means → the chips not only reached a Chinese company — the training itself partly happened on Chinese soil, pointing to a breakdown in export-control enforcement.
02

Where did the chips come from?

Moonshot AI tapped at least two Chinese cloud providers holding Blackwell chips. Those chips were likely obtained through channels that may violate U.S. export controls.
No single supplier had enough chips, so engineers linked multiple eight-GPU Blackwell servers across data centers — a setup far smaller than what leading U.S. AI companies typically use.
In plain terms = Moonshot didn't get one big cluster in one facility. It stitched together small pockets of GPUs scattered across different providers.
To compensate, engineers custom-optimized the network architecture, improving chip-to-chip communication across those separate data centers.
03

Is Moonshot the only one? What about Alibaba and DeepSeek?

Alibaba's Qwen3.8-Max (2.4 trillion parameters) was also trained on Nvidia chips, including Blackwell.
DeepSeek was previously reported to have used Blackwell chips smuggled into China to train a new model, since released as V4.
This reflects an industry-wide pattern: Chinese AI companies are challenging Silicon Valley with open-source models, yet their training pipelines remain heavily dependent on American chip technology.
04

Why can't domestic chips substitute?

A researcher at a major Chinese tech company noted that Nvidia's advantages in stability and high-speed interconnects make its chips very difficult to replace for training frontier-scale models.
For inference — running the trained model in production — Moonshot relies heavily on Nvidia's H20 chip, a model the U.S. permits for sale to China under current rules.
In plain terms = the top-tier chips needed for training have no legal supply channel, but the chips for day-to-day model serving can be bought openly. Two pipelines, one blocked, one open.
05

Are export controls actually working?

Blackwell chips are flowing into Chinese training pipelines via cloud providers, and some training is happening inside China. The real-world effectiveness of U.S. controls is now openly questioned.
This is unfolding as Chinese open-source models rapidly take market share from American models — making the next move on export-control policy a critical watch point.
This means → if controls can't stop the chips from getting in, and Chinese models keep closing the gap with those chips, Washington faces the awkward reality of a policy that restricts on paper but not in practice.

Content is for reference only, not financial advice.

Moonshot AI Seeks to Purchase More Nvidia Blackwell Chips to Train Next-Gen Models · nashnova