WAIC 2026: China's AI Competition Pivots Toward Low-Cost Token Production

Taylor Wilson
Published todayAbout 14 min read

At Shanghai's World Artificial Intelligence Conference, 'token factory' became the defining concept — China's AI race is shifting from model size and chip counts to per-token inference cost and profitability, signaling a new phase where the winner is whoever produces intelligence most cheaply.

01

Why is everyone suddenly talking about "token factories"?

Enterprise token consumption in China's model-as-a-service market surged from roughly 1.6 trillion per day in January 2025 to about 9.6 trillion per day by year-end — nearly a six-fold jump in under twelve months.
This means → demand has grown so large that "how good is the model" is no longer the only question. "How much does one inference run cost" is now the decisive metric.
In plain terms = the race used to be about who builds the biggest, smartest model. Now it is about who produces each answer at the lowest cost. That is the starting point of the "token factory" logic.
02

SenseTime's approach: different chips for different jobs — how does it work?

SenseTime's infrastructure arm SenseCore unveiled heterogeneous hybrid inference — splitting inference into two stages handled by different domestic chips: general-purpose processors tackle the compute-intensive context-understanding (prefill) phase, while high-end chips handle the memory-bandwidth-sensitive token-generation (decode) phase.
A unified networking, scheduling, and caching layer stitches chips of different brands and generations into one system. In plain terms = it does not matter whose chip it is — if it works, it joins the same production line.
SenseTime says the cluster already runs at positive gross margin, with daily throughput rising from roughly 400 billion tokens in early 2025 to about 2.42 trillion by late July, targeting 10 trillion by year-end.
03

Are SenseTime's performance claims credible?

The newly released SenseNova U1 Pro foundation model claims domestic-chip utilization (MFU) gains of 85 % to 152 %, inference cost-performance at 1.25× Nvidia's H-series, and 2.5× the token output at equivalent cost.
All figures are vendor-reported and have not been independently verified. This reflects a broader industry problem: performance benchmarks lack a common standard.
SenseTime also plans at least five domestic 10,000-card clusters serving over 200 AI startups, with deployments in Yancheng, Hong Kong, and Saudi Arabia.
04

What can Sugon's 100,000-card supercomputing cluster do?

Sugon's "Dawning 8000" — described as China's first fully domestic 100,000-card AI supercomputing cluster — runs on Hygon DCU chips linked via RDMA networking. It won the conference's "star exhibit" title.
Connected to the Zhengzhou national supercomputing node, the cluster can handle 5 % to 10 % of China's current token demand while supporting roughly 1.2 million concurrent AI conversations. It ran at full load in its first week, processing over 150,000 jobs per day.
This means → a single cluster can absorb a meaningful slice of national demand, a sign that compute infrastructure is crossing from lab scale to industrial scale. Henan province's AI sector has topped 100 billion yuan, with a capacity target of 120 EFLOPS by end-2026.
05

Beyond SenseTime and Sugon, who else is betting on this path?

Full-stack cloud provider Phancy named its platform outright "Token Factory"; Qingcheng AI launched a "front-shop, back-factory" token model — the concept has moved from a single vendor's label to an industry consensus.
On the hardware side, Huawei's Ascend 950 super-node — 1 EFLOPS per system — won the SAIL Award, while MetaX and Alibaba showcased in-house chips. This reflects a systematic effort across China's compute ecosystem to build inference capacity around Nvidia.
Tencent Cloud VP Wu Yunsheng told the conference that companies are now "doing token math": ambiguous tasks should go to token-heavy agents, while deterministic tasks should use cheaper fixed workflows. In plain terms = not every question deserves the most expensive answer — allocating compute by task type is what makes the economics work.
06

Can this path actually deliver?

The core of the "token factory" logic is cost discipline — squeezing more sellable tokens out of every watt and every chip. Chinese vendors now treat inference economics, not model scale, as the next competitive frontier.
The key proof point: whether domestic hybrid-inference clusters can sustain positive gross margin at much larger scale. SenseTime claims it has done so, but hitting the year-end 10-trillion-token target means multiplying capacity several times over.
This means → the "token factory" is not a proven endgame but a bet in progress. Whether the cost curve keeps falling as scale rises is the real suspense in this race.

Content is for reference only, not financial advice.

WAIC 2026: China's AI Competition Pivots Toward Low-Cost Token Production · nashnova