Chinese AI Firms Face Computing Power Constraints; High-Quality Inference Still Relies on Nvidia Chips
Nashnova编辑部
China's daily token calls now exceed 140 trillion, but high-quality inference still depends on Nvidia chips — domestic processors can only handle the low-value end, making the compute gap the biggest structural bottleneck in the country's AI supply chain.
How fast is inference demand growing?
National Data Administration figures: daily token calls in China topped 140 trillion in March 2026, up more than 1,000× from early 2024.
AI apps are shifting from "answer a question" to "agents" — AI programs that carry out real-world tasks. This means → each operation burns far more compute than a simple chatbot reply.
In plain terms = the AI used to give you a paragraph; now it runs errands for you. The compute bill is in a different league.
Why can domestic chips only take the low-end work?
Guan Jiawei, VP at inference-optimization startup Approaching.AI, says inference demand has split in two — high-quality token demand far outstrips supply, while low-quality tokens see weak demand and weak monetization.
Coding and similar tasks demand extreme chip performance; users will pay a premium. Domestic processors cannot reliably meet that bar.
This means → domestic chips are stuck with "cheap but unprofitable" inference jobs. The work that actually pays still goes to Nvidia.
Why does high-quality inference depend so heavily on Nvidia?
Guan put it bluntly: "If you rely solely on domestic chips for inference, you can only handle low-quality tiers — weak demand, weak monetization. That path is hard to sustain."
High-quality tokens — writing code, complex reasoning — require precision and speed that Nvidia chips still lead by a full generation.
This reflects a core tension: China's AI applications have raced ahead, but the underlying hardware is still bottlenecked by export controls.
Can software optimization close the hardware gap?
With Nvidia chips constrained by export controls, Chinese AI firms are using software optimization to stretch existing compute further.
The open question: can software-level gains close a generational hardware gap, especially in high-value tasks like coding and complex reasoning?
Put simply = software optimization is like tuning an old car to its limit — but the competitor drives a new one. Tuning narrows the gap; it rarely erases it.
Content is for reference only, not financial advice.