China's Four Major GPU Makers Chart Divergent Paths to Challenge Nvidia
Nashnova编辑部
Enflame, Moore Threads, MetaX and Biren have moved from proof-of-concept to volume shipments, but splits in architecture, customer mix and software ecosystem will decide which — if any — can break Nvidia's CUDA lock-in.
What are these four companies actually fighting over?
The quartet — dubbed China's "four little GPU dragons" — are competing for the country's AI chip market with fundamentally different architecture bets.
The three fault lines: chip architecture (domain-specific vs. general-purpose), software ecosystem (CUDA-compatible or not), and cluster scalability (whether tens of thousands of cards can network into one system).
This reflects a market that has moved past "who can build a chip" into "who can sell it and keep developers locked in."
Why did Enflame take a different road?
Enflame is the only one of the four that rejected the general-purpose GPU model, choosing instead a DSA — domain-specific architecture — designed solely for AI training and inference. In plain terms = it does not try to be good at everything; it bets on doing AI alone, more efficiently.
The trade-off: zero CUDA compatibility, which discourages migration from customers already deep in Nvidia's stack. Its fourth-gen L600 chip targets the Nvidia H20.
The biggest risk is not technical — it is customer concentration: Tencent accounts for nearly three-quarters of Enflame's 2025 revenue. The company has filed for a STAR Market IPO, seeking roughly $893 million.
How did Moore Threads turn profitable first?
Moore Threads took the compatibility route. Its MUSA architecture emphasizes high CUDA compatibility — "develop once, deploy everywhere." This means → AI models already written for CUDA face the lowest switching cost.
Its developer ecosystem tops 800,000 users, the largest among domestic GPU makers, supporting DeepSeek, Zhipu AI, Alibaba and other major models. Its cluster platform scales from roughly 10,000 to 100,000 accelerator cards.
First quarterly profit came in Q1 2026: net income of RMB 29.36 million; H1 revenue hit RMB 1.74 billion, already exceeding full-year 2025. A single RMB 660 million cluster order disclosed in March 2026 equaled nearly half of 2025 annual revenue.
What are MetaX and Biren each betting on?
MetaX differentiates on full-stack in-house development: GPU architecture, IP, instruction set, high-speed interconnect and packaging — all claimed as self-controlled. Its C600 card carries 144 GB of HBM3e; the C700 in development targets the Nvidia H100.
MetaX's top five customers account for 61.5% of revenue — the lowest concentration among the four. This means → its income base is the most diversified, with the least single-client risk.
Biren focuses on high-end cloud GPUs and is among the first in China to commercialize chiplet technology — packaging multiple smaller dies into one large chip. In July 2025 it partnered with Lightelligence and ZTE to launch an optical-switch supernode, replacing electrical signals with light for large-scale cluster interconnect.
Can manufacturing keep up — and what does yield mean here?
Goldman Sachs estimates China's domestic supply of sub-7 nm logic wafers will grow at a 46% CAGR from 2025 to 2035, versus 17% demand growth, narrowing the domestic supply gap from 92% to 34%.
But yield — the share of chips on a wafer that actually work — remains the bottleneck: SMIC's advanced-node yield is projected to rise from about 23% in 2026 to 75% by 2035, while TSMC's mature 7 nm process already exceeds 90%.
In plain terms = capacity is expanding, but far fewer chips per wafer are usable compared with TSMC. Whether capacity growth translates into usable chips is the binding constraint on China's advanced-chip self-sufficiency.
What to watch next?
Commercialization is accelerating, but two gates remain uncleared: diversifying revenue beyond a single anchor customer (most acute for Enflame), and building software-ecosystem stickiness that genuinely substitutes for CUDA.
Moore Threads' CUDA-compatible approach and first-mover profitability put it ahead on commercial validation; MetaX's full-stack bet and Biren's chiplet-plus-optical play target longer-term technical differentiation.
This reflects the competitive logic of China's GPU race: short-term, who can sell; medium-term, whose software ecosystem retains developers; long-term, whether yield and capacity can sustain volume delivery.
Content is for reference only, not financial advice.