Goldman Sachs: China's LLMs Now Dominate Global Token Usage, Market Underestimates the Shift

Miles Bennett
Published todayAbout 10 min read

Goldman Sachs says Chinese AI models already dominate token usage on global platforms, priced at roughly one-third of Western peers — and has lifted its 2026 China AI model ARR forecast to $13 billion, arguing the market is far behind in pricing this structural shift.

01

How did Chinese models capture global token share?

On platforms like OpenRouter, Chinese models now hold the dominant share of token usage, typically priced at one-third of comparable Western models.
In some use cases, costs drop to as low as one-hundredth of Western frontier models. This means → developers get far more tokens per dollar from Chinese models, so volume naturally migrates to the cheaper end.
In plain terms = Chinese models are competing on "more for less," and global developers are voting with their wallets.
02

Why is Goldman suddenly raising revenue forecasts?

Goldman lifted its 2026 combined ARR — annual recurring revenue, the steady subscription income a business collects each year — forecast for Chinese AI models from $10 billion to $13 billion.
Z.AI and MiniMax saw 2026 revenue estimates raised by 35% and 63%, respectively.
This reflects Goldman's view that the market underestimates a key inflection: users are migrating fast from free consumer apps to monetizable enterprise-agent workflows. In plain terms = AI is no longer just a chat toy — enterprises are starting to pay real money for it.
03

Which five developments made Goldman more bullish?

MiniMax H3 launched as an open-weight multimodal model, priced 30–50% below peers; ByteDance's Seedance 2.5 went live via API the same day, supporting up to 30-second video generation.
Alibaba's Qwen 3.8 Max debuted at 2.4 trillion parameters, climbing sharply in Arena.ai code rankings. This means → competition among 2–5-trillion-parameter coding and agent models will intensify visibly in the coming months.
DeepSeek V4 Flash launched formally, outperforming its earlier V4 Pro preview; Tencent's Hunyuan Hy3 production release also showed clear gains — Goldman credits SFT (supervised fine-tuning — a post-training step that calibrates models with human-labeled data) as the key driver.
04

Where do enterprise apps and compute supply stand?

Tencent's Workbuddy and Alibaba's QwenWork are accelerating enterprise-agent deployment; major internet platforms lead in workplace-app scenarios.
Domestic compute supply is expected to keep expanding through H2 2026, laying the groundwork for training next-generation models above 2 trillion parameters.
This means → the positive loop of "bigger models + more compute" has already started — it is no longer a slide-deck aspiration.
05

What is Goldman telling investors to buy?

Goldman maintains a Buy on MiniMax with an HK$800 target, citing H3's multimodal capabilities and aggressive pricing as catalysts for H2 2026 ARR growth.
Cloud computing and data centers are the top-pick sub-sector — Goldman favors GDS, VNET, Alibaba, and Kingsoft Cloud — reasoning that surging token demand plus continued hyperscaler capex will deliver margin expansion and pricing power.
This reflects Goldman's core thesis: the competitive battleground has shifted from "whose model is smartest" to cost-efficiency and the breadth of the agent ecosystem. In plain terms = it is no longer about benchmark scores — it is about who can sell AI cheaply and get enterprises to actually use it.

Content is for reference only, not financial advice.