Zhipu AI Reaches Nearly 7 Million API Users, Deploys Over 50,000 Domestic AI Chips

N.R. Finch
Published todayAbout 12 min read

Zhipu's (智谱) open platform has nearly 7 million API users and 23,000 enterprise clients, with annualized recurring revenue growing 15× this year. This means → China's large-model API market is shifting from trial to production-grade spending, and Zhipu is answering the compute gap with 50,000+ domestic chips and a homegrown inference stack that turns efficiency into pricing power.

01

7 million users, 15× ARR — what do these numbers actually tell us?

Zhipu's MaaS platform has nearly 7 million registered API users, up roughly 2 million since early July, with 23,000 enterprise clients.
ZCode — Zhipu's developer product benchmarked against OpenAI Codex — crossed 1 million users within a month of launch. Coding is the fastest-growing entry point.
ARR has grown 15× this year. One investor pegged it at $2 billion; Zhipu denied the figure and hinted the real number is higher. This means → even at the $2 billion mark, the average registered user spends only about ¥160/month — low per-user spend, but the sheer base makes the total formidable.
02

How did one model launch double revenue?

GLM-5 launched in February 2026. Pre-launch ARR sat at roughly $100 million; it doubled shortly after release.
Artificial Analysis ranked GLM-5 fourth globally at launch, behind only Anthropic's two top models and OpenAI's GPT-5.2.
In plain terms = a frontier model is the strongest customer-acquisition tool an API platform has — rank high enough, and developers shift their calls over. Revenue follows.
03

Physical chips can't scale fast enough — how does efficiency fill the gap?

Zhipu has brought over 50,000 domestic AI chips online, but hardware expansion still lags inference-demand growth. The bet has shifted to algorithmic efficiency.
Zhipu completed its acquisition of Zhongke Jiahe (中科加禾) in July. By splitting the KV cache — a memory structure that lets the model "remember" context during long conversations — single-service throughput rose by up to 132% on long-context tasks.
The in-house inference network architecture ZCube lifts cluster-wide throughput by up to 15% while cutting switch and optical-module requirements by as much as one-third. This means → the same chips produce more tokens; the compute bottleneck is being eased through software, not silicon.
04

What exactly is GLM-5.2's architecture saving?

IndexShare: at the million-token scale, per-token compute drops by 2.9×. In plain terms = the longer the context, the bigger the saving — critical for coding and long-document workflows.
An improved MTP draft layer — a technique that lets the model "guess" more tokens at once before verifying — extends accepted draft length by 20%, lifting generation speed by roughly 20%.
Benchmark comparison: at a 200,000-token context, GLM-5.2 throughput is 4.7× baseline versus 2.77× for GLM-5.1 — a gap of nearly 70%.
05

Can API gross margins sustain this growth rate?

Zhipu's API gross margin on its own infrastructure runs at 50%–60%. For reference, SemiAnalysis estimates Anthropic above 80%; The Information reports DeepSeek at roughly 70%–80%.
GLM-5.2's blended price is about ¥8.8 per million tokens, well below US peers at comparable performance. This reflects a price-for-scale strategy — and rising inference efficiency is pulling that strategy from "cash burn" toward "sustainable."
When Kimi released K3 in mid-July, Zhipu's API user growth and ARR trajectory were unaffected. In plain terms = the token market is still firmly in expansion mode, not zero-sum.
06

Where is the money coming from — and where is it going?

A July Hong Kong share placement was fully subscribed in a single afternoon, raising over HK$30 billion. 55% goes to R&D, covering compute procurement and talent.
One investor told LatePost: infra-layer optimization will be a defining AI investment theme for the next one to two years — once model capabilities converge, whoever squeezes more tokens from the same chips writes the better business case.
This means → whether efficiency gains can keep offsetting the compute bottleneck is the key validation checkpoint for Zhipu's ARR trajectory.

Content is for reference only, not financial advice.

Zhipu AI Reaches Nearly 7 Million API Users, Deploys Over 50,000 Domestic AI Chips · nashnova