Zhipu Claims GLM-5.3-Flash Runs Entirely on Domestic Chips; HK-Listed Stock Surges Over 8% in a Single Day

Nashnova编辑部
今天发布阅读约 11 分钟

Zhipu announced that GLM-5.3-Flash serves all online requests on over 100,000 domestically made chips, sending its Hong Kong shares up more than 8% to HK$1,115 — pushing the domestic-compute narrative from lab stage to production scale.

01

What kind of results did this model put up?

GLM-5.3-Flash launched anonymously as "Ox Alpha" on OpenRouter. In five days it generated over 50 trillion tokens of traffic; current usage is more than double DeepSeek's.
It ranks 10th on the Artificial Analysis Intelligence Index, above DeepSeek V4 Pro Max, and ties Anthropic's Claude Opus 4.8 at 57 points.
This means → Zhipu matched top-tier flagship scores with a model whose active parameters are just 18B — while pricing it at one-tenth of its own flagship.
02

"All domestic chips" — how solid is that claim?

Zhipu stated that every online inference request for GLM-5.3-Flash runs on more than 100,000 domestically made chips, moving domestic-compute substitution from proof-of-concept to production deployment.
CNBC noted it could not independently verify the chip claim. Zhipu declined to name its suppliers. LatePost previously reported the chips likely come from Huawei, Moore Threads, and Hygon; Zhipu offered no comment.
In plain terms = this is the largest public claim of "domestic chips running a frontier model" to date, but no third party has verified it, and the exact chip suppliers remain unconfirmed.
03

How do domestic chips handle a million-token context window?

The main bottleneck per domestic chip is memory capacity and bandwidth. Zhipu built a dedicated inference engine on top of SGLang — an open-source serving framework — using W8A8 quantization (compressing model weights and activations to 8-bit precision to save memory and speed up compute), mixed-cache quantization, and intra-node tensor parallelism.
The system also uses a production-grade EPD architecture — splitting "encode," "prefill," and "decode" into separate stages, each scheduled on its own chip resources. End-to-end serving performance improved 3× over the initial baseline on the same hardware.
This means → domestic chips still trail Nvidia on single-card performance, but Zhipu's engineering stack — quantization, disaggregation, parallelism — closes the gap enough to serve commercial traffic.
04

How aggressive is the pricing?

GLM-5.3-Flash costs RMB 0.8 per million input tokens and RMB 2.8 per million output tokensone-tenth of Zhipu's flagship GLM-5.3.
A launch promotion halves even that price for two weeks, putting the promotional rate at roughly 1/40th of Claude Opus 4.8.
This reflects Zhipu's playbook: grab traffic and developer ecosystem share at razor-thin margins first, then use scale to amortize domestic-chip unit costs.
05

How is rival MiniMax doing? What does the competitive picture look like?

MiniMax (00100.HK) rose about 3% the same day. It reported first-half revenue up 283% year-on-year, but adjusted net losses more than doubled to US$293 million.
Both companies listed in Hong Kong in January. Zhipu is up over 800% since IPO; MiniMax is up over 80% — a full order of magnitude apart.
In plain terms = MiniMax is growing revenue but bleeding more cash. Zhipu's stock trajectory shows the market currently values the "domestic chips + top benchmark scores" combination far more than revenue growth alone.
06

What comes next?

Zhipu is set to report first-half earnings next Monday. Whether commercial revenue can support the current valuation is the key test of whether this domestic-chip narrative holds.
This means → the market has already priced in a premium for "domestic compute substitution," but if the revenue numbers fall short, correction pressure will be significant.

市场有风险,内容仅供研究参考,不构成投资建议。

Zhipu Claims GLM-5.3-Flash Runs Entirely on Domestic Chips; HK-Listed Stock Surges Over 8% in a Single Day · nashnova