AI Token Price Deflation: Three-Way Structural Divergence Beneath the Surface

nashnova research
今天发布阅读约 10 分钟

Silicon Data's Token Spending Index has fallen 52.4% from its late-May peak, but Syz Group CIO Charles-Henri Monchau warns the headline number conflates three very different deflationary forces — mixing them up means trading on noise.

01

One index down by half — why does it "conceal more than it reveals"?

Silicon Data's Token Spending Index is a usage-weighted average price benchmark. It does not reflect total token spending. This means → a 52.4% price drop does not mean the industry is spending half as much.
Syz Group CIO Monchau breaks the price decline into three categories: technical deflation (models getting more efficient), structural deflation (enterprises routing low-priority tasks to cheaper open-source models), and competitive deflation (price wars).
In plain terms = all three show up as "lower prices," but the first is genuine progress, the second is demand re-routing, and the third is cash-burn for market share — each carries a completely different implication for profits and investment.
02

Frontier vs. open-source models — where does the price gap sit?

Seaport Research Partners tech analyst Jay Goldberg's data shows frontier LLM token prices fell from roughly $50 per million tokens in late 2023 to about $6 today.
The median price, however, has held broadly steady at just under $1 since mid-2024. The floor price — around $0.11, close to raw electricity cost — has barely moved.
This reflects a market that has split into two tiers: the top tier still has room to fall, while the bottom tier already runs at cost.
03

Costs are falling faster than prices — who is making money?

Goldberg estimates production costs are dropping much faster than market prices. This means → sellers' gross margins are actually widening, not shrinking.
That scissor gap creates starkly different profit structures: frontier models (mostly U.S. vendors) run at roughly 70% gross margin; open-source models (mostly Chinese vendors) sit at about 20%.
This split has remained relatively stable over three years. In plain terms = frontier players cut prices but still earn outsized margins; open-source players have been grinding through thin margins all along.
04

Why is training cost "the hottest variable"?

Every margin figure above covers only inference — the step where a model answers a query. Training capex is not included.
Goldberg calls training cost "the hottest variable" in the entire valuation framework — its scale and uncertainty both dwarf the inference side.
This means → the 70% inference gross margin looks attractive on paper, but once massive training outlays are factored in, the real return profile could look very different.
05

Can Anthropic and OpenAI make the IPO story work?

Goldberg's framework points to a core contradiction: high inference margins + fast growth make the top line compelling, but sustaining a reasonable return on investment requires two things at once.
First, maintain the technological moat. Second, slow open-source models from closing the performance gap. Lose either, and the moat collapses.
Regulatory safety narratives and within-border protectionism are floated as potential moat options — but their effectiveness remains unproven. Put simply = the tech barrier's durability is uncertain, the policy shield has not landed, and the core logic underpinning IPO valuations is still hanging in the air.

市场有风险,内容仅供研究参考,不构成投资建议。