Goldman Sachs Warns: AI Sector Token Prices Falling Faster Than Demand Is Growing

nashnova research
今天发布阅读约 9 分钟

Goldman's One-Delta trading desk warns AI stocks face a systematic de-rating — token prices fell 29% in August to ~$0.97 per million, more than halving from the May peak, as volume growth can no longer offset collapsing unit economics.

01

How far have token prices fallen?

Silicon Data's LLM Token Spend Index shows the market-clearing price per million tokens dropped to ~$0.97 in August — an all-time low.
That is down more than 50% from the May peak of ~$2.05, with 29% of the decline in August alone.
This means → The index tracks spend-weighted price, not raw demand. A halving signals that providers are cutting list prices, users are migrating to cheaper open-source models, or both — it does not necessarily mean people are using AI less.
02

Volume is surging — why isn't revenue keeping up?

JPMorgan's data-center report shows OpenRouter token routing volume rose ~47% month-on-month in August, but dollar spend grew only ~7%.
H100 GPU rental prices fell in tandem, contradicting the "demand boom" narrative.
In plain terms = Users are running more AI queries than ever, but each query costs less and less. Volume up nearly five-fold in percentage terms, revenue up single digits — the gap is profit eaten by the price war.
03

How is open-source competition accelerating the price war?

Meta's Muse Spark 1.3 went live on September 2. Independent benchmarks show performance comparable to GPT-5.6 Sol and Claude Opus 5, at a blended price of ~$0.78 per million tokens.
This means → When a "next-best" model is close enough in capability yet markedly cheaper, enterprise users route around the premium tier. The Token Spend Index is the aggregate of those routing decisions.
This reflects open-source weights systematically compressing closed-model pricing power — not because open-source is better, but because "good enough and cheaper" is a powerful force.
04

Can OpenAI's Astra reverse the trend?

On September 1, OpenAI said its next-generation model Astra had reached a "critical cybersecurity threshold" under its preparedness framework — the first model in that category — and would open to external users "soon," but no confirmed launch date was given.
Goldman's Rich Privorotsky is skeptical: if Astra ships with access restrictions, monitoring requirements, and usage friction, its impact on billable token volume will be limited.
In plain terms = A powerful model gated behind heavy guardrails may generate less revenue than an open one, because high-value tasks cannot enter the public metering system — and may even reduce billable volume.
05

What does this mean for the AI sector as a whole?

The report's core judgment: the cloud-inference pricing model — billing per million tokens — only works when models have pricing power.
Three pressures are converging simultaneously: open-source weight competition + on-device inference hardware + overcapacity from massive capex build-outs. Pricing power is migrating away from model providers.
This means → The capex-underwriting logic of the past two years — "more tokens = more revenue" — was broken by August's data. Whether Astra can deliver a capability leap users will pay a premium for is the only near-term variable that could disprove the overcapacity thesis.

市场有风险,内容仅供研究参考,不构成投资建议。

Goldman Sachs Warns: AI Sector Token Prices Falling Faster Than Demand Is Growing · nashnova