UBS Survey: Open-Source Model Token Share Surges, Inference Compute Faces Widespread Shortage
Taylor Wilson
A UBS survey finds enterprise open-source model token share rocketed from 0.1% to above 10% in six months, yet inference compute is already maxed out — Kimi 3 hit its GPU ceiling 48 hours after launch, and Databricks is GPU-short across multiple regions.
Open-source token share jumped 100× in half a year — where did the money go?
One AI-native company disclosed hard numbers: open-source models accounted for just 0.1% of its enterprise token volume in January, rose to 1% two months later, and now sit well above 10%. The firm projects the share will top 90% within 12 months.
This means → task volume running on open-source models is growing exponentially, but that does not mean open-source is capturing the revenue. Over 90% of compute spend still goes to frontier closed-source models.
In plain terms = open-source is winning the "bulk workload" game, but for high-value, high-difficulty core business tasks, customers still pay a premium for closed-source. Frontier model providers retain pricing power in those scenarios.
Third-party data backs this up: Together AI reports open-source token consumption rose from 10% to 30% over the past year. Open-source models lag closed-source leaders by roughly 90 days in capability, but run at only 10–30% of the cost.
How fast are enterprise AI bills climbing — and why do some companies not care?
UBS estimates roughly 60% of enterprises now list AI compute / token cost as a core pain point. Examples: one financial institution's Claude procurement order grew from $10 million to $60 million in three months; a bank spends hundreds of thousands of dollars per month on Opus 4.8 inference, much of it on low-complexity queries.
One AI-native company's monthly Anthropic spend jumped from $20,000 last December to nearly $1 million this July — but revenue grew even faster. Its response: cap usage, not cut deployment.
This means → soaring costs do not automatically trigger cost-cutting. When AI-driven revenue growth outpaces spend growth, companies choose to cap volume rather than slash contracts.
AI agents — programs that let AI complete multi-step tasks autonomously — use retry loops that can amplify a single workflow's token consumption by 10× to 100×. One client's single-department Claude bill exceeds that department's entire payroll, yet the department is still hiring.
The pay-per-token era has arrived — who feels the pain first?
Microsoft switched GitHub Copilot — its AI coding assistant — to full per-token billing on June 1. Summit attendees broadly agreed the move accelerated enterprise demand for cost optimization.
Yet roughly 60% of customers stick with the model they first deployed and never switch. This reflects strong inertia — migration costs and validation cycles keep most clients locked in.
Compute-routing platforms — middleware that automatically matches each task to the most cost-effective model — can cut client compute spend by 20–30% on average through scheduling and partial migration to cheaper in-house models.
This means → under pay-per-token rules, "saving your customer money" is itself a lucrative business. The value space for compute-routing products is wide open.
Just how tight is inference compute?
Moonshot's Kimi 3 hit its compute ceiling within 48 hours of launch and suspended new subscriber sign-ups. Databricks, which hosts Kimi and multiple open-source models, is GPU-short across several regions; its latest funding round is primarily earmarked for GPU procurement.
Fireworks' latest annualized revenue has topped $1 billion (up from $400 million in January), processing over 40 trillion tokens per day. It simultaneously closed a $1 billion round at a $17.5 billion post-money valuation. Baseten processes 30 trillion tokens daily; management says its traffic may already exceed the OpenAI official API.
On the closed-source side, Anthropic and Google recently locked in additional compute capacity from SpaceX. UBS supply-chain research indicates OpenAI and Anthropic will add hundreds of thousands of GW-scale compute reserves over the coming years.
In plain terms = open-source or closed-source, everyone is scrambling for GPUs. The inference compute gap has shifted from "might happen" to "happening now."
Why is Nvidia's Nemotron suddenly the talk of the summit — and what's still missing for open-source deployment?
Nvidia's Nemotron was the most frequently mentioned U.S.-origin open-source model at the summit; six months ago it was barely discussed. This reflects a rapid rise in Nvidia's open-source ecosystem presence after the December 2025 release of Nemotron 3.
But open-source deployment still faces hard constraints: one platform provider disclosed that 80% of its compute traffic still runs through Claude. Every client in finance, government, and other heavily regulated sectors refuses to deploy certain large models.
This means → the surge in open-source token share does not mean open-source is trusted in high-value scenarios. Regulatory compliance and model reliability remain closed-source moats.
Whether compute supply can keep pace with the exponential growth of token consumption will determine whether these inference platforms can sustain their lofty valuations — the central unresolved question in AI infrastructure today.
Content is for reference only, not financial advice.