IBM and Together AI Sign $240M Deal to Build NVIDIA AI Inference Cluster on IBM Cloud

Nashnova编辑部
Published 2026-08-11About 10 min read

IBM and Together AI signed a $240 million multi-year deal to build a large-scale AI inference cluster on IBM Cloud, powered by Nvidia's latest Blackwell chips and dedicated to open-source models — one of the biggest single commercial bets yet that the future of AI inference runs on open source.

01

What does this deal actually buy?

IBM and Together AI will jointly build a large-scale AI inference cluster — infrastructure purpose-built for running trained AI models and generating responses — hosted on IBM Cloud.
The hardware core is the Nvidia HGX B300 system, featuring Nvidia's latest-generation Blackwell processors with Spectrum-X Ethernet networking. This means → the cluster uses Nvidia's top-tier inference stack from chip to network.
In plain terms = IBM supplies the venue and cloud platform, Together AI brings the technology and customers, Nvidia provides the newest hardware — each party contributes its strongest card.
02

Who is Together AI, and why them?

Together AI, based in San Francisco, lets enterprises run AI workloads on open-source models — DeepSeek, MiniMax, Kimi and others — at lower cost than closed-source systems.
After its latest funding round in July, the company's valuation reached $8.3 billion. This reflects a rapid rise in capital-market confidence in the "open-source inference platform" positioning.
In plain terms = it operates as a cloud-based managed service for open-source AI models — enterprises pay to use them without buying GPUs or handling deployment themselves.
03

Why has inference suddenly become the biggest compute bottleneck?

Inference is the "use the model" stage of the AI pipeline — every answer generated, every image produced after training consumes inference compute. This means → the more users and the more frequent the calls, the faster inference demand scales.
Cloud providers and chipmakers have already poured billions of dollars into expanding inference infrastructure. Nvidia has stated explicitly that its Blackwell chips are optimized specifically for inference workloads.
In plain terms = training is "teaching the model"; inference is "putting the model to work." There are enough models now — the bottleneck has shifted to the working stage.
04

What went wrong with closed-source models that made open source an alternative?

Anthropic, OpenAI, and Meta have all experienced cybersecurity incidents with their closed-source models in recent months, heightening enterprise concerns over data security and vendor lock-in.
This reflects a structural trend: once trust in a closed-source provider's security is broken, enterprises actively seek auditable, self-deployable open-source alternatives.
This means → the open-source inference lane where Together AI operates is absorbing demand that flows out of the closed-source ecosystem.
05

Can this order really mark a milestone for open-source inference?

$240 million makes this one of the largest commercial contracts in the open-source inference space to date — but the contract is multi-year, and its ultimate value depends on Together AI's actual compute utilization rate over the contract term.
In plain terms = the money is signed, but whether it gets fully spent — and spent well — hinges on how many enterprise clients Together AI can attract to keep running models on the cluster.
This reflects a sector transitioning from "technically viable" to "commercially validated at scale" — this order is the entry ticket, not the finish line.

Content is for reference only, not financial advice.