OpenAI

OpenAI's $10B Cerebras Deal Adds 750MW of Inference Compute

OpenAI signed a Cerebras deal worth over $10 billion: 750MW of wafer-scale compute through 2028 for low-latency ChatGPT inference. What the terms, the G42 risk, and the supplier mix mean.

OpenAI's $10B Cerebras Deal Adds 750MW of Inference Compute — article cover
On this page6 SECTIONS
  1. The Terms: 750 Megawatts, Through 2028
  2. Why Inference, Not Training
  3. Cerebras’ Bet: Diluting G42
  4. OpenAI’s Multi-Vendor Compute Portfolio
  5. What It Means for Developers and Product Teams
  6. Sources

On January 14, 2026, OpenAI announced a multiyear partnership with Cerebras, the wafer-scale chip startup: OpenAI will buy up to 750 megawatts of computing capacity delivered on Cerebras silicon, with deployments running through 2028. CNBC, citing people close to the company, pegged the value of the agreement at more than $10 billion — the startup’s flagship customer win ahead of a long-planned IPO.

What makes this different from the usual infrastructure headline is where the megawatts go. All 750 MW point at inference. OpenAI is buying latency, not training throughput, and it is buying it from the one company whose architecture exists specifically to make single requests fast.

The Terms: 750 Megawatts, Through 2028

The agreement is sized in power rather than chip counts. 750 MW is campus-scale — the draw of a large AI data center, not a few racks — and the capacity is hosted on data centers Cerebras operates across the United States. CNBC reported that the two companies signed a term sheet just before Thanksgiving 2025 and went public on January 14; the $10 billion-plus figure comes from sources close to Cerebras rather than the official posts.

Three details worth noticing. The contract covers up to 750 MW, so the real deployment curve depends on how quickly Cerebras can expand. The relationship is not new: OpenAI evaluated Cerebras technology as early as 2017, and this deal converts a long courtship into contracted capacity. And the timing lands right before Cerebras’ IPO, where a multiyear OpenAI contract does obvious work for the valuation story.

Why Inference, Not Training

Sachin Katti, who leads compute infrastructure at OpenAI, put the motivation plainly in the announcement: “Cerebras adds a dedicated low-latency inference solution to our platform. That means faster responses, more natural interactions, and a stronger foundation to scale real-time AI to many more people.”

Cerebras’ pitch is the wafer-scale engine: an entire wafer fabricated as a single dinner-plate-sized processor, which sidesteps the interconnect bottlenecks of clustered GPUs and pushes per-request throughput very high. That is not the mainstream architecture for training frontier models, but it fits the “one user needs an answer now” shape of inference almost perfectly.

The integration is also pre-validated. OpenAI’s gpt-oss open-weight models already run on Cerebras silicon, alongside Nvidia and AMD chips. This is not a from-scratch port; it is working technology promoted into contracted production capacity.

Cerebras’ Bet: Diluting G42

For Cerebras, the strategic value rivals the dollar figure. CNBC reported that in the first half of 2024, roughly 87 percent of Cerebras’ revenue came from a single customer: G42, the Abu Dhabi AI group. A company positioning itself as the Nvidia challenger, with revenue concentrated in one large Middle Eastern client, is exactly the story capital markets punish.

CEO Andrew Feldman’s framing was pragmatic: “The way you have three very large customers is start with one very large customer, and you keep them happy, and then you win the second one.” OpenAI is that second customer. With the IPO pending, the contract attacks the revenue-concentration question and the growth narrative in a single move.

OpenAI’s Multi-Vendor Compute Portfolio

From OpenAI’s side, this is one more deal that thins out single-supplier dependence. Nvidia became the first company to reach a $5 trillion market cap in October 2025, a measure of how concentrated the AI silicon market remains. In December, Nvidia agreed to a nonexclusive licensing deal with Groq valued at $20 billion in cash, which CNBC called the company’s largest transaction to date. The supply side is consolidating even as the buy side scrambles to diversify.

Cerebras gives OpenAI an inference path that shares nothing with the Nvidia architecture. gpt-oss already runs on Nvidia and AMD chips; wafer-scale systems now enter the formal procurement mix. For the first time at OpenAI, inference-dedicated compute is its own line item rather than a byproduct of training clusters.

What It Means for Developers and Product Teams

Three practical effects. First, latency-sensitive products — voice, real-time agent interaction, streaming UIs — now have a technology route with OpenAI’s own endorsement; wafer-scale inference belongs in the evaluation set. Second, compute procurement is splitting along the training/inference line, and cost models should split with it: training is bought by cluster scale, inference by latency and throughput, and the two can come from different suppliers. Third, supplier diversification has moved from slogan to contract behavior. From our 2026 opening outlook to this deal, the first month of the year keeps sending the same signal: infrastructure choices are product decisions.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL