AI Infrastructure

Baseten's $1.5B Round Bets Big on Open-Source Inference

Baseten is reportedly raising $1.5B at up to a $13B valuation, five months after a $300M Series E at $5B. Inside the split-priced round and the open-source inference bet.

Baseten's $1.5B Round Bets Big on Open-Source Inference — article cover
On this page6 SECTIONS
  1. The Timeline Behind a 160% Jump in Five Months
  2. The Split-Priced Round: Two Prices for the Same Equity
  3. What Baseten Actually Sells: Inference as a Business
  4. The Risks Behind the Gold Rush
  5. What It Means for Engineering and Product Teams
  6. Sources

On June 18, the Wall Street Journal reported that Baseten, an AI inference infrastructure startup, is close to finalizing a $1.5 billion round at a valuation of up to $13 billion. That is barely five months after its $300 million Series E in January, which valued the company at $5 billion — a paper jump of about 160%.

The money is flowing into the hottest layer of the AI stack: inference. As model training concentrates in a handful of labs, “how do we run models fast and cheap in production” has become the daily question for every product team — and the wave of venture money chasing answers is what The Next Web dubbed the “inference gold rush.”

The Timeline Behind a 160% Jump in Five Months

Line up Baseten’s recent rounds and the velocity becomes obvious:

  • Mid-2025: a $150 million Series D, nine months before the Series E
  • September 2025: valued around $2.15 billion (per TNW)
  • January 2026: $300 million Series E at a $5 billion valuation, with IVP, CapitalG, and Nvidia participating
  • June 2026: reportedly raising $1.5 billion at up to $13 billion

In under a year, the valuation has gone from just over $2 billion to $13 billion — roughly six times.

The Split-Priced Round: Two Prices for the Same Equity

The WSJ describes a split-priced structure: some investors enter at $13 billion, others at $11 billion. TechCrunch’s explainer from March noted this tactic is often used to inflate headline valuations and flatter lead investors on paper — the true blended price lands somewhere between the two numbers. The round is co-led by Spark Capital, Sands Capital, Altimeter Capital, and Wellington Management; TNW’s roster also lists Conviction.

What Baseten Actually Sells: Inference as a Business

Founded in 2019 and based in San Francisco, Baseten sells inference infrastructure: it leases capacity from roughly 20 cloud providers and layers its own inference software on top, letting customers deploy and fine-tune models on their own data. The differentiation is routing — steering each request to a model that is good enough and cheap, with a deliberate bias toward capable open-source options over defaulting to OpenAI or Anthropic frontier models. The pitch is simple: as inference turns into a commodity input, whoever shaves the most off each token — by picking the right model, the right hardware, and the right cloud for each request — wins the account. Customers include Cursor, Mercor, and OpenEvidence; one customer reportedly ran a task at about 30% of what the proprietary-model alternative would have cost. CEO Tuhin Srivastava told the WSJ that “open-source models are getting very, very good,” and that customers increasingly mix open and proprietary models depending on task difficulty.

The Risks Behind the Gold Rush

The backdrop to all this is a brutal price war. Chinese open-source models from DeepSeek and Moonshot AI, plus Nvidia’s open Nemotron family, keep driving the per-token cost of inference down. In the same market, Cerebras has gone public and Fireworks AI competes directly, while the hyperscalers whose capacity Baseten leases also sell managed inference of their own. TNW flags the core tension: a $13 billion valuation assumes Baseten survives the margin squeeze that the price war creates — the company benefits as open models get better, and is equally exposed as they get cheaper. A broker that makes money on the spread between expensive and cheap models has a hard ceiling when everything converges toward cheap.

What It Means for Engineering and Product Teams

Three signals. First, multi-model routing is now the default architecture: deciding which request goes to a frontier model and which to an open one has become a product-level cost lever. Second, open-source models have been production-validated by the heaviest users — when an inference-intensive product like Cursor trusts them with its traffic, that is the real stress test. Third, the open-source supply side is not guaranteed: Meta’s Avocado frontier model has been reported to possibly go closed-source. Baking your cost structure onto a single open-source lineage is now no less risky than betting everything on a single proprietary model.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL