Generative AI

Decart Oasis 3: An API-Served World Model for AV Training

Decart ships Oasis 3, the first world model callable by API: 22 FPS three-camera driving scenes at $0.02/sec, a $300M raise behind it, and real caveats in hands-on testing.

Decart Oasis 3: An API-Served World Model for AV Training — article cover
On this page6 SECTIONS
  1. What Oasis 3 Is
  2. Performance Numbers and Pricing
  3. The Caveats Found in Testing
  4. The Business Bet
  5. Where It Lands in the World-Model Race
  6. Sources

On June 10, 2026, Decart, a startup founded in 2023, released Oasis 3: an interactive world model you can call over an API. Feed it a text prompt and it generates a photorealistic driving environment in real time, one that keeps evolving in a closed loop as a robot acts inside it. Decart bills it as “the first world model accessible by API, live from day one,” and CEO Dean Leitersdorf’s launch line goes further: “Every robot will be trained inside a world model. That era starts today.”

Why it matters: world models are graduating from lab demos to metered compute. For autonomous vehicle teams, a generator of long-tail scenarios — snowy nights, tunnels, obstructed cameras — now has a public price list for the first time.

What Oasis 3 Is

Oasis 3 sits on top of Decart’s real-time video foundation model, Lucy, and its inference stack, DOS (Decart Optimization Stack). The architecture is autoregressive “Live Stream Diffusion”: frames are generated one at a time, conditioned on prior frames and on robot actions. Where the original Oasis took keyboard and mouse input, Oasis 3 is robot-native — conditioned directly on action signals — and uses a geometry-aware architecture to output three synchronized camera views (one front-facing, two side-facing), matching how AV perception stacks are actually wired.

Environments are prompt-driven and generate without an upper bound; the company claims drift correction over long sessions. Training used NVIDIA’s Physical AI Open Dataset on Hugging Face, and the model runs on CoreWeave clusters of NVIDIA HGX B200 systems. Decart names four target uses: online reinforcement learning inside the generated world, policy evaluation before deployment, high-fidelity teleoperation to harvest expert training data, and simulating rare events that are unsafe or impractical to stage in reality.

Performance Numbers and Pricing

The published specs: 22 FPS at 512×768×3 resolution (three cameras), latency under 200 milliseconds, and a claimed ~100x efficiency gain from DOS. API pricing is $0.02 per second, with enterprise terms by use case. Leitersdorf gives a concrete sense of scale: every frame is roughly 8,000 tokens, so at tens of frames per second the system is pushing hundreds of thousands of tokens per second. Users can interact with an environment for hours, and Decart says its community already tops 100,000 developers, most of them building on Lucy.

The Caveats Found in Testing

TechCrunch’s launch-day hands-on was considerably cooler. A New York City prompt drifted into a generic Western city until the original intersection vanished entirely. Steering control went unresponsive for stretches. Cars passed through other cars — physics is not really simulated. The reviewer’s verdict: a “dream-like, disjointed stream of consciousness” that turns nonsensical over time. The autoregressive design also fills the context window fast; the team concedes that longer memory and token compression are open research, and the next version may let users start from a video rather than a single image.

Leitersdorf’s explanation of the physics gap is candid: there is drastically more data on good driving than on accidents, and that imbalance is, in his words, “a major research problem that we’re cracking now.”

The Business Bet

The real strategy is the ecosystem. Leitersdorf says the moment reminds him of the early days of LLMs, “when OpenAI invented the API for models,” and predicts “there’s going to be an entire developer community that emerges on top of this.” The war chest is stocked: a $300 million round announced in mid-May and led by Radical Ventures brought in NVIDIA, Toyota Ventures, Adobe Ventures, and eBay Ventures, with Sequoia and Benchmark returning; TechCrunch puts the valuation near $4 billion, total raised above $450 million. Leitersdorf adds the company has burned “drastically less” than $100 million in its lifetime.

Where It Lands in the World-Model Race

The 2026 world-model track is crowded: Google’s Genie 3, World Labs’ Marble, Runway, and Luma are all positioned here, a lane we surveyed in our opening-of-2026 outlook. Decart’s differentiation is vertical integration — its own models plus an inference stack that runs across NVIDIA, Google, and Amazon silicon — which Leitersdorf claims makes it “more than an order of magnitude cheaper” than rivals. Whether the API-ification of world models replays the LLM commercialization script will be visible within a few quarters. Until then, treat “programmable world” as a direction, not a destination: physical consistency remains the shared, unsolved homework for everyone on the track.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL