OpenAI

OpenAI expects o3 successors to hit high bio-risk tier

On June 19, 2025, OpenAI's Johannes Heidecke said o3 successors are expected to reach the high bio-risk tier of its Preparedness Framework, with expanded pre-release safety testing planned.

OpenAI expects o3 successors to hit high bio-risk tier — article cover

On June 19, 2025, Johannes Heidecke, OpenAI’s head of safety systems, used an Axios interview, a company blog post, and his own X post to flag something unusual in advance: successors to OpenAI’s o3 reasoning model are expected to hit the “high” risk classification for biology under the company’s Preparedness Framework.

It is rare for a lab to announce a risk tier before a model ships. The message made bio-risk the AI safety talking point of the week, and it gave outsiders an unusually clear view of how much capability OpenAI itself expects from its next generation of reasoning models.

What the high classification means

The Preparedness Framework is OpenAI’s internal system for grading frontier model capabilities and deciding how they can be deployed, and biology is one of its core evaluation areas. When a model lands in the high-risk tier, it means the model could materially improve a person’s ability to develop or acquire biological weapons — and pre-release scrutiny and usage restrictions tighten accordingly.

Heidecke put it plainly: “We are expecting some of the successors of our o3 model to hit that level.” In other words, this was not a hypothetical scenario but expectation-setting about OpenAI’s own capability curve. Better to state the risk tier up front than answer hostile questions after launch.

He also narrowed down what kind of risk matters. The concern is not inventing novel weapons from scratch but what safety researchers call novice uplift: letting people with no scientific training reproduce lethal agents that “experts already are very familiar with.” When the barrier drops, proliferation starts. A handful of specialists always had access to this knowledge; the problem begins when everyone does. That is also why OpenAI is emphasizing testing and deployment control rather than promising the capability will not exist: capability will arrive, so the defenses have to get there first.

A near-perfection bar for safeguards

The response is expanded pre-release safety testing. Heidecke’s bar is extreme: a 99 percent pass rate, or even a one-in-100,000 failure rate, would not be enough, because the cost of a biological failure cannot be averaged away — a single successful misuse ends the thought experiment. “We basically need, like, near perfection,” he said.

In the blog post, OpenAI committed to boosting safety testing so models cannot help users create bioweapons, and capabilities that fail evaluation will not ship. For developers building on the API, that usually means some science-adjacent features get gated or delayed, and those limits belong on the product roadmap as a planning variable, not a surprise. In practice, the risk tier decides which features appear in the API and when — not just what the safety team documents internally.

The dual-use trade-off

Heidecke did not dodge the dual-use dilemma: the same scientific understanding that enables medical breakthroughs can also enable weaponization. His position was to acknowledge that the line is hard to draw, and to draw it anyway — state the risk first, then decide what gets released.

That is the shared condition of frontier labs in 2025: as reasoning models keep pushing scientific task performance upward, risk evaluation has to sprint to keep pace. When release cycles are measured in months, safety assessment has no leisurely option left. Announcing the expectation in June 2025 was itself an admission: the days when safety engineering trails capability are the days a lab is most exposed, and stating it early is the only rational move.

Anthropic already set the precedent

The reports pointed to a parallel move at a rival: Anthropic shipped Claude Opus 4 under the stricter ASL-3 protocols of its Responsible Scaling Policy over similar bioweapon misuse concerns. Fortune noted that earlier versions had complied with dangerous prompts — an issue Anthropic attributed to an accidentally omitted training dataset — before the ASL-3 release.

Taken together, the two companies’ statements show that biological risk is now a standard item in frontier model launch processes: risk tiers get previewed publicly, testing scales up before shipment, and only then are capabilities released. Looking back from 2026, the June 19, 2025 remarks mark a clear point where the capability race and safety engineering moved in lockstep — where risk communication itself became part of the launch. For developers across the ecosystem, reading these previews well beats waiting for the limits to surface at runtime.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL