OpenAI

OpenAI launches o3-pro for ChatGPT Pro and the API

On June 10, 2025, OpenAI shipped o3-pro to ChatGPT Pro and Team users plus the API, replacing o1-pro at $20 per million input and $80 per million output tokens.

OpenAI launches o3-pro for ChatGPT Pro and the API — article cover

On June 10, 2025, OpenAI released o3-pro, a souped-up version of its o3 reasoning model with a one-line pitch: spend more compute, get more reliable answers. ChatGPT Pro and Team users could pick it from the model picker the same day, where it replaced o1-pro; Enterprise and Edu users followed the week after, and the developer API went live that same afternoon.

Some background for the unfamiliar: conventional models generate an answer directly, while reasoning models work through problems step by step, which makes them more dependable in domains like physics, math, and coding. OpenAI had already proven the category with o3 and o4-mini in April; o3-pro pushed the compute-for-quality trade to the top of the product line. TechCrunch described it as a souped-up o3 — and noted OpenAI’s claim that it is the company’s most capable model yet.

Pricing: a tier below o1-pro

The API costs $20 per million input tokens and $80 per million output tokens. The comparison is brutal: when o1-pro entered the API in March 2025 it cost $150 per million input and $600 per million output — TechCrunch then called it OpenAI’s most expensive model, twice the price of GPT-4.5 on input and ten times regular o1 on output. o3-pro cut the entry price for frontier-grade reasoning to roughly a seventh of that.

That is the most practical news here for cost-conscious teams: deep reasoning moved from “special budget only” to something you can route high-value tasks to without a procurement conversation. TechCrunch notes a million input tokens is roughly 750,000 words — longer than War and Peace. At that volume, the per-million-token rate is what decides whether a project pencils out at all.

Reliability over speed

The evidence OpenAI offered in its changelog is expert evaluation: “In expert evaluations, reviewers consistently prefer o3-pro over o3 in every tested category,” with particular strength in science, education, programming, business, and writing help. Reviewers also rated o3-pro consistently higher for clarity, comprehensiveness, instruction-following, and accuracy.

A separate “4/4 reliability” test makes the claim concrete: answering the same question correctly four times in a row, o3-pro beat both o1-pro and standard o3. The cost is latency. OpenAI acknowledged that o3-pro responses typically take longer than o1-pro’s — a frank statement that this model sells correctness, not reaction time. Unlike automated benchmarks, this expert-preference format asks which answer is actually better, closer to how the model gets judged in real use. InfoQ’s one-line summary of the positioning: a model for users who prioritize correctness and depth over speed.

Tool use, and the fine print

o3-pro ships with a full toolset: web search, file analysis, reasoning about visual inputs, Python execution, and personalized responses using memory. That makes it closer to a research assistant that can go find sources and run code than a chat window.

The limitations are stated just as clearly. At launch, temporary chats were disabled while OpenAI resolved a technical issue, Canvas was not supported, and the model cannot generate images. OpenAI also noted that o3-pro uses the same underlying model as o3, so safety details live in the existing o3 system card rather than a new document. The developer docs add one more flag: predicted outputs are not supported on o3-pro, a gap worth noting for long-output cost optimization.

The developer view: when to pick it

For developers, o3-pro is the correctness-first option: hard problems, high-value decisions, and batch jobs where waiting is acceptable belong in its lane; low-latency interactive features should stay on other models. The developer docs expose an o3-pro-2025-06-10 snapshot, pinning behavior to the launch-day version so production runs stay reproducible, and rate limits scale by usage tier — Tier 1 allows 500 requests and 30,000 tokens per minute, while Tier 5 reaches 10,000 requests and 30 million tokens per minute.

Early user feedback ran decidedly mixed. In the reactions compiled by InfoQ, some saw it finally crossing the threshold on tasks that had just missed — “not game-changing, but it might cross the threshold,” as one commenter put it, with real productivity gains attached. Others complained it was taking “an awfully long time,” reported frequent timeouts on the Android and macOS apps, and questioned whether hallucinations were actually addressed. Judged from 2026, o3-pro marked a turning point where flagship deep reasoning moved toward pragmatic pricing: correctness became priceable, and the trade between latency and reliability became a question every API selection had to answer.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL