Every agent loop spends a surprising share of its turns on tiny questions: is this shell command reversible, which tool comes next, does this page belong in the answer. A chat model handles each one by generating prose, which means a parser, a schema validator, and a retry path — before your actual policy code runs. On September 29, 2026 at DevDay, OpenAI named a fix: the Decisions API, a limited-preview endpoint that constrains GPT-6 Luna to developer-defined questions with finite pre-defined answers. Fourteen days earlier, TypeSafe AI had shipped Jev, a model built for exactly this job. The Firecrawl comparison of the two is worth studying because it shows two very different architectures pointed at the same bottleneck.
What a bounded decision endpoint actually changes
OpenAI’s entire public description is one paragraph in the DevDay recap: developers supply context as text or images, ask questions with finite answers, and get back selections they can use to classify content, route requests, or choose an agent’s next action. All three named jobs are control flow. Sam Altman demoed it driving a computer-use agent on stage, and the DevDay slide claims 150 ms versus 1.6 s for a standard GPT-6 Luna call — a 10x speedup, but against OpenAI’s own cheapest model doing the job the slow way.
The structural win isn’t just speed. Prompt-and-parse code tends to grow into a string-handling codebase nobody wants to own. Bounding the answer space before inference removes the parser, the retry branch, and the whole class of bugs where the model answers the right question in the wrong format. That’s distinct from structured outputs, which constrain the shape of text after the model has already decided what to say.
The uncomfortable detail: no public contract
As of September 30, 2026, OpenAI has published no request schema, response schema, endpoint path, SDK method, or pricing. Hiba Fathima at Firecrawl checked the API guides index, the endpoint reference, the combined docs export, the OpenAPI spec, and the Python, Node, and Go SDKs — none include it. The DevDay recap’s Decisions section is the only announcement with no “Learn more” link.
Practical implication: any code sample showing an OpenAI Decisions request body right now is invented. If you’re evaluating this for production, the honest move is to wait for the schema, the way you’d wait for published pricing before reworking how a KV cache change reshapes your agent bill — unit economics you can’t verify aren’t unit economics.
Jev, the incumbent that never learned to write
Jev launched September 15, 2026 from TypeSafe AI, founded by former OpenAI researcher Diogo Almeida. It’s a purpose-built “System One” model trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions — it architecturally cannot generate text, so output tokens bill at $0, against $0.042 per million input tokens. It exposes three documented primitives (Choice, Score, Noul) with per-option probabilities, mixable in a single call, and it’s generally available with full schemas and SDKs.
The tradeoffs are clear. Jev is text-only; OpenAI’s endpoint takes images, which matters if your decision point is a screenshot — the exact computer-use scenario Altman demoed. On latency, OpenAI’s 150 ms is a vendor claim under unstated conditions, while OpenRouter’s production telemetry puts Jev at a 0.21 s P50 and 0.34 s P95. Firecrawl’s takeaway is reasonable: treat them as the same ballpark until someone benchmarks both.
What the community made of it
Reception was muted compared with Jev’s launch. In the fourteen days between the two announcements, Hacker News front-paged six community reimplementations — including “Jev in 25 Lines of Python” at 691 points — so OpenAI’s version read as the seventh, and the first without a schema. Some argued the two-week gap proved there’s no durable moat in the category; others read OpenAI’s entry as validation that Jev’s category is real. Not everyone was sold on the framing at all: one Hacker News commenter pushed back that a bounded answer space only guarantees formatting, not correctness, and that grammars on any LLM get you the same thing.
How to decide while the dust settles
If your decisions run on screenshots, OpenAI’s image input is the differentiator — but you’re building against an undocumented preview. If your decisions are text and you need per-option probabilities today, Jev is the only one of the two with a contract you can code against, and its $0 output pricing is straightforward to model. Either way, keep authority in your own policy code: a 0.97 confidence score on “this looks safe” shouldn’t grant an account permissions it didn’t already have. The grounded next step is small — pick one high-frequency yes-or-no branch in your agent loop, measure the current round trip, and re-run the math once a schema and price are actually published.
