AGENTIC COMMONSAI industry briefings

繁中EN

TAGAI APIPUBLISHED 2026-10-10

All English articlesAI API

Scheduling Around Half-Price Tokens: DeepSeek-V4-Pro Ships

DeepSeek announced the GA launch of DeepSeek-V4-Pro on August 13, 2026, and the part that will change how you build is not the model itself — it is the pricing structure that arrives with it. Starting at 16:00 UTC on August 16, 2026, the API moves to peak and off-peak rates, with off-peak exactly 50% cheaper than peak. That is a real lever for anyone running batch jobs, eval sweeps, or scheduled agent workflows.

What actually shipped

According to the DeepSeek API docs announcement, V4-Pro is available on app and web through an “Expert Mode” toggle, and via the API with unchanged model names — so existing integrations should keep working without a rename. The announcement highlights three things:

  • Agent-focused upgrades, described by DeepSeek as “strong production gains” (their characterization, not a benchmark result I can verify here)
  • Flexible reasoning effort across V4-Pro and V4-Flash: low for simple tasks, high for daily agent workflows, max for complex ones
  • Native OpenAI Responses API support, optimized for Codex with one-click setup

The reasoning-effort dial is the quiet standout. If you already route between models by task difficulty, an in-model effort setting may let you collapse some of that routing logic into a single model choice.

The off-peak opportunity

Time-tiered pricing rewards builders who control when work runs. Jobs that map well to off-peak windows: nightly document processing, backfilling embeddings, replaying agent traces for evaluation, large-scale data extraction. Jobs that do not: anything user-facing where latency windows are dictated by traffic.

If your stack already queues heavy work, the practical move is to tag which pipeline stages tolerate a delay, then schedule those against off-peak rates. The supplied announcement does not specify the exact peak and off-peak hour boundaries or the base prices, so check the API docs before re-planning your budget.

This pairs naturally with the caching math I covered earlier — in what a 1/8-size KV cache means for your agent bill, the lever was shrinking per-request cost; here the lever is shifting when requests run. Combine both and your effective rate drops twice.

What to verify before committing

A few gaps matter for planning. The supplied RSS summary does not specify how the peak schedule is defined (which timezone, which hours), whether off-peak applies to both input and output tokens equally, or what the agent upgrades measure against. Treat “production gains” as a claim to test with your own workload, not a spec.

A sensible first step: point a low-risk batch job at V4-Pro during an off-peak window, compare output quality and token counts against your current setup, and only then decide whether the effort settings let you retire a smaller model from your rotation.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
SUPPORT US

Related reading

  1. What a 1/8-Size KV Cache Means for Your Agent Bill

    DeepSeek-V4.1-Flash cuts KV cache HBM needs to a quarter and starts retiring V4-Pro, changing the math on cache-heavy agent workloads.

    AI API

  2. Running the Subagent Math on Anthropic's Latest Small Model

    Claude Haiku 5.5 cuts Haiku-class costs by roughly 75% and adds an effort dial, reshaping which agent tasks stay on the cheap tier.

    Anthropic

  3. OpenRouter's $113M Series B Values Model Routing at $1.3B

    OpenRouter raised a $113M Series B led by CapitalG at a reported $1.3B valuation, processing 100 trillion tokens monthly across 400+ models — why routing now matters.

    OpenRouter