DeepSeek announced the GA launch of DeepSeek-V4-Pro on August 13, 2026, and the part that will change how you build is not the model itself — it is the pricing structure that arrives with it. Starting at 16:00 UTC on August 16, 2026, the API moves to peak and off-peak rates, with off-peak exactly 50% cheaper than peak. That is a real lever for anyone running batch jobs, eval sweeps, or scheduled agent workflows.
What actually shipped
According to the DeepSeek API docs announcement, V4-Pro is available on app and web through an “Expert Mode” toggle, and via the API with unchanged model names — so existing integrations should keep working without a rename. The announcement highlights three things:
- Agent-focused upgrades, described by DeepSeek as “strong production gains” (their characterization, not a benchmark result I can verify here)
- Flexible reasoning effort across V4-Pro and V4-Flash: low for simple tasks, high for daily agent workflows, max for complex ones
- Native OpenAI Responses API support, optimized for Codex with one-click setup
The reasoning-effort dial is the quiet standout. If you already route between models by task difficulty, an in-model effort setting may let you collapse some of that routing logic into a single model choice.
The off-peak opportunity
Time-tiered pricing rewards builders who control when work runs. Jobs that map well to off-peak windows: nightly document processing, backfilling embeddings, replaying agent traces for evaluation, large-scale data extraction. Jobs that do not: anything user-facing where latency windows are dictated by traffic.
If your stack already queues heavy work, the practical move is to tag which pipeline stages tolerate a delay, then schedule those against off-peak rates. The supplied announcement does not specify the exact peak and off-peak hour boundaries or the base prices, so check the API docs before re-planning your budget.
This pairs naturally with the caching math I covered earlier — in what a 1/8-size KV cache means for your agent bill, the lever was shrinking per-request cost; here the lever is shifting when requests run. Combine both and your effective rate drops twice.
What to verify before committing
A few gaps matter for planning. The supplied RSS summary does not specify how the peak schedule is defined (which timezone, which hours), whether off-peak applies to both input and output tokens equally, or what the agent upgrades measure against. Treat “production gains” as a claim to test with your own workload, not a spec.
A sensible first step: point a low-risk batch job at V4-Pro during an off-peak window, compare output quality and token counts against your current setup, and only then decide whether the effort settings let you retire a smaller model from your rotation.
