If you run long-context agents, the line item that quietly eats your budget is usually not the output tokens. It is the cache. DeepSeek’s announcement of V4.1-Flash on the DeepSeek API docs site puts that problem at the center of the release, and the numbers are specific enough to plan around.
The architecture bet: asymmetry over size
V4.1-Flash is a 552B-parameter mixture-of-experts model, but only a slice of that is active per request. The new Causal Encoder–Decoder setup uses 8B active parameters for input and 16B for output. DeepSeek says new pre-training methods plus larger-scale reinforcement learning post-training produce benchmark results ahead of its own flagship, including DeepSeek-V4-Pro.
That claim — a small model beating the flagship — is the interesting part for anyone doing model selection. It also comes with native visual understanding, which matters because DeepSeek retired the separate V4-Flash-Vision-Exp experiment alongside this release.
Why the KV cache numbers matter for cost
Compared with the previous generation, V4.1-Flash’s KV cache needs 1/4 the HBM and 1/8 the SSD storage. The announcement states plainly that cache-hit charges often account for a large share of agent costs, and that compressing the cache cuts those costs significantly.
If you have been tracking per-request economics on small models, this is the same discipline I ran through in my subagent cost math on Anthropic’s latest small model: the model price tag is the least of it. What determines whether an agent is viable is what a 200k-token conversation costs after the twentieth turn, and cache pricing is what that conversation mostly consists of.
What changed for anyone currently calling deepseek-v4-pro
A few migration facts worth putting in your calendar now:
- V4.1-Flash is live on the DeepSeek API as
deepseek-flash, with native multimodal support. - The old
deepseek-v4-flashanddeepseek-v4-flash-vision-expendpoints temporarily route to V4.1-Flash for compatibility. - Starting at 04:00 UTC on Sept 14, 2026, all
deepseek-v4-prorequests route to V4.1-Flash at V4.1-Flash rates, continuing until V4.1-Pro launches.
DeepSeek cites tests by multiple parties putting V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime, and says it is phasing out V4-Pro. If V4-Pro is in your production stack, you do not need to do anything for requests to keep working — but you should verify outputs, since you will be on a different model whether you planned for it or not.
Pricing mechanics worth exploiting
The tiered peak/off-peak structure carries over: off-peak rates are 50% of peak rates, and DeepSeek explicitly suggests scheduling flexible workloads off-peak. The new pricing takes effect at 04:00 UTC on Sept 10, 2026. For batch-style pipelines — document processing, overnight summarization, data extraction — the off-peak window is free money if your scheduler can move.
For self-hosting teams, DeepSeek says it will work with the open-source community on V4.1-Flash inference support and is inviting conversations about large-scale deployments of 2,000 GPUs plus a storage cluster. The model weights and tech report are on the DeepSeek Hugging Face page, so the details behind the 1/8 storage claim are checkable rather than marketing copy.
A practical next step
The takeaway for builders: when a provider compresses the KV cache this aggressively, the cheapest architecture for your agent may shift — longer retained contexts and multi-turn loops get relatively cheaper than aggressive context pruning. Before rewriting anything, though, rerun your cost benchmark on deepseek-flash and compare real bills. The savings only materialize if your workload is actually cache-heavy, and the announcement does not publish a per-token price table in the section summarized here, so check the current pricing page before committing.
