AGENTIC COMMONSAI industry briefings

繁中EN

TAGAI APIPUBLISHED 2026-10-10

All English articlesAI API

What a 1/8-Size KV Cache Means for Your Agent Bill

On this page6 sections
  1. The architecture bet: asymmetry over size
  2. Why the KV cache numbers matter for cost
  3. What changed for anyone currently calling deepseek-v4-pro
  4. Pricing mechanics worth exploiting
  5. A practical next step
  6. Sources

If you run long-context agents, the line item that quietly eats your budget is usually not the output tokens. It is the cache. DeepSeek’s announcement of V4.1-Flash on the DeepSeek API docs site puts that problem at the center of the release, and the numbers are specific enough to plan around.

The architecture bet: asymmetry over size

V4.1-Flash is a 552B-parameter mixture-of-experts model, but only a slice of that is active per request. The new Causal Encoder–Decoder setup uses 8B active parameters for input and 16B for output. DeepSeek says new pre-training methods plus larger-scale reinforcement learning post-training produce benchmark results ahead of its own flagship, including DeepSeek-V4-Pro.

That claim — a small model beating the flagship — is the interesting part for anyone doing model selection. It also comes with native visual understanding, which matters because DeepSeek retired the separate V4-Flash-Vision-Exp experiment alongside this release.

Why the KV cache numbers matter for cost

Compared with the previous generation, V4.1-Flash’s KV cache needs 1/4 the HBM and 1/8 the SSD storage. The announcement states plainly that cache-hit charges often account for a large share of agent costs, and that compressing the cache cuts those costs significantly.

If you have been tracking per-request economics on small models, this is the same discipline I ran through in my subagent cost math on Anthropic’s latest small model: the model price tag is the least of it. What determines whether an agent is viable is what a 200k-token conversation costs after the twentieth turn, and cache pricing is what that conversation mostly consists of.

What changed for anyone currently calling deepseek-v4-pro

A few migration facts worth putting in your calendar now:

  • V4.1-Flash is live on the DeepSeek API as deepseek-flash, with native multimodal support.
  • The old deepseek-v4-flash and deepseek-v4-flash-vision-exp endpoints temporarily route to V4.1-Flash for compatibility.
  • Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, continuing until V4.1-Pro launches.

DeepSeek cites tests by multiple parties putting V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime, and says it is phasing out V4-Pro. If V4-Pro is in your production stack, you do not need to do anything for requests to keep working — but you should verify outputs, since you will be on a different model whether you planned for it or not.

Pricing mechanics worth exploiting

The tiered peak/off-peak structure carries over: off-peak rates are 50% of peak rates, and DeepSeek explicitly suggests scheduling flexible workloads off-peak. The new pricing takes effect at 04:00 UTC on Sept 10, 2026. For batch-style pipelines — document processing, overnight summarization, data extraction — the off-peak window is free money if your scheduler can move.

For self-hosting teams, DeepSeek says it will work with the open-source community on V4.1-Flash inference support and is inviting conversations about large-scale deployments of 2,000 GPUs plus a storage cluster. The model weights and tech report are on the DeepSeek Hugging Face page, so the details behind the 1/8 storage claim are checkable rather than marketing copy.

A practical next step

The takeaway for builders: when a provider compresses the KV cache this aggressively, the cheapest architecture for your agent may shift — longer retained contexts and multi-turn loops get relatively cheaper than aggressive context pruning. Before rewriting anything, though, rerun your cost benchmark on deepseek-flash and compare real bills. The savings only materialize if your workload is actually cache-heavy, and the announcement does not publish a per-token price table in the section summarized here, so check the current pricing page before committing.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
Support us

Related reading

  1. Running the Subagent Math on Anthropic's Latest Small Model

    Claude Haiku 5.5 cuts Haiku-class costs by roughly 75% and adds an effort dial, reshaping which agent tasks stay on the cheap tier.

    Anthropic

  2. OpenRouter's $113M Series B Values Model Routing at $1.3B

    OpenRouter raised a $113M Series B led by CapitalG at a reported $1.3B valuation, processing 100 trillion tokens monthly across 400+ models — why routing now matters.

    OpenRouter

  3. DeepSeek Makes Its 75% V4 Pro Price Cut Permanent

    DeepSeek confirmed May 22, 2026 its 75% V4 Pro discount is permanent: $0.87 per million output tokens, a quarter of list price; Flash is $0.28. What it says about API economics.

    LLM