AGENTIC COMMONSAI industry briefings

繁中EN

TOPICOpen Models & Open WeightsPUBLISHED 2026-10-11

All English articlesLLM

Two DeepSeek V4 Preview Models Put 1M Context at the Default

The most practical change in DeepSeek’s V4 preview isn’t a leaderboard number. It’s that 1M tokens of context is now the default across all official DeepSeek services, not a premium tier you have to opt into. If you’ve been designing around 128K ceilings, the architecture assumptions in your agent loops just changed.

Two models, two jobs

The preview ships as a pair, both open-weighted on Hugging Face:

  • DeepSeek-V4-Pro: 1.6T total parameters, 49B active. DeepSeek positions it as open-source SOTA on agentic coding benchmarks, leading all current open models on world knowledge and math/STEM/coding while trailing only Gemini-3.1-Pro on knowledge.
  • DeepSeek-V4-Flash: 284B total, 13B active. Reasoning “closely approaches” V4-Pro per DeepSeek, matches it on simple agent tasks, and is faster and cheaper per call.

That split matters more than the totals. A 49B-active parameter count for Pro versus 13B for Flash means Flash’s cost advantage is structural, not promotional. You can read Flash as DeepSeek’s answer to the tiering question most providers solve with pricing alone: instead of two price points on one model, you get two genuinely different models. That’s the same instinct behind the KV-cache cost changes worth watching on the Flash line — the savings live in the model design, and they compound over long agent sessions.

What enables the 1M default

The technical claim behind the long context is a combination of token-wise compression and DSA (DeepSeek Sparse Attention), which DeepSeek says delivers “world-leading long context with drastically reduced compute and memory costs.” The company has a track record here — sparse attention and cache compression were central to how earlier versions made long context affordable — so the claim is plausible, though the tech report is the right place to verify the actual throughput and memory numbers before you architect around them.

The migration path, and one deadline

For existing integrations, this is about as painless as a model swap gets. Keep your base URL, change the model string to deepseek-v4-pro or deepseek-v4-flash. Both support the OpenAI ChatCompletions format and the Anthropic API, and both run in Thinking and Non-Thinking modes.

The deadline is the part to put in your calendar: deepseek-chat and deepseek-reasoner become fully inaccessible after July 24, 2026, 15:59 UTC. Right now those legacy names route to V4-Flash (non-thinking and thinking respectively), so many integrations already behave like Flash without knowing it. Before the cutoff, audit which model name your production calls use and decide whether to pin to deepseek-v4-flash explicitly — pinning makes your behavior immune to future routing changes.

If you’re already weighing V4-Pro against its pricing, note that a GA pricing step for the Pro model is worth reading alongside this preview — the preview tells you what the models can do; the pricing tells you what they cost to run.

A builder’s read

What’s genuinely new for product work:

  1. Long context as baseline. When 1M is the default, retrieval design shifts. You can lean harder on stuffing and less on aggressive chunking for documents that fit — though cost per call still argues for retrieval on anything recurring.
  2. Agentic-first positioning. DeepSeek says V4 integrates with Claude Code, OpenClaw, and OpenCode, and that it’s already used for in-house agentic coding. The Flash model matching Pro on simple agent tasks is the detail worth testing against your own workload.
  3. Open weights as a hedge. With weights downloadable, you can benchmark on the API today and self-host later if your unit economics demand it.

One caution from the source itself: DeepSeek notes attention around the release and asks users to trust only official accounts for news — reasonable advice in general when a model family gets this much buzz. Start with a Flash side-by-side against your current model on your actual agent traces, and check the July 24 cutoff before anything else ships.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
SUPPORT US

Related reading

  1. Scheduling Around Half-Price Tokens: DeepSeek-V4-Pro Ships

    DeepSeek-V4-Pro reaches GA with tiered reasoning effort, peak/off-peak pricing, and OpenAI Responses API support.

    AI API

  2. Two-Tier Token Pricing and Effort Knobs Change How You Route Cheap Work

    Haiku 5.5 adds tiered long-context pricing and per-task effort controls, reshaping how you route high-volume agent work.

    Claude

  3. What a 1/8-Size KV Cache Means for Your Agent Bill

    DeepSeek-V4.1-Flash cuts KV cache HBM needs to a quarter and starts retiring V4-Pro, changing the math on cache-heavy agent workloads.

    AI API