The headline from DeepSeek's September 10, 2026 release is a numbers inversion that does not usually happen. DeepSeek-V4.1-Flash, a model with 552 billion backbone parameters and only 8 billion active during input processing, now beats DeepSeek-V4-Pro across every published benchmark, and DeepSeek is routing all V4-Pro API traffic to it starting September 14. The official announcement and the model card on HuggingFace are the primary sources for the architecture and the numbers that follow. The Pro sits at 1.6 trillion total parameters and 49 billion active. The Flash beats a model with roughly one third the total parameters and one sixth the active parameters.
This is not a re-post-training of an existing architecture. The July 31 refresh kept the same decoder-only structure. V4.1 Flash is a ground-up redesign, the first in DeepSeek's new Causal Encoder-Decoder family, and its efficiency gains come from structural choices rather than sheer scale. For anyone who has ever paid to serve a long-context agent over a month, that distinction is the whole story.
The asymmetric split between input and output
Most modern language models are decoder-only Transformers. Every layer computes and stores its own key-value states during a forward pass, and those states pile up as context grows. DeepSeek-V4.1-Flash splits its 40 layers into two halves of 20: a causal encoder on the input side and a causal decoder on the output side. Both remain left-to-right attention, not bidirectional, so nothing changes about the autoregressive decoding path.
The important engineering move is where the decoder gets its KV cache. Instead of deriving key and value projections from each of its own 20 layers, the decoder pulls them from a single projection of the encoder's final hidden state. DeepSeek draws the parallel to a research design called YOCO (You Only Cache Once). During prefill the input tokens pass through all 20 encoder layers and the last layer produces a hidden state of shape $H_E$. That single tensor is linearly projected into a global KV cache of shape $H_{KV}$ that all 20 decoder layers attend to. During decode, each new token flows through the encoder once more, gets projected, and is appended to that same global cache.
This is where the activation asymmetry comes from. The prompt reading path activates only 8 billion parameters per token because it rides the lean encoder. The generation path activates 16 billion per token because it runs the full encoder plus all decoder layers. That split is deliberate. Agent workloads, code repositories, and document-reading pipelines are overwhelmingly input heavy. You read a large context once and generate a small response. Paying more per token to generate is fine when output is short.
Why the KV cache is the actual product
The benchmark numbers are the marketing copy. The KV cache is the reason the model exists. For an agentic session that holds a codebase, tool history, and screenshots in a single context, memory cost during inference dominates the bill, and that cost grows linearly with both sequence length and layer count.
DeepSeek's model card reports a global KV cache footprint of about 890 bytes per token. The previous V4 Flash sat at 3,456 bytes per token. That is a ratio of 0.257, or roughly one quarter. Over four generations the reduction reads as about 437-fold relative to DeepSeek's original V1 model. For a deeper layer-by-layer breakdown of every component, the technical writeup on Local AI Zone traces each piece back to the HuggingFace model card.
The compression stacks in layers, and each piece is worth understanding in isolation because they compound:
- Causal Encoder-Decoder projection removes the 20 separate decoder KV tensors down to one. This is the single largest lever.
- Compressed Sparse Attention 2 (CSA2) assigns each attention layer a fixed mode, either Full, Reindex, or Reuse. Reindex layers pick the top-K most relevant tokens using a two-stage hierarchical indexer, dropping attention cost from $O(n^2)$ to $O(K \cdot n)$. Reuse layers reuse the token indices chosen by the nearest previous Reindex layer, so the expensive indexer runs once and downstream layers inherit the result. This gives cross-layer KV sharing that the model card estimates reduces the effective number of unique KV caches from 40 to roughly 12.
- The KV values themselves are stored in FP4 using the E2M1 encoding, one sign bit, two exponent bits, one mantissa bit. That halves the footprint relative to FP8.
- SWA Bounded Replay reconstructs sliding-window attention KV states by replaying the forward pass for only the most recent $n_{win}$ tokens rather than persisting the whole window. That pushes persistent cache storage for those layers down to about one eighth of the V4 Flash requirement.
The quality cost is small and quantified. On GPQA Diamond the score drops from 91.1 under FP8 to 90.9 under FP4, a 0.2-point degradation. That is a deliberate engineering trade, and DeepSeek is publishing the number rather than hiding it, which is more than most vendors do.
The memory math that drives throughput
The practical consequence of 890 versus 3,456 bytes per token shows up as throughput on fixed hardware. Consider a single request with a 1 million token context. KV cache VRAM per request at the older footprint is about 3.2 gigabytes. At the new footprint it drops to roughly 850 megabytes. On the same GPU that is close to a 4x jump in how many concurrent long-context requests a server can hold, which maps almost directly to lower per-request serving cost.
The model card itself frames the trade with a concrete figure. Generating one new token against a 1 million token context at 890 bytes per token costs about 890 bytes of HBM traffic for the cache read. At the old 3,456 byte figure the same step moves 3.46 kilobytes. That difference is pure memory bandwidth, and long-context decoding is memory bound, not compute bound. When you are memory bound, smaller KV traffic is faster and denser, period.
Benchmarks: a mixed pattern, not a sweep
The honest reading of the benchmark table is that V4.1 Flash wins on the tasks that matter for agents and loses slightly on the knowledge-heavy tasks where a bigger model still has an edge. This mixed pattern is far more informative than a clean sweep.
| Model | Release | Architecture | Total params | Active | GPQA Diamond | HLE w/tools | Codeforces | DeepSWE v1.1 | Terminal-Bench 2.1 |
|---|---|---|---|---|---|---|---|---|---|
| V4 Flash | ~July 2026 | Decoder-only, CSA1 | 1.6T | 49B | 88.7 | 47.2 | 3,156 | 54.4 | 82.7 |
| V4 Pro | August 16, 2026 | Decoder-only, CSA1 | 1.6T | 49B | 92.4 | 60.0 | 3,348 | 62.7 | not listed |
| V4.1 Flash | September 10, 2026 | CED, CSA2 | 748B | 8B/16B | 90.9 | 63.9 | 3,471 | 74.2 | 90.6 |
Source: DeepSeek model card and changelog as compiled by independent technical writeups published September 10-11, 2026. All scores are vendor-reported.
Notice the shape. GPQA Diamond, a knowledge reasoning test, sits at 90.9, just under the V4-Pro 92.4. But HLE with tools jumps to 63.9 from the V4-Pro 60.0, a 3.9 point gain, and DeepSWE v1.1, an agentic coding task, jumps from 62.7 to 74.2, a 11.5 point gain that matches Claude Opus 5. Terminal-Bench 2.1 lands at 90.6, beating Opus 5 at 89.1. The model is competitive where it is meant to be, tool use, long context, agentic loops, and a touch weaker on pure knowledge recall. That is the signature of an architecture optimized for the input heavy, tool driven workflow rather than for trivia.
The pricing and the cost per completed task
The API pricing is where the architecture pays for itself. Off-peak rates are $0.15 per million uncached input tokens, $0.003 per million cached input tokens, and $0.60 per million output tokens. Peak pricing doubles those numbers. Peak hours are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC on weekdays. The 1 million token context window supports up to 384K tokens of output per request.
For comparison, the same DeepSeek pricing page still lists V4-Pro at $0.66 per million uncached input and $1.98 per million output off-peak. The gap between a cache hit and a cache miss is what really matters here. At peak pricing, one million uncached input tokens cost $0.30, while the same number of cached tokens costs $0.006, a 50x multiplier. Stable prompt prefixes, reusable repository context, and repeated agent instructions become economically decisive, because the cheapest token in the catalog is not the cheapest completed task if the model needs retries.
A worked example makes the economics concrete. For a single 1 million token context conversation with 10K output tokens, V4.1 Flash lands at roughly $0.156 total. V4-Pro is about $0.570, and Claude Opus 4.8 is around $5.250. On this long-context scenario the new Flash model is 3.7x cheaper than V4-Pro and roughly 34x cheaper than a top-tier closed frontier model.
Open weights do not make deployment lightweight
DeepSeek released the weights under an MIT license, which permits commercial self-hosting, and the repository includes prompt encoding guidance and inference notes for reproducing selected DeepSWE results. But the model is still large. The announcement itself cites a resource profile of roughly 2,000 GPU cards plus a storage cluster for large-scale deployment. Most teams should evaluate the API or a managed inference provider before considering self-hosting a 552 billion parameter multimodal MoE model.
Open weights still carry real value here for a different reason. They let inference frameworks inspect the artifacts, adapt their kernels, and reproduce parts of the evaluation without waiting on a closed API. Community frameworks already show support, with vLLM and SGLang picking up the model in mid-September 2026, though CSA2 needs custom kernels and speculative decoding, DeepSeek's DSpark feature, has its own support path. That is normal for a new architecture. The same pattern appeared with earlier DeepSeek releases, which is worth remembering when reading the current framework support list. The earlier DeepSeek V4 guide covering open weights and harness selection still tracks the same tokenizer and API ecosystem, so its workflow notes remain useful.
Our Read: why the encoder-decoder asymmetry is the signal
The framing most vendors push is parameter count. DeepSeek has just inverted it. A 552B model with 8B active beating a 1.6T model with 49B active is the first time a Flash-tier model in DeepSeek's lineup has entirely replaced a Pro-tier model on every benchmark. But the more interesting signal is what the benchmark split tells us about where frontier research is heading.
From a data science view, the HLE-with-tools and DeepSWE gains relative to GPQA are the tell. The model was trained on large-scale automated generation of agent tasks and environments, scaled progressively across data volume, task variety, and rollout count. The gains came from the data side, not a new algorithm. The architecture simply makes that data cheaper to serve at scale. When the quality gains come from synthetic agent-trajectory data rather than novel math, the real moat shifts to inference efficiency and to who can generate the most realistic trajectories cheaply.
From an engineering view, the KV cache work is a response to a systems problem that most benchmark tables ignore entirely. Long-context inference is KV-bound, and the cache hit rate is what determines whether an agent service is profitable at all. By collapsing the decoder's KV down to a single projected tensor and compressing it to FP4, DeepSeek has made the input-heavy workload that agents actually run dramatically cheaper per token. The 50x cache-hit to cache-miss ratio in pricing is not a pricing gimmick. It is an admission that for real agent deployments, most of the tokens a customer pays for are repeats, and the architecture is built to make those repeats nearly free.
There is a genuine caveat worth flagging. The legacy API names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work, but they now route to V4.1 Flash and bill at Flash rates. Anyone reading older benchmark logs that reference those names should treat them as the new model unless they recorded the request date and returned model metadata at the time. Storing that metadata with every evaluation is a small hygiene cost that prevents a large confusion later.
Outlook
DeepSeek's positioning is explicit. The new architecture is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models in the future. The 748B model with the projected global KV is meant to be the foundation, not the ceiling. V4.1-Pro is already referenced as the next stop.
For builders, the takeaways are concrete. If your workload is input heavy with long context and repeated tool calls, V4.1 Flash warrants a real evaluation against both its own V4-Pro and against closed frontier models on your own harness. Keep the task and harness fixed, compare cost per successful task rather than per token, include at least one long-context and one multimodal task, and test at more than one reasoning-effort setting, since DeepSeek's published scores use reasoning_effort at 100 with temperature 1.0. If your workload is short prompts and pure knowledge recall, the cost advantage shrinks and the larger Pro or a closed model may still win.
The deeper point is that V4.1 Flash marks a shift in open-weight development where efficiency now beats raw scale. Architectural work like projected global KV, cross-layer sparse attention, and FP4 caching is producing better quality per dollar than simply adding parameters. That is a meaningful inflection, and it is one the rest of the industry will have to respond to, especially on the serving economics that actually determine whether an AI product survives.
For context on how quickly this space moves, the earlier DeepSeek V4 guide covering open weights and harness selection still tracks the same tokenizer and API ecosystem, and the ongoing 2026 model release tracker (ai-model-releases-2026-tracker) documents the broader competitive field. Open weights made V4.1 Flash inspectable from day one, and that transparency is what lets this kind of architectural breakdown happen publicly within hours of release.