On September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash, a model that handles images and text. It supports a context of 1 million tokens and, compared with the previous-generation V4-Flash, reportedly cuts the KV cache that holds the full context to about one quarter, and the KV cache stored long-term for reuse to about one eighth. For AI agents that repeatedly fix code and run tests, the history grows as the work proceeds, and storing and reloading it costs money. How much does V4.1 Flash's new design lighten that burden, and how does it change what developers pay?
Why a lighter input pass helps agents
V4.1 Flash uses an asymmetric structure: 8 billion parameters are active when reading input, and 16 billion when generating output. According to the technical report, it is a "Causal Encoder-Decoder (CED)" that connects a 20-layer encoder to a 20-layer decoder. The KV the decoder uses to reference the full context is built from the encoder's final internal state, which lets the decoder skip processing over the entire long input.
For example, when an agent fixes code, runs tests, reads the logs, and fixes the code again, the model repeatedly receives new logs along with the conversation history. Unlike returning a single short answer, this means processing large amounts of input many times over. That is why lightening the input side adds up.
However, the figures of 8 billion and 16 billion do not describe the model's storage footprint. V4.1 Flash is a mixture-of-experts (MoE) model that selects which parts to use for processing, and its total parameter count has grown to 552 billion, up from 284 billion in the previous generation. It also has 196 billion parameters' worth of a conditional memory called "Engram," which retrieves information according to the token sequence. This is a mechanism for referencing learned information and is separate from the KV cache that holds conversation history.
In other words, the input is not lighter because the model got smaller; the design changes which parts of a large model run, and when.
An 890-byte cache and a design that avoids re-saving
In Figure 1 of the technical report, the full-context KV cache falls from 3,514 bytes per token to 890 bytes. The comparison is against V4-Flash, for the same context length. This portion, which sits in HBM (the GPU's high-speed memory), has shrunk to about a quarter.
The KV cache is computed information the model retains in order to refer back to past input. V4.1 Flash's "Compressed Sparse Attention 2 (CSA2)" shares across layers information that was previously kept separately in each layer. It uses some layers that re-select what to reference from the shared information and others that reuse even the reference targets, reducing duplicated storage and search. The main KV also uses the FP4 format, which saves capacity by lowering the precision at which values are stored.
Furthermore, for the cache stored long-term on SSDs and similar media, the handling of "SWA," the local attention that covers only a narrow recent window, has changed. According to Section 3.2.1 of the report, in the old configuration this local state took up nearly half of the persistent cache capacity. Yet local state is mainly reused within the short window while a conversation is ongoing. The policy of storing it long-term did not match how long it was actually needed.
So V4.1 Flash keeps the local state briefly in host-side memory and removes it from long-term storage. If the needed local state has been discarded, "SWA Bounded Replay" reprocesses a fixed recent range to rebuild it. Combined with the roughly 4x compression of the full-context KV, this reduces the persistent KV cache to about one eighth under the same processing load.
This reconstruction is an approximation, not a mathematically exact reproduction of the original state. DeepSeek says that in its own experiments the effect on response quality was small. The design trades reduced storage for limited recomputation and some approximation.
The "about one quarter in HBM" and "about one eighth on SSD" reductions apply to this KV cache. It cannot be said that the memory required by the whole server, including model weights, Engram, and compute workspace, falls by the same ratio.
Where scores rose, and where gaps with competitors remain
In DeepSeek's published DeepSWE v1.1 results, V4.1 Flash's issue resolution rate is 74.2%, a large gain from V4-Flash's 54.4%. It also exceeds Opus 5's 74.0% and GPT-5.6 Sol's 73.0% in the same table. But the rankings reverse on other tests.
| Benchmark | V4-Flash | V4.1 Flash | GPT-5.6 Sol | Opus 5 |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 82.7 | 90.6 | 88.8 | 89.1 |
| Terminal-Bench 3.0 | 7.6 | 30.0 | 34.4 | 43.3 |
| Terminal-Bench 4.0 | 7.0 | 31.2 | 39.9 | 51.8 |
| DeepSWE v1.1 | 54.4 | 74.2 | 73.0 | 74.0 |
| CyberGym | 76.7 | 88.1 | 84.5 | — |
| SEC-Bench Pro | 30.9 | 62.8 | 74.3 | — |
| ExploitGym | 1.8 | 15.3 | 33.7 | 22.1 |
The figures are percentages excerpted from Table 3 of DeepSeek's technical report, with each model at its maximum reasoning setting. DeepSWE is the issue resolution rate; the rest are Pass@1. "—" marks items with no value in the table. V4.1 Flash's coding evaluations mainly use DeepSeek Harness in Minimal mode with a 1-million-token context, but DeepSWE uses mini-SWE and SEC-Bench Pro uses Claude Code, so this cannot be treated as a comparison in which every item and every model was measured with the same execution software.
V4.1 Flash improves on its predecessor in every row. On the other hand, on newer Terminal-Bench versions, SEC-Bench Pro, and ExploitGym, it falls short of the competitors' scores in the table. It would be an overreach to say that CyberGym's 88.1% means it has surpassed competitors in security capabilities overall. These are evaluations from DeepSeek's own technical report, not the results of independent replication.
Differences in execution software also matter. According to the official model card, the same V4.1 Flash scores 74.2% on DeepSWE v1.1 with mini-SWE but 65.6% with Codex. Changing the mechanism around the model that manages tool execution and history changes the results. You cannot simply pick a model name and expect the same performance as the top score.
Cache hit rate and reasoning depth drive the bill
Under the official API's new pricing, per million tokens, cache-hit input costs $0.003 off-peak and $0.006 at peak. New input and output are billed at separate rates.
| Price per 1M tokens | Off-peak | Peak |
|---|---|---|
| Input: cache hit | $0.003 | $0.006 |
| Input: cache miss | $0.15 | $0.3 |
| Output | $0.6 | $1.2 |
Based on DeepSeek's official pricing page as revised on September 10. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday (10:00–13:00 and 15:00–19:00 Japan time); all other times are billed at off-peak rates.
Within the same time band, the unit price of cache-hit input is one fiftieth that of new input: 0.3 ÷ 0.006 at peak and 0.15 ÷ 0.003 off-peak both give 50. This is a comparison of input unit prices, not a claim that the total cost of the work falls to one fiftieth. The actual bill depends on how much input could be reused, how much was new, and the volume of output.
Raising reasoning depth also increases that output volume. The technical report says that raising the reasoning setting from 25 to 100 improved DeepSWE v1.1 from 66.0% to 74.2%, while output tokens grew roughly 2.5 times. Simply combining the low unit prices with the scores achieved at the maximum setting will not let you estimate the cost of completing a task.
Being able to handle repeatedly read history cheaply and the rising cost of thinking longer on hard problems can both be true at once. When adopting the model, you need to compare total token counts and billed amounts at each reasoning setting while checking the success rate your own work requires.
What to check before the September 14 API switch
For existing users, there is a practical change: the backend changes even if you don't change the model name. The official call name is deepseek-flash, and requests to the old deepseek-v4-flash and deepseek-v4-flash-vision-exp are already routed to V4.1 Flash.
In addition, from 13:00 on September 14 (Japan time), requests to deepseek-v4-pro will also be switched to V4.1 Flash and billed at that model's unit prices. DeepSeek says it will continue this arrangement until a future V4.1 Pro is released. Even if you thought you were pinned to the previous Pro, you will not keep calling the same model under the same name.
The weights and repository are released under the MIT license, and reference implementations for inference and prompt formatting are also provided. But it is premature to conclude from the cache reduction rates that it will run on your own hardware. Beyond the storage needed for the main model and Engram, you need to confirm a supported runtime environment.
Before the API switch, it is worth rechecking tool-call success rates and output quality on your existing code-fixing tasks and image-containing inputs. If quality holds there and total cost falls while the cache is reused, V4.1 Flash's new design becomes an option for expanding the length and volume of work you can hand to agents.
