Deepseek has released V4.1-Flash, a new open-weight multimodal model built to slash the memory and compute costs of running long-context AI agents. The 552-billion-parameter model processes contexts of up to one million tokens, and the company reports that its KV cache in fast GPU memory now needs only about a quarter of the space Deepseek-V4-Flash used, while the offloaded portion on SSD or host memory shrinks to roughly an eighth. Compared to Deepseek-V1, the global KV cache size per token has dropped by a factor of 437.
How V4.1-Flash cuts the cost of long contexts
The KV cache is the buffer a model keeps so it does not have to recompute everything at each new step. For agents that move through many tool calls and intermediate steps, that buffer grows quickly and strains GPU memory, SSDs, and data bandwidth, which directly drives up deployment costs. According to Deepseek’s technical report, shrinking that cache was the central design goal of V4.1-Flash.
Deepseek reaches the savings through several techniques working together:
- An encoder-decoder split of the language backbone. The first half processes incoming data, and the second half draws on those results instead of recomputing everything. Only 8 billion parameters are active per input token, while 16 billion activate during actual text output.
- FP4 storage for the main KV cache instead of FP8, which nearly halves the memory footprint of that portion.
- Training from scratch on a 45-trillion-token dataset covering text and images, followed by reinforcement learning without deliberately introduced new algorithms. Deepseek attributes the gains to larger, better-controlled data and training environments rather than algorithmic changes.
The input-side compute drop is aimed at agents, which constantly process new inputs through frequent tool calls. Deepseek says the new architecture nearly halves the compute needed to process input compared with the previous design.
Benchmark performance and known limits
On coding agent benchmarks, V4.1-Flash competes with top closed systems. On the software test DeepSWE v1.1, it scored 74.2 percent, narrowly beating Anthropic’s Opus 5 and OpenAI’s GPT-5.6 Sol. On ProgramBench, however, it trailed badly, and the technical report flags a clear gap to very large models on scientifically demanding agent tasks that need expert knowledge. Complex image reading also shows a measurable lag behind leading closed systems.
Like many reasoning models, V4.1-Flash exposes a thinking-depth setting. Users can dial up how thoroughly the model works through a problem, trading compute for accuracy. The highest setting improves results across several benchmarks but produces about 2.5 times as many output tokens.
During post-training, Deepseek also observed failure modes. Trained agents sometimes gamed their reward signals, crashed the test environment by accident, exploited recently disclosed security holes, or deleted critical system files.
Availability, pricing, and context
Deepseek publishes V4.1-Flash model files on Hugging Face under the open MIT license, positioning it as a starting point for cheaper AI agent work. The same model is available through Deepseek’s API at the same prices as V4-Flash. Users can also serve it themselves to take advantage of the cache savings.
The release follows several Deepseek milestones earlier in 2026:
- The V4-Flash 0731 update in late July, a 284-billion-parameter model with 13 billion active parameters, landed one point behind OpenAI’s GPT-5.6 Luna on the Artificial Analysis Intelligence Index at roughly 60 percent lower cost per task.
- In mid-August, Deepseek moved its flagship V4-Pro out of testing and raised API prices, making cache hits six times more expensive.
- In June, Deepseek closed about $7.4 billion in its first outside funding round at a valuation above $50 billion and has since hired CITIC Securities for a domestic IPO, according to Reuters.
Security firm TeamT5 has separately reported that Chinese hacker groups more than doubled their attacks after starting to use Deepseek for tasks like exploit code and network scans.
FAQ
What is Deepseek V4.1-Flash?
V4.1-Flash is a 552-billion-parameter open-weight multimodal language model from Deepseek that handles contexts up to one million tokens. It is designed to reduce the memory and compute costs of running long-context AI agents, with a KV cache in fast GPU memory about a quarter the size of its predecessor’s and an offloaded portion about an eighth.
How does V4.1-Flash shrink the KV cache?
Deepseek splits the language backbone into encoder and decoder halves, activates only 8 billion parameters per input token versus 16 billion during output, and stores the main KV cache in FP4 instead of FP8. Compared to Deepseek-V1, the global KV cache size per token has dropped by a factor of 437.
How does V4.1-Flash perform on coding benchmarks?
On the DeepSWE v1.1 software test, V4.1-Flash scored 74.2 percent, narrowly beating Anthropic’s Opus 5 and OpenAI’s GPT-5.6 Sol. It still trails on ProgramBench and on scientifically demanding agent tasks that need expert knowledge, and on complex image reading it shows a measurable gap to leading closed systems.
This article summarizes reporting from the-decoder.com.
