Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents
Deepseek released and open-sourced (MIT license, via Hugging Face) a new model, V4.1-Flash, with 552B total parameters (8B active on input/encoding, 16B active on output/decoding), supporting up to 1M token contexts. It reduces KV cache memory footprint substantially versus its predecessor V4-Flash and uses FP4 storage for the cache, and is also available via API at the same prices as V4-Flash.
Entities: Deepseek, V4.1-Flash, V4-Flash, V4-Pro, OpenAI, Anthropic
0 primary
What happened
Deepseek released and open-sourced V4.1-Flash under an MIT licence on Hugging Face, alongside a technical report. The model has 552B total parameters (8B active for input/encoding, 16B active for output/decoding), supports up to 1M token context windows, and uses FP4 storage to shrink its KV cache memory footprint compared with the prior V4-Flash. It is also live via API at the same prices as V4-Flash.
Why it matters
KV cache memory is one of the main cost bottlenecks for running long-context AI agents at scale, so a genuine reduction (Deepseek claims roughly a quarter of the GPU memory footprint, an eighth when offloaded) would lower the cost of deploying agentic workloads that hold large amounts of context. This matters most to developers and enterprises running or evaluating long-horizon agents, and to competitors, since an open-weights model claiming near-frontier coding performance at lower serving cost weakens the pricing power of closed labs like OpenAI and Anthropic. The practical upside is real but conditional: it only changes anyone's decisions once the memory and cost claims are verified on independent hardware, not just Deepseek's own benchmarks.
What is noise
The "matches OpenAI and Anthropic on coding and agent benchmarks" framing is vendor-reported and selectively presented; the source material itself notes the model trails badly on ProgramBench and performs weakly on science and image tasks, so this is not a clean win across the board. Comparisons like "437x smaller per-token than V1" are marketing arithmetic from a self-selected baseline and should not be read as a general efficiency multiplier for real deployments.
Watch next
- 01Independent benchmarking of V4.1-Flash's actual GPU memory usage and inference cost versus V4-Flash and competing open models, not just Deepseek's own figures
- 02Third-party replication of the coding/agent benchmark scores (e.g. DeepSWE, ProgramBench) to check whether the near-frontier claims hold up outside Deepseek's report
- 03Adoption signals over the next 4-8 weeks: developer uptake on Hugging Face, integration into agent frameworks, and whether API usage or pricing shifts in response from OpenAI or Anthropic
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Google DeepMind launches AlphaGenome Atlas, a free public database of predicted effects for 9 billion possible human genome variants8 Sept 202679
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679