DeepSeek releases V4.1-Flash, a new open-weight model using a novel causal encoder-decoder architecture with separate 8B prefill / 16B decode active parameters
DeepSeek released V4.1-Flash, replacing V4 Pro as its flagship open-weight model. It introduces a new causal encoder-decoder architecture with prefill/decode parameter separation (763B total parameters, 8B active for input/prefill, 16B active for output/decode), native text+image input, 1M-token context, MIT license, and techniques like Sliding-Window Attention Bounded Replay that reduce KV cache footprint to as low as 1/8 that of V4 Flash. It is priced at $0.30/1M input tokens and $1.20/1M output tokens (cached input $0.006/1M, with an additional 50% off-peak discount), and scored 40 on the Artificial Analysis Intelligence Index and took #1 on the Vals Index among open-weight models. Day-0 support was shipped by Baseten and Ollama began rolling it out to Max/Team/Pro accounts.
Entities: DeepSeek, DeepSeek V4.1-Flash, DeepSeek V4 Pro, Artificial Analysis, Vals AI, Baseten
0 primary
What happened
DeepSeek released V4.1-Flash, replacing V4 Pro as its flagship open-weight model. It uses a new causal encoder-decoder architecture that separates active parameters for prefill and decode (763B total, 8B active on input, 16B active on output), supports 1M-token context and native text plus image input, and ships under an MIT licence. Pricing is $0.30 per 1M input tokens and $1.20 per 1M output tokens, with day-0 hosting from Baseten and a rollout on Ollama's paid tiers. It scored 40 on the Artificial Analysis Intelligence Index and placed first among open-weight models on the Vals Index.
Why it matters
Developers and enterprises building long-running AI agents get a cheaper, MIT-licensed model with a claimed KV cache footprint as low as 1/8 of the previous version, which lowers the cost of maintaining long conversations or agent state. Because it is open-weight with immediate third-party hosting support, teams can test and deploy it today rather than waiting on an announcement. The benchmark score of 40 on the Intelligence Index is middling, so the practical case for switching rests more on cost and context efficiency than on raw capability gains.
What is noise
The "Return of the Whale" framing and commentator claims that this is effectively a "DeepSeek V5" undersold by naming are promotional spin, not evidence. The article leans on unfalsifiable claims that benchmarks fail to capture the architecture's real advance, which conveniently explains away the fact that the model trails some open competitors on standard measures. No primary source links (paper, docs, benchmark pages) are included despite claiming research-paper and benchmark evidence, so the specific numbers cannot be independently checked from this coverage alone.
Watch next
- 01Independent third-party benchmarks (e.g. Artificial Analysis, Vals AI, LMSYS) confirming or revising the Intelligence Index score of 40 and the #1 open-weight ranking over the next 2-4 weeks
- 02Whether the claimed 1/8 KV cache reduction and prefill/decode split hold up in independent inference benchmarks from hosting providers beyond Baseten and Ollama
- 03Developer and enterprise adoption signals: GitHub stars, Hugging Face downloads, and whether major inference platforms (Together, Fireworks, Replicate) add support within the next month
- 04Whether DeepSeek or independent researchers publish the underlying research paper with full architecture details, since none was linked in this coverage
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680