DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026, describing it as the smallest model in a new architecture family and the first Flash-tier DeepSeek model with native visual understanding. The model is a 552-billion-parameter Mixture-of-Experts system built on what DeepSeek calls a Causal Encoder-Decoder architecture: an asymmetric split that activates just 8 billion parameters to read input and 16 billion parameters to generate output. DeepSeek says the design cuts KV-cache memory to one-quarter the HBM and one-eighth the SSD storage of the prior generation, which is what lets a Flash-tier model take on multimodal input without the usual throughput penalty.
The headline claim is that V4.1-Flash beats DeepSeek’s own flagship V4-Pro on performance, cost, speed, and total runtime, not just on a narrower efficiency metric. DeepSeek’s Harness team framed the result as full surpassal of V4-Pro on key metrics, which is why the company is retiring V4-Pro API traffic into V4.1-Flash rather than running them in parallel: starting at 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash and are billed at Flash rates until a V4.1-Pro ships. Existing callers using the deepseek-v4-flash and deepseek-v4-flash-vision-exp model IDs are also transparently routed to the new build with no code change required.
The model is available immediately through DeepSeek’s standard API under the model name deepseek-flash, and the weights are published under the MIT license - continuing DeepSeek’s practice of shipping frontier-adjacent capability as freely licensed, redistributable weights rather than an API-only product. New pricing took effect at 04:00 UTC the same day, with DeepSeek’s existing peak/off-peak split preserved (off-peak rates run at 50 percent of peak).
For a market watching whether Chinese labs can keep closing the gap on efficiency rather than just raw capability, V4.1-Flash is a data point in DeepSeek’s specific style of pressure: it does not claim to beat GPT-6 Astra or Gemini 3.8 on capability, it claims to beat its own larger, more expensive model on a smaller compute budget while adding a new modality. That is a commodification argument aimed at API buyers’ unit economics, not a leaderboard argument, and it is the same playbook that made V3 and R1 consequential in 2025.