Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. - DeepSeek
Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, higher throughput, and scaling to larger models.

- Introducing the smallest model in our new architecture family, with native visual understanding.
- New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
- New pretraining methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, higher throughput, and scaling to larger models. Asymmetric architecture. More intelligence, less cost. New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. New pretraining methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro. Compared with the previous generation, V4.1-Flash’s KV cache needs just: Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly. V4.1-Flash is now live on the DeepSeek API with native multimodal support. V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. We’re phasing out V4-Pro. Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches. Official partners WorkBuddy (including CodeBuddy) & OpenCode now fully support V4.1-Flash. Try it today! More efficient architecture. Lower API prices. V4.1-Flash lets us serve more users at a lower cost. We’re passing the savings on to you. Peak/off-peak pricing continues to balance demand. Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save. New pricing takes effect at 04:00 UTC on Sept 10, 2026. Supporting open source. Expanding deployment options. We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options. Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk.
Sources
Related stories

Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses
Avid Artifacts readers know that we have been covering not only models but also their licenses for quite some time.

Controlling Reasoning Effort in LLMs
It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models.

Models & Pricing - DeepSeek
The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark.

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.