DeepSeek V4.1 Flash
Open-weight, very fast MoE built on DeepSeek’s new encoder-decoder architecture.
DeepSeek V4.1 Flash is a sparse mixture-of-experts model and the first built on DeepSeek’s Causal Encoder-Decoder (CED) architecture, activating 8B parameters on input and 16B on output.
- Provider
- DeepSeek
- Type
- Language model
- Released
- Sep 10, 2026
- Context window
- 1M tokens
- Max output
- 944K tokens
- Price
- $0.30 input / $1.20 output per 1M tokens
- Input
- text, image
- Output
- text
- Open weights
- Yes
- License
- MIT
Best for
- High volume / low cost
- Real-time / low latency
- Coding
- Local & on-device
- Agents
Strengths
- Over 200 tok/s
- Sub-1.5s first token
- Open weights
More from DeepSeek
- DeepSeek V4 Pro — DeepSeek. DeepSeek’s large open MoE, GA release (0813).