Models / DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash
DeepSeek-V4.1 family · Provider: deepseek · Status: active · Released: 2026-09-10
DeepSeek-V4.1-Flash is a DeepSeek-V4.1 model developed by deepseek, a Chinese AI company headquartered in Hangzhou, Zhejiang, China (derived from the official footer company name 杭州深度求索人工智能基础技术研究有限公司 and Zhejiang ICP / Hangzhou public-security filings, released 2026-09-10 with a 1,048,576-token context window and open weights.
| Key fact | Value |
|---|---|
| Model ID | deepseek-v4-1-flash |
| Architecture | 552B-parameter MoE; Causal Encoder-Decoder; 8B active parameters on input, 16B on output |
| Context window | 1,048,576 tokens |
| Max output | 393,216 tokens |
| Open weights | Yes |
| License | MIT |
| Self-hosting | Yes |
| API available | Yes |
| API pricing | $0.15 input / $0.6 output per 1M tokens (USD) · provider pricing page |
| Regions | unknown |
| Cloud providers | DeepSeek Platform |
Capabilities
| Capability | Supported |
|---|---|
| Reasoning | Yes |
| Coding | Yes |
| Math | Unknown |
| Chinese | Unknown |
| English | Unknown |
| Multilingual | Unknown |
| Vision | Yes |
| Audio | Unknown |
| Video | Unknown |
| Tool calling | Yes |
| Function calling | Yes |
| Structured output | Yes |
| Agent capability | Unknown |
| RAG | Unknown |
| Computer use | Unknown |
Benchmark results
| Benchmark | Version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| GPQA Diamond | — | 90.9 | accuracy | 2026-09-10 | vendor_reported | link |
| HLE | — | 36.8 | accuracy | 2026-09-10 | vendor_reported | link |
| Codeforces | — | 3471 | rating | 2026-09-10 | vendor_reported | link |
| Terminal-Bench 2.1 | — | 90.6 | accuracy | 2026-09-10 | vendor_reported | link |
| DeepSWE v1.1 | — | 74.2 | accuracy | 2026-09-10 | vendor_reported | link |
Benchmark scores are single data points, not universal rankings. Vendor-reported scores are labeled as such.
Known limitations
- Pricing is peak/off-peak: listed prices are off-peak; peak (01:00-04:00 and 06:00-10:00 UTC, Mon-Fri) is 2x
- Benchmarks are vendor-reported using DeepSeek Harness (minimal mode, max effort); not independently verified
- HLE score is on the pure-text subset (39.1 on that subset; 36.8 full)
- Legacy API names deepseek-v4-flash and deepseek-v4-flash-vision-exp route to V4.1-Flash
DeepSeek-V4.1-Flash is the current standard DeepSeek API model (deepseek-flash), released
2026-09-10 with MIT-licensed open weights. It is the smallest model of the “asymmetric architecture”
V4.1 family: a 552B-parameter MoE that activates only 8B parameters on input and 16B on output.
It has a 1M-token context window, 384K maximum output, and supports thinking (default) and non-thinking modes, native vision understanding, JSON output and tool calls. DeepSeek claims 1/4 the HBM and 1/8 the SSD KV-cache storage of the previous generation. API pricing is peak/off-peak: off-peak $0.15 input / $0.60 output per 1M tokens, with cache hits at $0.003 (as of 2026-09-20).
Provider
Pricing
Benchmarks with results for this model
Agents built on this model
API
Data interpretation
| Field | Value | Evidence type |
|---|---|---|
| Context window | 1,048,576 tokens | Official |
| Architecture | 552B-parameter MoE; Causal Encoder-Decoder; 8B active parameters on input, 16B on output | Official |
| Open weights | Yes | Official |
| License | MIT | Official |
| API pricing | $0.15 / $0.6 per 1M tokens (USD) | Official |
| GPQA Diamond | 90.9 accuracy | Vendor-reported |
| HLE | 36.8 accuracy | Vendor-reported |
| Codeforces | 3471 rating | Vendor-reported |
| Terminal-Bench 2.1 | 90.6 accuracy | Vendor-reported |
| DeepSWE v1.1 | 74.2 accuracy | Vendor-reported |
| Family position | 1 model in the DeepSeek-V4.1 family | China AI Hub analysis |
Evidence types: Official = vendor documentation, pricing pages or model cards. Vendor-reported = benchmark scores published by the vendor. China AI Hub analysis = derived from the database itself. See the sourcing policy.
Sources
Confidence and source hierarchy per the sourcing policy. Facts change; check the source before relying on this page.
What is the context window of DeepSeek-V4.1-Flash?
DeepSeek-V4.1-Flash has a 1,048,576-token context window and a maximum output of 393,216 tokens.
Is DeepSeek-V4.1-Flash open weight?
Yes — DeepSeek-V4.1-Flash weights are openly available under the MIT license.
How much does DeepSeek-V4.1-Flash cost through the API?
$0.15 per 1M input tokens and $0.6 per 1M output tokens (USD).
Where does China AI Hub get its DeepSeek-V4.1-Flash data?
From 4 sources (official pages first), last verified 2026-09-20. Benchmark scores are labeled by source type; see the sourcing policy for details.