China AI Hub China AI Hub

News / DeepSeek releases V4.1-Flash with multimodal vision and lower API prices

DeepSeek releases V4.1-Flash with multimodal vision and lower API prices

2026-09-10 · model_release

DeepSeek releases V4.1-Flash with multimodal vision and lower API prices
Image: mikemacmarketing / CC BY 2.0, via Wikimedia Commons · source

DeepSeek released V4.1-Flash on 2026-09-10, the smallest model of its new “asymmetric architecture” family, with MIT open weights and native multimodal vision. The API change log lists the release alongside a price reduction; V4-Flash and V4-Flash-Vision-Exp are retired, with their model names temporarily routed to V4.1-Flash for compatibility.

The model is a 552B-parameter mixture-of-experts that activates only 8B parameters per token on input (16B on output), with a 1M-token context window and 384K maximum output. DeepSeek reports GPQA Diamond 90.9, Terminal-Bench 2.1 90.6 and DeepSWE v1.1 74.2 (vendor-reported). Off-peak API pricing is $0.15 input / $0.60 output per 1M tokens with $0.003 cache hits.

The same change log entry reversed the earlier statement that DeepSeek-V4-Pro requests would route to V4.1-Flash after 2026-09-14; V4-Pro API service now continues with unchanged billing (see the related news item).

Database updates triggered

Sources