DeepSeek announces V4.1-Flash with native vision and lower API prices
DeepSeek says its KV cache, which lets the model reuse earlier calculations, needs a quarter of the high-bandwidth memory and an eighth of the SSD storage required by the previous generation's cache.
TLDR
DeepSeek announced V4.1-Flash on September 10, 2026, with native visual understanding and API access through the deepseek-flash model name. The company says its 552-billion-parameter model activates 8 billion parameters for input processing and 16 billion for output generation.
DeepSeek announced lower API prices effective at 04:00 UTC that day, with off-peak rates at half of peak rates. It also said V4-Pro was being phased out: all deepseek-v4-pro requests would route to V4.1-Flash at V4.1-Flash rates from 04:00 UTC on September 14, 2026, until V4.1-Pro launches.
SGLang and vLLM both announced launch-day support.
Combined views
8.3M
43 Sources, first seen 21d ago
