DeepSeek announced V4.1-Flash on September 10, 2026, with native visual understanding and API access through the deepseek-flash model name. The company says its 552-billion-parameter model activates 8 billion parameters for input processing and 16 billion for output generation. DeepSeek announced lower API prices effective at 04:00 UTC that day, with off-peak rates at half of peak rates. It also said V4-Pro was being phased out: all deepseek-v4-pro requests would route to V4.1-Flash at V4.1-Flash rates from 04:00 UTC on September 14, 2026, until V4.1-Pro launches. SGLang and vLLM both announced launch-day support.