DeepSeek Releases Experimental Multimodal Model on API
DeepSeek states its new experimental model adds vision to the V4-Flash API along with a Files API for image reuse.
DeepSeek posted that DeepSeek-V4-Flash-Vision-Exp is now live on its API platform. The model matches V4-Flash on text tasks and shows gains on multimodal agent benchmarks. Users set the model name to deepseek-v4-flash-vision-exp for Chat Completions, Messages, or Responses endpoints. Images count toward billing at up to 384 tokens each. A separate Files API lets developers upload an image once and reference it by file_id across requests at no charge. The company also noted that the model works with existing agent frameworks for combined text and visual tool use.
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major…

