• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology

DeepSeek announces V4.1-Flash with native vision and lower API prices

DeepSeek says its KV cache, which lets the model reuse earlier calculations, needs a quarter of the high-bandwidth memory and an eighth of the SSD storage required by the previous generation's cache.

DE
ZP
ZA
43 Sources, 21d ago, first seen 21d ago

TLDR

DeepSeek announced V4.1-Flash on September 10, 2026, with native visual understanding and API access through the deepseek-flash model name. The company says its 552-billion-parameter model activates 8 billion parameters for input processing and 16 billion for output generation.

DeepSeek announced lower API prices effective at 04:00 UTC that day, with off-peak rates at half of peak rates. It also said V4-Pro was being phased out: all deepseek-v4-pro requests would route to V4.1-Flash at V4.1-Flash rates from 04:00 UTC on September 14, 2026, until V4.1-Pro launches.

SGLang and vLLM both announced launch-day support.

Combined views

8.3M

43 Sources, first seen 21d ago

55.1K likes1.8K comments9.2K saves4.6K reposts
Featured Source

Combined views

8.3M

43 Sources, first seen 21d ago

55.1K likes1.8K comments9.2K saves4.6K reposts

Sentiment

Positive82.7%17.3%Negative

Summary

Many accounts praised DeepSeek-V4.1-Flash for its efficiency, low pricing, and controllable reasoning effort, while some replies called it a distillation of other models or claimed it underperforms the prior pro version.

Based on 141 sentiment-bearing replies from 127 accounts across 10 conversations.

Sentiment

Positive82.7%17.3%Negative

Summary

Many accounts praised DeepSeek-V4.1-Flash for its efficiency, low pricing, and controllable reasoning effort, while some replies called it a distillation of other models or claimed it underperforms the prior pro version.

Based on 141 sentiment-bearing replies from 127 accounts across 10 conversations.

43 Sources

@deepseek_ai🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
@zizhpanDeepSeek-V4.1-Flash is now publicly available on App, Web, and API. Our first flagship model with native multimodal support. Faster. More capable. Lower price. Open weights as usual.
@zainhasoh wow have not seen this before for any model... this complicates things > "DeepSeek-V4.1-Flash supports a continuously controllable reasoning effort from 1 to 100." reasoning_effort = [1 to 100]
@lmsysorgThrilled to see day-0 SGLang support for DeepSeek V4.1 Flash! Technical blog: http://lmsys.org/blog/2026-09-10-deepseek-v41 This is another leap for the DeepSeek family, with several new upgrades to the V4 stack: shared compressed KV, a two-stage sparse indexer, mHC, and Engram. SGLang supports the new architecture on launch and will keep pushing performance optimizations! We are also excited to see Miles support RL on day-0!
@vllm_project🐳 DeepSeek-V4.1-Flash is out, and vLLM serves it from day 0, verified on NVIDIA and AMD GPUs! 🎉 552B MoE backbone, native vision, 1M context. Built for agents: 8B active while it reads your prompt, 16B while it writes. If you already run DeepSeek-V4 on vLLM, most of this stack will feel familiar: the hyper-connections, the sliding-window plus compressed sparse attention, DSpark drafting, MXFP4 experts. vLLM has carried all of it since V4 landed. Two things are new, and both are worth a look: ✨ Engram: a quarter of the checkpoint is n-gram memory the model looks up instead of computes. 197B parameters of it. ✨ Only four layers write compressed KV now. The rest of the model shares it. Spin it up 👇 🔗 http://recipes.vllm.ai/deepseek-ai/DeepSeek-V4.1-Flash
@iScienceLuvrDeepSeek releases DeepSeek-V4.1-Flash!! 🔥 This model is wild... the benchmarks are showcasing it's GPT-5.6 Sol level, yet it's only 552B params?! This feels like it shouldn't be possible, what's the catch? Perhaps benchmarkmaxxed? Let's look into the architecture and training of the model: The key focus seems to be on more aggressive KV cache compression: "We adopt a Causal Encoder-Decoder (CED) architecture, in which decoder global KV is projected from the final encoder hidden states. This design enables the model to activate 8B parameters per token during prefill and 16B during decode, which is particularly cost-effective for input-heavy agentic scenarios." "DeepSeek-V4 can be viewed as an SWA-based local-processing backbone augmented with compressed global context." The idea behind CED is that the decoder's KV cache is constructed from the encoder output, bypassing full decoder computation. They also introduce Compressed Sparse Attention 2 (CSA2) that has three operating modes that differ in how they obtain main KV, indexer K, and Top-K indices. (frankly I don't understand this part very well 😭) DeepSeek-V4.1-Flash uses a variant of mHC called single-pass mHC, and they also incorporate a 196B Engram module to decouple memorization from computation. The model is natively multimodal: a vision embeddings generated from "DeepSeek-ViT" (a pretty standard ViT arch) are passed jointly with the text tokens into the model. This is trained first with SigLIP loss then with autoregressive loss for the combined vision encoder+LLM. "we train DeepSeek-V4.1-Flash on a large-scale multimodal corpus comprising 45T tokens." They use FP4 KV cache with quantization-aware training to further save storage. Regarding post-training: "In this release, we refrain from introducing novel post-training algorithms." "at the current stage, the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training." They utilize the model itself to construct its own training environments, based on data they are getting from model use internally. "As we transitioned from DeepSeek-V3 to V4, the rapidly growing number and diversity of agentic training environments motivated us to build DeepSeek Elastic Compute (DSec), a production-grade sandbox platform for large-scale agentic training and evaluation." "We therefore introduce a scalar effort level 𝑏 as an explicit conditioning signal during reinforcement-learning." max --> b=100, high --> b=75, low --> b=50. "As the last stage of post-training, the final full-vocabulary OPD task is trained on datasets from all domains using over 40 teacher models." Damn, this is a dense report, I've barely touched the surface tbh, very interesting!! model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
@epsilver_DeepSeek-V4.1-Flash is incredible. It has already replaced Gemini for everything I used it for. If China is the only one willing to accelerate now, then so be it. Congratulations DeepSeek!! 🐋
@gchampeauDeepseek a sorti un nouveau modèle open weights (DeepSeek V4.1-Flash) qu'ils comparent à GPT-5.6 Sol, mais le vrai intérêt est surtout là 👇Gros gains dans le ratio coût/qualité, qui font qu'exploiter ce modèle coûtera bien moins chère à qualité comparable, puisqu'il consomme beaucoup moins de ressources. Et ça confirme que le prix de l'IA baisse de génération en génération, il n'augmente pas. On fait aujourd'hui beaucoup plus à prix égal qu'on ne faisait il y a 2 ans.
@radixarkMiles brings Day-0 RL support to DeepSeek-V4.1-Flash. Miles keeps the trainer close to what @sgl_project samples. Parallelism and shared state let the new architecture scale intact across GPUs. Quantization-aware training mirrors SGLang's FP4/FP8 rounding, while routing replay reuses the rollout's expert choices. Numerical consistency comes from FP32 and deterministic reductions. Colocated training and rollout fit full-parameter RL on 16 GPUs. Over steps 0–80 of a DAPO run, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78. Blog and cookbook in the comments. 📚
@kimmonismusDeepSeek just released V4.1-Flash with a new architecture, six weeks after its July V4-Flash update. July’s release improved post-training while keeping the architecture unchanged. (Same with GLM-5.3/Flash) V4.1 introduces a Causal Encoder–Decoder architecture with native visual understanding. DeepSeek reports: - 552B MoE parameters, with 8B active during input processing and 16B during output generation. - KV-cache requirements cut to ¼ of the HBM and ⅛ of the SSD storage versus the previous generation. - Lower API prices. These are *significant* jumps in just a few weeks with post training. This is the new reality we have to adapt to: weekly releases with significant improvements. The company says Flash now beats V4-Pro on capability, cost and speed. Starting September 14, V4-Pro API requests will temporarily route to V4.1-Flash until V4.1-Pro arrives.
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    43 Sources

    @deepseek_ai🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
    @zizhpanDeepSeek-V4.1-Flash is now publicly available on App, Web, and API. Our first flagship model with native multimodal support. Faster. More capable. Lower price. Open weights as usual.
    @zainhasoh wow have not seen this before for any model... this complicates things > "DeepSeek-V4.1-Flash supports a continuously controllable reasoning effort from 1 to 100." reasoning_effort = [1 to 100]
    @lmsysorgThrilled to see day-0 SGLang support for DeepSeek V4.1 Flash! Technical blog: http://lmsys.org/blog/2026-09-10-deepseek-v41 This is another leap for the DeepSeek family, with several new upgrades to the V4 stack: shared compressed KV, a two-stage sparse indexer, mHC, and Engram. SGLang supports the new architecture on launch and will keep pushing performance optimizations! We are also excited to see Miles support RL on day-0!
    @vllm_project🐳 DeepSeek-V4.1-Flash is out, and vLLM serves it from day 0, verified on NVIDIA and AMD GPUs! 🎉 552B MoE backbone, native vision, 1M context. Built for agents: 8B active while it reads your prompt, 16B while it writes. If you already run DeepSeek-V4 on vLLM, most of this stack will feel familiar: the hyper-connections, the sliding-window plus compressed sparse attention, DSpark drafting, MXFP4 experts. vLLM has carried all of it since V4 landed. Two things are new, and both are worth a look: ✨ Engram: a quarter of the checkpoint is n-gram memory the model looks up instead of computes. 197B parameters of it. ✨ Only four layers write compressed KV now. The rest of the model shares it. Spin it up 👇 🔗 http://recipes.vllm.ai/deepseek-ai/DeepSeek-V4.1-Flash
    @iScienceLuvrDeepSeek releases DeepSeek-V4.1-Flash!! 🔥 This model is wild... the benchmarks are showcasing it's GPT-5.6 Sol level, yet it's only 552B params?! This feels like it shouldn't be possible, what's the catch? Perhaps benchmarkmaxxed? Let's look into the architecture and training of the model: The key focus seems to be on more aggressive KV cache compression: "We adopt a Causal Encoder-Decoder (CED) architecture, in which decoder global KV is projected from the final encoder hidden states. This design enables the model to activate 8B parameters per token during prefill and 16B during decode, which is particularly cost-effective for input-heavy agentic scenarios." "DeepSeek-V4 can be viewed as an SWA-based local-processing backbone augmented with compressed global context." The idea behind CED is that the decoder's KV cache is constructed from the encoder output, bypassing full decoder computation. They also introduce Compressed Sparse Attention 2 (CSA2) that has three operating modes that differ in how they obtain main KV, indexer K, and Top-K indices. (frankly I don't understand this part very well 😭) DeepSeek-V4.1-Flash uses a variant of mHC called single-pass mHC, and they also incorporate a 196B Engram module to decouple memorization from computation. The model is natively multimodal: a vision embeddings generated from "DeepSeek-ViT" (a pretty standard ViT arch) are passed jointly with the text tokens into the model. This is trained first with SigLIP loss then with autoregressive loss for the combined vision encoder+LLM. "we train DeepSeek-V4.1-Flash on a large-scale multimodal corpus comprising 45T tokens." They use FP4 KV cache with quantization-aware training to further save storage. Regarding post-training: "In this release, we refrain from introducing novel post-training algorithms." "at the current stage, the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training." They utilize the model itself to construct its own training environments, based on data they are getting from model use internally. "As we transitioned from DeepSeek-V3 to V4, the rapidly growing number and diversity of agentic training environments motivated us to build DeepSeek Elastic Compute (DSec), a production-grade sandbox platform for large-scale agentic training and evaluation." "We therefore introduce a scalar effort level 𝑏 as an explicit conditioning signal during reinforcement-learning." max --> b=100, high --> b=75, low --> b=50. "As the last stage of post-training, the final full-vocabulary OPD task is trained on datasets from all domains using over 40 teacher models." Damn, this is a dense report, I've barely touched the surface tbh, very interesting!! model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
    @epsilver_DeepSeek-V4.1-Flash is incredible. It has already replaced Gemini for everything I used it for. If China is the only one willing to accelerate now, then so be it. Congratulations DeepSeek!! 🐋
    @gchampeauDeepseek a sorti un nouveau modèle open weights (DeepSeek V4.1-Flash) qu'ils comparent à GPT-5.6 Sol, mais le vrai intérêt est surtout là 👇Gros gains dans le ratio coût/qualité, qui font qu'exploiter ce modèle coûtera bien moins chère à qualité comparable, puisqu'il consomme beaucoup moins de ressources. Et ça confirme que le prix de l'IA baisse de génération en génération, il n'augmente pas. On fait aujourd'hui beaucoup plus à prix égal qu'on ne faisait il y a 2 ans.
    @radixarkMiles brings Day-0 RL support to DeepSeek-V4.1-Flash. Miles keeps the trainer close to what @sgl_project samples. Parallelism and shared state let the new architecture scale intact across GPUs. Quantization-aware training mirrors SGLang's FP4/FP8 rounding, while routing replay reuses the rollout's expert choices. Numerical consistency comes from FP32 and deterministic reductions. Colocated training and rollout fit full-parameter RL on 16 GPUs. Over steps 0–80 of a DAPO run, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78. Blog and cookbook in the comments. 📚
    @kimmonismusDeepSeek just released V4.1-Flash with a new architecture, six weeks after its July V4-Flash update. July’s release improved post-training while keeping the architecture unchanged. (Same with GLM-5.3/Flash) V4.1 introduces a Causal Encoder–Decoder architecture with native visual understanding. DeepSeek reports: - 552B MoE parameters, with 8B active during input processing and 16B during output generation. - KV-cache requirements cut to ¼ of the HBM and ⅛ of the SSD storage versus the previous generation. - Lower API prices. These are *significant* jumps in just a few weeks with post training. This is the new reality we have to adapt to: weekly releases with significant improvements. The company says Flash now beats V4-Pro on capability, cost and speed. Starting September 14, V4-Pro API requests will temporarily route to V4.1-Flash until V4.1-Pro arrives.
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet