• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DeepSeek's V4.1 pricing has one user wondering about its margins

    The user claims V4.1 is cheaper to serve than V4-Flash but costs twice as much per output token even off-peak.

    HF
    RA
    AC
    51 Sources, ,

    TLDR

    One user argues that DeepSeek's V4.1 is cheaper to serve than V4-Flash, citing four times less cache and a better caching system, while claiming its output-token price is twice as high even off-peak. The user also puts DeepSeek's margins with V4 at roughly 80% at its cheapest and wonders where margins stand now.

    Combined views

    757.2K

    51 Sources, first seen 22d ago

    Combined views

    757.2K

    51 Sources, first seen 22d ago

    9.1K likes
    22d ago
    first seen 22d ago
    9.1K likes
    672 comments
    1K saves
    1.2K reposts
    672 comments
    1K saves
    1.2K reposts

    Sentiment

    Positive71.1%28.9%Negative

    Summary

    Sentiment

    Positive71.1%28.9%Negative

    Many accounts welcomed DeepSeek V4.1 Flash for its speed, 1M context, and lower cost on agentic benchmarks, while some replies called the architecture an inference nightmare or comprehension downgrade and questioned its origins.

    Based on 298 sentiment-bearing replies from 273 accounts across 11 conversations.

    Summary

    Many accounts welcomed DeepSeek V4.1 Flash for its speed, 1M context, and lower cost on agentic benchmarks, while some replies called the architecture an inference nightmare or comprehension downgrade and questioned its origins.

    Based on 298 sentiment-bearing replies from 273 accounts across 11 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    51 Sources

    @teortaxesTexvery interesting. V4-Flash-Vision [Intermediate] *New model arch*, *faster*; stronger; same price. But – 20 concurrent requests. For reference, it's 500 for Pro and 2500 for Flash now. So it's not about batch size fiddling, I guess. Deeper diffusion integration in DSpark2?
    @zephyr_z9No point in pushing Pro It's a bad pre train
    @kimmonismusDeepSeek Flash 4.1 is already being tested via API and rolling out. Faster and new architecture. Official release post probably very soon.
    @novita_labs🤗 Novita now supports DeepSeek-V4.1-Flash on @huggingface. • 552B backbone parameters • Native image and text input • Up to a 1M-token context window • Continuously controllable reasoning effort
    @nutlopeThis model is insane at landing pages. I asked DeepSeek V4.1 Flash & Claude Fable 5 to build me a landing page for a movie theater. Fable cost $1.21 while V4.1 Flash cost 2.6 cents, making it more than 40x cheaper at similar quality. Gave both the exact same prompt!
    @TekniumDeepseek V4.1 Flash is an incredibly powerful model!
    @NielsRoggeFor folks wondering what YOCO means, you can simply ask it on Papers with Code :) It powers the new DeepSeek-V4.1-Flash architecture. It's not an encoder-decoder, but rather a decoder-decoder architecture. It's all about reducing GPU memory and prefill latency. Find the chat here: https://paperswithcode.co/share/a943da6a-6b30-4f79-b20c-cc3970068d35
    @QuantumTransfdeepseek 这次事情我真的觉得挺讽刺的 一开始 pro 换 v4.1 flash:会影响生产,eval 指标更好不代表实际表现更好 后面听从用户意见,继续路由到原 pro,又变成:言而无信,朝令夕改,是不是发现新 flash 模型能力还不如 pro 只能说最好是不要把这种公众意见听进去,认定了怎么做就坚定做下去,隔壁 anthropic 就把这一点贯彻的很彻底 hhh
    @altryneChris Alexiuk (@llm_wizard): DeepSeek V4.1 Flash baked disaggregation into the model itself. Prefill and decode are different jobs. Whale put that split in the weights, not just the serving stack.
    @achowdheryRT @SemiAnalysis_: Congrats to @deepseek_ai on releasing DeepSeek-V4.1-Flash! > 552B backbone, with a causal encoder-decoder activating ju…

    51 Sources

    @teortaxesTexvery interesting. V4-Flash-Vision [Intermediate] *New model arch*, *faster*; stronger; same price. But – 20 concurrent requests. For reference, it's 500 for Pro and 2500 for Flash now. So it's not about batch size fiddling, I guess. Deeper diffusion integration in DSpark2?
    @zephyr_z9No point in pushing Pro It's a bad pre train
    @kimmonismusDeepSeek Flash 4.1 is already being tested via API and rolling out. Faster and new architecture. Official release post probably very soon.
    @novita_labs🤗 Novita now supports DeepSeek-V4.1-Flash on @huggingface. • 552B backbone parameters • Native image and text input • Up to a 1M-token context window • Continuously controllable reasoning effort
    @nutlopeThis model is insane at landing pages. I asked DeepSeek V4.1 Flash & Claude Fable 5 to build me a landing page for a movie theater. Fable cost $1.21 while V4.1 Flash cost 2.6 cents, making it more than 40x cheaper at similar quality. Gave both the exact same prompt!
    @TekniumDeepseek V4.1 Flash is an incredibly powerful model!
    @NielsRoggeFor folks wondering what YOCO means, you can simply ask it on Papers with Code :) It powers the new DeepSeek-V4.1-Flash architecture. It's not an encoder-decoder, but rather a decoder-decoder architecture. It's all about reducing GPU memory and prefill latency. Find the chat here: https://paperswithcode.co/share/a943da6a-6b30-4f79-b20c-cc3970068d35
    @QuantumTransfdeepseek 这次事情我真的觉得挺讽刺的 一开始 pro 换 v4.1 flash:会影响生产,eval 指标更好不代表实际表现更好 后面听从用户意见,继续路由到原 pro,又变成:言而无信,朝令夕改,是不是发现新 flash 模型能力还不如 pro 只能说最好是不要把这种公众意见听进去,认定了怎么做就坚定做下去,隔壁 anthropic 就把这一点贯彻的很彻底 hhh
    @altryneChris Alexiuk (@llm_wizard): DeepSeek V4.1 Flash baked disaggregation into the model itself. Prefill and decode are different jobs. Whale put that split in the weights, not just the serving stack.
    @achowdheryRT @SemiAnalysis_: Congrats to @deepseek_ai on releasing DeepSeek-V4.1-Flash! > 552B backbone, with a causal encoder-decoder activating ju…