• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    DeepSeek-V4.1-Flash on vLLM is claimed to run 1.9× faster at low concurrency

    The vLLM project says it delivers 5.3× the throughput at 150 tokens per second per user on SemiAnalysis AgentX.

    Woosuk KwonWK
    vLLMVL
    3 Sources, ,

    TLDR

    The vLLM project says that three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 tokens per second per user on SemiAnalysis AgentX. It links to a walkthrough with interactive figures.

    Combined views

    12.5K

    3 Sources, first seen 3h ago

    Combined views

    12.5K

    3 Sources, first seen 3h ago

    93 likes
    3h ago
    first seen 3h ago
    93 likes
    5 comments
    24 saves
    21 reposts
    5 comments
    24 saves
    21 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    vLLM@vllm_project1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on @SemiAnalysis_ AgentX. Here is how, with interactive figures you can step through 🧵 https://vllm.ai/blog/2026-10-07-deepseek-v41-flash3h
    Woosuk Kwon@woosuk_kRT @vllm_project: 1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughp…3h

    3 Sources

    vLLM@vllm_project1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on @SemiAnalysis_ AgentX. Here is how, with interactive figures you can step through 🧵 https://vllm.ai/blog/2026-10-07-deepseek-v41-flash3h
    Woosuk Kwon@woosuk_kRT @vllm_project: 1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughp…3h