Announcement
DeepSeek-V4.1-Flash on vLLM is claimed to run 1.9× faster at low concurrency
The vLLM project says it delivers 5.3× the throughput at 150 tokens per second per user on SemiAnalysis AgentX.
TLDR
The vLLM project says that three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 tokens per second per user on SemiAnalysis AgentX. It links to a walkthrough with interactive figures.
Combined views
12.5K
3 Sources, first seen ago