Single Modal deployment reportedly served up to half a trillion tokens in a day
A Modal team member says smarter routing reduced uneven server load and long waits for the first token, alongside broader gains in interactivity and throughput.
TLDR
A Modal team member credits optimizations and the company's autoscaling global GPU fleet with serving as much as half a trillion tokens in one day for a single deployment. The account highlights tiered KV caching and smarter routing, alongside speculative decoding. Performance gains on a single replica did not translate directly to production, the author said, with occasional spikes in time to first token and generally lower per-replica throughput. Routing that accounts for load and KV-cache state reduced uneven load and long waits for the first token, according to the account. The author described the routing as an experimental option for Modal Servers.
Combined views
1.2K
6 Sources, first seen 8h ago
Single Modal deployment reportedly served up to half a trillion tokens in a day
A Modal team member says smarter routing reduced uneven server load and long waits for the first token, alongside broader gains in interactivity and throughput.
TLDR
A Modal team member credits optimizations and the company's autoscaling global GPU fleet with serving as much as half a trillion tokens in one day for a single deployment. The account highlights tiered KV caching and smarter routing, alongside speculative decoding. Performance gains on a single replica did not translate directly to production, the author said, with occasional spikes in time to first token and generally lower per-replica throughput. Routing that accounts for load and KV-cache state reduced uneven load and long waits for the first token, according to the account. The author described the routing as an experimental option for Modal Servers.
