Databricks Claims Top Spot for Kimi K3 Inference Speed
Databricks engineer states the company leads Artificial Analysis for Kimi K3 inference speed.
TLDR
Yuchen Jin, a Databricks engineer, posted that the company leads Artificial Analysis rankings for Kimi K3 inference speed and latency at 239 tokens per second. He noted the model has 2.8 trillion parameters and is the largest open-source model Databricks has served. Matei Zaharia retweeted the post. Jonathan Frankle retweeted a separate post by Ali Ghodsi that linked to the Artificial Analysis page and stated Databricks has the fastest and lowest latency on Kimi K3. Other replies discussed the result but added no new performance data.
Combined views
93K
6 Sources, first seen 58d ago
