PhoneLLM Hits Sub-600ms P95 Latency at Low Cost
Alex Volkov posted performance figures for the open-weights model on X.
TLDR
Alex Volkov, an AI podcaster who hosts the ThursdAI podcast and works at Weights & Biases on Weave, posted on X that PhoneLLM delivers P95 inference latency under roughly 600 milliseconds. He added that the cost runs about a quarter cent per minute and attached a video. The packet records these statements as Volkov's claims about the open-weights model. No other sources or corroboration appear in the supplied lines.
Combined views
1.8K
1 Source, first seen 31d ago