PhoneLLM Hits Sub-600ms P95 Latency at Low Cost
Alex Volkov posted performance figures for the open-weights model on X.
Alex Volkov, an AI podcaster who hosts the ThursdAI podcast and works at Weights & Biases on Weave, posted on X that PhoneLLM delivers P95 inference latency under roughly 600 milliseconds. He added that the cost runs about a quarter cent per minute and attached a video. The packet records these statements as Volkov's claims about the open-weights model. No other sources or corroboration appear in the supplied lines.
Combined views
1.1K
1 post, first seen 7h ago
PhoneLLM Hits Sub-600ms P95 Latency at Low Cost
Alex Volkov posted performance figures for the open-weights model on X.
Alex Volkov, an AI podcaster who hosts the ThursdAI podcast and works at Weights & Biases on Weave, posted on X that PhoneLLM delivers P95 inference latency under roughly 600 milliseconds. He added that the cost runs about a quarter cent per minute and attached a video. The packet records these statements as Volkov's claims about the open-weights model. No other sources or corroboration appear in the supplied lines.
