Local AI’s efficiency gains came more from hardware than models, a user says
Discussing a Stanford University and Together AI paper, the user reports an 18-fold rise in accuracy per joule over 16 months: 5.9-fold from better accelerators versus 3-fold from better models.
TLDR
A user summarizing “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” described as a Stanford University and Together AI paper, credits hardware with the biggest efficiency gains. In a follow-up reply, they report that local AI’s accuracy per joule improved 18-fold in 16 months, with gains of 5.9-fold from better accelerators versus 3-fold from better models. Their summary also says hybrid local-cloud routing reduced energy, compute and cost by 60%–80% against the paper’s batched-cloud baseline. The gains had limits: according to the summary, about 95% of problems in the paper’s hardest reasoning slice remained unsolved by local models.
Combined views
24.6K
2 Sources, first seen 20d ago