Local AI’s “intelligence per watt” improved 5.3× from 2023 to 2025, a user reports
Summarizing a Stanford University and Together AI paper, a user says routing queries between local and cloud models cut energy, compute and cost by 60%–80% versus the paper’s batched-cloud baseline.
TLDR
A user sharing findings from “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI” describes substantial gains for local models. According to their summary of the Stanford University and Together AI paper, locally serviceable query coverage rose from 23.2% to 71.3% between 2023 and 2025, alongside a 5.3× improvement in intelligence per watt. The summary also reports that hybrid local-cloud routing reduced energy, compute and cost by 60%–80% against the paper’s batched-cloud baseline. Hard reasoning remained a weakness: the user says about 95% of problems in the paper’s hardest reasoning slice were still unsolved by local models.
Combined views
10.1K
3 Sources, first seen 19d ago