• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Local AI’s “intelligence per watt” improved 5.3× from 2023 to 2025, a user reports

    Summarizing a Stanford University and Together AI paper, a user says routing queries between local and cloud models cut energy, compute and cost by 60%–80% versus the paper’s batched-cloud baseline.

    RP
    3 Sources, 19d ago, first seen 19d ago

    TLDR

    A user sharing findings from “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI” describes substantial gains for local models. According to their summary of the Stanford University and Together AI paper, locally serviceable query coverage rose from 23.2% to 71.3% between 2023 and 2025, alongside a 5.3× improvement in intelligence per watt. The summary also reports that hybrid local-cloud routing reduced energy, compute and cost by 60%–80% against the paper’s batched-cloud baseline. Hard reasoning remained a weakness: the user says about 95% of problems in the paper’s hardest reasoning slice were still unsolved by local models.

    Combined views

    10.1K

    3 Sources, first seen 19d ago

    Combined views

    10.1K

    3 Sources, first seen 19d ago

    189 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    189 likes
    21 comments
    125 saves
    70 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    21 comments
    125 saves
    70 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    @rohanpaul_aiAbsolutely beautiful new paper from Stanford Univ + Together AI "Intelligence per Watt: Measuring Intelligence Efficiency of Local AI" - It found that hybrid local-cloud routing reduced energy, compute, and cost by 60% to 80% against its batched-cloud baseline. - From 2023 to 2025, local AI’s intelligence-per-watt improved 5.3×, while locally serviceable query coverage jumped from 23.2% to 71.3%. - An iPhone 16 Pro achieved roughly 7× higher intelligence-per-watt than workstation GPUs on the same model and precision, showing how power-efficient mobile AI can be for lightweight queries. - A diverse pool of 20+ local models actually beat the paper’s 3 frontier cloud models on 3 of 4 benchmarks when each query was routed to the best model, showing how much model diversity can matter. - Dropping precision from FP16 to FP4 cut inference energy by 3X–3.5X, while costing roughly 2.5 percentage points of accuracy per precision step; in one test, a larger FP4 model even beat a smaller FP16 model. - The biggest remaining weakness is concentrated at the hard end: on the paper’s hardest reasoning slice, about 95% of problems were still unsolved by local models, even as easier and medium-difficulty tasks improved rapidly.

    3 Sources

    @rohanpaul_aiAbsolutely beautiful new paper from Stanford Univ + Together AI "Intelligence per Watt: Measuring Intelligence Efficiency of Local AI" - It found that hybrid local-cloud routing reduced energy, compute, and cost by 60% to 80% against its batched-cloud baseline. - From 2023 to 2025, local AI’s intelligence-per-watt improved 5.3×, while locally serviceable query coverage jumped from 23.2% to 71.3%. - An iPhone 16 Pro achieved roughly 7× higher intelligence-per-watt than workstation GPUs on the same model and precision, showing how power-efficient mobile AI can be for lightweight queries. - A diverse pool of 20+ local models actually beat the paper’s 3 frontier cloud models on 3 of 4 benchmarks when each query was routed to the best model, showing how much model diversity can matter. - Dropping precision from FP16 to FP4 cut inference energy by 3X–3.5X, while costing roughly 2.5 percentage points of accuracy per precision step; in one test, a larger FP4 model even beat a smaller FP16 model. - The biggest remaining weakness is concentrated at the hard end: on the paper’s hardest reasoning slice, about 95% of problems were still unsolved by local models, even as easier and medium-difficulty tasks improved rapidly.