• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Power, hardware failures and specialized TPUs in Google's AI buildout

    A post about Google's AI infrastructure says something fails multiple times an hour at a scale of 100,000 accelerators.

    Sonya Huang 🐥SH
    1 Source, ,

    TLDR

    A post about Google's AI infrastructure buildout says Google forecasts $200 billion in 2026 capital spending. It describes “goodput” per watt as a key metric and says something fails multiple times an hour at a scale of 100,000 accelerators. It also discusses separate TPUs for training and inference, rising demand from long-horizon agents and Google's preference for grid-connected power.

    Combined views

    2.5K

    1 Source, first seen 2h ago

    Combined views

    2.5K

    1 Source, first seen 2h ago

    6 likes
    2h ago
    first seen 2h ago
    6 likes
    4 comments
    9 saves
    2 reposts
    4 comments
    9 saves
    2 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Sonya Huang 🐥@sonyatweetybirdAI is driving the biggest infrastructure build-out in history, with @Google alone forecasting $200B of capex in 2026E. What are the real constraints on a system at this scale? Amin Vahdat is the chief in charge of Google's buildout. As SVP AI Infrastructure (AI^2), he role spans @GoogleDeepMind, @googlecloud, and the TPU/accelerator hardware roadmap, planning chips 2-5 years out and overseeing a data center buildout of epic proportions. Takeaways: — Why "goodput" per watt is the north star metric and designing around failures: At the 100,000-accelerator scale, something fails multiple times an hour, capacity has to double every six months, and as much of that comes from software and model optimization as from silicon. Hardware is a multiplier that lifts everyone above it — Why Google split TPU into two chips (8I for inference, 8T for training), and the calculus for when a workload is big and durable enough to specialize for — The "bitter lesson of chips," and why the TPU architecture hasn't fundamentally changed since v1 — How DeepMind and the hardware team co-design: delaying a tape-out by two weeks for the right model optimization, and using Gemini to design chips for future Geminis — Long-horizon agents as a new workload shape: no human in the loop to rate-limit requests, and CPU, networking and storage demand going "through the roof" alongside accelerators — Optical circuit switching: mirrors redirecting light to swap a failed TPU rack in milliseconds, without touching a fibe — Why Google prefers grid-connected over behind-the-meter power — Why 7- and 8-year-old TPUs are still at 100% utilization — Open standards, and why IP won the internet as the narrow waist of the hourglass — Orbital data centers: 1.4x the solar power, near-100% sunlight in sun-synchronous orbit, free-space optics between satellites, and "no fundamental showstoppers" — The 2036 rack: multiple megawatts, wheeled in, plug in water, power and fiber Timestamps: 0:00 – Introduction 1:47 – What makes a data center an AI data center 5:30 – Goodput, not FLOPS: holding yourself accountable when something fails every hour 11:52 – Doubling token capacity every six months, and where the gains actually come from 15:32 – The TPU bet: from a contrarian call in 2013 to splitting 8i and 8t 23:30 – The case for and against co-design 26:11 – Shoulder to shoulder with DeepMind: intercepting chips mid-flight 34:16 – Long-horizon agents change the shape of the data center 37:50 – Optical circuit switching and the state of networking 42:35 – Power is the binding constraint: utilities, gigawatts, and sizing a data center 49:23 – Training vs. serving clusters, seven-year-old TPUs, and open standards 58:29 – Orbital data centers and the supercomputer of 20362h

    1 Source

    Sonya Huang 🐥@sonyatweetybirdAI is driving the biggest infrastructure build-out in history, with @Google alone forecasting $200B of capex in 2026E. What are the real constraints on a system at this scale? Amin Vahdat is the chief in charge of Google's buildout. As SVP AI Infrastructure (AI^2), he role spans @GoogleDeepMind, @googlecloud, and the TPU/accelerator hardware roadmap, planning chips 2-5 years out and overseeing a data center buildout of epic proportions. Takeaways: — Why "goodput" per watt is the north star metric and designing around failures: At the 100,000-accelerator scale, something fails multiple times an hour, capacity has to double every six months, and as much of that comes from software and model optimization as from silicon. Hardware is a multiplier that lifts everyone above it — Why Google split TPU into two chips (8I for inference, 8T for training), and the calculus for when a workload is big and durable enough to specialize for — The "bitter lesson of chips," and why the TPU architecture hasn't fundamentally changed since v1 — How DeepMind and the hardware team co-design: delaying a tape-out by two weeks for the right model optimization, and using Gemini to design chips for future Geminis — Long-horizon agents as a new workload shape: no human in the loop to rate-limit requests, and CPU, networking and storage demand going "through the roof" alongside accelerators — Optical circuit switching: mirrors redirecting light to swap a failed TPU rack in milliseconds, without touching a fibe — Why Google prefers grid-connected over behind-the-meter power — Why 7- and 8-year-old TPUs are still at 100% utilization — Open standards, and why IP won the internet as the narrow waist of the hourglass — Orbital data centers: 1.4x the solar power, near-100% sunlight in sun-synchronous orbit, free-space optics between satellites, and "no fundamental showstoppers" — The 2036 rack: multiple megawatts, wheeled in, plug in water, power and fiber Timestamps: 0:00 – Introduction 1:47 – What makes a data center an AI data center 5:30 – Goodput, not FLOPS: holding yourself accountable when something fails every hour 11:52 – Doubling token capacity every six months, and where the gains actually come from 15:32 – The TPU bet: from a contrarian call in 2013 to splitting 8i and 8t 23:30 – The case for and against co-design 26:11 – Shoulder to shoulder with DeepMind: intercepting chips mid-flight 34:16 – Long-horizon agents change the shape of the data center 37:50 – Optical circuit switching and the state of networking 42:35 – Power is the binding constraint: utilities, gigawatts, and sizing a data center 49:23 – Training vs. serving clusters, seven-year-old TPUs, and open standards 58:29 – Orbital data centers and the supercomputer of 20362h