Frontier AI infrastructure is not a contest to maximize FLOPS. Amin Vahdat, Google’s Chief Technologist for AI Infrastructure, explains why the more useful measure is goodput: how much productive work a system completes after accounting for failures, communication bottlenecks, and idle resources. That framing makes the episode a practical guide to where large-scale compute is actually won or lost.

Vahdat describes the split between Google’s TPU 8i and 8t lines as an architectural response to different workloads, rather than a one-size-fits-all chip strategy. He also outlines how Google and DeepMind co-design future chips while systems are still being built, allowing workload needs to inform decisions before tape-out. At the cluster level, optical circuit switching can reroute light to spare racks in milliseconds, helping keep a training job productive when hardware or network paths fail.

The discussion extends beyond silicon. Power availability, utility negotiations, and site constraints increasingly determine whether capacity can be brought online—and when. Meanwhile, long-horizon agents are changing the compute mix: persistent, tool-using workloads can raise demand for CPUs and storage alongside accelerators.

For infrastructure builders and investors, the takeaway is to evaluate the whole system: accelerator fit, network resilience, usable power, and workload composition. The scarce resource is not simply peak compute. It is dependable, well-matched capacity that can deliver useful work at scale.

Key Insights

  • Goodput is a more decision-useful measure than peak FLOPS: it captures productive work delivered after accounting for idle capacity, failures, and communication overhead.
  • Google’s TPU 8i and 8t split reflects workload-specific architectural tradeoffs; frontier infrastructure increasingly requires distinct designs rather than one accelerator optimized for every task.
  • Google and DeepMind co-design chips while systems are still in flight, feeding workload requirements into design choices before tape-out instead of treating hardware as a fixed input.
  • Optical circuit switching can reroute light to spare racks in milliseconds, providing a way to preserve cluster productivity when a rack or network path becomes unavailable.
  • Power is a deployment constraint negotiated with utilities, not just a data-center engineering input; access and timing can shape capacity as much as chip supply.
  • Long-horizon agents can increase CPU and storage requirements alongside accelerator demand, so infrastructure planning must account for the full execution stack, not just GPUs or TPUs.

Who should listen: AI infrastructure leaders, accelerator and networking architects, and investors assessing the cost and scalability of frontier-model compute.

Why This Matters

As AI workloads shift from training runs toward persistent, tool-using agents, infrastructure economics are broadening beyond accelerator counts. This episode shows why power access, network resilience, and CPU-and-storage capacity are becoming strategic inputs to frontier-scale deployment.

Listen to the full episode →