Most discussions about AI scaling bottlenecks stop at the obvious: Nvidia's order books and TSMC's advanced packaging capacity. Dylan Patel goes three layers deeper, tracing the constraint to its true origin. This conversation is a masterclass in semiconductor supply chain physics and economics, mapping the entire value chain from gigawatt-scale data centers down to the fundamental limits of high-NA EUV lithography. Patel explains why the bottlenecks have shifted from logic to memory and, critically, why they will ultimately settle on ASML by the end of the decade. You'll walk away with a precise understanding of why an H100 appreciates over time, how hyperscaler capital allocation is reshaping the foundry landscape, and why the memory industry is heading for a crunch that will rival the logic shortage. There is no hype here, just a rigorous, first-principles breakdown of the capital expenditure, lead times, and physical constraints that will define the next five years of AI compute scaling.

Key Takeaways

  • The ultimate bottleneck for AI compute by 2028-2030 is not TSMC's wafer capacity but ASML's ability to ship high-NA EUV tools, which are required for the next node shrinks and have a decade-long supply chain inertia.
  • Nvidia's H100 GPUs are one of the few technology assets that appreciate after purchase because the underlying wafer and advanced packaging supply is so constrained that spot market access commands a premium over contracted pricing.
  • The memory industry is structurally underinvested and faces an enormous incoming crunch, as HBM (High Bandwidth Memory) consumes a disproportionate share of wafer capacity relative to standard DRAM, squeezing supply for the broader market.
  • Scaling power in the US for gigawatt-class data centers is not the primary bottleneck it is often portrayed to be, with the real lead time constraints residing in the semiconductor fabrication equipment supply chain rather than utility infrastructure.
  • TSMC's strategic allocation of leading-edge nodes like N2 is a zero-sum game, forcing companies like Apple and Google to compete directly for capacity, with hyperscaler demand now large enough to displace traditional smartphone volume commitments.

Who should watch: Infrastructure engineers, hardware procurement leads, and technical founders who need to model the true cost and availability curves for AI compute over a 3-7 year horizon.

Why This Matters

Patel's analysis reframes AI scaling as a supply chain physics problem rather than a software or capital problem, implying that the companies with the deepest ASML and TSMC relationships, not the largest GPU clusters, will control the pace of frontier model development.

Watch the full video →