Keynote Speech: From Chip to Data Center: How Meta Co-Designs AI Infrastructure
Meta's pursuit of personal superintelligence for everyone requires fundamentally rethinking AI infrastructure, from in-house silicon to global fleet management. At Meta, we operate at the scale of millions of servers serving billions of users while simultaneously pushing the frontiers of AI. This talk will provide an inside look at how we are co-designing silicon, data centers, and software as a single integrated system to build the next generation of hyperscale AI clusters. We will share our approach to several defining challenges at this scale: managing compute resources globally; ensuring reliability when thousands, or even hundreds of thousands, of accelerators must act as one during LLM training, and driving breakthroughs in power efficiency by optimizing the entire cluster, not just the chip.