One giant cluster buys scale but creates a single point of failure
Asianometry, Dylan Patel@Asianometry & Dylan Patel — How the semiconductor industry actually worksat 17:00
From the conversation
He's trying to deal with all these different things. Suddenly, you have a single point of failure, and that's bad. On the flip side, there are obviously immense gains from being centralized because of the scaling laws. The flip side is compute efficiency which is obviously going to be hurt because you can't experiment and have different people lead and try their efforts as much if you're more centralized. There is a balancing act there. That is actually really interesting, the fact that they can centralize. I didn't think about this. Even if America as a whole is getting millions of GPUs a year, the fact that any one company is only getting hundreds of thousands or fewer means that there's no one person who can do a single training run as big in America as if China as a whole decides to do one together. The 10 gigawatts you mentioned near the Three Gorges Dam, how widespread is it? Is it a state? Is it like one wire? Would you do a sort of distributed training run? It’s not just the dam itself, but also all of the coal. There's some nuclear reactors there as well I believe. Between all of that and renewables like solar and wind, in that region there is an absurd amount of concentrated power that could be built. I'm not saying it's like one button, but it's more like, "hey within X mile radius." That's more of the correct way to frame it.…
Summary
Centralizing AI compute into a single large cluster creates a single point of failure while simultaneously enabling gains from scaling laws. There is a trade-off between centralization's scaling benefits and the compute efficiency and experimentation advantages that come from distributing resources across multiple independent efforts.
Watch the clip on YouTubeStarts at 17:00