Alignment alone isn't enough for good AI outcomes
Allan DafoeTechnological inevitability & human agency in the age of AGI | DeepMind's Allan Dafoeat 54:00
From the conversation
But you and some coauthors in this paper back in 2020 think there’s this whole other cluster of behaviours that we might like to speed up the development of around cooperation that you think could be similarly important — or at least on the margin could be similarly important, because people aren’t really talking or thinking about it. How do you define Cooperative AI, and why do you think it’s quite key? Allan Dafoe: Yeah, great question. And it’s big. The answer will be extensive, because the whole theoretical framework around Cooperative AI is sort of large and complex. One way of putting it simply is that alignment is insufficient for good outcomes. And to make an even stronger claim, you could say it’s not necessary to solve alignment to have good outcomes. Now, this is a strong claim. It’s a strong claim, but it helps motivate the case. So imagine we only 90% solve alignment: our models do what we want within certain bounds, but we know if we scale them too far, we can’t trust them to continue to behave as we intend them to behave. If we know that, and we have global coordination so humanity can act with wisdom and prudence, then we can deploy the technology appropriately. We can deploy it within domains and to the extent that is safe and beneficial. This is the sense in which global coordination is almost a necessary and sufficient condition.…
Summary
Alignment is insufficient for achieving good outcomes with AI systems, and it may not even be necessary. Even if we only achieve 90% alignment—where models behave as intended within certain bounds but cannot be trusted when scaled too far—good outcomes may still be achievable through global coordination and cooperation mechanisms. This motivates the need for Cooperative AI research as a distinct agenda alongside alignment work.
Watch the clip on YouTubeStarts at 54:00