If humans can't settle alignment debates, they can't verify AI's proposals

Eliezer YudkowskyEliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationalityat 44:40

From the conversation

So in alignment, the thing hands you a thing and says “this will work for aligning a super intelligence” and it gives you some early predictions of how the thing will behave when it’s passively safe, when it can’t kill you. That all bear out and those predictions all come true. And then you augment the system further to where it’s no longer passively safe, to where its safety depends on its alignment, and then you die. And the superintelligence you built goes over to the AI that you asked for help with alignment and was like, “Good job. Billion dollars.” That’s observation number one. Observation number two is that for the last ten years, all of effective altruism has been arguing about whether they should believe or Paul Christiano, right? That’s two systems. I believe that Paul is honest. I claim that I am honest.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

For ten years, the effective altruism community has been unable to determine whether to believe Yudkowsky or Christiano on alignment — and this is between two honest humans. The problem of verifying AI-generated alignment proposals will be vastly harder.

Watch the clip on YouTubeStarts at 44:40

Concept

More from Eliezer Yudkowsky

Related clips