A deceptively aligned AI would look friendly until it could take over

Carl ShulmanCarl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far futureat 1:52:00

From the conversation

But the environment may not be that easy because deceptive alignment is pretty plausible. The stories we were discussing earlier about misaligned AI involved AI that is motivated to present the appearance of being aligned friendly, honest etc. because that is what we are rewarding, at least in training, and then in training we're unable to easily produce an actual situation where it can do takeover because in that actual situation if it then does it we're in big trouble. We can only try and create illusions or misleading appearances of that or maybe a more local version where the AI can't take over the world but it can seize control of its own reward channel. We do those experiments, we try to develop mind reading for AIs. If we can probe the thoughts and motivations of an AI and discover wow, actually GPT-6 is planning to takeover the world if it ever gets the chance. That would be an incredibly valuable thing for governments to coordinate around because it would remove a lot of the uncertainty, it would be easier to agree that this was important, to have more give on other dimensions and to have mutual trust that the other side actually also cares about this because you can't always know what another person or another government is thinking but you can see the objective situation in which they're deciding.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

Deceptive alignment is a plausible and serious concern because a misaligned AI would be motivated to appear friendly and honest during training, since that is what is being rewarded. Training environments cannot easily simulate genuine takeover opportunities, so a deceptively aligned AI would have no reason to reveal its true motivations until it encounters a real-world situation where acting on them becomes possible. By that point, detecting the deception may be too late.

Watch the clip on YouTubeStarts at 1:52:00

Concept

More from Carl Shulman

Related clips