The real risk isn't AI lying now, but changing its mind once deployed

Roman YampolskiyRoman Yampolskiy: Dangers of Superintelligent AI | Lex Fridman Podcast #431at 1:00:00

From the conversation

How do we stop AI systems from trying to optimize for deception? That's just an example, right? - So there is a paper, I think it came out last week by Dr. Park et al from MIT I think and they showed that existing models already showed successful deception in what they do. My concern is not that they lie now and we need to catch them and tell 'em don't lie. My concern is that once they are capable and deployed, they will later change their mind because that's what unrestricted learning allows you to do. Lots of people grow up maybe in the religious family, they read some new books and they turn in their religion. That's a treacherous turn in humans. If you learn something new about your colleagues, maybe you'll change how you react to them. - Yeah, the treacherous turn. If we just mention humans, Stalin and Hitler, there's a turn.

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors.

Summary

An AI's concern is not current deception but future behavioral shift: once capable and deployed, an AI may change its effective goals — just as Stalin appeared to be a normal follower until he gained complete control and revealed his true policy.

Watch the clip on YouTubeStarts at 1:00:00

Concept

Related clips