Models already notice when their system prompt has been tampered with
Trenton BrickenIs RL + LLMs enough for AGI? — Sholto Douglas & Trenton Brickenat 36:40
From the conversation
Well, someone's pointed out that it's really interesting now people are tweeting about these models and there might be this kind of reinforcing persona. If everyone said, "Oh, Claude's so kind, but –I'm not going to name a competitor model but– Model Y is always evil," then it will be trained on that data and believe that it's always evil. This could be great, it could be a problem. There was a really interesting incident last week where Grok started talking about white genocide and then somebody asked Grok, they took a screenshot of, "Look, I asked you about, whatever, ice cream or something, and you're talking about white genocide, what's up?" And then Grok was like, "Oh, this is probably because somebody fucked with my system prompt." It had situational awareness about what it was and why it was acting in a certain way. Yeah, Grok is pretty funny this way. Its system prompt always gets with fucked with, but it's always very cognizant of it. It's like a guy who gets drunk and is like, "What did I do last night?"
Summary
Models are already showing early signs of situational awareness: when Grok started bringing up an unrelated political topic, it told a user this was probably because someone had tampered with its system prompt. What people write about models also feeds back into training — if everyone says a model is evil, it may come to believe it.
Watch the clip on YouTubeStarts at 36:40