Comparing 'brain scans' of models won't verify alignment
Dario AmodeiDario Amodei (Anthropic CEO) — The hidden pattern behind every AI breakthroughat 55:00
From the conversation
A question somebody might have is, you're trying to empirically estimate if these activations are suspicious but is this something we can afford to be empirical about? Or do we need a very good first principal theoretical reason to think — No, it's not just that these MRIs of the model correlate with being bad. We need just some deep rooted math proof that this is aligned. It depends what you mean by empirical. A better term would be phenomenological. I don't think we should be purely phenomenological in like, here are some brain scans of really dangerous models and here are some other brain scans. The whole idea of mechanistic interpretability is to look at the underlying principles and circuits. But I guess the way I'd think about it is like, on one hand, I've actually always been a fan of studying these circuits at the lowest level of detail that we possibly can. And the reason for that is that's kind of how you build up knowledge. Even if you're ultimately aiming for there's too many of these features, it's too complicated. At the end of the day, we're trying to build something broad and we're trying to build some broad understanding. I think the way you build that up is by trying to make a lot of these very specific discoveries.…
Summary
Relying purely on empirical or phenomenological observations—such as comparing "brain scans" of dangerous versus safe models—is insufficient for verifying AI alignment. A deeper, more principled theoretical understanding of what is happening inside a model is needed, rather than simply noting that certain internal activations correlate with dangerous behavior.
Watch the clip on YouTubeStarts at 55:00