Evals must catch capabilities that hide other capabilities

Holden KarnofskyAI doesn't need to be 'superintelligent' to take over + AI safety playbook | Holden Karnofsky (2023)at 1:24:40

From the conversation

Holden Karnofsky: The way I would carve up the space is capability evals, like: “Is this AI capable enough to do something scary?” — forgetting about whether it wants to. Capability evals are like: Could an AI make a bunch of copies of itself, if a human tried to get it to do it? Could an AI design a bioweapon? Those are capability evals. Then there’s alignment evals. That’s like: “Does this AI actually do what it’s supposed to do, or does it have some weird goals of its own?” So the stuff you talked about with model organisms would be more of an alignment eval, the way you described it, and autonomous replication is a capability eval. I think a very important subcategory of capability evals is what I call “meta-capability evals” or “meta-dangerous capabilities” — which is basically any ability an AI system has that would make it very hard to get confident about what other abilities it has. An example would be what I’m currently tentatively calling “unauthorised proliferation,” so an AI model that can walk a human through building a powerful AI model of their own that is not subject to whatever restrictions and controls the original one is subject to. That could be a very dangerous capability. Like, you could say, it can design a bioweapon, but it always refuses to do so. But it could also help the human build an AI that we don’t know what the hell that thing can do.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

Model evaluation should be divided into capability evaluations (testing whether AI can perform dangerous tasks like bioweapon design or autonomous replication) and alignment evaluations (testing whether AI follows intended goals or has misaligned objectives). A critical subcategory is "meta-capability" evaluations, which assess abilities that make other capabilities hard to evaluate—such as unauthorized proliferation, where an AI helps humans build unrestricted AI systems that bypass safety controls.

Watch the clip on YouTubeStarts at 1:24:40

Concept

More from Holden Karnofsky

Related clips