Red-team AI systems to learn whether they can fool your reviewers

Adam GleaveCan Defense in Depth Work for AI? (with Adam Gleave)at 47:20

From the conversation

How can you be confident that PR does not have some adversarial attack that fools your automated reviewer and also is really persuasive to human beings? I think it's much harder to have like, you know, confidence in that. Um, it is a meaningfully harder thing for a system to pull off. So maybe one area I would be sort of more optimistic than speed is I I just don't see where these sort of AI systems get really smart and collude with each other to like not try and pull off an attack before they're smart enough to know that they can definitely succeed. Especially if you just fine-tune the systems to be a little bit um you know adversarial and see what you can they can do. You should at least be able to get a warning sign that these systems are capable of fooling you and breaking your security even if you're unsure whether your current system is is actually just honest and aligned or is is scheming. um that that could be hard to distinguish, but at least eliciting these things don't seem that hard. Um but it's yeah, in in most kind of like unconstrained deployment scenarios, I think you could definitely run into some problems, but I I think if you spoke to the kind of AI control people, and I I don't work in AI control, so I'll do a bad job representing it.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

We put an AI-powered firewall on the output coming out of these systems and review the PR. But how can you be confident the PR does not have some adversarial attack that fools your automated reviewer and is also really persuasive to humans? If you just fine-tune systems to be a little adversarial and see what they can do, you should at least get a warning sign that these systems are capable of fooling you and breaking your security.

Watch the clip on YouTubeStarts at 47:20

Concept

Related clips