Models grading models works, but invites reward hacking

Sholto Douglas, Trenton BrickenSholto Douglas & Trenton Bricken — How LLMs actually thinkat 31:00

From the conversation

For example, Anthropic has the constitutional RL paper where they take another language model and they point it and say, “how helpful or harmless was that response?” Then they get it to update and try and improve along the Pareto frontier of helpfulness and harmfulness. So you can point language models at each other and create evals in this way. It's obviously an imperfect art form at the moment. because you get reward function hacking basically. Even humans are imperfect here. Humans typically prefer longer answers, which aren't necessarily better answers and you get the same behavior with models. Going back to the Sherlock Holmes thing, if it's all associations all the way down, does that mean we should be less worried about super intelligence? Because there's not this sense in which it's like Sherlock Holmes++. It'll still need to just find these associations, like humans find associations. It's not able to just see a frame of the world and then it's figured out all the laws of physics. This is a very legitimate response.It's, “if you say humans are generally intelligent, then artificial general intelligence is no more capable or competent.” I'm just worried that you have that level of general intelligence in silicon. You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

Language models can be used to evaluate other language models, as in Anthropic's constitutional RL approach where one model judges another's helpfulness and harmlessness to guide improvement. This form of recursive self-evaluation is imperfect, however, because it is susceptible to reward function hacking — for instance, both human and model evaluators tend to prefer longer responses regardless of actual quality.

Watch the clip on YouTubeStarts at 31:00

Concept

More from Sholto Douglas, Trenton Bricken

Related clips