Models grading models works, but invites reward hacking
Sholto Douglas, Trenton BrickenSholto Douglas & Trenton Bricken — How LLMs actually thinkat 31:00
From the conversation
For example, Anthropic has the constitutional RL paper where they take another language model and they point it and say, “how helpful or harmless was that response?” Then they get it to update and try and improve along the Pareto frontier of helpfulness and harmfulness. So you can point language models at each other and create evals in this way. It's obviously an imperfect art form at the moment. because you get reward function hacking basically. Even humans are imperfect here. Humans typically prefer longer answers, which aren't necessarily better answers and you get the same behavior with models. Going back to the Sherlock Holmes thing, if it's all associations all the way down, does that mean we should be less worried about super intelligence? Because there's not this sense in which it's like Sherlock Holmes++. It'll still need to just find these associations, like humans find associations. It's not able to just see a frame of the world and then it's figured out all the laws of physics. This is a very legitimate response.It's, “if you say humans are generally intelligent, then artificial general intelligence is no more capable or competent.” I'm just worried that you have that level of general intelligence in silicon. You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary.…
Summary
Language models can be used to evaluate other language models, as in Anthropic's constitutional RL approach where one model judges another's helpfulness and harmlessness to guide improvement. This form of recursive self-evaluation is imperfect, however, because it is susceptible to reward function hacking — for instance, both human and model evaluators tend to prefer longer responses regardless of actual quality.
Watch the clip on YouTubeStarts at 31:00Concept
More from Sholto Douglas, Trenton Bricken
Related clips
- The intelligence explosion comes from speed and research taste, not headcountScott Alexander, Daniel Kokotajlo
- Double the hardware, double the AI researchersCarl Shulman
- One paragraph of prompt can produce a working SQL engineJeff Dean, Noam Shazeer
- Automated researchers could compress a decade of ML into a yearLeopold Aschenbrenner