Open-ended goals like 'make money online' are powerful and risky
Sholto Douglas, Trenton BrickenIs RL + LLMs enough for AGI? — Sholto Douglas & Trenton Brickenat 44:00
From the conversation
It might be like “achieve some goal”. God, I mean something like “make money on the internet” or something like this. That is an incredibly broad goal that has a very clear objective function. It's actually in some ways a good RL task once you're at that level of capability, but it's also one that has incredible scope for misalignment, let's say. Totally. I feel like we optimize humans for specific objectives all the time. Sometimes it goes off the rails, obviously, but I don't know… You could make a theoretical argument that you teach a kid to make a lot of money when he grows up and a lot of smart people are imbued with those values and just rarely become psychopaths or something. But we have so many innate biases to follow social norms. I mean Joe Heinrich's The Secret of our Success is all about this. Even if kids aren't in the conventional school system, I think it's sometimes noticeable that they aren't following social norms in the same ways. The LLM definitely isn't doing that. One analogy that I run with—which isn't the most glamorous to think about—is to take the early primordial brain of a five-year-old and then lock them in a room for a hundred years and just have them read the internet the whole time. It's already happening. No, but they're locked in a room, you're putting food through a slot and otherwise they're just reading the internet.…
Summary
Broad goals like "make money on the internet" could serve as effective RL objectives for highly capable AI systems, but they also carry significant potential for misalignment due to their open-ended scope. A parallel can be drawn to how humans are often optimized for specific objectives, which sometimes produces unintended outcomes. This suggests that the same mechanisms that make goal-directed training powerful also introduce alignment risks.
Watch the clip on YouTubeStarts at 44:00Concept
More from Sholto Douglas, Trenton Bricken
Related clips
- Misalignment likely turns catastrophic before most other destructive technologiesPaul Christiano
- A deceptively aligned AI would look friendly until it could take overCarl Shulman
- Who writes the moral constitution of powerful AI?Dwarkesh Patel
- Let AI hunt for vulnerabilities only inside an air-gapped boxEliezer Yudkowsky