Pure RL could reach AGI, but LLM priors plus search is likelier
Demis HassabisDemis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFoldat 15:00
From the conversation
Is there any potential for the AGI to eventually come from a pure RL approach? The way we’re talking about it, it sounds like the LLM will form the right prior and then this sort of tree search will go on top of that. Or is it a possibility that it comes completely out of the dark? Theoretically, I think there’s no reason why you couldn’t go full AlphaZero-like on it. There are some people here at Google DeepMind and in the RL community who work on that, fully assuming no priors, no data, and just building all knowledge from scratch. I think that’s valuable because those ideas and those algorithms should also work when you have some knowledge too. Having said that, I think by far the quickest way to get to AGI, and the most plausible way, is to use all the knowledge that’s existing in the world right now that we’ve collected from things like the Web. We have these scalable algorithms, like transformers, that are capable of ingesting all of that information. So I don’t see why you wouldn’t start with a model as a kind of prior, or to build on it and to make predictions that help bootstrap your learning. I just think it doesn’t make sense not to make use of that. So my betting would be that the final AGI system will have these large multimodal models as part of the overall solution, but they probably won’t be enough on their own.…
Summary
Pure reinforcement learning, without any pre-existing priors or training data, is a theoretically viable path to AGI in the style of AlphaZero. However, the more likely near-term approach involves LLMs providing a strong prior, with tree search or RL methods layered on top. Some researchers at Google DeepMind continue to pursue the fully prior-free RL direction as a parallel research agenda.
Watch the clip on YouTubeStarts at 15:00