GPT-3, not GPT-2, was when language models became serious
John SchulmanJohn Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGIat 49:00
From the conversation
But when GPT-2 was done, I wasn't completely sold on it being revolutionizing everything. It was really after GPT-3 that I pivoted what I was working on and what my team was working on. After that, we got together and said, "oh yeah, let's see what we can do here with this language model stuff." But after GPT-2, I wasn't quite sure yet. Let’s say the stuff we were talking about earlier with RL starts working better with these smarter models. Does the fraction of compute that is spent on pre-training versus post-training change significantly in favor of post-training in the future? There are some arguments for that. Right now it's a pretty lopsided ratio. You could argue that the output generated by the model is higher quality than most of what's on the web. So it makes more sense for the model to think by itself rather than just training to imitate what's on the web. So I think there's a first principles argument for that. We found a lot of gains through post-training. So I would expect us to keep pushing this methodology and probably increasing the amount of compute we put into it. The current GPT-4 has an Elo score that is like a hundred points higher than the original one that was released. Is that all because of what you're talking about, with these improvements that are brought on by post-training? Yeah, most of that is post-training.…
Summary
GPT-3, not GPT-2, was the inflection point that convinced Schulman to shift his and his team's focus to large language models. He expects the share of compute spent on post-training to keep growing: model outputs can be higher quality than most web text, and most of GPT-4's Elo gain since its original release came from post-training.
Watch the clip on YouTubeStarts at 49:00