GPT-3's few-shot results proved scaling was the main force

GwernGwern — Anonymous writer who predicted AI trajectory on $12K/year salaryat 16:00

From the conversation

And I'm like, “Holy shit, we are living in the scaling world. Legg and Moravec and Kurzweil were right!” And then I turned to Twitter and everyone else was like, “Oh, you know, this shows that scaling works so badly. Why, it's not even state-of-the-art!” That made me so angry I had to write all this up. Someone was wrong on the Internet. I remember in 2020, people were writing bestselling books about AI. It was definitely a thing people were talking about, but people were not noticing the most salient things in retrospect: LLMs, GPT-3, scaling laws. All these people who are talking about AI but missing this crucial crux, what were they getting wrong? I think for the most part they were suffering from two issues. First, they had not been paying attention to all of the scaling results before that which were relevant. They had not really appreciated the fact that, for example, AlphaZero was discovered in part by DeepMind doing Bayesian optimization on the hyperparameters and noticing that you could just get rid of more and more of the tree search as you went and you got better models. That was a critical insight, which could only have been gained by having so much compute power that you could afford to train many, many versions and see the difference that that made.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

Seeing the few-shot learning results from GPT-3 convinced Gwern that scaling was the dominant force in AI progress, validating earlier predictions from Legg, Moravec, and Kurzweil. The broader public and commentators at the time dismissed or misread these results, failing to recognize LLMs and scaling laws as the most important developments happening in AI. This disconnect between what the evidence showed and how it was received motivated Gwern to write extensively correcting the misinterpretation.

Watch the clip on YouTubeStarts at 16:00

Concept

Scaling

Related clips