GPT-3's few-shot results proved scaling was the main force
GwernGwern — Anonymous writer who predicted AI trajectory on $12K/year salaryat 16:00
From the conversation
And I'm like, “Holy shit, we are living in the scaling world. Legg and Moravec and Kurzweil were right!” And then I turned to Twitter and everyone else was like, “Oh, you know, this shows that scaling works so badly. Why, it's not even state-of-the-art!” That made me so angry I had to write all this up. Someone was wrong on the Internet. I remember in 2020, people were writing bestselling books about AI. It was definitely a thing people were talking about, but people were not noticing the most salient things in retrospect: LLMs, GPT-3, scaling laws. All these people who are talking about AI but missing this crucial crux, what were they getting wrong? I think for the most part they were suffering from two issues. First, they had not been paying attention to all of the scaling results before that which were relevant. They had not really appreciated the fact that, for example, AlphaZero was discovered in part by DeepMind doing Bayesian optimization on the hyperparameters and noticing that you could just get rid of more and more of the tree search as you went and you got better models. That was a critical insight, which could only have been gained by having so much compute power that you could afford to train many, many versions and see the difference that that made.…
Summary
Seeing the few-shot learning results from GPT-3 convinced Gwern that scaling was the dominant force in AI progress, validating earlier predictions from Legg, Moravec, and Kurzweil. The broader public and commentators at the time dismissed or misread these results, failing to recognize LLMs and scaling laws as the most important developments happening in AI. This disconnect between what the evidence showed and how it was received motivated Gwern to write extensively correcting the misinterpretation.
Watch the clip on YouTubeStarts at 16:00