Memorization isn't intelligence, and ARC-AGI is built to show it
Francois CholletFrancois Chollet — Why the biggest AI models can't solve simple puzzlesat 0:00
From the conversation
LLms are very good at memorizing static programs If you scale up the size of your database, you are not increasing the intelligence of the system one bit. I feel like you're using words like memorization that we would never use for human children. If they could solve any arbitrary algebra problem they wouldn’t say they memorized algebra, you’d say they learned algebra. So I’ve got a million dollar prize pool and there’s a 500,000 for the first team to get to the 85% benchmark. If ARC survives 3 months from here, we’ll up the prize. OpenAI basically set back progress to AGI by five to ten years. They caused this complete closing down of frontier research publishing and now LLMs have essentially sucked the oxagen out of the room, like everyone is doing LLMs. Today I have the pleasure to speak with François Chollet, who is an AI researcher at Google and creator of Keras. He’s launching a prize in collaboration with Mike Knoop, the co-founder of Zapier, whom we’ll also be talking to in a second. It’s a million dollar prize to solve the ARC benchmark that he created. First question, what is the ARC benchmark? Why do you even need this prize? Why won’t the biggest LLM we have in a year be able to just saturate it? ARC is intended as a kind of IQ test for machine intelligence.…
Summary
ARC-AGI serves as a benchmark specifically designed to resist memorization, and Chollet has backed this claim with a million- dollar prize pool to incentivize teams to reach 85% performance. Chollet distinguishes between memorizing static programs and genuine learning, arguing that scaling up a model's training database does not increase its intelligence. He contends that conflating the two — as large LLM developers tend to do — has set back meaningful progress toward AGI by years.
Watch the clip on YouTubeStarts at 0:00Concept
Related clips
- Measuring AGI takes a broad battery of tests against human baselinesShane Legg
- Automated AI research won't be starved of ground truthScott Alexander, Daniel Kokotajlo
- Chained tasks only fail when the base success rate is too lowSholto Douglas, Trenton Bricken
- AI advances fastest where tasks can be scored digitallyCarl Shulman