Measuring AGI takes a broad battery of tests against human baselines
Shane LeggShane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architecturesat 1:00
From the conversation
To be an AGI that's the bar you need to meet. So if we want to test whether we're meeting the threshold or we're getting close to the threshold, what we actually need is a lot of different kinds of measurements and tests that span the breadth of all the sorts of cognitive tasks that people can do and then to have a sense of what human performance is on these sorts of tasks. That then allows us to judge whether or not we're there. It's difficult because you'll never have a complete set of everything that people can do because it's such a large set. But I think that if you ever get to the point where you have a pretty good range of tests of all sorts of cognitive things that we can do, and you have an AI system which can meet human performance and all those things and then even with effort, you can't actually come up with new examples of cognitive tasks where the machine is below human performance then at that point, you have an AGI. It may be conceptually possible that there is something that the machine can't do that people can do but if you can't find it with some effort, then for all practical purposes, you have an AGI. Let's get more concrete. We measure the performance of these large language models on MMLU and other benchmarks. What is missing from the benchmarks we use currently?…
Summary
Measuring progress toward AGI requires a broad battery of tests spanning the full range of cognitive tasks humans can perform, benchmarked against human-level performance on each. No evaluation suite can ever be fully complete, given the vast scope of human cognitive abilities, but having wide coverage and calibrated human baselines allows for a reasonable judgment of whether a system is approaching the AGI threshold.
Watch the clip on YouTubeStarts at 1:00Concept
Related clips
- Automated AI research won't be starved of ground truthScott Alexander, Daniel Kokotajlo
- Chained tasks only fail when the base success rate is too lowSholto Douglas, Trenton Bricken
- Memorization isn't intelligence, and ARC-AGI is built to show itFrancois Chollet
- AI advances fastest where tasks can be scored digitallyCarl Shulman