AI advances fastest where tasks can be scored digitally
Carl ShulmanCarl Shulman (Pt 1) — Intelligence explosion, primate evolution, robot doublings, & alignmentat 1:26:00
From the conversation
You can do unit tests, you can prove theorems, you can do all sorts of operations entirely in the confines of a computer, which is one reason why programming has been benefiting more than a lot of other areas from LLMs recently whereas robotics is lagging. And considering they are getting better at things like the GRE, math, at programming contests, and some people have forecasts and predictions outstanding about doing well on the informatics olympiad and the Math Olympiad and in the last few years when people tried to forecast the MMLU benchmark which has a lot of sophisticated, graduate student level science kind of questions, AI knocked that down a lot faster than AI researchers and students who had registered forecasts on it. If you're getting top-notch scores on graduate exams, creative problem solving, it's not obvious that that area will be a relative weakness of AI. In fact computer science is in many ways especially suitable because of getting up to speed with new areas, being able to get rapid feedback from the interpreter at scale. But do you get rapid feedback if you're doing something that's more analogous to research? Let's say you have a new model and it’s like, if we put in 10 million dollars on a mini-training run on this this would be much better. Yeah for very large models those experiments are going to be quite expensive. You're going to look more at can you build up this capability by generalization?…
Summary
AI capabilities are advancing more rapidly in programming and formal reasoning than in areas like robotics because software tasks can be evaluated entirely within a digital environment, avoiding the costs and risks of physical interaction. Benchmarks such as the GRE, math competitions, and programming contests show measurable progress, and specific predictions exist about AI achieving strong performance on the Informatics Olympiad and Math Olympiad.
Watch the clip on YouTubeStarts at 1:26:00Concept
More from Carl Shulman
Related clips
- Measuring AGI takes a broad battery of tests against human baselinesShane Legg
- Automated AI research won't be starved of ground truthScott Alexander, Daniel Kokotajlo
- Chained tasks only fail when the base success rate is too lowSholto Douglas, Trenton Bricken
- Memorization isn't intelligence, and ARC-AGI is built to show itFrancois Chollet