Pre-training is a lecture flying by; real learning needs reasoning
Leopold AschenbrennerLeopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of historyat 18:00
From the conversation
I don’t know, between six months and three years. But it's possible. It’s also very related to the issue of the data wall. Here’s one intuition on learning by yourself. Pre-training is kind of like the teacher lecturing to you and the words are flying by. You’re just getting a little bit from it. That's not what you do when you learn by yourself. When you learn by yourself, say you're reading a dense math textbook, you're not just skimming through it once. Some wordcels just skim through and reread and reread the math textbook and they memorize. What you do is you read a page, think about it, have some internal monologue going on, and have a conversation with a study buddy. You try a practice problem and fail a bunch of times. At some point it clicks, and you're like, "this made sense." Then you read a few more pages. We've kind of bootstrapped our way to just starting to be able to do that now with models. The question is, can you use all this sort of self-play, synthetic data, RL to make that thing work. Right now, there's in-context learning, which is super sample efficient. In the Gemini paper, it just learns a language in-context. Pre-training, on the other hand, is not at all sample efficient. What humans do is a kind of in-context learning. You read a book, think about it, until eventually it clicks. Then you somehow distill that back into the weights.…
Summary
Pre-training on internet text is an inefficient form of learning, analogous to passively listening to a lecture where information "flies by" without deep absorption. A more effective approach — analogous to how humans genuinely learn — involves pausing, reasoning through material, and engaging actively rather than simply ingesting large volumes of data repeatedly. This distinction points toward alternative training paradigms that could help address the data wall constraining current scaling approaches.
Watch the clip on YouTubeStarts at 18:00