YouSaid · the spoken record
Leopold Aschenbrenner
- lines on the record
- 332
- first
- 2024-06-04
- most recent
- 2024-06-04
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Companies, right? Because, like, they're the ones that are going to be like the rapper companies are betting on stagnation, right? The rapper companies are betting, like, you have these intermediate models and take so much stuff to integrate them. And I'm kind of like, I'm really bearish because I'm like, we're just going to Sonic boom you, you know, and we're going to get the unhoblins. We're going to get the drop-in remote worker. And then, you know, your stuff is not going to matter. Okay, sure, sure.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“But, you know, I think a lot more of the world is going to start feeling it. And I think that's going to start being kind of intense.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“I think 2023 was the sort of moment for me where it went from kind of AGIs, the sort of theoretical abstract thing, and you'd make the models to like, I see it, I feel it. And I see the path. I see where it's going. I think I can see the cluster where it's trained on, like the rough combination of algorithms, the people, like how it's happening. And I think, you know, most of the world is not right here.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Things have been quiet since then, but you know, the next thing has been in the oven, and I sort of expect sort of every generation, these kind of like ge forces to intensify, right? It's like people see the models. There's people haven't counted them, so they're going to be surprised and it'll be kind of crazy. And then revenue is going to accelerate. Suppose you do hit the 10 billion end of this year. Suppose just continues on this sort of doubling trajectory of every six months of revenue doubling. It's like you're not actually that far from 100 billion. Maybe that's like 26. And so at some point, what happened to Nvidia is going to happen to big tech. That's going to explode. And I mean, I think a lot more people are going to feel it, right? I mean, I think the”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, OpenAI. Oh, okay. Yeah, yeah, yeah. And It kind of went, you know, I mean, you know, I was been thinking about this and, you know, like talking to a lot of people in the years before there was this kind of weird thing, you know, you almost didn't want to talk about AI or AGI. It was kind of a dirty word, right? And then 2023, you know, people saw ChatGPT for the first time, they saw GP4, and it just exploded, right? It triggered this kind of like, you know, a huge sort of capital expenditures from all these firms and the explosion in revenue from NVIDIA and so on.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so Know, I think 2023 was kind of a really interesting year to experience as somebody who was like, you know, really following the eye stuff, where, you know, before that, it was.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“And the scaling, right? So it's like you have this baseline, just enormous force of scaling, right? Where it's like GPT-2 to GP4, you know, GP2, it could kind of like, it was amazing, right? It could string together plausible senses, but it could barely do anything. It was kind of like preschooler. And then GPT-4 is, you know, it's writing code. It like, you know, can do hard math. So it's like smart high schooler. And so this big jump. And, you know, in sort of the essay series, I go through and kind of count the orders of magnitude of compute scale up, of algorithmic progress. And so sort of scaling alone by 27-28 is going to do another kind of preschool to high school jump on top of GPD4. And so that will already be just like out of per token level, just incredibly smart. That'll get you somewhere reliability. And then you add these on hobblings that make it look much less like a chatbot, more like this agent, like a drop-in remote worker. And that's when things really get going.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“But what humans do is they kind of like they do in context learning, you read a book, you think about it until eventually it collects, but then you somehow distill that back into the weights. And in some sense, that's sort of like what RL is trying to do. And like, when RL is super finicky, but when RRL works, RL is kind of magical because it's sort of the best possible data for the model. It's like when you try a practice problem and then you fail and at some point you kind of figure it out in a way that makes sense to you. That's sort of like the best possible data for you because like the way you would have solved the problem. And that's sort of that's what RL is rather than just you kind of read how somebody else solved the problem and doesn't usually click.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Translate trade, translating in context, like, right? Like right now, there's in context learning super sample efficient. In the Gemini paper, right? It just learns language in context. And then you pre-training, not at all sample efficient”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“You reread and reread the math textbook, and then they memorize it. If you just repeated the data, then they memorize. What you do is you kind of like you read a page, you kind of think about it, you have some internal monologue going on, you have a conversation with the study buddy, you try a practice problem, you fail a bunch of times, and at some point it clicks, and you're like, this made sense. Then you read a few more pages. And so we've kind of bootstrapped our way to being able to do that now with models or just starting to be able to do that. And then the question is being able to read it, think about it, try problems. And the question is, can you, you know, all this sort of self-place synthetic data URL is kind of like making that thing work?”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Between six months and three years, you know. But I think it's possible. And I think there's also very related to the sort of issue of the data wall. But I mean, I think one intuition on the learning by yourself is sort of pre-training is kind of the words are flying by, or it's like the teacher is lecturing to you. And the model is, you know, the words are flying by. They're just getting a little bit from it. But that's sort of not what you do when you learn from yourself, right? When you learn by yourself, so you're reading a dense math textbook. You're not just kind of like skimming through at once. You wouldn't learn that much from it. I mean, some word cells just skim through.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“College, you know, if you're smart, you can kind of teach yourself. And in sort of models, you're just starting to enter that regime, right? And so it's sort of like, it's a little bit, it's probably a little bit more scaling. And then you got to figure out what goes on top. And it won't be trivial, right? So a lot of. Of deep learning is sort of like, you know, it sort of seems very obvious in retrospect. And there's sort of some obvious cluster of ideas, right? There's sort of some kind of thing that seems a little dumb, but this kind of works. But there's a lot of details you have to get right. So I'm not saying this, we're going to get this next month or whatever. I think it's going to take a while to really figure out the details”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“And again, there's sort of this advantage of bootstrapping, right? So your Twitter bio is being pre-trained, right? But you're actually not being pre-trained anymore. You're not being pre-trained anymore. You are pre-trained in grade school and high school. At some point, you transition to being able to learn by yourself, right? You weren't able to do that in elementary school. I don't know, middle school, probably high school is maybe when sort of started, you need some guidance”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“In sort of the unhobbling we've done over sort of like GP2 to GP4 was you kind of took the sort of like raw mass and then you like RHF it into really good chatbot and that was a huge win right like going you know in the original I think in Struct GPT paper you RLHF versus non-RLHF model it's like 100x model size win on sort of human preference ratings you know Yeah, I think people used to say it was a hardware problem, but I think the hardware stuff is getting solved. But the thing we have right now is you don't have the huge advantage of being able to brotstrap yourself with pre-training. You don't have all this sort of unsupervised learning you can do. You have to start right away with the sort of RL self-play and so on. All right, so now the question is why Why might some of this unhubbling and RL and so on work?”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Of all, free chaining is magical, right? And it gave us this huge advantage for models of general intelligence because you just Predict the next token, but predicting the next token, I mean, this is sort of a common misconception. But what it does is it lets this model learn these incredibly rich representations, right? Like these sort of representation learning properties are the magic of deep learning. You have these models, and instead of learning just kind of like whatever, Cisco artifacts or whatever, it learns sort of these models of the world. That's also why they can kind of generalize, right? Because it learned the right representations. And so you pre-train these models and you have this sort of like raw bundle of capabilities that's really useful and sort of this almost unformed raw mass.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“But sometimes you hit a weird construction zone or a weird intersection. And then I sometimes' seat, my girlfriend, I'm kind of like, ah, be quiet for a moment. I need to figure out what's going on. And that's sort of like, you know, you go from autopilot to the system too is jumping in. And you're thinking about how to do it. And so the scaling is improving that system one autopilot. And I think it's sort of, it's the brute force way to get to kind of agents. If you just improve that system, but if you can get that system to working”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Learn, right? You need to kind of learn error correction tokens, the tokens where you're like, ah, I think I made a mistake. Let me think about that again. You need to learn the kind of planning tokens that's kind of like, I'm going to start by making a plan. Here's my plan of attack. And then I'm going to write a draft. And I'm going to like, now I'm going to critique my draft. I'm going to think about it. And so it's not things that models can do right now. But the question is, how hard is that? And in some sense, also, there's sort of two paths to agents, right? When Cholto was on your podcast, he talked about kind of scaling leading to more nines of reliability. And so that's one path. I think the other path is this sort of like unhobbling path where it needs to learn this kind of like system two process. And if it can learn this sort of system two process, it can just use kind of millions of tokens and think for them and be cohesive and be coherent. One analogy, so when you drive, here's an analogy. When you drive, right? Okay, you're driving. And most of the time you're kind of an autopilot, right? You're just kind of driving and you're doing well. And then”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“If I thought for like 100 tokens a minute, you know, it's like what you before does. Maybe it's like, you know, it's equivalent to me thinking for three minutes or whatever, right? Suppose GPT 4 could think for millions of tokens, right? That's sort of plus four rooms, plus four orders of magnitude on test time compute, just like on one problem. It can't do it right now. It kind of gets stuck, right? It writes some code, even if you can do a little bit of iterative debugging, but eventually just kind of like it kind of gets stuck in something. It can't correct its errors and so on. In a sense, there's this big overhang, right? And like other areas of ML, there's this great paper on AlphaGo, right? Where you can trade off train time and test on compute. And if you can use four ooms more test on compute, that's almost like a three and a half oom bigger model. Just because, again, like if 100 tokens a minute, a few million tokens, that's a few months of sort of working time. There's a lot more you can do in a few months of working time than right now. So the question is, how hard is it to unlock that? And I think the sort of short timelines AI world is if it's not that hard. And the reason that might not be that hard is that, you know, there's only really a few extra tokens you need to.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think a really key question for sort of AI progress in the next few years is sort of how hard is it to do sort of unlock the test time compute overhang? So right now, GPD4 answers a question. And it kind of can do a few hundred tokens of kind of chain of thought. And that's already a huge improvement, right? Sort of like, this is a big unhobbling before answer a math question is just shotgun. If you try to kind of like answer a math question by saying the first thing that came to mind, you wouldn't be very good. So GP4 thinks for a few hundred tokens. And if I thought for a few hundred, you know, if I think at like 100 tokens a minute and I thought.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, yeah. So basically, the intermediate models could do it, but it would take a lot of slap. I see. And so then, you know, the actually it's just the drop in remote worker kind of AGI that can automate cognitive tasks that actually just ends up kind of like, you know, basically it's your intermediate models would have made the software engineer more productive, but will the software engineer adopt it? And then the 27 model is, well, you just don't need the software engineer. You can literally interact with it like a software engineer and it'll do the work of a software engineer.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“But I think in some sense, the way a lot of these systems want to be integrated is you kind of get this sort of sonic boom where it's the sort of intermediate systems could have done it, but it would have taken schlap. And before you do the schlap to integrate them, you get much more powerful systems, much more powerful systems that are sort of unhobbled. And so they're this agent and there's this drop-in remote worker. And then you're kind of interacting with them like a coworker, right? You can take Zoom calls with them and you're slacking them and you're like, ah, can you do this project? And then they go off and they go away for a week and write a first draft and get feedback on them and run tests on their code and then they come back and you see and you tell them a little bit more things or that'll be much easier to integrate. And so it might be that actually you need a bit of overkill to make the sort of transition easy and really harvest the gains.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“And then I think by 2078, if you extrapolate the trends, and we'll talk about that more later, and I talk about it in the series, I think we hit basically as smart as the smartest experts. I think the hobbling trajectory kind of points to, it looks much more like an agent than a chatbot.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think probably the sort of 10 gigawatt ish range is sort of my best guess for when you get the sort of true AGI. I mean, yeah, I think it's sort of like one gigawatt data center. And again, I think actually compute is overrated, and we're going to talk about that, but we will talk about compute right now. So, you know, I think 25, 26, we're going to get models that are basically smarter than most college graduates. I think sort of the practice, a lot of the economic usefulness, I think, really depends on sort of unhobbling. Basically, it's, you know, the models are kind of, they're smart, but they're limited, right? There's this chat bot and things like being able to use a computer, things like being able to do kind of like a Gentic long horizon tasks”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“But it's, you know, for the average knowledge worker, it's like a few hours of productivity a month. And it's kind of like you have to be expecting pretty lame AI progress to not hit some few hours of productivity a month of, yeah.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“You know, really, one sort of big tack, it'll be gangbusters is $100 billion a year. And so the question is sort of how feasible is $100 billion a year from AI revenue. And it's a lot more than right now, but I think if you sort of believe in the trajectory of the AI systems as I do, and which we'll probably talk about, it's not that crazy, right? So there's, I think there's like 300 million-ish Microsoft Office subscribers, right? And so they have copilot now, and I know what they're selling it for, but suppose you sold some sort of AI add-on for $100 a month and you sold that to a third of Microsoft Office subscribers subscribe to that. That'd be $100 billion right there. $100 a month is, you know, that's a lot. It's a lot. It's a lot.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Eventually, you got to get to like 100 billion a year. I think this is where it gets really interesting for the big tech companies, right? Because like their revenues are on order, you know, hundreds of billions, right? So it's like 10 billion, fine, you know, and it'll pay off the 2024 size training cluster.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“It's not just sort of my crazy take. I mean, AMD, AMD, I think, forecasted a $400 billion AI accelerator market by 27. I think it's an AI accelerators are only part of the expenditures. It's sort of, I think sort of a trillion dollars of sort of like total AI investment by 2027 is sort of like, we're very much in track on it. I think the trillion dollar cluster is going to take a bit more sort of acceleration. But we saw how much sort of chat GPT unleashed, right? And so every generation, the models are going to be kind of crazy and people, it's going to shift the overton window. And then obviously the revenue comes in, right? So these are forward-looking investments. The question is, do they pay off? And so if we sort of estimated the GPT-4 cluster at around 500 million, by the way, that's sort of a common mistake people make is they say people say like $100 million, but that's just the rental price, right? They're like, ah, you rent the cluster for three months, but it's, you know, if you're building the biggest cluster, you got to like, you got to build the whole cluster. You got to pay for the whole cluster. You can't just rent it for three months. Once you're trying to get into these sort of hundreds of billions.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, I don't know, but if you try to map out how expensive with the 10 gigawatt cluster be, that's maybe a couple hundred billion. So it's sort of on that scale and they're planning it. They're working on it. So”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“You know, I don't know. I think to send you like six months ago, you know, 10 gigawatts was the taco town. I mean, I think, I feel like now, you know, people have moved on, 10 gigawatts is happening. I mean, I don't know, there's the information report on OpenAI and Microsoft planning a hundred billion dollar cluster. So you got to, you know, if you're a gigawatt.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“It's like the power of the Hoover Dam. That costs tens of billions of dollars. That's like a million H100 to Kovalens. 2028. That's a class jet's 10 gigawatts, right? That's more power than kind of like most US states. That's like 10 million H100's equivalents costs hundreds of billions of dollars. And then 2030 trillion dollar cluster, 100 gigawatts, over 20% of US electricity production, 100 million H100 equivalents. And that's just the training cluster, right? That's like the one largest training cluster. And then there's more inference GPUs as well, right? Once there's products, most of them are going to be inference GPUs. And so Know U.S. power production has barely grown for decades, and now we're really in for a ride.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Of straight lines on a graph, right? There's this kind of like long run trend basically almost a decade of sort of training compute of the sort of largest AI systems growing by about half an order of magnitude, 0.5 booms a year. And you just kind of play that forward, right? GPT-4 reported to have finished pre-training in 2022, the sort of cluster size there was rumored to be about 25,000 H100s, sorry, A100s on semi-analysis. That's roughly, you know, if you do the math on that, it's maybe like a $500 million cluster. It's very roughly 10 megawatts. And just play that forward half noon a year, right? So then 2024, that say, you know, that's a cluster that's 100 megawatts. That's like 100,000 H100 equivalents. Cost in the billions, you know, play it forward two more years, 2026. That's a cluster. That's a gigawatt. That's sort of a large nuclear reactor side.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so unlike basically most things that have come out of Silicon Valley recently, AI is kind of this industrial process. The next model doesn't just require some code. It's building a giant new cluster. Now it's building giant new power plants. Pretty soon it's going to be building giant new fabs. Since ChatGPT, this kind of extraordinary sort of techno capital acceleration has been set into motion. I mean, basically exactly a year ago today, NVIDIA had their first kind of blockbuster earnings call, right? Where like why not 25% after hours and everyone was like, oh my God, AI, it's a thing. You know, I mean, I think within a year, you know, Nvidia, NVIDIA data center revenue has gone from like, you know, a few billion a quarter to like, you know, $25 billion a quarter now. And continuing to go up like big tech CapEx is skyrocketing. And it's funny because it's both, there's this sort of this kind of crazy scramble going on. But in some sense, it's just the sort of continuity.”
2024-06-04 · Dwarkesh Podcast · Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history · IDENTIFIED FROM THE TRANSCRIPT · source