YouSaid · the spoken record

Sergey Levine

lines on the record
110
first
2025-09-12
most recent
2025-09-12
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Makes a lot more sense to me for an embodied system to have parallel processes. Now, mathematically, you can actually make close equivalences between parallel and sequential stuff. Like transformers aren't actually fundamentally sequential, like you kind of make them sequential by putting in position embeddings. Transformers are fundamentally very paralyzable things. That's what makes them so great. So I don't think that actually

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Compared to the brain. Yeah, that's a really good question. So I definitely don't know the answer to this. I am not by any means well versed in neuroscience, but if I had to guess and also provide an answer that leans more on things I know, it's something like this, that the brain is extremely parallel. It kind of has to be just out of just because of the biophysics. But it's even more parallel than your GPU. If you think about how a modern multimodal language model processes the input, if you give it some images and some text, like first it reads in the images, then it reads in the text and then proceeds one token at a time to generate the output.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Yeah, how we represent both contexts, both what happened in the past, and also plans or reasoning, as you can call it in LLM world, which is what we would like to happen in the future, or intermediate processing stages in solving a task. I think... Doing that in a variety of modalities, including potentially learned modalities that are suitable for the job, is something that has, I think, enormous potential to overcome some of these challenges.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Representing your context in the right form that captures what you really need to achieve your goal and otherwise kind of discards all the unnecessary stuff. I think that's a really important thing. And I think what we're seeing the beginnings of that with multimodal models, but I think that multimodality has so much more to it than just like image plus text. And I think that that's a place where there's a lot of room for really exciting innovation.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  5. As a person, there are certainly some things where you keep track of them very symbolically, like almost in language. I have my checklist. I'm going shopping and at least for me, I can literally visualize in my mind my checklist, like, you know, pick up the yogurt, pick up the milk, pick up whatever. And I'm not picturing the milkshelf with the milk sitting there. I'm just thinking like milk, right? But then there's other things that are much more spatial, almost visual. When I was trying to get to your studio, I was thinking what the street looks like, here's what that street looks like, here's what I expect the doorway to look like.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Well, that's a very big question. Yeah, let's try to unpack this a little bit. I think there's a lot going on in there. One thing that I would say is a really interesting technical problem, and I think it's something where we'll see perhaps a lot of really interesting innovation over the next few years, is the question of representation for context. So if you imagine some of the examples you gave, if you have a home robot that's doing something and needs to keep track,

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  7. In the moment, right? Like it's like you've practiced that so much, you've baked it into your neural network and your brain that You don't have to think carefully about keeping all that context. So it really is just more of its paradox manifesting itself. But that doesn't mean that we don't need the memory. It just means that if we want to match the level of dexterity and physical proficiency that people have, there's other things we should get right first, and then

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  8. And I think this memory stuff is actually more of a paradox in disguise, where we think that the cognitively demanding tasks that we do, that we find hard that kind of cause us to think like, oh, man, I'm sweating, I'm working so hard, those are the ones that require us to keep lots of stuff in memory, lots of stuff in our minds. Like if you're solving some big math problem, if you're having a complicated technical conversation on a podcast, those are things we have to keep all those pieces, all those puzzle pieces in your head. If you're doing a well-rehearsed task, if you are an Olympic swimmer and you're swimming with perfect form and you're like right there in the zone, people even say like it's in the moment.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  9. But the reason why it's not the most important for the kind of skills that you saw when you visited us, at some level, I think it comes back to more of its paradox. So more of X paradox is basically that it's like, you know, if you know one thing about robotics, it's like that's the thing, more of Exparadox says that basically in AI, the easy things are hard and the hard things are easy, meaning like the things that we take for granted, like picking up objects, seeing perceiving the world, all that stuff, those are all the hard problems in AI and the things that we find challenging, like playing chess and doing calculus, actually are often the easier problems.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  10. I mean, it's not that there's something good about having less memory, to be clear. I think That adding memory, adding longer context, all that stuff, adding higher resolution images. I think those things will make the model better.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  11. But it's just you have this kind of compositionality that emerges when you do learning at scale. And that's really where all these remarkable capabilities come from. And now you put that together with language, you put that together with all sorts of chain of thought reasoning, and there's a lot of potential for the model to compose things in new ways.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  12. And we're like, we didn't know we'd do that. Holy crap. And then we tried to play around with it. And it's like, yep, it does that every time. Drop in, you know, it's doing its work, drop something else on the table, just pick it up, put it back. Okay, that's cool. Shopping bag, it starts putting things in the shopping bag. The shopping bag tips over, it picks it back up and stands it upright. We didn't tell anybody to collect data for that. I'm sure somebody accidentally at some point or maybe intentionally picked up the shopping bag.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Because of this, in principle, if we have a sufficient diversity of behaviors, the model should figure out that those behaviors can be composed in new ways, as the situation calls for it. We've actually seen things even with our current models, which I should say, I think they're in the grand scheme of things like looking back five years from now, we'll probably think that these are tiny in scale, but we've already seen what I would call emergent capabilities. When we were playing around with some of our laundry folding policies, actually we discovered this by accident, the robot accidentally picked up two t-shirts out of the bin instead of one, starts folding the first one, the other one gets in the white, picks up the other one, throws it back in the bin.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Like, holy crap, that is definitely not something that it has ever seen because IPAs only ever used for heading down pronunciations of individual words. So that's compositional generalization. It's putting together things you've seen. Like that, in new ways. And it's like, you know, arguably there's nothing profoundly new here because, yes, you've seen different words written that way, but you've figured out that now you can compose the words in this other language the same way that you've composed words. In English. So that's actually where the emergent capabilities come from

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Yeah. So there's a subtlety here. Emergent capabilities don't just come from the fact that internet data has a lot of stuff in it. They also come from the fact that generalization, once it reaches a certain level, becomes compositional. There was a acute example that one of my students really liked to use in some of his presentations, which is you know what international phonetic alphabet is, IPA. So if you look in a dictionary, They'll have the pronunciation of a word written in like kind of funny letters. That's basically international phonetic alphabet. So it's an alphabet that is pretty much exclusively used for writing down pronunciations of individual words and dictionaries. And you can ask an LLM to write you a recipe for making some meal in International Phonetic Alphabet. And it will do it. And that's like,

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Okay, that's like pretty dumb, right? Whereas if I told you first, like you're going to be playing tennis and then I let you study up, right? Now you really know what you're looking for. So I think that actually there's a very real challenge here. I don't want to understate the challenge, but I do think that there's also a lot of potential for foundation models that are embodied, that learn from interaction, from controlling robotic systems to actually be better at absorbing the other data sources because they know what they're trying to do. I don't think that that by itself is like a silver bullet. I don't think it solves everything. But I think that it does help a lot. And I think that we've already seen the beginning of that, where we can see that including web data in training for robots really does help with generalization. And I actually have the suspicion that in the long run,

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Well, let me put it this way let's say that I gave you lots of videotapes or lots of recordings of different sporting events and gave you a year to just watch sports and then after that year I told you okay now your job you're going to be playing tennis

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And its perception is in service to fulfilling that purpose. And that is like a really great focusing factor. We know that for people, this really matters. Like literally what you see is affected by what you're trying to do. There's been no shortage of psychology experiments showing that people have almost a shocking degree of tunnel vision where they will literally not see things right in front of their eyes if it's not relevant to what they're trying to achieve. And that is tremendously powerful. There must be a reason why people do that because certainly if you're out in the jungle, seeing more is better than seeing less. So if you have that powerful focusing mechanism, it must be darn important for getting it to achieve your goal. And I think robots will have that focusing mechanism because they're trying to achieve a goal.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  19. If you want to really predict everything that's going on in that scene, there's just so much stuff that even if you're doing a really great job and capturing 100% of something, by the time you get to everything else, ages will have passed. Whereas with text, it's sort of been abstracted into those bits that we as humans care about. So the representation is already there, and they're not just good representations, they actually focus in on what really matters. Okay, so that's the bad news. Here's the good news. The good news is that we don't have to just get everything out of like pointing a camera outside this building because when you have a robot, that robot is actually trying to do a job. So it has... A purpose

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Well, there's clouds moving around. Let me understand everything about water molecules and ice particles in the air. And you can go super deep on that. If you want to fully understand all, you know, down to the subatomic level, everything that's going on. As a person, you could spend like decades just thinking about that, and you'll never even get to the pedestrians or the water, right?

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  21. But it's not like just generating videos and images has already resulted in systems that have this kind of deep understanding of the world where you can ask them to do stuff beyond just generating more images and videos. Whereas with language, clearly it hasn't. And I think that this point about Representations is really key to it. One way we can think about it is this imagine pointing a camera outside this building. There's the sky, there's the clouds are moving around, the water, cars driving around, people, if you want to predict everything that will happen in the future, you can do so in many different ways. You can say, okay, there's people around, so let me get really good understanding the psychology of how people behave in crowds and predict the pedestrians. But you could also say, like,

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Yeah, so I have maybe two things I can say there. I have some bad news and some good news. So, the bad news is what you're saying is really getting at The core of a long running challenge with Video and image generation models. In some ways, the idea of getting intelligent systems by predicting video is even older than the idea of getting intelligent systems by predicting text. But the text stuff. Turn into practically useful things earlier than the video stuff did. I mean, the video stuff is great. Like you can generate cool videos, and I think that the work there that's been done recently is amazing.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  23. The big benefit that recent innovations in AI give to robotics is really that the ability to leverage prior knowledge. I think the fact that the model is the same model, that's kind of always been the case in deep learning, but it's that ability to pull in that prior knowledge, that abstract knowledge that can come from many different sources. That's really powerful.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Yeah, so one theme here that I think is important to keep in mind is that The reason that those building blocks are so valuable is because the AI community has gotten a lot better at leveraging prior knowledge. And a lot of what we're getting from the pre-trained LLMs and VLMs is prior knowledge about the world. And it's kind of like, it's a little bit abstracted knowledge. You can identify objects, you can figure out roughly where things are in image, that sort of thing. But I think if I had to summarize in one sentence,

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Dream is going on. That's right, with the exception of the actions are actually not represented as discrete tokens. It actually uses a flow matching kind of diffusion because they're continuous and you need to be very precise with your actions for extras control.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  26. To be a different module because the actions are continuous, they're high frequency, so they have a different data format than text tokens. But structurally, it's still an end-to-end transformer. And roughly speaking, technically, it corresponds to a kind of mixture of experts architecture.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Yeah, so the current model that we have basically is a vision language model that has been adapted for motor control. So to give you a little bit of like a fanciful brain analogy, a VLM, a vision language model is basically an LLM that has had a little pseudo-visual cortex grafted to it, a vision encoder. So our models, they have a vision encoder, but they also have an action expert, an action decoder, essentially. So it has like a little visual cortex and notionally a little motor cortex. And the way that the model actually makes decisions is it reads in the sensory information from the robot. It does some internal processing, and that could involve actually outputting intermediate steps like you might tell it, clean up the kitchen, and it might think to itself like, hey, to clean up the kitchen to pick up the dish, and I need to pick up the sponge and I need to put this and this. And then eventually it works its way through that chain of thought generation down to the action expert, which actually produces continuous actions.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Doing something like actually real. Yeah. I mean, ideally, I would like it to be RL because you can get away with the robot acting autonomously. Which is easier. But it's not out of the question that you can have. Nick's autonomy. As I mentioned before, robots can learn from all sorts of other signals. I described how we can have a robot that learns from a person talking to it. So there's a lot of middle ground in between fully teleoperated robots and fully autonomous robots

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Learning on the job or acquiring data in a way that the process of acquisition of that data itself is useful and valuable.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Well, that's the thing that we don't know that. It's certainly very reasonable to infer that robotics is a tough problem and probably it requires as much experience as the language stuff. But because we don't know the answer to that, to me a much more useful way to think about it is not How much data do we need to get before we're fully done? But how much data do we need to get before we can get started, meaning before we can get A data flywheel that represents a self sustaining and ever-growing data collection.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  31. That's right. It's very hard to do because robotic experience consists of time steps that are very correlated with each other. So, like the raw byte representation is enormous, but probably the information density is comparatively low. Maybe a better comparison as to the data sets that are used for multimodal training. And there it's, I believe last time we did that count, it was like between one and two orders of magnitude.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  32. So, answering those questions more thoroughly will give us a greater clarity on the axes on those dependent variables, on the things that we need to scale. And we don't fully know right now what that will look like. I think we'll figure it out. Soon, and something we'll work on actively. But we want to really get that right so that when we do scale it up, it'll directly translate into capabilities that are very relevant to practical use.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  33. That's a really good question. So the challenge here is in understanding which axes of scale contributes to which axis of capability. So, if we want to expand capability horizontally, meaning like the robot knows how to do 10 things now, and I'd like it to do 100 things later, that can be addressed by just directly horizontally scaling what we already have. But we want to get robots to a level of capability where they can do practical useful things in the real world, and that requires expanding along other axes too. It requires, for example, getting to very high robustness. It requires getting them to perform tasks very efficiently, quickly. It requires them to recognize edge cases and respond intelligently. And those things, I think, can also be addressed with scaling, but we have to identify the right axes for that, which means figuring out what kind of data to collect, what settings to collect it in, what kind of methods consume that data, how those methods work.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Was very much framed as a fundamental research effort. And that's good, like the fundamental research is really important, but it's not enough by itself. You need the fundamental research, and you also need the impetus to make it real and make it real means like actually put the robots out there, get data that is representative of the kind of tasks that they need to do in the real world, get that data at scale, build out the systems, get all that stuff right. And that requires a degree of focus, a singular focus on really nailing the robotic foundation model for its own sake, not just as a way to do more science, not just as a way to publish a paper, and not just as a way to kind of have a research lab.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  35. That's a really good question. So I'll start out with maybe a slight modification to your comment is I think they've made a lot of progress. And in some ways, a lot of the work that We're doing now at physical intelligence is built on the backs of lots of other great work that was done, for example, at Google. Like many of us were actually at Google before. We're involved in some of that work, some of it is work that we're drawing on that others did. So there's definitely been a lot of progress there. But to make robotic foundation models really work, it's not just a laboratory science kind of experiment. It's also, it also requires kind of industrial scale. Building effort. It's more like the Apollo program than it is like a science experiment. The excellent research that was done in the past industrial research labs.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  36. That's tremendously important, and that's something that we basically had no idea how to do about five years ago. But now we can actually use LLMs and VLMs, ask them questions, and they will make reasonable guesses. Like they will not give you expert behavior, but you can say, like, hey, there's a sign that says slippery floor. What's going to happen when I walk over that? It's kind of pretty obvious, right? And no autonomous car in 2009 would have been able to answer that question. So common sense plus the ability to make mistakes and correct those mistakes. That's sounding like an awful lot like what a person does when they're trying to learn something. All of that doesn't make robotic manipulation easy necessarily, but it allows us to get started with a smaller scope and then grow from there.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  37. But you would probably be okay with a child trying to do the dishes without somebody constantly sitting next to them with a break, so to speak. For a lot of tasks that we want to do with robotic manipulation, there's potential to make mistakes and correct those mistakes. And when you make a mistake and correct it, well, first you've achieved the task because you've corrected it, but you've also gained knowledge that allows you to avoid that mistake in the future. With driving because of the dynamics of how it's set up, it's very hard to make a mistake, correct it, and then learn from it because the mistakes themselves have significant ramifications. Now, not all manipulation tasks are like that. There are truly some very safety-critical stuff. And this is where the next thing comes in, which is common sense. Common sense, meaning the ability to make inferences about what might happen that are reasonable guesses, but that do not require you to experience that mistake and learn from it in advance.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  38. But there's also other things about robotics that are a bit different than driving. In some ways, robotic manipulation is a much, much harder problem, but in other ways it's a problem space where it's easier to get rolling to start that flywheel with a more limited scope. So to give you an example, If you're learning how to drive, you would probably be pretty crazy to learn how to drive on your own without somebody helping you. You would not trust your teenage child to learn to drive just on their own, just drop them in the car and say, like, go for it. And that's like a... A 16 year old who's had a significant amount of time to learn about the world. He would never even dream of putting a five-year-old in a car and tell him to get started. But if you want somebody to clean the dishes, dishes can break too.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  39. That's a really good question. So, one of the big things that is different now than it was in 2009 actually has to do with The technology for machine learning systems that understand the world around them. Principally, for autonomous driving, this is perception. For robots, it can mean a few other things as well. And perception certainly was not in a good place in 2009. The trouble with perception is that it's one of those things where you can nail a really good demo with a somewhat engineered system, but hit a brick wall when you try to generalize it. Now at this point in 2025, we have much better technology for generalizable and robust perception systems. And more generally, generalizable and robust systems for understanding the world around us. Like when you say that the system is scalable and machine learning scalable really means generalizable. So that gives us a much better starting point today. So that's not an argument about robotics being easier than autonomous driving. It's just an argument for 2025 being a better year than 2009.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Now, imagine what this implies for the human plus robot dynamic. Like now, basically learning for these systems is not just learning from raw actions. It's also learning from words, eventually be learning from observing what people do, from the kind of natural feedback that you receive when you're doing a job together with somebody else. And this is also the kind of stuff where that prior knowledge that comes from these big models is tremendously valuable because that lets you understand that that attraction dynamic So, I think that there is a lot of potential for these kind of human plus robot deployments to make the model better.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Also, because the human can help, the human can give hints. You know, let me tell you the story. When we were working on the PIO 5 project, this was the Paper that we released last April. We initially control our robots with teleoperation in a variety of different settings. And then at some point we actually realized that we can actually make significant headway once the model was good enough by supervising it not just with low-level actions, but actually literally instructing it through language. Now, you need a certain level of competence before you can do that, but once you have that level of competence, just standing there and telling the robot, okay, now pick up the cup, put the cup in the sink. Put the dish in the sink just with words. Already actually gives the robot information that it can use to get better.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  42. And that just makes total sense. It also makes it much easier to get all the technology bootstrapped because when it's robot plus human, now there's a lot more potential for the robot to actually learn on the job, acquire new skills. It's just like, you know,

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  43. That's a very hard question to answer I'm probably not prepared to tell you what percentage of all labor work can be done by robots because I don't think right now off the cuff I have a sufficient understanding of What's involved in That big of a cross section of all physical labor. I think what I can tell you is this, that I think it's much easier to get effective systems rolled out gradually in a human-in-the-loop setup. And again, I think this is exactly what we've seen with coding systems, and I think we'll see the same thing with automation, where basically robot plus human is much better than just human or just robot.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  44. It's a very subtle question. I think what it probably will come down to is this question of scope. The reason that LLMs are doing all software engineering is because they're good within a certain scope, but there's limits to that. Those limits are increasing, to be clear every year. And I think that there's no reason that we wouldn't see the same kind of thing with robots, that the scope will have to start out small because there will be certain things that these systems can do very well and certain other things where more human oversight is really important. And the scope will grow. And what that will translate into is increased productivity. And some of that productivity will come from the robots themselves being valuable and some of it will come from the people using the robots are now more productive in their role.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  45. So, I think there's a nuance here, and the nuance is it becomes more obvious if we consider the analogy to the coding assistance, right? It's not like the nature of coding assistance today is that there's a switch that flips and suddenly instead of writing software, suddenly like all software engineers get fired and everyone's using LMs for everything. And that actually makes a lot of sense that the biggest gain in productivity comes from experts, which is software engineers, whose productivity is now augmented by these really powerful tools.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  46. I think it's intellectually a very deep and profound problem, and figuring that out is going to be very exciting. But I think we kind of know roughly the puzzle pieces, and it's something that we need to work on. And I think if we work on it, and we're a bit lucky and everything kind of goes as planned, I think single digit is reasonable.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  47. I mean, my sense there, too, is that this is probably a single digit thing rather than a double digit thing, but the reason it's so hard to really pin down is because as with all research, it does depend on figuring out a few question marks. And I think my answer in terms of the nature of those question marks is I don't think these are things that require profoundly deeply different ideas, but it does require the right synthesis of the kinds of things that we already know. And sometimes synthesis, to be clear, is just as difficult. It's coming up with profoundly new stuff, right?

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  48. So I think it'll be the same thing. That will see an increase in the scope that we're willing to give to the robots as they get better and better, where initially the scope might be like there is a particular thing you do. You're making the coffee or something. Whereas as they get more capable as their ability to have common sense and a broader repertoire of tasks increases, then we'll give them greater scope. Now you're running the whole coffee shop.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Well, I think it's actually not that different than what we've seen with LMS in some ways, that it's a matter of scope. Like if you think about coding assistance, right? Initially, the best tools for coding they could do a little bit of completion. You give them a function signature and they'll try their best to type out the whole function and they'll maybe get half of it right. And as that stuff progresses, then you're willing to give these things a lot more agency so that the very best coding systems now, like if you're doing something relatively formulaic, maybe it can put together most of a PR for you for something fairly accessible.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source

  50. And I think that when you're doing physical things in the real world, that kind of stuff just happens more often than it does if you're an AI assistant answering a question. Like if you answer a question, you just answered it wrong. Well, it's not like you can just like go back and tweak a few things. Like the person you told the answer to might not even know that it's wrong. Whereas if you're like folding the t-shirt and you messed up a little bit, like, yeah, it's pretty obvious. You can reflect on that, figure out what happened.

    2025-09-12 · Dwarkesh Podcast · Fully autonomous robots are much closer than you think – Sergey Levine · IDENTIFIED FROM THE TRANSCRIPT · source