YouSaid · the spoken record
Pieter Abbeel
- lines on the record
- 55
- first
- 2018-12-16
- most recent
- 2018-12-16
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“The kind of it seems like the level of reasoning a dog has is pretty sophisticated, but then it's still not yet at the level of human reasoning. And so it seems like we don't even need to achieve human level reasoning to get very strong affection with humans. And so my thinking is why not, right? Why couldn't within AI, couldn't we achieve the kind of level of affection that humans feel among each other? friendly animals and so forth. The question, is it a good thing for us or not? That's another thing, right? Because, I mean. But I don't see why not”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Interesting question. Maybe I'll answer it with another question, right? But I'll come back to it. So another question can have is, okay, I mean, how close does some people's happiness get from interacting with just a really nice Dog. I mean, dogs, you come home. That's what dogs do. They greet you. They're excited. Makes you happy when you come home to your dog. You're just like, okay, this is exciting. They're always happy when I'm here. I mean, if they don't greet you, because maybe whatever your partner took them on a trip or something, you might not be nearly as happy when you get home, right? And so.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“It seems like that's the kind of space we converge on to. I mean, I'm not an expert in anthropology, but it seems like we're very kind of good within our own tribe, but need to be taught to be nice to other tribes. Well, if you look at”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, there was kind of two optimizations happening for humans, right? So for humans, there's kind of the very long-term optimization, which evolution has done for us. And we're kind of predisposed to like certain things. And that's in some sense what makes our learning easier because, I mean, we know things like pain. And hunger and thirst. And the fact that we know about those is not something that we were taught. That's kind of innate. When we're hungry, we're unhappy. When we're thirsty, we're unhappy. When we have pain, we're unhappy. And ultimately, evolution built that into us to think about those things. And so I think there is a notion that it seems somehow humans evolved in general to prefer to get along in some ways. But at the same time also to be very Territorial and kind of centric to their own tribe.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I feel like we don't have this kind of unit test or proper tests for robots. And I think there's something very interesting to be thought about there, especially as you update things. Your software improves. You have a better self-driving car suite. You update it. How do you know it's indeed more capable on everything than what you had before that you didn't have any... Things creep into it. So I think that's a very interesting direction. Except that's somehow for humans, we do. Because we say, okay, you have a driving test, you passed. You can go on the road now. And you must have accents every like million or 10 million miles, something pretty phenomenal compared to that short test that is being done.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Did that? Would you trust it that it can thrive? I'd be like, no, that's not enough for me to trust it. But somehow for humans, we figured out that somebody being able to do that is representative of them being able to do a lot of other things. And so I think somehow for humans we figured out representative tests of what it means if you can do this, what you can really do. Of course, testing humans don't want to be tested at all times. Self-driving cars or robots could be tested more often probably. You can have replicas that get tested and they're known to be identical because they use the same neural net and so forth. But still.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“You do a stop sign successfully and then, you know, you pull over again and you're pretty much done. And you're like, okay, if a self driving car”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“What do we do to test somebody for driving, right? To get a driver's license. What do they really do? I mean, you fill out some test and then you drive. And I mean, in suburban California, that driving test is just you drive around the block, pull over.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“So When a robot is doing things, you kind of have a few notions of safety to worry about. One is that the robot is physically strong and of course could do a lot of damage. Same for cars, which we can think of as robots do in some way. And this could be completely unintentional. So it could be not the kind of long term AI safety concerns that, okay, AI is smarter than us and now what do we do? But it could be just very practical. Okay, this robot, if it makes a mistake. Whether the results going to be, of course, simulation comes in a lot there to test in simulation It's a difficult question, and I'm always wondering let's say you look at, let's go back to driving, because a lot of people know driving well, of course.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Ensemble of simulators, not any single one of them is sufficiently representative of the real world such that it would work if you train in there, but if you train in all of them. Then there was something that's good in all of them. The real world will just be another one of them. That's not identical to any one of them, but just another one of them.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“So the power of simulation, right? Simulators get better and better, of course, becomes stronger and we can learn more simulation. But there's also another version, which is where you say assimilator doesn't even have to be that precise. As long as it's... Somewhat representative, and instead of trying to get one simulator that is sufficiently precise to learn in and transfer really well to the real world, I'm going to build many simulators.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Is still a very different thing. For autonomous driving, I think there is still the question imitation versus RL. So imitation gives you a lot more signal. I think where imitation is lacking and needs some extra machinery is it doesn't, in its normal format, doesn't think about goals or objectives. And of course you are versions of imitation learning. Inverse reinforcement learning type imitation learning, which also thinks about goals. I think then we're getting much closer. But I think it's very hard to think of a Fully reactive car generalizing well. If it really doesn't have a notion of objectives to generalize well to the kind of generality that you would want, you want more than just that reactivity that you get from just behavioral cloning slash supervised learning.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“So I think the distinction between third person and first person is not a very important distinction for autonomous driving. They're very similar because the distinction is really about who turns the steering wheel. Or maybe let me put it differently how to get from a point where you are now to a point, let's say, a couple meters in front of you. And that's a problem that's very well understood. And that's the only distinction between third and first person there. Whereas with the robot manipulation, interaction forces are very complex.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“So for autonomous driving, I would say it's Third person is slightly easier. And the reason I'm going to say it slightly easier to do with third person is because the car dynamics are very well understood.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“That's a skill, no matter where the bottle starts, maybe it always goes onto a target or something. That's fairly easy to teach you about with Teleop. Now, what's even more interesting if you can now teach a robot through third-person learning, where the robot watches you do something and doesn't experience it, but just watches it and says, okay, well, if you're showing me that, that means I should be doing this. And I'm not going to be using your hand because I don't get to control your hand, but I'm going to use my hand. I do that mapping. And so that's where I think One of the big breakthroughs has happened this year. This was led by Chelsea Finn here. It's almost like learning a machine translation for demonstrations where you have a human demonstration and the robot learns to translate it into what it means for the robot to do it. And that was a meta-learning formulation. Learn from one to get the other. And that, I think, opens up a lot of opportunities to learn a lot more quickly.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Because why not just show the robot? And now the question is, how do you show the robot? One way to show is to tally operate the robot. And then the robot really experiences things. And that's nice because that's really high signal-to-noise ratio data. And we've done a lot of that. And you teach your robot skills in just 10 minutes. You can teach a robot a new basic skill. Like, okay, pick up the bottle, place it somewhere else.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Leapfrogging, where somebody figures out a formalism to say, okay, any RL problem by playing this and this idea, you can turn it into a self-play problem where you get signal a lot more easily. Reality is many problems we don't know how to turn into self play. And so either we need to provide detailed reward that doesn't just reward for achieving a goal, but rewards for making progress. And that becomes time consuming. And once you're starting to do that, let's say you want a robot to do something, you need to give all this detailed reward. Well, why not just give it demonstration?”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“You're on both sides, so one of you succeeds. And the beauty is also one of you fails. And so you see the contrast, you see the one version of me that did better than the other version. And so every time you play yourself, you get signal. And so whenever you can turn something into self-play, you're in a beautiful situation where you can naturally learn much more quickly than in most other reinforced learning environments. So I think if somehow we can turn more reinforcement learning problems into self-play formulations, that would go really, really far. So far, self-play has been largely around games where there's natural opponents. But if we could do self-play for other things, let's say, I don't know, a robot learns to build a house. I mean, that's a pretty advanced thing to try to do for a robot, but maybe tries to build a hut or something. If that can be done through self-play, it would learn a lot more quickly if somebody can figure it out. And I think that would be something where it goes closer to kind of the mathematical.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“When we look at self-play. What's so beautiful about it is goes back to kind of the challenges in reinforcement learning. So the challenge in reinforcement learning is getting signal. And if you never succeed, you don't get any signal. In self-play,”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Making gradual progress one step at a time, a new experiment here, a new experiment there that gives us new insights and gradual building up, but not getting to something yet where we're just, okay, here's an equation that now explains how that would have been two years of experimentation to get there, but this tells us what the result is going to be. Unfortunately, not so much yet.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Maybe I would prefer we could make the progress with mathematics. And the reason maybe I would prefer that is because often if you have something you can mathematically formalize, you can leapfrog a lot of experimentation. And experimentation takes a long time to get through and a lot of trial and error, kind of reinforcement in your research process. But you need to do a lot of trial and error before you get to a success. So if you can leapfrog that, to my mind, that's what the math is about. And hopefully once you do a bunch of experiments, you start seeing a pattern. You can do some derivations that leapfrog some experiments. But I agree with you. I mean, in practice, a lot of the progress has been such that we have not been able to find the math that allows us to leapfrog ahead. And we are”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“And even after some kind of, if people get rewired in some way, they might be able to reuse parts of their brand for other functions. And so what that suggests is some kind of modularity. And I think it is a pretty natural thing to strive for to see can we find that modularity? Can we find this thing? Of course, it's not every part of the brain is not exactly the same. Not everything can be rewired arbitrarily. But if you think of things like the neocortex, which is a pretty big part of the brain, that seems fairly modular from what defining so far. Can you design something equally modular? And if you can just grow it, it becomes more capable probably. I think that would be the kind of interesting underlying principle to shoot for that is not unrealistic.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I think that's a really interesting pursuit. And in the following sense, in that there is a lot of evidence The brain is pretty modular, and so I wouldn't maybe think of it as the theory maybe, the underlying theory, but more kind of the principle. There have been findings where Who are blind will use the part of the brain usually used for vision, for other functions.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Can I make this even simpler? How simple can I get this? What's the simplest equation I can explain everything, right? Master equation for the entire dynamics of the universe, we haven't really pushed that direction as hard in deep learning, I would say. Not sure if it should be pushed, but it seems a kind of generalization you get from that that you don't get in our current methods so far.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“A little different from what we've achieved so far. And it's not clear. If just, you know, regularizing more and forcing it to come up with a simpler, simpler, simpler experience and say, look, this is not simple. But that's what physics researchers do, right? They say.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I think I might have heard this. I might have heard it somewhere else, and I think it might have been one of your interviews, maybe the one with Yosha Benjamin. I'm not 100% sure I like the example. I'm not sure who it was, but the example was essentially if you use current deep learning techniques, what we're doing to predict, let's say, the relative motion of our planets, it would do pretty well. But then now if a massive new mass enters our solar system, it would probably not predict what will happen. And that's a different kind of generalization. That's a generalization that relies on the ultimate simplest, simplest explanation that we have available today to explain the motion of planets. Whereas just pattern recognition could predict our current solar system motion pretty well. No problem. And so I think that's an example of kind of generalization that”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“And then they get reused for other tasks. And so I think there is something there where somehow if you train a big enough model on enough things, it seems to transfer some deep mind results that I thought were very impressive, the unreal results, where it was learning to navigate mazes in ways where it wasn't just doing reinforcement learning, but they had other objectives, was optimizing for. So I think there's a lot of interesting results already. Where it's hard to wrap my head around is to which extent or when do we call something generalization or the levels of generalization involved in these different tasks?”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“That was a huge breakthrough. And then recently, I feel like similar Of by scaling things up, it seems like this has been expanded upon. Like people training even bigger networks, they might transfer even better. If you looked at, for example, some of the OpenAI results on language models and some of the recent Google results on language models. They are learned for just prediction.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I would say when even with the initial kind of big breakthrough in 2012 with AlexNet, right, the initial thing is, okay, great. This does better on ImageNet, hence image recognition, but then, immediately thereafter, there was of course the notion that wow, what was learned on ImageNet and you now want to solve a new task, you can fine-tune AlexNet for new tasks. And that was often found to be the even bigger deal, that you learn something that was reusable, which was not often the case before. Usually machine learning, you learn something for one scenario and that was it.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I'm pretty torn on this in that I think there are some very impressive, there's just some very impressive results already, right? I mean.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Which at a time Rocky Duan led. And that's exactly the meta-learning approach where you say, okay, we don't know how to design hierarchy. We know what we want to get from it. Let's just end-to-end optimize for what we want to get from it and see if it might emerge. And we saw things emerge. The maze navigation had consistent motion down hallways Which is what you want. A hierarchical control should say, I want to go down this hallway. And then when there is an option to take a turnack and decide whether to take a turn or not and repeat, even had the notion of where have you been before or not to not revisit places you've been before? It still didn't scale yet to the real-world candidates I think you had in mind, but it was some sign of life that maybe you can metalarn these hierarchical concepts.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“We had a very hard time getting success with. Not saying it's a dead end necessarily, but we had a lot of trouble getting that to work. And then we start revisiting the notion of what are we really trying to achieve? What we're trying to achieve is not necessarily hierarchy per se, but you can think about what does hierarchy give us? We hope it would give us is better credit assignment, what does better credit assignment is giving us, it gives us faster learning. Right. And so faster learning is ultimately maybe what we're after. And so that's what we ended up with, the RL squared paper on learning to reinforcement learn”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Another direction we've been thinking about for a long time and didn't make any progress on was more information theoretic approaches. So the idea there was that what it means to take high level action is to choose a latent variable now that tells you a lot about what's going to be the case in the future. Because that's what it means to take a high level action. Say, okay, I decide I'm gonna navigate to the gas station because I need to get gas for my car. Well, that'll now take five minutes to get there. But the fact that I get there, I could already tell that from the high level action I took much earlier.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Really see with sensors, process that and understand what's in the world. And so it's a good time to try to bring these things together. I see a few ways of getting there. One way to get there would be to say deep learning can get bolted on somehow to some of these more traditional approaches. Now, bolted on would probably mean you need to do some kind of end-to-end training where you say my deep learning processing somehow leads to a representation that in turn uses some kind of traditional underlying dynamical systems that can be used for planning. And that's, for example, the direction Aviv Tamar and Thenard Kuritach here have been pushing with causal infogan and of course other people too. That's one way. Can we somehow force it into the form factor that Is amenable to reasoning?”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so maybe let me highlight what I think the limitations are of what already was done. Twenty-thirty years ago, in fact, you'll find reasoning systems that reason over relatively long horizons, but the problem is that they were not grounded in the real world. So people would have to hand design. Some kind of logical dynamical descriptions of the world that didn't tie into perception. And so that didn't tie into real objects and so forth. And so that was a big gap. Now with deep learning, we start having the ability to”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I think where things get really tricky in the real world compared to the things we've looked at so far with great success in reinforcement learning is Time scales, which takes us to an extreme. So when you think about real world, I mean, I don't know, maybe some student decided to do a PhD here, right? Okay, that's a decision, that's a very high level decision. But if you think about their lives, I mean, any person's life, it's a sequence of muscle fiber contractions and relaxations, and that's how you interact with the world. And that's a very high frequency control thing, but it's ultimately what you do and how you affect the world until I guess we have brain readings. You can maybe do it slightly differently, but typically that's how you affect the world. And the decision of doing a PhD is like so abstract relative to what you're actually doing in the world. And I think that's where credit assignment becomes just completely beyond what any current RL algorithm can do. And we need hierarchical reasoning at a level that is just not available at all.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“And so it's benefiting from this linear control aspect, it's benefiting from the tiling, but it's somehow tiling it one dimension at a time. Because if let's say you have a two-layer network, in that hidden layer, You make a transition from active to inactive or the other way around, that is essentially one axis, but not axis aligned, but one direction that you change. And so you have this kind of very gradual tiling of the space where you have a lot of sharing between the linear controllers that tile the space. And that was always my intuition as why to expect that this might work pretty well. It's essentially leveraging the fact that linear feedback control is so good, but of course not enough. And this is a gradual tiling of the space with linear feedback controls that share a lot of expertise across them.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Non stationing, but the stationer flight regime like hover, you can use linear feedback control to stabilize the helicopter very complex dynamical system, but the controller is relatively simple. And so I think that's a big part of it is that if you do feedback control, even though the system you control can be very, very complex, often relatively simple control architectures can already do a lot, but then also just linear is not good enough. And so one way you can think of this neural networks is that in some sense they tile the space, which people were already trying to do more by hand or with finite state machines. Say, this linear controller here, this linear controller here. Neural network learns to tile the spin state linear controller here and not a linear controller here, but it's more subtle than that.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Some intuition behind it. So I think There's a few ways to think about this. The way I tend to think about it mostly originally when we started working on deep reinforcement learning here at Berkeley, which was maybe 2011, 12, 13 around that time, John Schulman was a PhD student initially kind of driving it forward here. The way we thought about it at the time was if you think about rectified linear units or kind of rectifier type neural networks, what do you get? You get something that's piecewise linear feedback control. And if you look at the literature, linear feedback control is extremely successful, can solve many, many problems surprisingly well. I remember, for example, when we did helicopter flight, if you're in a stationary flat regime, not a”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“The neural network to make the actions that were kind of present when things are good more likely and make the actions that are present when things are not as good less likely.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“The counterpart of that is why is RL, so many experiences to learn from? Because really what's happening is when you have a sparse reward, you do something maybe for like, I don't know, you take 100 actions and then you get a reward. Or maybe you get like a score of three. And I'm like, okay, three, not sure what that means. You go again and now you get two. And now you know that that sequence of 100 actions that you did the second time around somehow was worse than the sequence of 100 actions you did the first time around. But that's tough to now know which one of those were better or worse. Some might have been good and bad in either one. And so that's why I needed so many experiences. But once you have enough experiences, effectively RLS teasing that apart is starting to say, okay, what is consistently there when you get a higher reward and what's consistently there when you get a lower reward? And then kind of the magic of in some sense the policy grant update is to say now let's update”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“Person having a mind in their own mind, they I wanted to do a backflip, but the robot didn't know what it was supposed to be doing. It just knew that sometimes the person said, this is better, this is worse. Then the robot figured out what the person was actually after was a backflip. And I'd imagine the same would be true for things like more interactive robots that the robot would figure out over time, oh, this kind of thing apparently is appreciated more than this other kind of thing.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“There was no reason it couldn't emerge through learning. And maybe one way to formulate as an objective, you wouldn't have to necessarily score it explicitly. So standard rewards are numbers. Numbers are hard to come by. This is a 1.5 or 1.7 on some scale. It's very hard to do for a person. But much easier is for a person to say, okay, what you did the last five minutes was much nicer than what you did the previous five minutes. And that now gives a comparison. And in fact, there have been some results on that. For example, Paul Cristiano and collaborators at OpenAI had the Hopper, Majoko Hopper, a one-legged robot. Backflips. From feedback, I like this better than that. That's kind of equally good. And after a bunch of interactions, it figured out what it was the person was asking for, namely a backflip. And so I think the same thing.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“The robot being fun to be around, and why wouldn't it then naturally become more and more attractive and more and more maybe like a person or like a pet? I don't know what it would exactly be, but more and more have those features and acquire them automatically.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“That's a question, I would say a lot of people ask. And I think part of why they ask it is they're thinking about Unique are we really still as people like after they see some results they see a computer play go they say computer do this that they're like okay but can it really have emotion can it really interact with us in that way and then Once you're around robots, you already start feeling it. And I think the kind of maybe methodologically the way that I think of it is. If you run something like reinforcement learning, it's about optimizing some objective. No reason that the Objective couldn't be tied into how much does a person like interacting with this system and why could not the reinforcement learning system optimize for”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“But then I was at an event organized by This was by Fidelity and they had scripted pepper to help moderate some sessions. And they had scripted pepper to have the personality of a child a little bit. And it was very hard to not think of it as its own person in some sense because it was just kind of jumping. It would just jump into conversation, make it very interactive. Moderate would be saying Pepper would just jump in. Hold on. How about me? Can I participate in this too? You're just like, okay, this is like a person. And that was 100% scripted. And even then it was hard not to have that sense of somehow there is something there.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“It's very hard with Brett here. We give him a name, Berkeley Robot for the Elimination of Tedious Task. It's very hard to not. Think of. Robot as a person, and it seems like everybody calls them a heave for whatever reason, but that also makes it more a person than if it was a it. It seems pretty natural to think of it that way. This past weekend really struck me. I've seen Pepper many times on.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“at the Mars event that Jeff Beasos organizes. They brought it out there and it was nicely falling around Jeff. When Jeff left the room, they had it follow him along, which is pretty impressive.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“So physically for me it's Boston Dynamics videos always just ring home and I'm just super impressed Recently, the robot running up the stairs doing the parkour type thing. I mean, yes, we don't know what's underneath. They don't really write a lot of detail. But even if it's hard-coded underneath, which it might or might not be, just the physical abilities of doing that parkour, that's a very impressive. Right there”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it's learnable. I think if you set up a ball machine, let's say on one side and then a robot with a tennis racket on the other side, I think it's learnable and maybe a little bit of pre-training and simulation. Yeah, I think that's feasible. I think the swing the racket is feasible. It'd be very interesting to see how much precision it can get. I mean, that's. That's where, I mean, some of the human players can hit it on the lines, which is very high precision.”
2018-12-16 · Lex Fridman Podcast · Pieter Abbeel: Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source