YouSaid · the spoken record
Anca Dragan
- lines on the record
- 98
- first
- 2020-03-19
- most recent
- 2020-03-19
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“It's frustrating. I'm frustrated by inability to comprehend. It feels very frustrating. It's like there's some stuff that we should time, blah, blah, blah, that we should really be understanding. And I definitely don't understand it. But, you know, the amazing physicists of the world have a much better understanding than me, but it still is epsilon and the grand scheme of things. So it's very frustrating. It also feels like our brain don't have some fundamental capacity yet well yet or ever. I don't know.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I was in a planetarium once. And they show you the thing, and then they zoom out and zoom out, and this whole speck of dust kind of thing. I think I was conceptualizing. There were kind of, you know, what our humans were just on this little planet, whatever. We don't matter much in the grand scheme of things. And then my mind got really blown because they talked about this multiverse theory where they kind of zoomed out and were like, this is our universe. And then like, there's a bazillion other ones and it just pop in and out of existence. So like our whole thing that we can't even fathom how big it is was like a blimp that went in and out. And at that point, I was like, okay, I'm done. This is not. There is no meaning. And clearly what we should be doing is try to impact whatever local thing we can impact. Our communities leave a little bit behind there, our friends, our family, our local communities. And just.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Sort of told me that I should take the SATs and apply to go to college abroad and do better on my English and all of that. And when it came to, well, financially, I couldn't. My parents couldn't really afford to do all these things. She started tutoring me on physics for free. And on top of that, sitting down with me to kind of train me for SATs and all that jazz that she had experience with.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“When I was in high school, my friends, my classmates did some tutoring. We were gearing up for our baccalaureate exam and they did some tutoring on, well, someone Math, someone whatever, I was comfortable enough with some of those subjects, but physics was something that I hadn't focused on in a while. And so. They were all working with this one teacher and I started working with that teacher. Her name is Nicole Becano. And she was the one who kind of opened up this whole world for me because she”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“tried to accomplish this is the thing I Thinking about this question and have a pretty joyous moment because I don't know that I would. Change much I'm trying to make some contributions to how we understand human AI interaction. I don't think I would change that. Maybe I'll take, you know, I take more trips to the Caribbean or something, but I tried some of that already from time to time. So yeah, I mean, I try to do the things that bring me joy and thinking about these things bring me joy, the Merry Condo thing, you know, don't do stuff that doesn't spark joy. For the most part, I do things that spark joy. Maybe I'll do like less service in the department or something. I'm not dealing with admissions anymore. But no, I mean, I think I have amazing colleagues and amazing students and amazing family and friends and kind of spending time and some balance with all of them is what I do. And that's what I'm doing already. So I don't know.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“You're not allowed to die. Yeah. Now I'll say that I want to live forever, but I watch this show. It's very silly. It's called The Good Place. And they reflect a lot on this. And, you know, the moral of the story is that you have to make the afterlife be finite too, because otherwise people just kind of just like wally. It's like whatever. So I think the finalist helps. I'm not a religious person. I don't think that there's something after. And so I think it just ends and you stop existing. And I really like existing. It's such a great privilege to exist that, yeah, it's just, I think that's the scary part”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“This is like my biggest nightmare, by the way. I really like living. So I'm actually, I really don't like the idea of being told that I'm going to die.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“A lot of what we do in my lab is write math about human behavior, combine it with data and learning, put it all together, give it to robots to plan with and hope that instead of writing rules for the robots, writing heuristics, designing behavior, they can actually autonomously come up with the right thing to do around people. That's kind of our, you know, that's our signature move. We wrote some math and then instead of kind of handcrafting this and that and that, the robot figured stuff out. And isn't that cool? And I think that is the same enthusiasm that I got from the robot figured out how to reach that goal in that graph. Isn't that cool?”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“We can write math about human behavior, right? Yeah. So that's, and I think that stuck with me because, you know, a lot of what I do.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“When I was in, it's kind of a personal story. When I was in 12th grade, I got my hands on a PDF copy in Romania of Russell Norvig AI modern approach. I didn't know anything about AI at that point. I was, you know, I had watched the movie The Matrix was my exposure. And so I started going through this thing and you know you were asking in the beginning, what are industries, you know, it's math and it's algorithms, what's interesting, it was so captivating this notion that you could just have a goal and figure out your way through a kind of a messy complicated situation. So what sequence of decisions you should make autonomously to achieve that goal? That was so cool. Bias, but that's a cool book.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Right? Because people don't have the bandwidth to do everything. So just because you know my house is messy doesn't mean that I want it to be messy, right? But that just because I, you know, I didn't put the effort into that. I put the effort into something else. So the robot should figure out, well, that something else was more important, but it doesn't mean that, you know, the house being messy is not. So it's a little subtle, but yeah, we really think of it, the state itself is kind of like a choice that People implicitly made about how they won their world.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“One really, really crazy one is the environment itself. Like our world. It's not, you know, you observe our world and the state of it, and it's not that you're seeing behavior and you're saying, oh, people are making decisions that are rational, blah, blah, blah. But our world is something that we've been acting in according to our preferences. So I have this example where like the robot walks into my home and my shoes are laid down on the floor, kind of in a line, right? It took effort to do that. So even though the robot doesn't see me doing this, you know, actually aligning the shoes, it should still be able to figure out that I want the shoes aligned because there's no way for them to have magically instantiated themselves in that way. Someone must have actually”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“But I'm gonna interpret it not as this is the universal reward function that I shall always optimize, always and forever, but as this is good evidence about what the person wants. And I should interpret that evidence in the context of these situations that it was specified for because ultimately that's what the designers thought about. That's what they had in mind. And really them specifying reward function that works for me in all these situations is really kind of telling me that whatever behavior that incentivizes must be good behavior with respect to the thing that I should actually be optimizing for. And so now the robot kind of has uncertainty about what it is that it should its reward function is. And then there's all these additional signals we've been finding that it can kind of continually learn from and adapt its understanding of what people want. Every time the person corrected, maybe they demonstrate, maybe they stop, hopefully not, right?”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Obey what, right? Because we just talked about how you try to say what you want, but you don't always get it right and you want these machines to do what you want not necessarily exactly what you literally. So you don't want them to take you literally. You want to take what you're saying and interpret it in context. And that's what we do with the specified rewards. We don't take them literally anymore from the designer. Not we as a community, we as some members of my group. And some of our collaborators like Peter Beale and Stuart Russell We sort of say, okay, the designer specified this thing.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it should constantly be adjusting kind of thing. The issue with the laws is, I don't even, you know, there are words and I have to write math. And have to translate them into math. What does it mean to”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“It moves in a way where it's easy if you correct it in some other way or not, and then kind of actually plans its motion so that it can disambiguate and collect information about what you want. Anyway, so that's one way that's kind of sort of leaked information, maybe even more subtle leaked information is if I just press the e-stop, right? I'm doing it out of panic because the robot is about to do something bad. There's again information there, right? Okay, the robot should definitely stop, but it should also figure out that whatever was about to do was not good. And in fact, it was so not good that stopping and remaining stopped for a while was better, a better trajectory for it than whatever it is that it was about to do. And that, again, is information about what are my preferences, what do I want.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. Yeah. So That's true. Again, it comes back to Can you structure a little bit your assumptions about how human behavior relates to what they want? And you can't, one thing that we've done is literally just treated this external torque that they applied as when you take that and you add it with what the torque the robot was already applying, that overall action is probably relatively optimal in respect to whatever it is that the person wants. And then that gives you information about what it is that they want. So you can learn that people want you to stay further away from them. Now you're right that there might be many things that explain just that one signal and that you might need much more data than that for the person to be able to shape your reward function over time. You can also do this info gathering stuff that we were talking about. Now that we've done that in that context, just to clarify, but it's definitely something we thought about where you can have the robot start acting in a way, like if there are a bunch of different explanations, right?”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Part of realization has been that that is signal that communicates about the reward because if my robot was moving in an optimal way and I intervened, that means that I disagree with his notion of optimality, whatever it thinks is optimal is not actually optimal. And sort of optimization problems aside, that means that the costs of the reward function is incorrect, or at least not what I want it to be.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“So I push away because it's a reaction to what the robot is currently doing. And this is what we call physical human robot interaction. And now there's a lot of interesting work on how the heck do you respond to physical human robot interaction? What should the robot do if such an event occurs? And there's sort of different schools of thought. Well, you know, you can sort of treat it to control theoretically and say this is a disturbance that you must reject. You can sort of treat it more kind of heuristically and say, I'm going to go into some like gravity compensation mode so that misery maneuverable around or I'm going to go in the direction that the person pushed me. And to us.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Would do, yeah, absolutely. So I think maybe some surprising bits, right? So we were talking before about I'm a robot arm in needs to move around people, carry stuff, put stuff away, all of that. And now Imagine that the robot has some initial objective that the programmer gave it so they can do all these things functionally. It's capable of doing that. And now I noticed that it's doing something and maybe it's coming too close to me. And maybe I'm the designer. Maybe I'm the end user and this robot is now in my home. And I push it away.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“And maybe leaving a lot of other information on the table, like what are other things we could do to actually communicate to the robot about what we want them to do besides attempting to specify a reward?”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Right, collaborating on the reward design. And so, what does it mean, right? What is it when we think about the problem not as someone specifies all of your job is to optimize and we start thinking about urine, this interaction and this collaboration? And the first thing that comes up is when the person specifies a reward, it's not, you know, gospel, it's not like the letter of the law. It's not the definition of the reward function you should be optimizing because they're doing their best, but they're not some magic perfect oracle. And the sooner we start understanding that, I think the sooner we'll get to more robust robots that function better in different situations. And then you have kind of say, okay, well, it's almost like robots are overlearning, over putting too much weight on the reward specified by definition.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, it stopped doing his job. So there's such a thing as offer optimizing for things and failing to. Think ahead of time of all the possible things that might be important. And so that's interesting because historically I work a lot on reward learning from the perspective of customizing to the end user, but it really seems like it's not just the interaction with the end user that's a problem of the human and the robot collaborating so that the robot can do what the human wants, right? This kind of back and forth, the robot probing, the person being informative, all of that stuff might be actually just as applicable to this kind of maybe new form of human robot interaction, which is the interaction between the robot and the expert programmer roboticist designer in charge of actually specifying what the heck should do, specifying the task for the robot.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“That's the thing Right, so there is this thing called Good Heart's Law, which is you set a metric for an organization, and the moment it becomes a target that people actually optimize for.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Suboptimal behavior that is actually optimal. I mean, this, I guess, the idea of unintended consequences, you know, it's optimal with respect to what you specified, but it's not what you want. And there's a difference between those”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Don't know, kind of, I guess I closed my eyes to it for a while because I've been tuning cost functions for 10 years now. But it really strikes me that, yeah, we've moved the tuning and the designing of features or whatever from the... Behavior side into the reward side. And yes, I agree that there's way less of it, but it still seems really hard to anticipate any possible situation and make sure you specify a reward function that went optimized will work well in every possible situation.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“What I can do is I can sample, I can think of some situations that I think are representative of what the robot will face. And I can tune and add and tune some reward function until the optimal behavior is what I want on those situations, which first of all is super frustrating because through the miracle of AI, we don't have to specify rules for behavior anymore, right? We were saying before, the robot comes up with the right thing to do. You plug in this situation. It optimizes Bragnet situation. It optimizes. But you have to spend still a lot of time on actually defining what it is that that criterion should be. Make sure you didn't forget about 50 bazillion things that are important and how they all should be combining together to tell the robot what's good and what's bad and how good and how bad. And so I think this is a lesson that”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“If you write it out, or maybe it's a set depending on the situation or whatever it is, if you write it out and then you deploy the agent, you'd want to make sure that whatever you specified incentivizes the behavior you want from the agent in any situation that the agent will be faced with, right? So I do motion planning on my robot arm. I specify some cost function, like this is how far away should try to stay so much amount to stay away from people and this is how much it matters to be able to be efficient and blah blah blah, right? I need to make sure that whatever I specify those constraints or trade-offs or whatever they are. That when the robot goes and solves that problem in every new situation, that behavior is the behavior that I want to see. And what I've been finding is that we have no idea how to do that. That basically”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Really good question because the answer is we don't. This was. You know, I used to think about how, well, it's actually really hard to specify rewards for interaction because, you know, it's really supposed to be what the people want. And then you really, you know, we talked about how you have to customize what you want to do to the end user. I kind of realized that even if you take the interactive component away, Still really hard to design reward functions. So, what do I mean by that? I mean, if we assume this survey paradigm in which there's an agent and its job is to optimize some objective, some reward, utility, loss, whatever, cost.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Agree. I think that is actually really possible. I guess mainly I'm pointing out that if you do it naively, you're and implicitly assuming something, that assumption might actually really be wrong. But I do think that if you explicitly think about What the agent should do such that the person still stays engaged, what you essentially empower the person do more than they could, that's really the goal, right? You still have a driver, so you want to empower them to be so much better than they would be by themselves. And that's Difference, very different mindset than I want them to basically not drive. And be ready to sort of take over.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“And the assumption that the person functions just as well there as they function in the states that they would normally encounter is a little questionable. Now another part is the kind of the human factor side of this, which is that I don't know about you, but I think I definitely feel like I'm experiencing things very differently. Versus when I'm a passive observer. Like, even if I try to stay engaged, right, it's very different than when I'm actually actively making decisions. And you see this in life in general. Like you see students who are actively trying to come up with the answer, learn the thing better than when they're passively told the answer. I think that's somewhat related. And I think people have studied this in human factors for airplanes. And I think it's actually fairly established that these two are not the same.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Be just as safe as when I'm driving, like that I will, you know, if I wouldn't get into some kind of accident, if I'm driving, I will be able to avoid that accident when I'm supervising too. And I think I'm concerned about this assumption from a few perspectives. So from a technical perspective, it's that when you let something kind of take control and do its thing, and it depends on what that thing is, obviously, and how much it's taking control and how what things are you trusting it to do. But if you let it do its thing and take control, it will go to what we might call off policy from the person's perspective states. So states to the person wouldn't actually find themselves in if they were the ones driving.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think what we have to be careful about there is, it seems like some of these systems, not all, are making this underlying assumption that we If so, I'm a driver and I'm now really not driving but supervising, and my job is to intervene, right? And so we have to be careful with this assumption that when if I'm supervising”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“So when you're sort of treating agents as having these objectives, these incentives, humans or artificial, you're kind of implicitly modeling that they'd like to stick around so that they can accomplish those goals. So I think in a sense maybe that's what draws me so much to the rationality framework, even though it's so broken. We've been able to It's been such a useful perspective and like we were talking about earlier, what's the alternative? I give up and go home. Or, you know, I just use complete black boxes, but then I don't know what to assume out of distribution. I come back to this. It's just, it's been a very fruitful way to think about the problem and a very more positive way, right? It's just people aren't just crazy. Maybe they make more sense than we think. But I think we also have to somehow be ready for it to be.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it might be simpler, and I'm no person who likes to complicate this. I think it might be simpler than that. Because it turns out, for instance, if you say model people in the very, I'll call it traditional way, I don't know if it's fair to look at as a traditional way, but calling people as, okay, they're rational somehow, the utilitarian perspective. That once you say that, you automatically capture that they have an incentive to keep on being. Stuart likes to say, you can't fetch the coffee if you're dead. Stuart Russell, by the way.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Who knows? Who knows how much you need, right? In terms of if your state is really like the positions of everything or whatnot and velocities, who knows how much you need. And then there's this, there's so many mappings. And so now you're talking about how do you regularize that space? What priors do you impose? Or what's the inductive bias? There's all very related things to think about it. Basically, what are assumptions that we should be making? Such that these models actually generalize outside of the data that we've seen. And now you're talking about, well, I don't know. What can you assume? Maybe you can assume that people actually have intentions and that's what drives their actions. Maybe that's the right thing to do when you haven't seen data very nearby that tells you otherwise. I don't know. It's a very open question.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“So, and while we're talking about here is how do we reason about what other people do in situations where we haven't seen them? And somehow we just I can anticipate what will happen in situations that are even novel in many ways. And I have a pretty good intuition for it. I always get it right, but I might be a little uncertain and so on. But I think it's this that if you just rely on data, there's just too many possibilities. There's too many policies out there that fit the data. And by the way, it's not just state, it's really kind of history of state because to really be able to anticipate what the person will do. It kind of depends on what they've been doing so far because that's the information you need to kind of at least implicitly, so to say, this is the kind of person that this is. This is probably what they're trying to do. So anyways, like you're trying to map history states to actions, there's many mappings. And history”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“All the possible policies, and then you take only the ones that are consistent with the human data that you've observed, that still leads a lot of things could happen outside of that distribution where you're confident and you know what's going on”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think they have to be a combination because if you get some human data and then you say, this is going to be my model of the person, what are four simulation and training or for just deployment time? And that's what I'm planning with as my model of how people work, regardless. If you take some data and you don't assume anything else and you just say, okay, this is some data that I've collected. Let me fit a policy to help people work based on that. What tends to happen is you collected some data and some distribution and then now your robot sort of computes a best response to that, right? It's sort of like, what should I do if this is how people work and easily goes off of distribution where that model that you've built of the human completely sucks because out of distribution you have no idea, right?”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Simulation, and then that means, well, guess what? You need to model things at least to model people, model the world enough that you, you know, whatever policy you get of that is like actually fine to roll out in the world and do some additional learning there.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Figure out on its own how to achieve its goal without hitting stuff and all that stuff here all the good stuff for emotion planning 101 I think of that as very much AI not this is some rule or some here there's nothing rule based around that right it's just you're you're searching through a space and figuring out are you optimizing through a space and figure out what seems to be the right thing to do um and i think it's hard to just do that because you need to learn models of the world and i think it's hard to just do the learning part where you don't you know you don't bother with any of that because then you're saying well i could do imitation but then when i go off distribution i'm really screwed or you can say i can do reinforcement learning which adds a lot of robustness but then you have to do either reinforcement learning in the real world which sounds a little challenging or that trial and error you know um or you have to do reinforcement learning”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think on the one hand that Learning is inevitable here, right? I think on the other hand that when people characterize the problem as it's a bunch of rules that some people wrote down versus it's an end-to-end DRL system or imitation learning, then maybe there's kind of something missing from maybe that's more So for instance I think a very, very useful tool in this sort of problem both in how to generate the car's behavior and robots in general and how to model human beings is actually planning search optimization right so robotics is the sequential decision making problem and when when a robot can”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“And I wouldn't say that it's just, you know, now it's just pure engineering, and it's probably the, I mean, by the way, I'm speaking kind of very generally here as hypothesizing, but I think that there are successes and yet no one is everywhere out there. So that seems to suggest that things can be expanded and can be scaled and we know how to do a lot of things, but there's still probably, you know, new algorithms or modified algorithms that you still need to put in there as you learn more and more about new challenges that you get faced with.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it A sense it really depends. I think that we were talking about how, well, look, it's really hard because anticipatory people do is hard. And on top of that, playing the game is hard. But I think we sort of have the fundamental some of the fundamental understanding for that. And then you already see that these systems are being deployed. You know, even driverless, like there's, I think now a few companies that don't have a driver in the car In some small areas”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“With Glare exactly what you were saying, stuff that people would see that you don't. I think I have more of my intuition comes from systems that can actually use lighter as well”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Really, just don't know enough to say, all vision alone, what, you know, what's like there's a lot of how many cameras do you have? Is it how you're using them? I don't know. There's all sorts of details. I imagine there's stuff that's really hard to actually see how do you deal with.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I'm assuming LIDAR here, by the way. I think it's kind of irresponsible to not use LIDAR. That's just my personal opin Depending on your use case, but I think if you have the opportunity to use LIDAR, in a lot of cases you might not.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, it might be. I mean, you know, and I picked Dantern San Francisco. Adapting to well, now it's knowing, now it's no longer snowing, now it's slippery in this way, now it's the dynamics part. I could imagine being still somewhat challenging.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“I think the perception problem, I mean, and by the way, a bunch of years ago, this would not have been true. And a lot of issues in the space were coming from the fact that, oh, we don't really, you know, we don't know what's where. But I think it's fairly safe to say that at this point, although you could always improve on things and all of that, you can drive through downtown San Francisco if there are no people around. There's no really perception issues standing in your way there. Perception is hard, but yeah, it's we've made a lot of progress on the perceptions, and I had to undermine the difficulty of the problem. I think everything about robotics is really difficult, of course. I think that, you know, the planning problem, the control problem, all very difficult. But I think what makes it really...”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah good question I have many opinions on this imagine downtown San Francisco Yah crazy busy everything. Okay, now take all the humans out. No pedestrians, no human driven vehicles, no cyclists, no people on little electric scooters zipping around, nothing. Think we're done. I think driving at that point is done. We're done. There's nothing really that still needs to be solved about that.”
2020-03-19 · Lex Fridman Podcast · #81 – Anca Dragan: Human-Robot Interaction and Reward Engineering · IDENTIFIED FROM THE TRANSCRIPT · source