YouSaid · the spoken record

Leslie Kaelbling

lines on the record
79
first
2019-03-12
most recent
2019-03-12
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Just how do we, what's the best way to make this work out? And I think it's clear, it's a combination of learning, to me, it's clear. It's a combination of learning and not learning. And what should that combination be and what's the stuff we build in? So to me, that's the most compelling question.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Well, so there's the story I've been telling you about. How to engineer intelligent robots. So that's what we want to do. We all kind of want to do, well, I mean, some set of us want to do this. And the question is, what's the most effective strategy? And we've tried, and there's a bunch of different things you could do at the extremes, right? One super extreme is we do introspection and we write a program. Okay, that has not worked out very well. Another extreme is we take a giant bunch of neural goo and we train it, train it up to do something. I don't think that's going to work either. So the question is, what's the middle ground? And again, this isn't a theological question or anything like that. It's just like

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Designing objectives. I mean, they're mixed together because, as we also know, as machine learning people, right, when you're designing, in fact, this is the lecture I gave in class today. When you design an objective function, you have to wear both hats. the hat that says what do I want and there's the hat that says ah but i know what my optimizer can do to some degree and i have to take that into account So it's always a trade off, and we have to kind of be mindful of that. The part about taking people's jobs, I understand that that's important. I don't understand. Sociology or economics are people very well. So I don't know how to think about that.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Pragmatically, in the shortish term, it seems to me that those are really interesting and critical questions. And the idea that we're going to go from being people who engineer algorithms to being people who engineer objective functions. I think that's definitely going to happen. And that's going to change our thinking and methodology and stuff.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Want to be sure as robots or software systems get more competent that their objectives are aligned with your objectives or that our objectives are compatible in some way or we have a good way of mediating when they have different objectives and so I think it is important to start thinking in terms like you don't have to be freaked out by the robot apocalypse to accept that it's important to think about objective functions of value alignment. And that you have to really everyone who's done optimization knows that you have to be careful what you wish for, that sometimes you get the optimal solution and you realize, man, that was that objective was wrong.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Oh, here's a hypothesis class. Maybe it's a space of plans, or maybe it's a space of classifiers or whatever, but there's some set of answers in an objective function. And then we work on some optimization method that tries to optimize a solution in that class. We don't know what solution is going to come out, right? So I think it's important to communicate that. So, I mean, of course, probably people who listen to this, they know that lesson. But I think it's really critical to communicate that lesson. And then lots of people are now talking about, you know, the value alignment problem.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Well, actually, let me talk a little bit about utility. Actually, I had an interesting conversation with some military ethicists who wanted to talk to me about autonomous weapons. And they were interesting, smart, well-educated guys who didn't know too much about AI or machine learning. And the first question they asked me was, has your robot ever done something you didn't expect? And I like burst out laughing because anybody who's ever done something other robot knows that they don't do much. And what I realized was that their model of how we program a robot was completely wrong. program a robot was like Lego mindstorms like oh go forward a meter turn left take a picture do this do that and so if you have that model of programming then it's true it's kind of weird that your robot would do something that you didn't anticipate but the fact is and and actually so now this is my new educational mission if I have to talk to non-experts I try to teach them the idea that we don't operate we operate at least one or maybe menu levels of abstraction above that and we say

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  8. So it's clear that the deep learning stuff has made deep and important improvements. And so the high watermark is now higher. There's no question. But of course, I think people are overselling and eventually investors, I guess, and other people look around and say, well, you're not quite delivering on this grand claim and that wild hypothesis. So probably it's going to crash some of them out. And then it's okay. I mean, but I don't, I can't imagine that there's like some awesome monotonic improvement from here to human level AI.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Think the cycles are inevitable, but I think each time we get higher, right? I mean, so, you know, it's like climbing some kind of landscape with a noisy optimizer.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Back in my day, when we worked on our theses, we did not publish papers. You did your thesis for years. You picked a hard problem and then you worked and chewed on it and did stuff and wasted time and for a long time. And when it was roughly when it was done, you would write papers. And so I don't know how to insert, and I don't think that everybody has to work in that mode, but I think there's some problems that are hard enough that it's important to have a longer research horizon. And I'm worried that we don't incentivize that at all at this point.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  11. No, it's a big investment to do a good review of a paper. And the flood of papers is out of control, right? So, you know, there aren't 3,000 new, I don't know how many new movies are there in a year. I don't know, but there's probably going to be less than how many machine learning papers there are in a year now. And I'm worried, you know, I, right, so I'm like an old person. So of course, I'm going to say things are moving too fast. I'm a stick in the mud. So I can say that. But my particular flavor of that is I think the horizon for researchers has gotten very short, that students want to publish a lot of papers and it's exciting. In that, and you get patted on the head for it, and so on. But, and some of that is fine, but I'm worried that we're driving out. People who would spend two years thinking about something

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  12. But if I felt inclined to do something more in the publication direction, I would do this other thing, which I thought about doing the first time, which is to get together some set of people whose opinions I value and who are pretty articulate. And I guess we would be public, although we could be private. I'm not sure. And we would review papers. We wouldn't publish them and you wouldn't submit them. We would just find papers and we would write reviews and we would make those reviews public. And maybe if you know so we're Leslie's friends who review papers and maybe eventually if we, our opinion was sufficiently valued, like the opinion of JMLR is valued, then you'd say on your CV that Leslie's friends gave my paper a five-star reading, and that would be just as good as saying I got it, you know, accepted into this journal. So I think we should have good public commentary and organize.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  13. That's good. But I think there's still. Value in careful reading and commentary of things. And it's hard to tell when people are upvoting and downvoting or arguing about your paper on Twitter and Reddit whether they know what they're talking about, right? So then I have the second order problem of trying to decide whose opinions I should value and such. So I don't know. If I had infinite time, which I don't, and I'm not going to do this because I really want to make robots work.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So, actually, when I started Jamlar, I wanted to do something completely different. And I didn't because it felt like we needed a traditional journal of record. And so we just made JamlR be almost like a normal journal, except for the open access parts of it, basically. Increasingly, of course, publication is not even a sensible word. You can publish something by putting it in an archive so I can publish everything tomorrow. So making stuff public is there's no barrier. We still need curation and evaluation. I don't have time to read all of Archive. And you could argue that Kind of social thumbs upping of articles suffices, right? You might say, oh, heck with this, we don't need journals at all. We'll put everything on archive and people will upvote and downvote the articles and then your CV will say, oh man, he got a lot of upvotes.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  15. I paid for the lawyer to incorporate and the IP address. And it just didn't cost a couple hundred dollars a year to run. It's a little bit more now, but not that much more. But that's because I think computer scientists are competent and autonomous in a way that many scientists in other fields aren't. I mean, at doing these kinds of things, we already types out our own papers. We all have students and people who can hack a website together in the afternoon. So the infrastructure for us was like not a problem, but for other people in other fields, it's a harder thing to do.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Yeah, right. So it's completely open. It's open access. Actually, I had a postdoc, George Conadaris, who wanted to call these journals free for all. Because there were, I mean, it both has no page charges and has no access restrictions. And the reason, and so lots of people, I mean, there were people who are mad about the existence of this journal who thought it was a fraud or something. It would be impossible, they said, to run a journal like this with basically, I mean, for a long time, I didn't even have a bank account.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Artificial intelligence. Yeah, okay, good. So it came about because there was a journal called Machine Learning, which still exists, which was owned by Kluwer. And there was, I was on the editorial board and we used to have these meetings annually where we would complain to Cluer that it was too expensive for the libraries and that people couldn't publish and we would really like to have some kind of relief on those fronts and they would always sympathize but not do anything. I just decided to make a new journal. And there was the journal of AI research which was on the same model which had been in existence for maybe five years or so and it was going along pretty well. We just made in our journal. It wasn't, I mean, I don't know. I guess it was work, but it wasn't that hard. So basically the editorial board probably 75% of the editorial board of

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  18. But I think what I mean, I think you get an interesting cycle where for a contest, a bunch of smart people get super motivated and they hack their brains out and much of what gets done is just hacks, but sometimes really cool ideas emerge. And then that gives us something to chew on after that. So it's not a thing for me, but I don't regret that other people do it.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  19. I think you can get in the way. I mean, some people, many people find it motivating, and so that's good. I find it anti motivating personally.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  20. No, I resist. I mean, I think all the time that we spend arguing about those kinds of things could be better spent just making their robots work better.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Well, to me, I guess the thing that seems most compelling to me at the moment is this question of what to build in and what to learn. I think We're missing a bunch of ideas. And we, you know, people, you know, don't you dare ask me how many years it's going to be until that happens because I won't even. Participate in the conversation I think we're missing ideas, and I don't know how long it's going to take to find them.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Well, but I, okay. But then what does self awareness mean? I mean, that you need to have some part of the system that can observe other parts of the system and tell whether they're working well or not. That seems critical. So does that count as, I mean, does that count as self-awareness or not? Well, it depends on whether you think that there's somebody at home who can articulate whether they're self-aware. But clearly, if I have like, you know, some piece of code that's counting how many times this procedure gets executed. That's a kind of self awareness, right? So there's a big spectrum. It's clear you have to have some of it.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Don't think much about consciousness. Even most philosophers who care about it will give you that you could have robots that are zombies, right, that behave like humans but are not conscious. And I at this moment would be happy enough with that. So I'm not really worried one way or the other.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  24. So, I think the pressing question is what kinds of structure can we build in that are like the moral equivalent of convolution that will make a really awesome superstructure that then learning can kind of progress on efficiently?

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  25. That's right. And people are still like the graph convolutions are an idea that are related to relational representations. So, I think there are, so you, I've come far afield from perception, but I think the thing that's going to make perception, that kind of the next step is actually understanding better what it should produce, right? So what are we going to do with the output of it? It's fine when what we're going to do with the output is steer. It's less clear when we're just trying to make one integrated intelligent agent. What should the output of perception be? We have no idea. And how should that hook up to the other stuff? We don't know.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  26. It's a very strong bias and it's a very critical bias So, my own view is that we should look for more things that are like convolution but that address other aspects of reasoning, right? So convolution helps us a lot with a certain kind of spatial reasoning that's quite close to the imaging. I think there's other ideas like that. Some amount of forward search, maybe some notions of abstraction, maybe the notion that objects exist. Actually, I think that's pretty important. And a lot of people won't give you that to start with, right?

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  27. We shouldn't build in any modularity. We should make a giant, gigantic neural network, train it end-to-end to do the thing. And that's the best way forward. And it's hard to argue with that, except on a sample complexity basis, right? So you might say, oh, well, if I want to do end-to-end reinforcement learning on this giant, giant neural network, it's going to take a lot of data and a lot of black broken robots and stuff. Then the only answer is to say, okay, we have to build something in, build in some structure or some bias we know from theory of machine learning the only way to cut down the sample complexity is to kind of cut down, somehow cut down the hypothesis space. You can do that by building in bias. There's all kinds of reason to think that nature built bias into humans. Convolution is a bias.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Well, I think the big question. Is representational hugely the question is representation. So Perception has made great strides lately, right? And we can classify images and we can. Play certain kinds of games and predict how to steer the car and all this sort of stuff. I don't think we have a very good idea of. What perception should deliver, right? So if you believe in modularity, okay, there's a very strong view which says.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Right. And so learning to do algebra manipulations for some reason is, I mean, that's probably going to want naturally a sort of a different representation than riding a unicycle. The time constraints on the unicycle are serious, the space base is maybe smaller. I don't know.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  30. But it's just as much of a policy as the other policy. It's just I've made, I think, the way I see it is it's a time space trade-off in computation. A more overt policy representation. Maybe it takes more space, but maybe I can compute quickly what action I should take. On the other hand, maybe a very compact model of the world dynamics plus a planner, let's compute what action to take too just more slowly. There's no, I don't, I mean, I don't think there's no argument to be had. It's just like a question of what form of computation is best for us.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Or maybe styles of. Problems. I mean, there's probably some reasoning that needs to go on in image space. I think, again, There's this model based versus model free idea, right? So on reinforcement learning, people talk about oh, should I learn? I could learn a policy just straight up a way of behaving. I could learn it's popular to learn a value function. That's some kind of weird intermediate ground. Or I could learn a transition model, which tells me something about the dynamics of the world. If I take a, imagine that I learn a transition model and I couple it with a planner and I draw a box around that. I have a policy again. It's just stored a different way.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Well, that would be a mistake. I mean, it's not all a planning problem, right? I think it's really, really important that we understand that you have to put together pieces and parts that have different styles of reasoning and representation and learning. I think it seems probably clear to anybody that it can't all be this or all be that. Brains aren't all like this or all like that, right? They have different pieces and parts and substructure and so on. So I don't think that there's any good reason to think that there's going to be like one true algorithmic thing that's going to do the whole job.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  33. You say, if you can reason at this other level and say, here's what I'm hoping to achieve, what could I do to make that true, that somehow the branching is smaller? Now, what's interesting is that like in the AI planning community, that hasn't worked out. In the class of problems that they look at and the methods that they tend to use, it hasn't turned out that it's better to go backward. It's still kind of my intuition that it is, but I can't prove that to you right now.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yeah, I mean, it's interesting. Simon, Herb Simon back in the early days of AI talked a lot about means, ends, reasoning, and reasoning back from the goal. There's a kind of an intuition that people have that The number of state space is big, the number of actions you could take is really big. So if you say, here I sit and I want to search forward from where I am, what are all the things I could do? That's just overwhelming.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Give me an estimate and it wouldn't be crazy. And you have to have an estimate of that in order to make plans that involve walking through the Kuala Lumpur Airport, even if you don't need to know it in detail. So I'm really interested in these kinds of abstract models and how do we acquire them. But once we have them, we can use them to do hierarchical reasoning, which is I think is very important.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  36. So, there's no point in planning in detail. But you have to have, you have to make a leap of faith that you can figure it out once you get there. And it's really interesting to me how you arrive at that. How do you, so you have learned over your lifetime to be able to make some kinds of predictions about how hard it is to achieve some kinds of sub-goals? And that's critical. Like you would never plan to fly somewhere if you couldn't didn't have a model of how hard it was to do some of the intermediate steps. So one of the things we're thinking about now is how do you do this kind of very aggressive generalization? To situations that you haven't been in, and so on, to predict how long will it take to walk through the koala lumpur airport.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Can plan to go to New York and arrive at the airport and then find yourself in an office building later. You can't even tell me in advance what your plan is for walking through the airport. Partly because you're too lazy to think about it maybe, but partly also because you just don't have the information. You don't know what gate you're landing in or what people are going to be in front of you or anything.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  38. And abstractions sort of temporal abstractions, then you can make plans at a high level. And you can say, I'm going to go to town, and then I'll have to get gas, and then I can go here and I can do this other thing. And you can reason about the dependencies and constraints among these actions. Again, without thinking about the complete details. What we do in our hierarchical planning work is then say, all right, I make a plan at a high level of abstraction, I have to have some reason to think that it's feasible without working it out in complete detail. And that's actually the interesting step. I always like to talk about walking through an airport.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  39. People since probably reasoning began have thought about hierarchical reasoning, the temporal hierarchy in particular. Well, there's spatial hierarchy, but let's talk about temporal hierarchy. So you might say, oh, I have this long execution I have to do, but I can divide it into some segment abstractly, right? So maybe I have to get out of the house, I have to get in the car, I have to drive, so on. And so... You can plan. If you can build abstractions, so this we started out by talking about abstractions and we're back to that now. If you can build abstractions in your state space.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  40. You have to trade off the fact that you're not seeing in front of you and you're looking behind you, and how valuable is that information, and so on. And so to make choices about information gathering, you have to reasonably leave space. Also to just. Take into account your own uncertainty before trying to do things. So you might say if I understand where I'm standing relative to the door jam pretty accurately, then it's okay for me to go through the door. But if I'm really not sure where the door is, then it might be better to not do that right now.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Well, any problem that requires deliberate information gathering, right? So if in some problems Like chess, there's no uncertainty. Or maybe there's uncertainty about the opponent. There's no uncertainty about the state. And some problems, there's uncertainty, but you gather information as you go, right? You might say, oh, I'm driving my autonomous car down the road and it doesn't know perfectly where it is, but the LIDARs are all going all the time. So I don't have to think about whether to gather information. But if you're a human driving down the road, you sometimes look over your shoulder to see what's going on behind you in the lane, and you have to decide whether you should do that now.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Control problem is actually the problem of controlling my beliefs. So I think about taking actions, not just what effect they'll have on the world outside, but what effect they'll have on my own understanding of the world outside. And so that might compel me to ask a question or look somewhere to gather information, which may not really change the world state, but it changes my own belief about the world

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Yeah, okay. It sounds good. So, believe space That is instead of thinking about what's the state of the world and trying to control that as a robot, I think about what is the space of beliefs that I could have about the world. If I think of a belief as a probability distribution over ways the world could be, a belief state is a distribution. And then my control problem, if I'm reasoning about how to move through a world I'm uncertain about,

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Need to know what the principles are. Why does this work? Why does that not work? I mean, for a while, people built bridges by trying, but now we can often predict whether it's going to work or not without building it. Can we do that for learning systems or for robots?

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  45. I hope there's I hope science will, I mean, there's engineering and there's science. I think that they're not exactly the same. And I think right now we're making huge engineering leaps and bounds so that engineering is running away ahead of the science, which is cool and often how it goes, right? So we're making things and nobody knows how and why they work, roughly. We need to turn that into science

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Right. So very bad job, very big problems. I like to do that. But I wish I could say something of a formal solution concept that I could use to say, oh, this algorithm actually, it gives me something. Like I know what I'm going to get. I can do something other than just run it and get out seven.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  47. And if you look at like computer science theory, so people talked for a while, everyone was about solving problems optimally or completely, and then there were interesting relaxations. So people look at, oh, are there regret bounds or can I do some kind of approximation? Can I prove something that I can approximately solve this problem or that I get closer to the solution as I spend more time and so on? What's interesting, I think, is that we don't have good approximate solution concepts for very difficult problems, right? I like to, you know, I like to say that I'm interested in doing a very bad job of very big problems.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  48. That's interesting. I think we have a little bit of a methodological crisis, actually, from the theoretical side. I mean, I do think that theory is important and that right now we're not doing much of it. So, there's lots of empirical hacking around and training this and doing that and reporting numbers, but is it good? Is it bad? We don't know. It's very hard to say things.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So we have to make approximations. Approximations in modeling, approximations in solution algorithms, and so on. And so I don't have a problem with saying, yeah, my problem actually is PalmDP in continuous space with continuous observations. And it's so computationally complex. I can't even think about it's big O whatever. But that doesn't prevent me from. It helps me, gives me some clarity to think about it that way. And to then take steps to make approximation after approximation to get down to something that's like computable in some reasonable time.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Optimal planning for even discrete Pom DPs can be undecidable depending on how you set it up. Lots of people say I don't use PalmDPs because they are intractable. And I think that that's kind of a very funny thing to say because. Problem you have to solve is the problem you have to solve. So if the problem you have to solve is intractable, that's what makes us AI people, right? So we solve, we understand that the problem we're solving is complete wildly intractable, that we can't, we will never be able to solve it optimally. At least I don't. Yeah, right. Later we can come back to an idea about bounded optimality and something. But anyway, I don't.

    2019-03-12 · Lex Fridman Podcast · Leslie Kaelbling: Reinforcement Learning, Planning, and Robotics · IDENTIFIED FROM THE TRANSCRIPT · source