YouSaid · the spoken record

Michael Littman

lines on the record
118
first
2020-12-13
most recent
2020-12-13
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Yeah, I'll send you a video. I'm not great, but I managed. And so balance. Yeah. So my wife has a really good one that she sticks to and is probably pretty accurate. And it has to do with healthy relationships with people that you love and working hard for good causes. But to me, yeah, balance in a word. That works for me. Not too much of anything because too much of anything is iffy.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Hitchhiker's Guide to the Galaxy that that is the meaning of life. So when I turned 42, I had a meaning of life party. Where I invited people over and everyone shared their meaning of life. We had slides made up. And so we all sat down and Did a slide presentation to each other about the meaning of life. And mine is great. Mine was balanced. I think that life is balanced. And so the activity at the party for 42-year-old, maybe this is a little bit non-standard. But I found all the little toys and devices that I had where you had to balance on them. You had to stand on it and balance or pogo stick I brought, a rip stick, which is like a weird two-wheeled skateboard. I got a unicycle, but I didn't know how to do it. I now can do it.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I mean, I think reinforcement learning researchers maybe think about this from a science perspective more often than a lot of other people, right? As a supervised learning person, you're probably not thinking about the sweep of a lifetime, but reinforcement learning agents are having little lifetimes, little weird little lifetimes. And it's hard not to project yourself into their world sometimes. But as far as the meaning of life. So when I turned 42, you may know from that is a book I read, the

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  4. So, sci-fi can be good at that too. So one sci-fi book to recommend is exhalations by Ted Chang, a bunch of short stories. Ted Chang is the guy who wrote the short story that became the movie Arrival. And all his stories just from a, he was a computer scientist, actually. He studied at Brown, they all have this sort of really insightful bit of science or computer science that drives them. And so it's just a romp, right, to just like he creates these artificial worlds with these by extrapolating on these ideas that we know about but hadn't really thought through to this kind of conclusion. And so his stuff is, it's really fun to read. It's mind warping.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  5. You know, that's probably a possible thing. He's maybe thought a thing or two about how to explain AI to people. Yeah. Yeah, that's a really good point. This book so far has been remarkably good at telling the story of sort of the history, the recent history of some of the things that have happened. I'm in the first third. He said this book is in three-thirds. The first third is essentially AI fairness and implications of AI on society that we're seeing right now. And that's been great. I mean, he's telling the stories really well. He went out and talked to the frontline people whose names were associated with some of these ideas and it's been terrific. He says the second half of the book is on reinforcement learning. So maybe that'll be fun. And then the third half, third, third, is on superintelligence alignment problem. And I suspect that that part will be less fun for me to read.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  6. In that same general neighborhood. I mean, they have different. Emphases that they're concentrating on. I think Stuart's book did a remarkably good job, like just a celebratory good job at describing AI technology and sort of how it works. I thought that was great. It was really cool to see that in a book.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Programming is power. Right. Yeah. It's like magic. It's like magic spells. And it's not out of reach of everyone. But at the moment, it's just a sliver of the population who can commune with machines in this way. So I don't know. So that book had a big impact on me. Currently, I'm reading The Alignment Problem, actually, by Brian Christian. So I don't know if you've seen this out there yet.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Do the things that that person once done. And as software people, we know how to do that when we have a problem. We're like, okay, I'll just hack up a Perl script or something and make it so. If we lived in a world where everybody could do that, that would be a better world. And computers would be have, I think, less sway over us and other people's software would have less sway over us as a group. Yeah.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  9. The analogy that he makes is that programming is a similar kind of thing that We need to have a say in right so being a reader, being literate, being a reader means you can receive all this information, but you don't get to put it out there And programming is the way that we get to put it out there. And that was the argument he made. I think he specifically has now backed away from this idea. He doesn't think it's happening quite this way. And that might be true that it didn't, society didn't sort of play forward quite that way. I still believe in the premise. I still believe that at some point we have the relationship that we have to these machines and these networks has to be one of each individual can has the wherewithal to make the machines help them.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Oh, yeah, I mean, yeah, I've read a lot of really fun stuff. In terms of books that I find myself thinking back on that I read a while ago, like that have stood the test of time to some degree. I find myself thinking of program or B programmed a lot by Douglas Roshkop, which basically put out the premise that we all need to become programmers in one form or another. And it was in analogy to once upon a time we all had to become readers. We had to become literate. And there was a time before that when not everybody was literate. But once literacy was possible, the people who were literate had more of a say in society than the people who weren't. And so we made a big effort to get everybody up to speed. And now it's not 100% universal, but it's quite widespread. Like the assumption is generally that people can read.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  11. What I think is going to happen is we're going to have a lot more, like a very gradual kind of rollout where people have these cars in closed communities, where it's somewhat realistic, but it's still in a box, right? So that we can really get a sense of what are the weird things that can happen, how do we have to change the way we behave around these vehicles. It's obviously requires a kind of co-evolution that you can't just plop them in and see what happens. But of course, we're basically popping them in and see what happens. So I was wrong, but I do think that would have been a better plan.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So many years ago, before self-driving cars were an actual thing, you could have a discussion about, somebody asked me, like, what if we could use that robotic technology and use it to drive cars around? Are people going to be killed? And then it's, you know, blah, blah, blah. I'm like, that's not what's going to happen, I said, with confidence incorrectly, obviously.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I know that you know that I know kind of planning. And the last time I spoke with her, she was very articulate about the ways in which self-driving cars are not solved. Like what's still really, really hard.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Pedestrians, you're practically talking to an octopus at that point. They've got all these weird degrees of freedom. You don't know what they're going to do. They can turn around any second.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Theories of mind of the other cars. Yeah, yeah, which I just hadn't heard discussed in the self driving car talks that I've been to. Since then, there's some people who do.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  16. And that blew me away. I was completely unaware of it until I watched my son learning to drive. And I was realizing that he was sending signals to all the cars around him. And those, in his case, he's always had social communication challenges. He was sending very mixed, confusing signals to the other cars. And that was causing the other cars to drive weirdly and erratically. And there was no question in my mind that he would have an accident because. They didn't know how to read him. There's things you do at the speed that you drive, the positioning of your car that you're constantly like in the head of the other drivers and seeing him not knowing how to do that and having to be taught explicitly. Okay, you have to be thinking about what the other driver is thinking. Was a revelation to me. I was stunned.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  17. So I sit in the passenger seat and it's really scary. Have wishes to live, and they're figuring things out now. They start off very Very much better than I imagine, like a neural network would, right? They get that they're seeing the world, they get that there's a road, that they're trying to be on, they get that there's a relationship between the angle, the steering, but it takes a while to not be very jerky So that happens pretty quickly. Like the ability to stay in lane at speed happens relatively fast. It's not zero shot learning, but it's pretty fast. The thing that's remarkably hard, and this is, I think, partly why self-driving cars are really hard, is the degree to which driving is a social interaction activity.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Have you taught anyone to drive? So I have two children. I learned a lot about car driving because my wife doesn't want to be the one in the car while they're learning. So that's my job. Yeah.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  19. In development. And I'm not that guy. In fact, I used to joke that when I got into college that it was on kind of help out the illiterate kind of program because I got to college. In my house, I wasn't a particularly bad or good reader. But when I got to college, I was surrounded by these people that were just voracious in their reading appetite. And they would like, have you read this? Have you read this? Have you read this? And I'd be like, no, I'm clearly not qualified to be at this school. Like there's no way I should be here. Now I've discovered books on tape, like audiobooks. And so I'm much better. I'm more caught up. I read a lot of books.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  20. So, I don't read very, I mean, obviously, I read and I've read plenty of books, but like some people, like Charles, my friend Charles and others, like a lot of people in my field, a lot of academics, like reading was really a central topic to them.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  21. And I don't read. He's like, you must. I'm like, I don't think I do. I mean, like, stop signs. I definitely read stop signs. But like reading books is not a thing that I do a lot. You really, though?

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  22. That called latent semantic analysis that would take words of English and embed them in multi-hundred dimensional space and then use that as a way of assessing similarity and basically doing reinforcement learning, not sorry, not reinforcement, information retrieval, sort of pre-Google information retrieval. And he was trained as an anthropologist, but then became a cognitive scientist. So I was in the cognitive science research group. Like I said, I'm a cognitive science groupie. At the time, I thought I'd become a cognitive scientist, but then I realized in that group, no, I'm a computer scientist, but I'm a computer scientist who really loves to hang out with cognitive scientists. And he said, he studied language acquisition in particular. He said, you know, humans have about this number of words of vocabulary, and most of that is learned from reading. And I said, that can't be true because I have a really big vocabulary.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  23. I'd like to tell you a story. So, my first job out of college was at Belcore. I mentioned that before, where I worked with Dave Ackley. The head of the group was a guy named Tom Landauer. And I don't know how well known he's known now. But arguably, he's the inventor and the first proselytizer of word embeddings. So they developed a system shortly before I got to the group.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Focus. So you're right in that broader system We're also a part of and can have some influence on, it is much more complicated and much more powerful. Yeah, I agree with that.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  25. It's a system that includes human companies and corporations, right? Because corporations are funny organisms in and of themselves that really do seem to have self-preservation built in. And I think that's at the design level. I think they're designed to have self-preservation be.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  26. And they arguably already are manipulating human behavior. But not for self Which I think is a big, that would be a big step. Like, if they were trying to manipulate us to convince us not to shut them off, I would be very freaked out. But I don't see a path to that from where we are now. They don't have any of those abilities. That's not what they're trying to do. They're trying to keep people on the site.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  27. That means we're going to discover life on other planets. No, it doesn't. It means that we're in a sigmoid curve on the front half, which looks a lot like an exponential. The second half is going to look a lot like diminishing returns.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  28. That's got to end, right? And there's hints now that Moore's Law is starting to feel some friction, starting to, the world is pushing back a little bit. One thing, I don't know, lots of people know this. I didn't know this. I was trying to write an essay. And yeah, Moore's Law has been amazing and it's enabled all sorts of things. But there's also a kind of counter Moore's law, which is that the development cost for each successive generation of chips also is doubling. So it's costing twice as much money. So the amount of development money per cycle or whatever is actually sort of constant. And at some point, we run out of money, or we have to come up with an entirely different way of doing the development process. So I guess I always, always a bit skeptical of the, look, it's an exponential curve. Therefore, it has no end. Soon the number of people going to Neurops will be greater than the population of the Earth.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Motivated part based models that, you know, that just feel like obviously the right thing that you have to have, or we can throw a lot of data at it, and guess what? We're doing better with a lot of data. Hadn't thought about it until this moment in this way, but what I believe, well, I've thought about what I believe. What I believe is that, you know, Compositionality The right way to say it. The complexity grows rapidly as you consider more and more possibilities, like explosively. And so far, Moore's law has also been growing explosively, exponentially. And so it really does seem like, well, we don't have to think really hard about the algorithm design or the way that we build the systems because the best benefit we could get is exponential and the best benefit that we can get from waiting is exponential.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Yes, that's right. Just waiting for them to do their job so that we can pretend to have done ours. So, I mean, the argument reminds me a lot of, I think it was a Fred Jellinek quote, early computational linguist who said, you know, we're building these computational linguistic systems. And every time we fire a linguist, performance goes up by 10%, something like that. And so the idea of us building the knowledge in that case was much less, he was finding it to be much less successful than get rid of the people who know about language as a you know from a kind of scholastic academic kind of perspective and replace them with more compute. And so I think this is kind of a modern version of that story, which is, okay, we want to do better on machine vision. You could build in all these, you know,

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  31. As an academic discipline, that we're really just a subfield of computer architecture. We're just kind of weighing around for them to do it.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Be pushed back at. I think that conversations, even people who are pretty smart, maybe the smartest thing that we know, maybe not the smartest thing we can imagine, but we get so much benefit out of talking to each other and interacting. That's presumably why you have conversations live with guests is that there's something in that interaction that would not be exposed by, oh, I'll just write you a story and then you can read it later. And I think because these systems are just learning from our stories, they're not learning from being pushed back at by us that they're fundamentally limited into what they could actually become on this route. They have to get, you know, shut down. We have to have an argument that they have to have an argument with us and lose a couple times before they start to realize, oh, okay, wait, there's some nuance here that actually matters.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  33. I don't believe that you can get something that really Can do language and use language as a thing that doesn't interact with people. Like I think that it's not enough to just take everything that we've said written down and just say, that's enough. You can just learn from that and you can be intelligent. I think you really need to.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  34. But it's really hard to see in the statistics because so much of what we're saying is kind of wrote. And so our metrics that we use to measure how these systems are doing don't reveal that because it's in the interestes that is very hard to detect.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Right. So I've had a similar thought, which was that The stories that GPT 3 spits out are amazing and very human like. And it doesn't mean that computers are smarter than we realize necessarily. It partly means that people are dumber than we realize or that much of what we do day to day is not that deep. Like we're just kind of going with the flow. We're saying whatever feels like the natural thing to say next. Not a lot of it is. Is creative or meaningful or intentional? But enough is that we actually get by, right? We do come up with new ideas sometimes and we do manage to talk each other into things sometimes. And we do sometimes vote for reasonable people sometimes.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Yeah, yeah, yeah. So, first of all, the whole transformer network. Family of things is really cool. It's really, really cool. I mean, if you've ever Back in the day, you played with, I don't know, Markov models for generating text, and you've seen the kind of text that they spit out and you compare it to what's happening now, it's amazing. It's so amazing. Now, it doesn't take very long interacting with one of these systems before you find the holes, right? It's not smart in any kind of general way. It's really good at a bunch of things and it does seem to understand a lot of the statistics of language extremely well. And that turns out to be very powerful. You can answer many questions with that. But it doesn't make it a good conversationalist, right? And it doesn't make it a good storyteller. It just makes it good at imitating of things that it is seen in the past.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  37. And so it could be that what David is seeing is a kind of asymptoting that you can keep getting better, but with diminishing returns. And at some point, you hit optimal play. Like in theory, all these finite games, they're finite. They have an optimal strategy. There's a strategy that is the mini-max optimal strategy. And so at that point, you can't get any better. You can't beat that strategy. Now, that strategy may be from an information processing perspective intractable, right? You need all the situations are sufficiently different that you can't compress it at all. It's this giant mess of hard-coded rules. And we can never achieve that. But that still puts a cap on how many levels of improvement that we can actually make.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  38. So the question is, okay, so I don't know his analysis on that. From talking to Go experts, the depth, the strategic depth of Go seems to be substantially greater than that of chess, that there's more kind of Steps of improvement that you can make getting better and better and better and better. But there's no reason to think that it's infinite.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Right. And so using the information about how. These things called Elo scores, this sort of notion of how strong a player are you? There's kind of a range of possible scores. And you increment and score basically if you can beat another player of that lower score 62% of the time or something like that. Like there's some threshold of if you can somewhat consistently beat someone, then you are of a higher score than that person. And there's a question as to how many times can you do that in chess, right? And so we know that there's a range of human nobility levels that cap out with the best in humans. And the computers went a step beyond that. And computers and people together have not gone, I think, a full step beyond that. It feels the estimates that they have is that it's starting to asymptote, that we've reached kind of the maximum, the best possible chess playing. And so that means that there's kind of a finite strategic depth, right? At some point, you just can't get any better at this.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Right. So that, I think, is a really good question. And I think that we don't, I don't think we as a, I don't know, community really know the answer to this. So, okay, so I went to a talk by some experts on computer chess. So in particular, computer chess is really interesting because for, of course, for a thousand years, humans were the best chess playing things on the planet. And then computers like edge ahead of the best person. And they've been ahead ever since. It's not like people have overtaken computers. But computers and people together have overtaken computers. So, at least last time I checked, I don't know what the very latest is, but last time I checked that there were teams of people who could work with computer programs to defeat the best computer programs.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  41. But notice it's not getting better at other things. It's getting better at Go. And I think that's a big leap to say, okay, well, therefore it's... Better at other things

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  42. The system. I believe that the system can take that kind of leap. Yeah, no, and also I think that beginner knowledge. Like you can start to get a feel really quickly for the idea that Certain parts of the being in certain parts of the board seem to be more associated with winning. Because it's not stumbling upon the concept of winning. It's told that it wins or that it loses. Well, it's self-play. So it both wins and loses. It's told which side won. And the information is kind of there to start percolating around to make a difference as to, well, these things have a better chance of helping you win and these things have a worse chance of helping you win. And so, you know, it can get to basic play, I think, pretty quickly. Then once it has basic play, well, now it's kind of forced to do some search to actually experiment with, okay, well, what gets me that next increment of improvement?

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Humongous amount of information. But we know that it went beyond that. We know that it somehow got away from that information because it was learning strategies. I don't think AlphaGo is just better at implementing human strategies. I think it actually developed its own strategies that were more effective. And so from that perspective, okay, well, so it made at least one. Quantum leap in terms of strategic knowledge. Okay, so now maybe it makes three. Like, okay, but that first one is the doozy, right? Getting it to. Work reliably and for the networks to hold on to the value well enough. That was a big step.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Riches in Alberta and Satender would be in England, but I think he's in England from Michigan at the moment. But he was, yes, he was much more impressed with Alpha Go Zero, which is didn't get a kind of a bootstrap in the beginning with human-trained games. Just was purely self-play. Though the first one, Alpha. Was also a tremendous amount of self play. They started off, they kickstarted the action network that was making decisions, but then they trained it for a really long time using more traditional temporal difference methods. So as a result, it didn't seem that different to me. Like it seems like, yeah, why wouldn't that work? Like once it works, it works. So he found that removal of that extra information to be breathtaking. That's a game changer. To me, the first thing was more of a game changer.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Right. So then alpha zero. So I think I may have a slightly different opinion on this than some people. So I talked to Sitinder Singh in particular about this. So Satinder was like Rich Sutton, a student of Antibarto. So they came out of the same lab, very influential machine learning reinforcement learning researcher now at DeepMind, as is rich, though different sites, the two of them.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Yeah. And so a couple things seemed like they were really promising. None of them made it into production that I'm aware of. And neural nets as a whole started to kind of implode around then. And so there just wasn't a lot of air in the room for people to try to figure out, okay, how do we get this to work in the RL setting?

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Some work that Satinder Singh et al. did on handoffs with cell phones, deciding when should you hand off from this cell tower to this cell tower. Communication now.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  48. I mean, like I said, the students who I worked with, we tried to get, basically apply that architecture to other problems. And we consistently failed. There were a couple. A couple really nice demonstrations that ended up being in the literature. There was a paper about controlling elevators where it's like, okay, can we modify the heuristic that elevators use for deciding like a bank of elevators for deciding which floors we should be stopping on to maximize throughput, essentially? And you can set that up as a reinforcement learning problem and you can have a neural net represent the value function so that it's taking where all the elevators, where the button pushes, this high dimensional, well, at the time high dimensional input, a couple dozen dimensions, and turn that into a prediction as to, oh, is it going to be better if I stop at this floor or not? And ultimately, it appeared as though for the standard simulation, distribution for people trying to leave the building at the end of the day, that the neural net learned a better strategy than the standard one that's implemented in elevator controllers. So that was nice.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  49. But with enough skepticism that you're looking for where the problems are and fighting through them. Because you know there's got to be a way out of this thing.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  50. These techniques that were so good at playing chess and that could beat the world champion in chess couldn't beat your typical go-playing teenager in Go. So the fact that in a very short number of years we kind of ramped up to trouncing people in Go just blew me away.

    2020-12-13 · Lex Fridman Podcast · #144 – Michael Littman: Reinforcement Learning and the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source