YouSaid · the spoken record

David Silver

lines on the record
114
first
2020-04-03
most recent
2020-04-03
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Well, let's just take that argument and pursue it to its natural conclusion. So the next level indeed is for how can our learning brain achieve its goals most effectively? Well, maybe it does so by Us as learning beings, building a system which is able to solve for those goals more effectively than we can. And so when we build a system to play the game of go, when I said that I wanted to build a system that can play Go better than I can, I've enabled myself to achieve that goal of playing Go better than I could by directly playing it and learning it myself. And so now a new layer has been created, which is systems which are able to achieve goals for themselves. And ultimately, there may be layers beyond that where they set sub-goals to parts of their own system in order to achieve those and so forth.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  2. They lead to the outcome of that system is not in contradiction with the fact that it's also a decision making system that's optimizing for some goal and purpose.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Another level of discovery, which was learning systems, you know, parts of the brain which were able to learn for themselves and learn how to program themselves to achieve any goal. And presumably, there are parts of the brain where goals are set to parts of that system and provides this very flexible notion of intelligence that we as humans presumably have, which is the ability to kind of, the reason we feel that we can achieve any goal. So it's a very long-winded answer to say that I think there are many perspectives and many levels at which intelligence can be understood. And each of those levels, you can take multiple perspectives. You can view the system as something which is optimizing for a goal, which is understanding it at a level by which we can maybe implement it and understand it as AI researchers or computer scientists. Or you can understand it at the level of the mechanistic thing which is going on, that there are these atoms bouncing around in the brain.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  4. What's evolution then? Well, it's now got its own goal within that, which is to actually reproduce as effectively as possible. And now, how does reproduction, how is that made as effective as possible? Well, you need entities within that that can survive and reproduce as effectively as possible. And so it's natural that in order to achieve that higher level goal, those individual organisms discover brains, intelligences, which enable them to support the goals of evolution. And those brains, what do they do? Well, perhaps the early brains, maybe they were controlling things at some direct level. Maybe they were the equivalent of pre-programmed systems which were directly controlling what was going on and setting certain things in order to achieve these particular goals. But that led to a

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  5. How can that be done by a particular system? And maybe evolution is something that the universe discovered in order to kind of dissipate energy as efficiently as possible. And by the way, I'm borrowing from Max Tegmark for some of these metaphors, the physicist. But if you can think of evolution as a mechanism for dispersing energy, then evolution you might say then becomes a goal, which is if evolution disperses energy by reproducing as efficiently as possible.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  6. And say, I think that there are many levels at which you can understand a system and you can understand something as optimizing for a goal at many levels. And so you can understand, let's start with the universe. Does the universe have a purpose? Well, it feels like it's just one level just following certain mechanical laws of physics and that that's led to the development of the universe. But at another level, you can view it as actually there's the second law of thermodynamics that says that this is increasing in entropy over time forever and now there's a view that's been developed by certain people at MIT that this you can think of this as almost like a goal of the universe that the purpose of the universe is to maximize entropy so there's multiple levels at which you can understand a system the next level down you might say well if the goal is to maximize entropy well

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  7. And we choose those because we think it will lead us to something later on. We think that that's helpful to us to achieve some ultimate goal. Now I don't want to speculate whether or not humans as a system necessarily have a singular overall goal of survival or whatever it is. But I think the principle for understanding and implementing intelligences has to be that if we're trying to understand intelligence or implement our own, there has to be a well-defined problem. Otherwise If it's not I think it's like an admission of defeat that for there to be hope for understanding or implementing intelligence we have to know what we're doing. We have to know what we're asking the system to do otherwise if you don't have a clearly defined purpose you're not going to get a clearly defined answer

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  8. System may choose to create its own motivations and subgoals that help the system to achieve its ultimate goal. And that may indeed be a hugely important mechanism to achieve those ultimate goals. But there is still some ultimate goal, I think the system needs to be measurable and evaluated against. And even for humans, I mean, humans, we're incredibly flexible. We feel that we can, you know, any goal that we're given, we feel we can master to some degree. But if we think of those goals really, you know, like the goal of being able to pick up an object or the goal of being able to communicate or influence people to do things in a particular way or whatever those goals are, really, their sub-goals really that we set ourselves, you know, we choose to pick up the object. We choose to communicate. We choose to influence someone else.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  9. So I think when we think about intelligence, it's really important to be clear about the problem of intelligence. And I think it's clearest to understand that problem in terms of some ultimate goal that we want the system to try and solve for. And after all, if we don't understand the ultimate purpose of the system, do we really even have a clearly deprived problem that we're solving at all? Now, within that, as with your example for humans,

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Showed that in quantum computation, one of the big questions is how to understand the nature of the function in quantum computation and a system based on alpha beat the state of the art by quite some distance there again. So these are just examples. And I think the lesson which we've seen elsewhere in machine learning time and time again is that if you make something general, it will be used in all kinds of ways. You know, you provide a really powerful tool to society. And those tools can be used in amazing ways. And so I think we're just at the beginning. And for sure, I hope that we see all kinds of outcomes.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  11. So I absolutely do hope and imagine that we will get to the point where ideas just like these are used in all kinds of different domains. In fact, one of the most satisfying things as a researcher is when you start to see other people use your algorithms in unexpected ways. So in the last couple of years, there have been a couple of nature papers where different teams unbeknownst to us took alpha zero and applied exactly those same algorithms and ideas to real world problems of huge meaning to society. So one of them was the problem of chemical synthesis and they were able to beat the state of the art in finding pathways of how to actually synthesize chemicals, retrochemical synthesis. And the second paper actually actually just came out a couple of weeks ago in Nature.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Favour of its own Giuseki that humans didn't know about. It starts to say, Oh, well, you thought that the knights moved Pinza Joseki was a great idea. But here's something different you can do there, which makes some new variation that humans didn't know about. And actually now the human go players study the Gaseki that Alpha Go played and they become the new norms that are used in today's top-level Go competitions.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Timeline of discovery where what we saw was that there are these opening patterns that humans play called Giuseki. These are like the patterns that humans learn to play in the corners and they've been developed and refined over literally thousands of years in the game of Go. And what we saw was in the course of the training Alpha Go Zero over the course of the 40 days that we trained this system, it starts to discover exactly these patterns that human players play. And over time, we found that all of the Gusecki that humans played were discovered by this system through this process of self-play and this sort of essential notion of creativity. But what was really interesting was that over time it then started to discard some of these.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Here's something great. I'm going to start using that. And then that process, it's like a micro discovery that happens millions and millions of times over the course of the algorithm's life, where it just discovers some new idea, oh, this pattern, this pattern's working really well for me. I'm going to start using that. Oh, now, oh, here's this other thing I can do. I can start to connect these stones together in this way, or I can start to sacrifice stones or give up on pieces or play shoulder hits on the fifth line or whatever it is. The system's discovering things like this for itself continually, repeatedly all the time. And so it should come as no surprise to us then, when if you leave these systems going, that they discover things that are not known to humans, to the human norms are considered creative. And we've seen this several times. In fact, in alpha go zero, we saw this beautiful.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  15. So let me start by Think saying that we should ask what creativity really means. So to me, creativity. Means discovering something which wasn't known before, something unexpected, something outside of our norms. And so in that sense, the process of reinforcement learning or self-play approach that was used by AlphaZero is the essence of creativity. It's really saying at every stage you're playing according to your current norms and you try something and if it works out you say, hey.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  16. The full reinforcement learning problem needs to deal with worlds that are unknown and complex. And the agent needs to learn for itself how to deal with that. And so Musira was a step, a further step in that direction.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Showing that even without the rules, the system can learn for itself just by trial and error, just by playing this game of Go. And no one tells you what the rules are, but you just get to the end and someone says, you know, win or loss. You play this game of chess and someone says win or loss, or you play a game of breakout in Atari, and someone just tells you, you know, your score at the end. And the system for itself figures out essentially the rules of the system, the dynamics of the world, how the world works. not in any explicit way, but just implicitly enough understanding for it to be able to plan in that system in order to achieve its goals.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  18. This led us to the most recent step in the story of Alphago, which was a system called MuZero. And muzero is a system which learns for itself even when the rules are not given to it. It actually can be dropped into a system with messy perceptual inputs. We actually tried it in some Atari games, the canonical domains of Atari that have been used for reinforcement learning. And this system learned to build a model of these Atari games that was sufficiently rich and useful enough for it to be able to plan successfully. And in fact, that system not only went on to beat the state of the art in Atari, but the same system without modification was able to reach the same level of superhuman performance in Go, Chess, and Shoghi that we'd seen in AlphaZero.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  19. With this stream of observations coming in, rich sensory input coming in, actions going out in a way that allows it to reason in the way that AlphaGo or AlphaZero can reason, in the way that these Go and chess playing programs can reason. But in a way that allows it to take actions in that messy world to achieve its goals.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  20. One of the important steps is to acknowledge that the world is a really messy place. It's this rich, complex, beautiful, but messy environment that we live in. And no one gives us the rules. No one knows the rules of the world. At least maybe we understand that it operates according to Newtonian or quantum mechanics at the micro level or according to relativity at the macro level. But that's not a model that's useful for us as people to operate in it. Somehow the agent needs to understand the world for itself in a way where no one tells it the rules of the game, and yet it can still figure out what to do in that world.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  21. And we also beat the world's strongest programs and reached superhuman performance in that game, too. And it was the very first time that we'd ever run the system on that particular game was the version that we published in the paper on AlphaZero. It just worked out of the box literally. No touching it. We didn't have to do anything. And there it was, superhuman performance, no tweaking, no twiddling. I think there's something beautiful about that principle that you can take an algorithm and without twiddling anything, it just works. Go beyond alpha zero what's required. Alpha zero is just a step, and there's a long way to go beyond that to really crack the deep problems of AI.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Think the remarkable observation which we saw with alpha zero was that actually without modifying the algorithm at all, it was able to play and crack some of AI's greatest previous challenges, in particular we dropped it into the game of chess, and unlike the previous systems like Deep Blue, which had been worked on for years and years, we were able to beat the world's strongest computer chess program convincingly using a system that was fully discovered by its own from scratch with its own principles. And in fact, one of the nice things that we found was that we also achieved the same result in Japanese chess, a variant of chess where you get to capture pieces and then place them back down on your own side as an extra piece. So much more complicated variant of chess.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Essentially, progress will always lead you to if you have sufficient representational resource, like imagine you could represent every state in a big table of the game, then we know for sure that a progress of self-improvement will lead all the way in the single agent case to the optimal possible behavior and in the two-player case to the minimax optimal behavior. That is the best way that I can play knowing that you're playing perfectly against me. And so for those cases, we know that even if you do open up some new error, that in some sense you've made progress. You're progressing towards the best that can be done.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  24. The Game of Go would set the ceiling, but the Game of Go has 10 to the 170 states in it. So the ceiling is unreachable by any. Computational device that can be built out of the 10 to the 80 atoms in the universe. You asked a really good question, which is Do you not open up other errors when you correct your previous ones? And the answer is yes, you do. And so it's a remarkable fact about this class of two-player game and also true of single agent games.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Let me back this up. I think science always should make falsifiable hypotheses. Let me back up this claim with a falsifiable hypothesis, which is that if someone was to in the future take alpha zero as an algorithm and run it on with greater computational resources that we had available today, then I would predict that they would be able to beat the previous system 100 games to zero, and that if they were then to do the same thing a couple of years later, that previous system 100 games to zero, and that that process would continue indefinitely throughout at least my human lifetime.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  26. And understand well, what's that doing wrong? And it takes you on to the next level and the next level. And this progress can go on indefinitely. And indeed, what would have happened if we'd carried on training alpha go zero for longer, we saw no sign of it slowing down its improvements, or at least it was certainly carrying on to improve. Presumably, if you had the computational resources, this could lead to better and better systems that discover more and more.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Now, if you take that same idea and trace it back all the way to the beginning, it should be able to take you from no knowledge, from completely random starting point, all the way to the highest levels of knowledge that you can achieve in a domain. And the principle is the same, that if you bestow a system with the ability to correct its own errors, then it can take you from random to something slightly better than random because it sees the stupid things that the random is doing and it can correct them. And then it can take you from that slightly better.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  28. When it's doing something wrong and correct for it. And so it seemed to me that the way to correct delusions was indeed to have more iterations of reinforcement learning. No matter where you start, you should be able to correct those errors until it gets to play out and understand, oh, well, I thought that I was going to win in this situation, but then I ended up losing that suggest that I was misevaluating something. There's a hole in my knowledge and now the system can correct for itself and understand how to do better.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Then lo and behold, it actually ended up outperforming the previous version of Althago and indeed was able to beat it by 100 games to zero. So what's the intuition as to why? I think the intuition to me is clear that whenever you have errors in a system, as we did in AlphaGo, AlphaGo suffered from these delusions. Occasionally it would misunderstand what was going on in a position and misevaluate it. How can you remove all of these errors? Errors arise from many sources. For us, they were arising both from starting from the human data, but also from the nature of the search and the nature of the algorithm itself. But the only way to address them in any complex system is to give the system the ability to correct its own errors. It must be able to correct them. It must be able to learn for itself.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  30. I don't see much value in running an experiment where you're 95% confident that you will succeed. And so we could have tried maybe to take AlphaGo and do something which we knew for sure it would succeed on. But much more interesting to me was to try it on the things which we weren't sure about. And one of the big questions on our minds back then was could you really do this with self-play alone? How far could that go? Would it be as strong? And honestly, we weren't sure. It was 50-50, I think. I really, if you'd asked me, I wasn't confident that it could reach the same level as these systems, but it felt like the right question to ask. And even if it had not achieved the same level, I felt that that was an important direction to be studying. And so...

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Let me first say why we tried it. So we tried it both because I feel that it was the deeper scientific question to be asking to make progress towards AI. And also because in general in my research, I don't like to do research on questions for which we already know the likely outcome.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Bing this, like the algorithm for alpha zero just appeared. And in its full form. And this was actually before we played against Lisa Dahl, but we just didn't. I think we were so busy trying to make sure we could beat the world champion that it was only later that we had the opportunity to step back and start examining that sort of deeper scientific question of whether this could really work.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  33. In fact, I mean, Just for fun, I could tell you exactly the moment where the idea for Alpha Zero occurred to me because I think there's maybe a lesson there for researchers who are kind of too deeply embedded in their research and working 24-7 to try and come up with the next idea, which is it actually occurred to me on honeymoon. And I was like, my most fully relaxed state, really enjoying myself. And just...

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  34. First step was to try and take all of the knowledge out of AlphaGo in such a way that it could play in a fully self-discovered way, purely from self-play. And to me, the motivation for that was always that we could then plug it into other domains. But we saved that until later.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Such as that, which no matter what the goal is, no matter what goal we set to the system, we can come up with, we have an algorithm which can be placed into that world, into that environment, and can succeed in achieving that goal. And then that to me is almost the essence of intelligence, if we can achieve that. And so alpha zero is a step towards that. And it's a step that was taken in the context of two-player perfect information games like Go and Chess. We also applied it to Japanese chess

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Elegant principle by which a system can learn for itself all of the knowledge which it requires to play, to play a game such as Go. Importantly, by taking knowledge out you not only make the system less brittle in the sense that perhaps the knowledge you were putting in was just getting in the way and maybe stopping the system learning for itself, but also you make it more general. knowledge you put in, the harder it is for a system to actually be placed taken out of the system in which it's kind of been designed and placed in some other system that maybe would need a completely different knowledge base to understand and perform well. And so the real goal here is to strip out all of the knowledge that we put in to the point that we can just plug it into something totally different. And that to me is really the promise of AI is that we can have systems.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Let me start with self play. So the idea of self play is something which is really about systems learning for themselves, but in the situation where there's more than one agent. And so if you're in a game, a game is played between two players, then self-play is really about understanding that game. Just by playing games against yourself rather than against any actual real opponent. And so it's a way to kind of, um, Discover strategies without having to actually need to go out and play against any particular human player, for example. The main idea of alpha zero was really to try and Step back from any of the knowledge that we'd put into the system and ask the question, is it possible to come up with a single

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Any action that they might take within that world and any state they might find themselves in, and in order to do that. We make progress towards AI

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  39. There's a ceiling on how well it can do, but maybe more importantly, it means that the same idea cannot be applied in other domains where we don't have access to the kind of human grandmasters and that ability to kind of encode exactly their knowledge into an evaluation function. And the reality is that the story of AI is that most domains turn out to be of the second type, where when knowledge is messy, it's hard to extract from experts or it isn't even available. And so we need to solve problems in a different way. And I think AlphaGo is a step towards solving things in a way which puts learning as a first class citizen and says systems need to understand for themselves how to understand the world, how to judge the value of

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  40. So I think we should not demean the achievements of what was done in previous eras of AI. I think that Deep Blue was an amazing achievement in itself and that heuristic search of the kind that was used by DeepBlue had some powerful ideas that were in there. But it also missed some things. So the fact that the evaluation function function, the way that the chess position was understood was created by humans and not by the machine is a limitation which means that

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  41. To Gary Kasparov, that moment in computer chess was more profound than Deep Blue. And the reason he believes it mattered more was because it was done with learning and a system which was able to discover for itself new principles, new ideas, which were able to play the game in a way which he hadn't always known about or anyone. And in fact, one of the things I discovered at this panel was that the current world champion, Magnus Carlsen, apparently recently commented on his improvement in performance and he attributes it to AlphaZero that he's been studying the games of AlphaZero and he's changed his style Play more like Alpha Zero, and it's led to him actually increasing his rating to a new peak.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  42. So Gary Kasparoff has an incredible respect for. What we did with AlphaGo, and you know, it's an amazing Tribute coming from him of all people that he really appreciates and respects what we've done. Think he feels that the progress which has happened in computer chess, which Later after AlphaGo, we built the Alpha Zero system Which defeated the world's strongest chess programs.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  43. After that, we increased the power of the system, and the next version of AlphaGo beats the other strong human players 60 games to nil. So what a great moment for him and something to be remembered for.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Then his personal loss in that moment. Lisa Dole, I think, was a much more cognizant of that, even at the time. So in his closing remarks to the match, he really felt very strongly that what had happened in the AlphaGo match was not only meaningful for AI, but for humans as well. And he felt as a go player that it had opened his horizons and meant that he could start exploring new things it had brought his joy back for the game of Go because it had broken all of the conventions and barriers and meant that suddenly anything was possible again. And so, you know, I was sad to hear that he'd retired, but he's been a great world champion over many, many years. And I think he'll be remembered for that ever more. He'll be remembered as the last person to beat Alphago. I mean, after...

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Let me take you back first of all to the first part of your comment about Gary Kasparov, because actually, at the panel yesterday, he specifically said that when he first lost a deep blue, he viewed it as a failure. He viewed that this had been a failure of his. But later on in his career, he said he'd come to realize that actually it was a success. It was a success for everyone because this marked a transformational moment for AI. And so even for Gary Kasparov, he came to realize that that moment was pivotal.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Final game. For some reason, we as a team were convinced having seen Alpha Go in the previous game suffer from delusions. We as a team were convinced that it was suffering from another delusion. We were convinced that it was misevaluating the position and that something was going terribly wrong. And it was only in the last few moves of the game that we realized that actually although it had been predicting it was going to win all the way through, it really was. And so somehow, you know, it just taught us yet again that you have to have faith in your systems when they exceed your own level of ability and your own judgment, you have to trust in them to know better than you that design now. You've bestowed in them the ability to judge better than you can, then trust the system to do so.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  47. And so for us, this was a real challenge. Like, would Alpha go be able to deal with this or would it just kind of crumble in the face of this situation? And fortunately, it dealt with it perfectly. The fourth game was amazing in that Lisa Dole appeared to be losing this game. AlphaGo thought it was winning. And then Lisa Dole did something which I think only a true world champion can do, which is he found a brilliant sequence in the middle of the game, a brilliant sequence that led him to really just transform their position. He found just a piece of genius, really. And after that, AlphaGo, its evaluation just tumbled. It thought it was winning this game and all of a sudden it tumbled and said, oh, now I've got no chance. And it started to behave rather oddly at that point.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  48. These players are amazing. Lisa Doll was a true champion, 18 time world champion, and had this amazing ability to probe Alpha Gopher for weaknesses of any kind. And in the third game, he was losing and we felt we were sailing comfortably to victory, but he managed to, from nothing, stir up this fight and build what's called a double co, these kind of repetitive positions. And he knew that historically no computer GO program had ever been able to deal correctly with double co-positions. And he managed to summon one out of nothing.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Anticipated and computers discovered this idea. They were the ones to say actually, you know, here's a new idea, something new, not in the domains of human knowledge of the game. And now the humans think this is a reasonable thing to do and it's part of Go Knowledge now. The third game, something special happens when you play against a human world champion, which again I hadn't anticipated before going there, which is...

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  50. The second game became famous for a move known as Move 37. This was a move that was played by Alphago that broke all of the conventions of Go, that the Go players were so shocked by this. They thought that maybe the operator had made a mistake. They thought there was something crazy going on. And it just broke every rule that Go players are taught from a very young age. They're just taught this kind of move called a shoulder hit. You can only play it on the third line or the fourth line. And AlphaGo played it on the fifth line. And it turned out to be a brilliant move and made this beautiful pattern in the middle of the board that ended up winning the game. And so this really was a clear instance where we could say computers exhibited creativity, that this was really a move that was something humans hadn't known about.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source