YouSaid · the spoken record

David Silver

lines on the record
114
first
2020-04-03
most recent
2020-04-03
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. So I remember all of these games vividly, of course, you know, moments like these don't come too often in the lifetime of a scientist. The first game was magical because it was the first time that a computer program had defeated a world champion in this grand challenge of Go. And there was a moment where Where AlphaGo invaded Lisa Dole's territory towards the end of the game. And that's quite an audacious thing to do. It's like saying, hey, you thought this was going to be your territory in the game, but I'm going to stick a stone right in the middle of it and prove to you that I can break it up. And Lisa Dole's face just dropped. He wasn't expecting a computer to do something that audacious.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Funnily enough, there was So just before the match, we weren't betting on anything concrete, but we all held out a hand. Everyone in the team held out a hand at the beginning of the match. And the number of fingers that they had out on that hand was supposed to represent how many games they thought we would win against Lisa Dahl. And there was an amazing spread in the team's predictions. Have to say, I predicted 4-1.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Also, lucky that, you know, in a sense, I was insulated from the knowledge of, I think it would have been harder to focus on the research if the full kind of reality of what was going to come to pass had been known to me and the team. I think it was, you know, we were in our bubble and we were working on research and we were trying to answer the scientific questions. And then bam, you know, the... The public sees it, and I think it was better that way in retrospect.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  4. An adjustment process to realize that this Was something which the world really cared about and which was a watershed moment. And I think there was that moment of realization. It's also a little bit scary because if you go into something thinking it's going to be maybe of interest and then discover that 100 million people are watching, it suddenly makes you worry about whether some of the decisions you'd made were really the best ones or the wisest. We're going to lead to the best outcome, and we knew for sure that there were still imperfections in AlphaGo. Which we're going to be exposed to the whole world watching. And so, yeah, it was, I think, a great experience. And I feel privileged to have been part of it, privileged to have led that amazing team. I feel privileged to have been in a moment of history, like you say

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Well, thank you again. I mean It's funny because I guess I've been working on Computer Go for a long time, so I've been working at the time of the AlphaGo match on Computer Go for more than a decade. And throughout that decade, Had this dream of what would it be like to actually be able to build a system that could play against the world champion? And I imagine that that would be an interesting moment that maybe some people might care about that and that this might be a nice achievement. But I think when I arrived in Seoul and discovered the legions of journalists that were following us around and the 100 million people that were watching the match online, Life, I realized that I had been off in my estimation of how significant this moment was by several orders of magnitude. And so there was definitely...

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  6. That it was something which was extremely useful to us. It helped us to understand the system, it helped us to build deep learning representations which were clear and simple and easy to use. And so really I would say it served a purpose not just as part of the algorithm, but something which I continue to use in our research today, which is trying to break down a very hard challenge into pieces which are easier to understand for us as researchers and develop. So if you use a component based on human data, it can help you to understand the system such that then you can build the more principled version later that does it for itself.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  7. And so the reason that we did it that way was at that time we were exploring separately the deep learning aspect from the reinforcement learning aspect. That was the part which was new and unknown to me at that time was how far could that be stretched. Once we had that, it then became natural to try and use that same representation and see if we could learn for ourselves using that same representation. And so right from the beginning actually our goal had been to build a system using self-play. And to us, the human data right from the beginning was an expedient step to help us for pragmatic reasons to go faster towards the goals of the project than we might be able to starting solely from self-play. And that might not be the long holy grail of AI.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  8. In the early days of AlphaGo, we used human data to explore the science of what deep learning can achieve. And so when we had our first paper that showed that it was possible to predict the winner of the game, that it was possible to suggest moves, that was done using human data.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Where we had a clear goal, which was to try and crack this outstanding challenge of AI to see if we could beat the world's best players. And this led within the space of not so many months to playing against the European champion Fan Hui in a match which became memorable in history as the first time a Go program had ever beaten a professional player. And at that time, we had to make a judgment as to whether when and whether we should go and challenge the world champion. And this was a difficult decision to make, again, we were basing our predictions on our own progress and had to estimate based on the rapidity of our own progress when we thought we would exceed the level of the human world champion. And we tried to make an estimate and set up a match. And that became the govert. Is Lisa Dahl match in 2016.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So we scaled up. This was something where I had lots of conversations back then with Demis Asabes at the head of DeepMind, who was extremely excited. And we made the decision to scale up the project, brought more people on board. And so AlphaGo became something

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  11. I tend to be an optimist with the power of deep learning and reinforcement learning. So the system won, and we were able to beat this human down level player. And for me, that was the moment where it was like, okay, something special is afoot here. We have a system which without search is able to already just look at this position and understand things as well as a strong human player. And from that point onwards, I really felt that reaching the top levels of human play, you know, professional level, world champion level, I felt it was actually an inevitability. And if it was inevitable outcome, I was rather keen that it would be us that achieved it.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  12. I found this to be profoundly surprising. In fact, it was so surprising that we had a bet back then. And like many good projects, bets are quite motivating. And the bet was whether it was possible for a system based purely on deep learning, no search at all to beat a DAN level human player. And so we had someone who joined our team who was a DAN level player. He came in and we had this first match against him.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  13. And so that was our scientific question, which we were probing and trying to understand. And as we started to look at it, we discovered that we could build a system. So in fact, our very first paper on AlphaGo was actually a pure deep learning system which was trying to answer this question. And we showed that actually a pure deep learning system with no search at all was actually able to reach human level, master level, at the full game of Go, 19 by 19 boards. And so without any search at all, suddenly we had systems which were playing at the level of the best Monte Carlo tree step systems, the ones with randomized rollouts.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Is there another fundamentally different approach to this key question of Go, the key challenge of how can you build that intuition in? How can you just have a system that could look at a position and understand what move to play or how well you're doing in that position, who's going to win? And so the deep learning revolution had just begun. There are systems like ImageNet had suddenly been won by deep learning techniques back in 2012. And following that, it was natural to ask, well, you know, if deep learning is able to scale up so effectively with images to understand them enough to classify them, well, why not go? Why not take the black and white stones of the Go board and build a system which can understand for itself what that means in terms of what move to pick or who's going to win the game, black or white?

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Yeah, so programs based on Monte Carlo research were a first revolution in the sense that they led to suddenly programs that could play the game to any reasonable level, but they plateaued. It seemed that no matter how much effort people put into these techniques, they couldn't exceed the level of amateur DAN level go players. So strong players, but not anywhere near the level of professionals, never mind the world champion. And so that brings us to the birth of AlphaGo, which happened in the context of a startup company known as DeepMind. I heard of them. Project was born and the project was really a scientific investigation where myself and Ajah Huang and an intern, Chris Madison, were exploring a scientific question and that scientific question was really what

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  16. And so the intuition is that randomization captures something about the nature of the search tree from a position that you're understanding the nature of the search tree from that node onwards by using randomization. And this was a very powerful idea.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Worked with him a little bit in those days out of my PhD thesis, and Mogo was a first step towards the later successes we saw in Computer Go. But it was still missing a key ingredient. Mogo was evaluating purely by random rollouts against itself. And in a way, it's truly remarkable that random play should give you anything at all. Why in this perfectly deterministic game that's very precise and involves these very exact sequences? Why is it that randomization is helpful?

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Search and a particular form of Monte Carlo search that became very effective and was developed in Computer Go first by Remy Coulomb in 2006 and then taken further by others was something called Monte Carlo Tree Search, which basically takes that same idea and uses that insight to evaluate every node of a search tree is evaluated by the average of the random playouts from that node onwards. And this idea was very powerful and suddenly led to huge leaps forward in the strength of computer go playing programs. And among those the strongest of the GoPlaying programs in those days was a program called MOGO. to actually reach human master level on small boards, nine by nine boards. And so this was a program by someone called Sylvain Gelli, who's a good colleague of mine, but I

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Position could be evaluated or a state in general could be evaluated not by humans saying whether that position is good or not or even humans providing rules as to how you might evaluate it. But instead by allowing the system to randomly play out the game until the end multiple times and taking the average of those outcomes as the prediction of what will happen. So for example if you're in the game of Go, the intuition is that you take a position and you get the system to kind of play random moves against itself all the way to the end of the game and you see who wins. And if black ends up winning more of those random games than white, well you say, hey, this is a position that favors white. And if white ends up winning more of those random games than black, then it favors white. was known as Monte Carlo.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Back during my PhD on ComputerGo around about that time, there was a major new development which actually happened in the context of Computer Go. It was really a revolution in the way that heuristic search was done. And the idea was essentially that

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  21. And sometimes that requires putting together more complex systems where we don't have the full answers yet as to what those minimal ingredients might be.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  22. 100% agree. I think that If we try to anticipate what will generalize well into the future, I think it's likely to be the case that it's the simple, clear ideas which will have the longest legs and which will carry us furthest into the future. Nevertheless, we're in a situation where we need to make things work today.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  23. I often wonder when we one day do have AIs which are superhuman in their abilities to understand the world. Will they think of the algorithms that we developed back now? Will it be looking back at these days and thinking that will we look back and feel that these algorithms were naive first steps or will they still be the fundamental ideas which are used even in 100,000, 10,000 years?

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Essentially, the AI winter where people gave up on neural networks. I think it's really down to that lack of ability to generalize from low dimensions to high dimensions because back then we were in the low dimensional case. People could only build neural nets with 50 nodes in them or something. And to imagine that it might be possible to build a billion dimensional neural net and it might have a completely different qualitatively different property was very hard to anticipate. And I think even now we're starting to build the theory to support that. And it's incomplete at the moment, but all of the theory seems to be pointing in the direction that indeed this is an approach which truly is universal both in its representational capacity, which was known, but also in its learning ability, which is surprising.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  25. We're very tuned to working within a three-dimensional environment. To start to visualize what a billion dimensional neural network surface that you're trying to optimize over, what that even looks like is very hard for us. And so I think that really, if you try to account for the

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Perform so well despite the fact that these neural networks that they're representing have these incredibly nonlinear kind of bumpy surfaces, which to our kind of low dimensional intuitions Make it feel like surely you're just going to get stuck and learning will get stuck because you won't be able to make any further progress. And yet, the big surprise is that learning continues and these what appear to be local optima turn out not to be because in high dimensions when we make really big neural nets, there's always a way out and there's a way to go even lower and then still not in a local optima because there's some other pathway that will take you out and take you lower still. And so no matter where you are, learning can proceed and do better and better and better without bound. And so that is a surprising and beautiful property of neural nets, which I find elegant and beautiful and somewhat shocking that it turns out to be the case.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Think, let me Two parts to that question. I think It's not surprising to me that the idea of reinforcement learning works because in some sense. Think it's, I feel it's the only thing which can ultimately. And so I feel we have to address it. And there must be successes possible because we have examples of intelligence. And it must at some level be able to possible to acquire experience and use that experience to do better in a way which is meaningful to environments of the complexity that humans can deal with. It must be. I surprised that our current systems can do as well as they can do. One of the big surprises for me and a lot of the community. Really, the fact that deep learning can continue to

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  28. As we start to put more resources into the system, more memory and more computation, and more data, more experience, more interactions with the environment, that these are systems that can just get better and better and better at doing whatever the job is they've asked them to do, whatever we've asked that function to represent, it can learn a function that does a better and better job of representing that knowledge, whether that knowledge be estimating how well you're going to do in the world, the value function, whether it's going to be choosing what to do in the world, the policy, whether it's understanding the world itself, what's going to happen next, the model.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Amongst all the approaches for reinforcement learning, deep reinforcement learning is one family of solution methods that tries to utilize powerful representations that are offered by neural networks to represent any of these different components of the solution, of the agent, like whether it's the value function or the model or the policy. The idea of deep learning is to say, well, here's a powerful toolkit that's so powerful that it's universal in the sense that it can represent any function and it can learn any function. And so if we can leverage that universality, that means that whatever we need to represent for our policy or for our value function for our model, deep learning can do it. So deep learning is one approach that offers us a toolkit that has no ceiling to its performance.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Like a value function or a model, are you learning something to perform well, like a policy, and the form of that objective is kind of giving the semantics to the system. And so it really is at the next level down a fundamental choice. And we have to make those fundamental choices as system designers or enable our algorithms to be able to learn how to make those choices for themselves.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Wouldn't do a very good job of it. So, learning is required because it's the only way to achieve good performance in any sufficiently large and complex environment. So that's the first step. And so that step gives commonality to all of the other pieces. Because now you might ask, well, what should you be learning? What does learning even mean? In this sense, learning might mean, well, you're trying to update the parameters of some system which is then the thing that actually picks the actions. And those parameters could be representing anything. They could be parameterizing a value function or a model or a policy. And so in that sense, there's a lot of commonality in that whatever is being represented there is the thing which is being learned and it's being learned with the ultimate goal of maximizing rewards. But the way in which you decompose the problem is really what gives the semantics to the whole system. Are you trying to learn something to predict well?

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  32. I think the fundamental idea is maybe at the higher level, the fundamental idea is the first step of the decomposition is really to say, well, how are we really going to solve any kind of problem where you're trying to figure out how to take actions and just from this stream of observations, you've got some agent situated in its sensory motor stream and getting all these observations in, getting to take these actions? And what should it do? How can you even broach that problem? Maybe the complexity of the world is so great that you can't even imagine how to build a system that would understand how to deal with that. And so the first step of this decomposition is to say, well, you have to learn. The system has to learn for itself. And so note that the reinforcement learning problem doesn't actually stipulate that you have to learn. You could maximize your rewards without learning.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  33. A representation of a policy that means something which is deciding how to pick actions, is that decision making process explicitly represented? And is there a model in the system? Is there something which is explicitly trying to predict what will happen in the environment? And so those three pieces are, to me, some of the most common building blocks. And I understand the different choices in RL as choices of whether or not to use those building blocks when you're trying to decompose the solution. Should I have a value function represented? Should I have a policy represented? Should I have a model represented? And there are combinations of those pieces and of course other things that you could add into the picture as well. But those three fundamental choices give rise to some of the branches of RL with which we're very familiar.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yeah. So now if we think about, okay, so there's this ambitious problem definition of RL. It's truly ambitious. It's trying to capture and encircle all of the things in which an agent interacts with an environment and say, well, how can we formalize and understand what it means to crack that? Now let's think about the solution method. Well, how do you solve a really hard problem like that? Well, one approach you can take is to decompose that very hard problem into pieces that work together to solve that hard problem. And so you can kind of look at the decomposition that's inside the agent's head, if you like, and ask, well, what form does that decomposition take? And some of the most common pieces that people use when they're kind of putting this solution method together, some of the most common pieces that people use are whether or not that solution has a value function. That means, is it trying to predict explicitly trying to predict how much reward it will get in the future? Does it have a

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Reinforcement learning is the study and the science and the problem of intelligence in the form of an agent that interacts with an environment. So the problem you're trying to solve is represented by some environment, like the world in which that agent is situated. And the goal of RL is clear, that the agent gets to take actions. Those actions have some effect on the environment. And the environment gives back an observation to the agent saying, you know, this is what you see or sense. And one special thing which it gives back is called the reward signal. How well it's doing in the environment. And the reinforcement learning problem is to simply take actions over time so as to maximize that reward signal.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Methods. I would say that if we have such a thing, it would be a solution to the RL problem. Now, what particular methods have been used to get there? Well, we should keep an open mind about the best approaches to actually solve any problem. And the things we have right now for reinforcement learning, maybe I believe they've got a lot of legs, but maybe we're missing some things. Maybe there's going to be better ideas. I think we should keep, you know, let's remain modest and we're at the early days of this field and there are many amazing discoveries ahead of us.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Let me say it this way. I think it's helpful to separate out the problem from the solution. So I see the problem of intelligence. I would say it can be formalized as the reinforcement learning problem. And that formalization is enough to capture most, if not all of the things that we mean by intelligence, that they can all be brought within this framework and gives us a way to access them in a meaningful way that allows us as scientists to understand intelligence and us as computer scientists to build them. And so in that sense, I feel that it gives us a path, maybe not the only path, but a path towards AI. And so do I think that any system in the future that's solved AI would have to have RL within it? Well, I think if you ask that, you're asking about the solution.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Really, an opportunity to frame what intelligence means, like what are the goals of AI in a clear, single, clear problem definition such that if we're able to solve that clear, single problem definition, in some sense we've cracked the problem of AI.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Asked him if he would be interested in supervising me on a PhD thesis in ComputerGo. And he basically said that if he's still alive, he'd be happy to. But unfortunately, he'd been struggling with very serious cancer for some years. And he really wasn't confident at that stage that he'd even be around to see the end event. But fortunately, that part of the story worked out very happily. I found myself out there in Alberta. They've got a great games group out there with a history of fantastic work in board games as well, as Rich sat in the father of RL, so it was the natural place for me to go in some sense to study this question. And the more I looked into it, the more strongly I felt that this wasn't just the path to progress in computer go. Really, you know, this was the thing I'd been looking for. This was

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  40. So I had just finished working in the games industry at this startup company, and I took a year out to discover for myself exactly which path I wanted to take. I knew I wanted to study intelligence, but I wasn't sure what that meant at that stage. I really didn't feel I had the tools to decide on exactly which path I wanted to follow. So during that year, I read a lot. And one of the things I read was Saturn and Barto, the sort of seminal textbook on an introduction to reinforcement learning. And when I read that textbook, I just had this resonating feeling that this is what I understood intelligence to be. And this was the path that I felt would be necessary to go down to make progress in AI. So I got in touch with Rich Saturn and.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Actually, captures a strong level of play happen very rarely, which means that at any moment in the game, you've got the same number of whitestones and black stones. And the only thing which differentiates how well you're doing is this intuitive sense of where are the territories ultimately going to form on this board. And if you look at the complexity of a real go position, it's mind-boggling that kind of question of what will happen in 300 moves from now when you see just a scattering of 20 white and black stones intermingled. And so that challenge is the reason why position evaluation is so hard in Go compared to other games. In addition to that, has an enormous search space. So there's around 10 to the 170 positions in the game of Go. That's an astronomical number. And that search space is so great that traditional heuristic search methods that were so successful.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Yeah, so in the game of Go, you know, you find yourself in a situation where both players have played the same number of stones.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Where it's an important part of the culture, so much so that it's considered one of the four ancient arts that was required by Chinese scholars. So there's a deep history there.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  44. And the only nuance to the game is that if you fully surround your opponent's piece, then you get to capture it and remove it from the board and it counts as your own territory. Now from those very simple rules, immense complexity arises. There's kind of profound strategies in how to surround territory, how to kind of trade off between making solid territory yourself now compared to building up influence that will help you acquire territory later in the game, how to connect groups together, how to keep your own groups alive, which patterns of stones are most useful compared to others. There's just immense knowledge. And human go players have played this game for, it was discovered thousands of years ago, and human go players have built up this immense knowledge base over the years. It's studied very deeply and played by something like 50 million players across the world, mostly in China, Japan, and Korea.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  45. So, the Game of Go has remarkably simple rules. In fact, so simple that people have speculated that if we were to meet alien life at some point, that we wouldn't be able to communicate with them, but we would be able to play a game. They probably have discovered the same rule set. So the game is played on a 19 by 19 grid, and you play on the intersections of the grid, and the players take turns. And the aim of the game is very simple. It's to surround as much territory as you can, as many of these intersections with your stone.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Although not all of the pieces were handcrafted, the overall effect was nevertheless still brittle and it was hard to make all these pieces work well together. And so really what I was pressing for and the main innovation of the approach I took was to go back to first principles and say, well, let's back off that and try and find a principled approach where the system can learn for itself just from the outcome. Learn for itself. If you try something, did that help or did it not help? And only through that procedure can you arrive at knowledge which is verified. The system has to verify it for itself, not relying on any other third party to say this is right or this is wrong. And so that principle was already very important in those days that unfortunately we were missing some important pieces back then.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  47. In a sense, that was the state of the art back then. So if you look at the Go programs which had been competing for this prize I mentioned, they were an assembly of different specialized systems, some of which used huge amounts of human knowledge to describe how you should play the opening, how you should, all the different patterns that were required to play well in the game of Go endgame theory combinatorial game theory, and combined with more principled search-based methods which we're trying to solve for particular subparts of the game, like life and death, connecting groups together, all these amazing sub-problems that just emerge in the game of Go, there were different pieces all put together into this collage, which together would try and play against a human.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  48. The only way you're going to be able to get a system which has sufficient knowledge in it, millions and millions of pieces of knowledge, billions, trillions, of a form that it can actually apply for itself and understand how those billions and trillions of pieces of knowledge can be leveraged in a way which will actually lead it towards its goal without conflict or other issues.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So I strongly felt that learning would be necessary, and that's why my PhD topic back then was trying to apply reinforcement learning to the Game of Go, and not just learning of any type, but I felt that the only way to really have a system to progress beyond human levels of performance wouldn't just be to mimic how humans do it, but to understand for themselves and how else can a machine hope to understand what's going on except through learning. If you're not learning, what else are you doing? Well, you're putting all the knowledge into the system. And that just feels like something which decades of AI have told us is maybe not a dead end, but certainly has a ceiling to the capabilities. It's known as the knowledge acquisition bottleneck. The more you try to put into something, the more brittle the system becomes. And so you just have to have learning. You have to have learning.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Many, many more problems in AI. So for me, that was the moment where, like, okay, this is not just about playing the game of Go, this is about something profound. And it was back to that bug which had been itching me all those years. This is the opportunity to do something meaningful and transformative. And I guess a dream was born.

    2020-04-03 · Lex Fridman Podcast · #86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning · IDENTIFIED FROM THE TRANSCRIPT · source