YouSaid · the spoken record

Eric Jang

lines on the record
188
first
2026-05-15
most recent
2026-05-15
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. But somehow, because maybe we're playing on a lot of games where the bots just resign instead of playing all the way to the Trump Taylor resolution, they kind of forget how to evaluate those kind of late stage plants. Like in the case that we showed with the corner play, maybe 100% of our trading data in our replay buffer has lost examples of how to evaluate the value function at those states. So you might end up in a scenario where your terminal value is very bad. And if the terminal values of the leaves are not good, then this will actually propagate all the way up and cause your pucked selection criteria and your backups to be off. And then you end up visiting a very, very different distribution than what your policy initially recommended. Also, if your number of sims is low, then you might also have a variance issue where you just don't explore enough, right? It's only guaranteed to converge.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  2. In practice, it is a heuristic, and it does work also in practice, but let me illustrate an example where MCTS can give you a worse distribution than your policy network. And this can often happen if your self-play algorithm has trained to a good point, but then somehow it collapses because it's not trained on diverse data or something. So let's say we have a board state where the policy recommendations here are very good. So pi of AS is like great.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  3. And this is very related to Dagger in robotics and imitation learning, where you want to collect an intervention here. And even if you're in a not great state, for example, like a self-driving car that veers off the side of the road, there is still a valid action that kind of corrects you and brings you back

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  4. So there is a family of algorithms that basically take trajectories and relabel the actions to better trajectories. So maybe a better action here would have been to take A0 prime. A better action here would have been to take A1 prime, and yet another one like A2 prime So What MCTS is doing is basically saying, like, you play this game where you eventually lost, but on every single action, I'm going to give you a strictly better action that you should take instead. It does not guarantee that you are going to win, but it does guarantee that if you take these tuples as training data so that you retrain your policy network to predict these ones instead of these ones, you're going to do better.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  5. And so instead of starting here, we get to now have a neural network start here, and then the play gets stronger once we then apply another thousand steps on top of it. And you can keep going, right? So the training algorithm for AlphaGo is to basically take the games where you've applied the search on every move that the policy encountered, whether you won or lost. And that's quite important. And you're just going to train the model to imitate the search process. So there's an analogy to robotics, actually, which is the dagger algorithm. First, I'm going to draw like a schematic of like, let's say, you know, the states, right? So S0, S1, S2, S3. So let's say we took a series of actions in an MDP to Get a trajectory And these actions may be suboptimal, right? Maybe we lost at the end of this game

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  6. To be honest, I actually don't know the test time scaling behavior of MCTS simulations, and I believe it might actually be quite sensitive to how strong this one is in practice. I'm just drawing a monotonically increasing function that gets to one. So don't pay too much attention to the shape of the curve. Just know that it's monotonic with respect to symmetry. Okay, so the idea of MCTS is very brilliant, which is like we're gonna, we got something better by applying search. And we're going to now on our next iteration of updating this network, just train this to approximate the outcome of a thousand steps of search.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  7. That gets you to a policy here that gets you to here, which is great. But if you were to distill this MCTS policy network back into your sort of shoot from the hip policy network, then you could actually start here. Let's say this was zero by distillation, then if you spend another 1000 sim steps, then you actually kind of get to here. It's almost like if you could just amortize the first 1,000 steps actually into the policy network instead of the search process, then you could begin at a much better starting point and then get a much better result for the number of sims that you put.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Let's say at zero simulations, you're sort of implicit win rate is like, I don't know, here without any simulation, if you just take this raw action, this is what your win rate is. And let's say as we increase the number of sims, maybe you kind of have a win rate that looks like this. When you search for, let's say, a thousand simulation steps.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  9. And then so on and so forth. And then the game ends. And one person wins, and one person loses. The beauty of how AlphaGo trains itself is that it actually can take this final search process, the outcome of the search process, and tell the policy network, hey, instead of having MCTS do all this legwork to arrive here, why don't you just predict that from the get-go? Why don't you not use this guess and just predict this to begin with? And if you have this guess to begin with in your policy network, then MCTS has to do a lot less work to get things to work. And so if we draw like a sort of test time scaling plot, so let's say this is like number of simulations.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  10. I'm sorry, that's correct, yes. To something that looks like So on every move, you have your initial guess from your policy network. And then the search process that combines your policy network and your value network arrives at a more confident action that you take.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Of a given s, right? So after applying MCTS process, your policy recommended distribution looks like this. It's a bit more peaky than the previous one. And so then you take the arcmax, or maybe you just sample from this. It doesn't have to be the arcmax. And then you make your move. And then you throw away the tree, and then you begin anew on the next move, right? Again, you compute a new distribution. So initially, maybe your guess looks like this, and then you refine it through MCTS

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  12. It gets more confident about one of these actions. And so maybe the distribution looks a bit more peaky like this based on the search. Now, of course, you can tune the search process so that it ends up more diffuse, but that's probably not a good idea. MCTS should get more confident about specific actions than others. But it, of course, might place a lot of weight on other actions initially. And then as you increase this number of sims, it should converge to a very peaky distribution. So this is your new, let's call this like pi. Let's wrap this in like a MCTS operator.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Okay, so we now talk about the RL part of how this thing gets stronger by playing itself. Let's say we play a game where the AI, so you make a move. AI will kind of compute the search, and then this is this sort of visit count distribution. Let's say this is your policy, your policy, initial policy recommendation at this node. And then after MCTS,

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Okay, so now that we have the search algorithm that applies the value function as well as the policy function, we can now talk about how the Monte Carlo tree search algorithm can actually act as an improvement operator on top of these guys here.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  15. I'm drawing a very low dimensional version of this. Of course, in the real game, it's much more high dimensional. But you'll end up with basically a tree structure that has A lot of leaves that kind of terminate and are not visited again because their value is deemed to be too low. But then along one path, there will be a set of actions with very, very high visit counts that kind of gravitate towards that one set of decisions as you increase n. So this is kind of like the mental picture of what the tree in Monte Carlo research looks like. And you should contrast this with an exhaustive tree like in tic-tac-toe where you could say there's nine actions and then eight and then seven and six. And so it's sort of like nine factorial sized tree. The Monte Carlo tree search in Go is very, very sparse. It only considers the paths that you've expanded children nodes on.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  16. You can usually batch these somewhat efficiently. So it probably is not a huge computational burden in practice. But yes, you would have to pass 361, like up to 361 boards into a single mini-batch update to evaluate all the values here, then normalize them. Now, there's actually a more important reason why we still do this, which is how Monte Carlo tree search is used to feed back on itself. And sort of recursively improve its own predictions and search capabilities. And that's where this one, having this as an explicit entity you're modeling rather than an implicit normalization over your value is a good idea. Okay, so we talked about the simulations and basically what you end up with as you roll out the number of simulations is a tree that kind of looks like

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  17. There is a sort of duality here. It would be weird if, let's say, the policy recommended an action that disagreed with the value, right? If let's say the policy said this was very high probability, but this one said it was low value, then there's actually something kind of fundamentally wrong between your policy head and your value head. So they are linked, and you probably could get rid of this if you came up with a different way to recover this from just the value evaluations.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  18. As induced through this search process. And so the visit count that we store in the node earlier actually becomes the sort of vote for which way we should finally select an action here. So, as a sort of test of understanding, it's worth thinking a little bit about whether we could make this even simpler, right? Could we actually maybe even get rid of this one and still make the thing work? So recall that when you do an expansion and then an evaluation at, let's say, this node, you are checking the sort of win probability of each of the child nodes, right? And so this one is one and these are zero, you do kind of know something about which action might be better to take. And so why would you still need this, right? Like why not just normalize this one into some distribution and call that your policy distribution? This is fine. You can do this and this probably does work, but in practice, having a single forward pass that gives you a pretty good guess is how the breadth is pruned out.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  19. So that's what is known as the backup step. And once you evaluate this, you can actually kind of recursively go back. So if you know the action value of this node, you can then take the average on its parent and so on and so forth. So you have this kind of four-step process where you are choosing the best action that you know of so far, then you may run into a node where you haven't been to before. So you need to grow the tree a bit. And then you run it through the network to guess whether you're going to win or not. And then you walk all the way back up to the root node to update your values on what the best moves are. So as you do this iteratively, this selection criteria will cause you to visit the, because you're always selecting according to this criteria, you're always going to be selecting the best action you think at any given branch. So the final visit counts of like how often you chose these things will reflect your correct policy distribution.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  20. So let me just finish up with actually the last step, which is the backup. So once you've scored these things, you basically take the mean, the value, the Q value assigned to the node here for taking this action is now just the average across your evaluated values here. You take a running mean over all of the simulations that you've taken, and they average the values of the children

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Almost. So a simulation is easy to think about when the whole tree already exists. You just walk down the tree using the puck selection criteria and then you keep going. Now in AlphaGo, the data structure is such that we begin with a tree that has basically only depth one, which is its only children, and you want to iteratively build out the tree as you're also selecting actions down the tree. So that's the kind of core thing here is that because Go is such a combinatorily complex game, you cannot afford to build the tree in advance and then search it. You must search while building the tree.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  22. And you might be wondering, okay, well, how do they play this out? It would be very, very costly to do another search on this playout, like almost like a tree within a tree. So they don't do this. Instead, they just take the policy network and play it against itself. So they just take this as both players and they just play it all the way to the end. And this is something that helps ground the estimates here in reality because you can get a single sample estimate of whether you win or not. You can think about in the end game where the board is almost resolved that this one actually becomes quite useful because the random, the play according to the policy will most likely decide a pretty reasonable guess of the game. And so you're not facing a problem where this one kind of becomes untethered from reality. It turns out this is totally unnecessary. So in all subsequent papers after AlphaGo Lee, they just got rid of this. And so in my implementation, I also did the same and it speeds things up a lot because you don't have to roll these games out on every single simulation.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Of a full board. And so this is like a zero or one, right? And so they took this value and they just averaged it with this one here. So the formula they did was like alpha times v theta of some node plus sort of like 1 minus alpha of a true randomly sampled plant.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  24. And so this is the expansion step. You've taken a non leaf node and expanded it and evaluated the value. And this is essentially a quick guess as to if I were to play to the end, am I going to win or not? So you can almost think about the v theta as a shortcut for searching to the end of the tree for any given simulation. And then this is essentially the evaluation step. We're evaluating the quality of each of these boards. In original Alpha Go Lee, they actually did something kind of interesting, which is that they took this value and they averaged it with the value of a real go play out. So they actually played a real game from here all the way to the end. So I'm just going to draw this squiggly line to indicate some path. And they kind of like play this all the way to Trump Taylor resolution.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So then this one has possible actions that we could take. And we expand basically the leaf nodes here. So for each of these nodes that we could arrive at, we're going to now check how good those nodes are. So maybe... From here, like the human could play here, the human could play here, or human could play here. And we're going to store essentially the v theta for each of these things. So v theta of node one. Or like node one prime, the theta node one. And so we're basically using our neural network to make an intuitive guess of how good is this board from the perspective of this player. And fortunately, because it's a zero-sum game, it's easy to deduce that the value for this player at this step is just 1 minus the value from this perspective. So it's easy to flip the search process depending on which player you're at. Um,

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Let's suppose P1 was the highest probability node. So you selected this one here. Now you got to this node and you realize that it's not a leaf node, right? It's not a terminal game, so you cannot resolve the final resolution. So the next step that you do is expansion. So you will then run this node, this board state, through the policy network. Note that this is the AI's move, right? AI is making this move. And so when we expand this tree, we're now thinking about what the human might do, or any opponent might do, right? So this is like your opponent. The tree expansion process actually is completely. So when we evaluate the node here, we're going to now evaluate the node from the perspective of this player.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  27. So for each of these, we're going to, you know, NA is zero for all the actions initially. N is zero. And so we're going to basically just pick according to this. Initially What is going to be the chosen action here is most likely going to be biased towards the highest likelihood action here, because these are sort of uniform for every bound.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  28. And each of these has an associated probability of taking that action. So there's P8, P1, P2, et cetera. Okay, so at the beginning of our Monte Carlo tree search, we have our root node, and we can initialize it with some children, right? Because we know the policy network evaluated on the root node gives us on a 3x3 board with one existing stone placed eight possible children that this AI could take. So with each of the children, our policy network also gives us the probability of selecting that child. So the first step is to do the selection of the tree. And again, this is a very shallow tree. All we have so far is a tree of depth one, essentially, right? So our first move is to select by maximizing or arg maxing the pucked criteria, which is basically C pucked times P of A divided by n over 1 plus na.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  29. So at the beginning of our Monte Carlo tree search, our tree is very basic. It only has the root node, our current board that our AI wants to play at. And so we're going to basically select the best action for this. When this root node is created, we also know that we can evaluate this under our neural network and get the quantities v theta, as well as our probability over actions. And I'm going to say So, for all of the actions here, we can create a bunch of children, right? So, this one has, well, in this case, I'm drawing a 3x3 board with one board missing. So basically there are eight possible children associated with this root node.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  30. So, we're going to pick this kind of numulations thing. And for every simulation, we're going to basically do several things simultaneously. We're going to see which moves are the best in the current tree. We're going to add extra leaves to the tree if we get to a point where we need to add a leaf. And we're going to update the action values for the tree. So that's what every simulation involves these kind of like four-step process. So the four-step process is basically selection. Expansion Evaluation

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Fair enough. Interestingly enough, modern Go bots don't need that much compute at test time. And what we'll actually find out as we talk about how the MCTS policy improvement works is that over time the raw network actually takes all of the burden of that big TPU pod and just pushes it into the network. And you can do all of that work with one neural network forecast. But the TPU pod will always add the extra OOMF on top. And so that's what they wanted for the match.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  32. And we now have a four-step iterative process to do MCTS. So this tripped me up when I was first reading the paper and trying to understand it, but essentially what we're going to do is we're going to choose a number of simulations, so like num. Simulations And this number varies. This can be somewhere between 200 to 2048. I believe in the AlphaGo Lee match, they use tens of thousands of simulations per move because they really wanted to boost the strength of the model as much as possible But in training, you don't actually need too many. And Katago, I think, uses something on this order as well

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And so you can start this way. And it's important when implementing this to kind of just verify that this is probably true. It's good to verify that your Go rules are implemented correctly, that you can run these simulations relatively quickly. And just as almost like a sort of a checkpoint that you want to make sure that you can actually do this basic step before you try to layer on more complex things like search. But yeah, we can do a lot better than taking the raw neural network and playing the moves. And this is how we can apply it to Monte Carlo Research. So let's apply the neural network to improve Monte Carlo tree search. So we start with our root n

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So, if you just do this, imagine you're vibe coding alpha go and you gather some expert data sets from LikataGo online, or you have a data set of human players and you train this model, actually it turns out this model is already a pretty good Go player. It'll most likely beat most human players, right? So if you just take this policy recommendation and take the arg max over its... If this is the probabilities, if you take the arg max and you just take this action as your Go play, it'll be a very, very fast go player that doesn't think in terms of reasoning steps. It just kind of shoots from the hip, and it'll be a very strong go player. Which is already quite miraculous if you think about 10 neural network layers, maybe under like 3 million parameters can already do something that impressive.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  35. It is not relevant to the expert data. It's true for any data that you traded on. Yeah, so if you were to learn Tableau Raza, you would also expect this to fall out.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  36. And so, as you get hundreds of steps into the game, it becomes much more clear who is more likely to win or who's more likely to lose under your expert data distribution.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  37. So, you might be wondering, okay, well, some of the early boards where basically only one stone has been put down, how could you possibly know who the winner of this game is? Well, if you have hundreds of thousands of games, then on average you'll probably see that boards that start like this have a sort of half of the games that branch off from this will win and half of the games that branch off from this will lose. So that'll actually be fine. When you train this model to predict those, the logit will sort of converge to 0.5. And so for these things, it's sort of expected that once you train the model, a starting board state will look like 0.5. And then as you progress towards the end of the game, it'll actually look something like if this is 0.5, the win probability will sort of either go like this or it'll go like this. And this is sort of your move number.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Initialize your research project to something as close to success as possible, especially if you're doing something new that you haven't done before. Like always pick something that works and then get it to do something better rather than start from something that doesn't work at all and then try to make it work. So under that philosophy, it's a great idea to start from something that has a good initialization. So we're going to take human expert place and train this model to predict good actions. So we're going to take all of the winning games, all the moves in which a human won and sorry, an expert won and then predict those actions. And then regardless of board state, like whether you won or lost, you're going to predict the outcome.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  39. So, the OG AlphaGo paper, or called Alphago Lee, initialized this network with a supervised learning data set of expert human play. Later, they removed this restriction by having the model teach itself how to play well, but I find it actually from a matter of implementation for your audience. Super, super nice to always kind of initialize your experiments to something that's easy and then get the problem working before trying to bite off the whole thing and learn a tabular resin. You generally want to kind of initialize, just as in deep learning, initializations everything, right? You always want to...

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Empties or maybe like a masked region if you want to train on multiple board sizes. I'm actually not going to talk about multiple board sizes for now. That's a little bit too complicated. So we'll just say we've got this two or three channel RGB-like image. And then we go into a resnet. And then we have two branching heads. One head predicts the value function, and this is like a single logit. So it's like an And then we have the policy, which is R361. So This is the architecture, and we're going to basically train this to predict the outcomes of games given the board state. And we're also going to train this to predict what are good moves.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  41. And I think that's a place where things get very, very exciting in terms of self play or diplomacy style. Yeah, interesting. Okay, so returning back to the neural network, the architecture, again, is not super important. You can get it to work with transformers, you can get it to work with ResNets. I found that for low budget experiments, resonets work a little better. You can also use kind of a Carpathy-style auto research hyperparameter tuning to make your architecture pretty good. And so you don't have to worry too much about that. You just need to sort of set up the problem so that you have a sort of target optimization. Okay, so we're going to pick just a somewhat arbitrary

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Encourage people to kind of fork my repo and try these things out, which is if you were to play, let's say, 2v2 go. Then you actually need to model your partner's behavior. And you may not have information on how they play. So you need to aggregate some information on how they play so that you can respond accordingly. Situations where it's no longer a perfect information game. And then in those cases, in games of imperfect information or partial observability, then you do need some context to build a model.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Right, great question. So Go is a perfect information game. And in perfect information games, there does exist a Nash equilibrium strategy for which you can do no worse than any other strategy. So if you know that your opponent has a particular bias, like they love to play aggressively, you can actually in principle counter that specific strategy better than a national equilibrium policy, but to counter any given strategy, it does exist a single national equilibrium that can be decided solely using the current state. So that is a design choice that most Go agents, AlphaGo chose to do, which in hindsight turned out to work very well because the Nash equilibrium seems to be superhuman. Like no human strategy seems to be able to beat it. Now, there are variations of this where you would actually need to consider temporal history. And this is a very exciting research area that I would

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  44. A lot of those tricks, but try as I might, I actually haven't figured out a way to make transformers better than Resonance for now.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Yeah, so if you have a very large 19 by 19 goat board. And you've got some sort of battles going on here, and you've got some battles going on here. When you pass this through a convolutional neural network, the receptive fields of the convolutional network are going to be good at computing local things and making that invariant, but they won't be able to kind of connect these two features easily, right? They need to sort of be pooled together and attend to each other somehow. So the argument about why transformers are good for computer vision tasks with vision transformers and so forth is that because they have a sort of global attention across the whole thing, they can more easily draw these kind of predictions. But you do need more data there so that you can kind of learn through data the sort of invariant local features. I've tried very hard to make transformers work for this problem because I was kind of curious if transformers would present some sort of breakthrough in Go and just remove.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  46. They provide the inductive bias of local convolutions, and generally, transformers start to outperform residual convolutional networks when you want more global context. So, one interesting finding from the Catago paper was that they found it actually quite useful to pull together global features together, at aggregate global features throughout the network to kind of give the network a global sense of how to connect value from one side of the board to another side of the board.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  47. So I'm going to draw a one dimensional flattened move distribution, but this is really like a square kind of grid, right? So maybe it thinks actions are like, these are the kind of probability distribution over good actions. And both of these are categorical classification problems, right? So you can train this like any classifier with deep learning, cross-entropy loss, that kind of stuff. So the specific architecture does not actually matter too much. I tried a few different architectures. Transformers work, resonance work. For small data regimes, my experience is that resonance still kind of outperform transformers and kind of give you more bang for the buck at lower budgets. But this may not be true.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Okay, so now we have a basic intuition of how moves are made with search. We're going to talk about how neural networks can speed this up by providing an analog to the human intuition. So there's two networks. There is the value network Which takes in a state and it predicts, you know, am I going to win or lose? It's a binary classification problem. Then we're gonna have a policy network which induces a distribution over good actions to take.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So I've used two to denote the AIs playing as white and one to denote the human playing as black, and zero as empty. And then now on the AI's turn, it does the MCTS tree search all over again, from scratch, right? So it throws away this old tree that is searched last round, and now there's a new root node, and it begins to search a new. And then so and so forth. So MCTS is basically a, you can think about it like a search algorithm that is deciding what moves to play best, aided by neural networks. And it's done on every move. Okay, great. So let's talk about the neural network part.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Let's go back before we talk about neural nets. Let's just go back to how this play out works. We've only talked about making one move, right? So the AI looks at this encoded Go board. It has a tree. It searches for deeply into the tree to find out which of its actions might be the best. And then it takes that action. And then now it goes back to the human. So maybe now the human sees a go board that looks like this. And then they make their move. So maybe they put their stone here And then now we go back to the AI. Which now looks at a new encoded board.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source