YouSaid · the spoken record

Eric Jang

lines on the record
188
first
2026-05-15
most recent
2026-05-15
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. That's essentially the kind of core idea behind it. Okay, so we take this idea that humans can glance at a board and instantly predict whether we win. And maybe that gives us the opportunity to really truncate how deep we search. And then we also know that humans can look at a board. And decide, you know. What Intuitively at a glance what moves might be good on a Go board. So these are kind of two things that we can use deep neural networks for to accelerate this search process.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Yeah, the important intuition at a high level just to step back about where we're going with all this is that classically, for games like Go, you could build a tree, but we don't have computers powerful enough for that. And Estimating the value of every action that you could possibly take is also hard because you don't know until the end of the game. You could take averages by playing them to the end, but that's also hard because you don't know which actions to take to sample these averages. So conceptually, there's kind of two problems. There's the breadth of the tree, and then there's the depth of the tree. And AlphaGo gives us a way to basically shrink both of those to be very tractable.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Function to look at a board and quickly resolve the game without playing out all of these trees into a very deep search depth.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  4. And so the human glances at the board and they know I'm probably going to lose. And they're essentially running a neural network that looks at a board and implicitly they are amortizing a huge number of possible game playouts and taking that average and then deciding whether the board is winnable or not and then whether they should concede or keep playing or not. And this is remarkable. If you think about the beauty of something like this, it's like a neural network in a human can somehow do all of this simulation at a glance and then just know within a few seconds without actually playing every single game logically based on just kind of like crystallized knowledge and experience that like they can do this and so this gives us a hint that like in games like Go, there are ways to basically radically speed up the search process. And this is one of the fundamental intuitions behind why AlphaGo works is that you can train a value

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Yeah, so we talked about this U value being your final resolution of whether you want or lost. And this is a terminal leaf node condition. Now humans don't play all the way to the sort of edges of the leaves of the tree, right? They kind of stop some dozens of moves before, maybe even hundred moves before in sort of high-level play. So how do they know, right? Like you can think about humans as implicitly having a neural network called a value function that basically takes in a board state and then kind of evaluates

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  6. OK, so if you were to do this without neural networks, it would still be intractable. You would have trouble finding which actions to sample, a lot of the actions would contribute very low value, especially if you're trying to fight your way out of a losing position, and only a few actions give you high value. So the search in practice is still very, very expensive. But the idea is that if you can, because Go follows a tree structure, you can actually inform a very good estimate of the value of this node based on the values downstream, assuming they're all correct and assuming you've searched deep enough.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  7. This action here is just the average of whether you won or lost at the leaf nodes. And correspondingly, you can kind of walk up the chain and say, like, well, the mean action value of this node, let's call this like QB, and this is action B, is just the average of a weighted average of these ones here. And the weighted average is it could be dependent on if you have a different sampling distribution or not, but the basic intuition is that you want to resolve the game where you have a deterministic win or lose, and then you can kind of go backwards. This is called the backup step and assign values to these nodes or actions corresponding to the averaged over the final terminal leaf.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Okay, so this is the action selection criteria for how you decide which moves to move down. Now, as you move down in TreeSearch, you will eventually run into a node where it's quite clear you've won or lost, right? At the very, very end of the game, when there are no valid moves to play left under Trump Taylor's scoring, you can decide whether you won or lost, right? So you either win or you lost. And so this is basically the final return of the whole game. And so the question here is we can assign a value u to a terminal leaf node of the tree, but how do we assign the values for nodes prior to that, the parents? And it turns out what you simply do is you just take the

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  9. And so, where does the random search process come in? That's where P of action comes in. So, if we assume a very naive algorithm where you have a uniform probability of taking any valid action, then this would just be one over the number of valid moves in this setup. And you would be kind of taking this average over this very diffuse tree. And this is a valid integral you can take, but it's very slow because you're going to consider a lot of trees that have very low value. And it's essentially almost like an important sampling problem where you want to, there's only a few actions and sort of paths that can contribute high value and almost everything else is low value. So this is sort of a tricky problem here.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So one small clarification to make is that you talked a little bit about simulations on improbabilities and so forth. We should remember that Go fundamentally is a deterministic game. So the notion of like where does a notion of probability come from here, right? If you had a very powerful computer, There is no probabilities. You can just compute the true average of what the mean action value is. So where does the probability come in? Well, it turns out that as in computer go before alpha go, we've always done some sort of Monte Carlo method where we have, we take the expected Q value averaged over a randomly selected tree. And that randomly selected tree is where probabilities come in. So the interpretation of Q is What is the expected action value under the random distribution induced by some random search process?

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Yes, that's right. So the motivation for UCB was to come up with an algorithm where if you don't know the payoff of the different actions you can select to begin with, this strategy basically was given some exploration term here bounds your regret In terms of how wrong you can possibly be. I don't know the proof. I don't also know if this one is proved to have a logarithmically or like square root bounded regret or anything, but I think the algorithm was just derived to look something like this. And you can tell that these terms, they grow a little bit differently, and this is actually just to account for the fact that Go has many more actions in every given move compared to your standard banded problem.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  12. An action, let's just say call it action A. Initially, it's zero. And so as n is increasing, let's say we've already made 10 action selections from that root node, but we haven't picked A yet, then this term actually starts to become quite large for A. And conversely, if we have chosen a 10 times out of 10, then now this term is quite small. It diminishes very quickly. And the same thing is actually true here.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  13. So, the equation and forms are actually pretty similar. These are both scoring criteria, right? Like you want to arg max this quantity and you want to arg max this quantity to determine which action to take. So let's break down the intuition of how you select actions here. This is the mean action value. So how good is a given child on average? And if you actually knew the whole tree, then this is all you need to select the best action. You don't really need to do more than that. But if you're interactively building this tree as you're figuring out what the Q values should be, then what you have to do is occasionally try some other actions as a sort of explore versus exploit trade-off. So in both UCB and PUCT, there's this term here that basically rewards taking actions that you haven't taken before. So as we mentioned before, each node stores the vis account of taking that specific action, right? So everything is initialized to zero. And so for a given,

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  14. From the parent. Yes. Like, what are the odds that we sampled this one And this will become relevant later. We've talked about a deterministic tree for now. So I'll bring probabilities into this later. And then finally, we have a sort of dictionary of children Which is just like more of these nodes in a sort of classic linked list style reference tree. So this is the basic data structure to implement a tree. And in AlphaGo, they use a slightly different action selection criteria called Puct. And it's short for predicted upper confidence with trees. And this is basically when you select which child to take, you do argmax a of Q of SA Plus constant.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  15. So then Q represents the mean Of this action. And I'll use a subscript A to denote that this kind of corresponds to taking a specific action to get here from the root node. So if we have root basically taking A gets us to this note here. And then we're going to also

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Yes, and we'll call this an action. So, one thing that is easy to trip on is if you come from robotics or other kinds of reinforcement learning is, where are the actions, right? I'm only talking about nodes. Nodes here represents states. And because this is a perfectly deterministic game with no randomness, you can actually just infer the action based on the child. So, if I go here, that implies an action. And this is the state that we resolved. So, the LLMs, if you ask to vibe code MCTS implementation, it'll most likely design the right data structure here. But it's sort of a chef's choice. You can actually rewrite the tree structure however you like. This was what Claude 4.6 wrote for me when I asked it, and it was a very reasonable choice.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Looks like on every move, we're going to take the best action, or the arg max over A that maximizes. Q of A, and I'll explain what Q of A is in a moment, plus some sort of exploration bonus. So on every node, we're going to track a few quantities. So let's consider each of these a node. This is the root node. Where you're making decisions from. And these are the children of the root node. And we're going to say each node is basically a data structure. That is, it stores a visit count Of this node, this child node.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  18. So alpha goes kind of core conceptual breakthrough was using neural nets to make this search problem tractable. So before we get into how neural networks are involved, let's talk a little bit about how we can, you know, assuming we had a powerful enough computer search this tree to find the best move. So in the beginning, you're not going to build out the whole tree because storing that tree would be very expensive. Instead, you might do something like interactively figure out which leaves of this tree are worthy of exploring and expanding into the future to see what else is there. So there are some early algorithms in bandit literature like UCB1, which is not exactly appropriate for a sequential game like Go, but very much inspired the action selection algorithm used in AlphaGo. So UCB1.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  19. And the branching factor decreases by one each time. Yes, yes. But in any case, this is a very, very, very large tree. And this is also why computer scientists for many years thought that Go was not a tractable problem this century because the amount of compute you would need to exhaustively search every possible possibility is just too large. If you could, Go is actually deterministic game. So on any given state, you can actually compute what the best possible strategy you can make is. In order to win the game, you can search all the possible futures where you win, and then just make sure you always stay in that set of futures.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Right, let me use the sport here. So if we start here and then you play here and then I play here and then you play here, that is equivalent to I start here, you play here, I play here. And then you play here, right So both of them arrived at the same spot, but through different paths. So, this child node can be thought about as a shared ancestor.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So, like 300 depth of the tree. So, if we keep on expanding possible moves here, in this move, the AI is going, and then here the human would go. And then there's some. And so forth. You can find that essentially what you end up with is an enormous explosion in the possible game outcomes originating from just this one state. So this is something to the order of like, you know, 361 to 300 power of 300, which is far more than the number of atoms in the universe. It's just, and of course, actually, there are redundancies in symmetries, so it's not actually 300, but that's sort of the, if you were to do a naive tree where there were no merging of children, then actually you end up with a tree about this big.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  22. So, the AI can choose let's just pick three possible random moves that can go, and I just drew these at random. So which move is best here, right? Well, we don't know until the game ends. Go does not have any kind of local reward of which move here is good. And this is what makes Go a very difficult game, is that you don't actually know who won until you really get to the end of the game. So how deep is this tree, right? Well, in a 19 by 19 Go board, there are roughly to the order of 361 moves on any given move. And of course, as it fills up, you have less moves. And the number of steps in the game can be somewhere from 250 to 300 moves. And maybe experts might decide to end the game well before that. But under Trump Taylor scoring, you actually have to play things all the way to the end. So this could be like 300 moves or something.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Great, yeah. Let's start with kind of an intuition about the underlying search process used to make moves. And we'll layer on ideas from deep learning to make it much more efficient and tractable. So, Go is a game where there's just two players. We're going to draw a person here, and we're going to draw an AI here. And let's say this person is playing black, so they go first. And then now the AI is going to make a move based on what it sees here. So there's a question of how you encode these inputs into the AI. Maybe you could use ones and zeros, but you want to represent black, white, and empty. So you would need at least three different values here. So maybe you could use zero, ones, and twos or something. So the AI might see something like 0, 0, 0. So, this is the input to the AI on its turn.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  24. The game ends when either a player chooses to resign. Or both players pass consecutively. Cool. Yep. So that's the rules.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So, in Trump Taylor scoring, it's perfectly unambiguous. So it can be decided algorithmically by a computer. So if let's say you have this at the end game, the way you score this is that you first count how many stones you control, and that's unambiguous. Then you count how many empty intersections that are not touched by your opponent's stones. So these intersections would not count for either player because both of these intersections are connected to both Whitestone's and Blackstones. If this were like this, then White would get three points. Now this is a little odd because a human would know that white is actually losing these points. But Trom Taylor's scoring would consider white to have all of these points as well as these points.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Humans basically say, I think the game is done, and then you have to also say, I think the game is done, and then we'll say, like, I think these are milestones, and then you have to agree. If you don't agree, then we keep playing. So essentially, once two humans, their so-called value function agree on a consensus, then the Chinese rules result that

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  27. However, if you have a way of breaking this formation and connecting Y to something outside of it, then it can flip. And so this is where it's a little bit hard for a computer to decide these kind of things. So how do humans do it? It's worth thinking a little bit about how humans resolve this because this will actually map later to how we think about the deep neural network.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Yeah, great question. So, this is where different rule sets have different ways of scoring. And so we should talk a little bit about how you resolve scores between humans and how you resolve scores between computer code. Because there's actually some ambiguity in how humans evaluate this. So most humans would look at this board configuration and conclude that black has kind of totally surrounded white. And so white has no chance of life. We could play out more here, but then at the end, I would capture everything.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  29. It's actually black because I have actually surrounded this whole area And assuming I have other black stones here, it's actually very hard for you to break this out of the control of these stones.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  30. And so now I would capture this entire group And they said this would be mine. Okay, there's one more case that I want to demonstrate, which actually I had a bug in my code recently, which is the following situation. So let's consider a formation like this, right? And then we have other pieces on the board in play or whatever. And so Let's talk a little bit about how the game ends. In this territory, who controls these areas? Is it white or is it black?

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  31. The clock is gone. Yes. So in Go, it's actually okay to let an opponent capture some stones. If, for example, it allows you to position, to capture more stones in somewhere else on the board. And this is what makes Go a very beautiful game is that you can kind of lose the battle but win the war. And as the board size increases, the complexity of these kind of like micro versus macro dynamics gets more interesting.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  32. And then if you think through what happens if you were to respond here, you can probably search into the future and deduce what I'll do in response once you do that.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  33. The cross section, not the diagonals. So, this one is surrounded on three sides. So you're at threat of losing that stone if you don't play one immediately there. Now you can see that I'm starting to pressure you because by putting a stone here, now you are forced to put one here.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So, this move basically exposes one empty neighbor for your whitestone, and it's very akin to a check in chess, where if you don't respond immediately by putting in one here, then I can immediately capture this.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  35. This white stone down here. It would be instant suicide. In Tromp Taylor, it's actually fine. You put it down and then it immediately resolves to death. So the outcome is sort of the same. Let's go ahead and start over and play a few stones, and then I'll explain some. I'll just start

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  36. So, the Game of Code is a very simple one that can be implemented quickly and easily in a computer. The objective of the game is basically to put down black and white stones and try to occupy as much territory in the game as possible. So I might start by putting down a black stone. Black always goes first. And so the way you capture an opponent's stones is that for every intersection, if you can surround all four of its neighbors with your stones, then this one is sort of cut off from oxygen, if you will, and then it is a dead stone. So then now I control these four stones as well as this empty intersection here. So there's like slight variations between Chinese, Japanese, and what is called Trump Taylor rules. Trump Taylor rules are designed to be completely unambiguous for Go, so this is what all Go AIs train against and resolve against. In typical Go, like what a humans play, you're not actually not allowed to put

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  37. So, if you plot out how much compute it took to build various iterations of strong GoBots over the years, you can see that in 2020, there was an open source project called Katago by David Wu from Jane Street, who basically achieved a 40x reduction in compute needed to train a really strong gobot table. I'm not certain if it's stronger than AlphaGo Zero or Alpha Zero Mu zero, but it's very, very strong. And this is what most Go practitioners today train against when they're playing an AI. And thanks to LLM coding, what took a whole team of research scientists at DeepMind and millions of dollars of research and compute can now be done for a few thousand dollars of rented compute.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Sure, yeah. I like making things. And Alpha Go and Go AI is one of those things that really got me into the field when I saw the kind of early breakthroughs off ago in 2014, 2015, 2016, and so forth. It was just profound to see how smart AI systems could become. And the kind of computational complexity class that they could tackle with deep learning. This is a problem that has long been understood to be kind of intractable for search, and yet it was solved through deep learning. And so that was quite mysterious to me. And I've always wanted to understand that phenomena a little bit better. My training is often in deep neural nets for robotics, where it's the decisions made by the neural networks are a bit more intuitive, but alpha go is a sort of problem where the decisions are actually the result of a very, very deep search. And it's always been very mysterious to me how a tenant

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source