YouSaid · the spoken record

Eric Jang

lines on the record
188
first
2026-05-15
most recent
2026-05-15
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. No way is a strong word. I think lots of people have thought about how to try to apply MCTS or its kind of successors like New Zero to continuous control spaces. And I'm sure very cool research work is still ongoing to try to crack that problem. But yes, the seeming challenge right now is that most problems in much higher dimensional action spaces or something that's combinatorily much bigger like language, they don't seem as amenable to the kind of discrete action selection heuristics as well as kind of game evaluation type stuff that Go does. But that's not to say the idea of like, you know, thinking into the future along multiple parallel tracks might not give you some information about which way to search. If you think about mathematics, I think mathematics often occupies a little bit more of like a logical search kind of procedure where you kind of can back up, you can kind of see which paths seem good or not. There's more of a rigid structure there.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So here's an example where LMs break down. The CPUT rule involves square root of n over 1 plus na. In an LLM, you're most likely never going to sample the same child more than once. So if you have, let's say, multi-steps of thinking, because language is so broad and open-ended, it's a sort of... Discrete set of actions is not really an appropriate choice for an LLM, even though they're a discrete tokens. It's just such a large number that this type of exploration heuristic is probably not the right thing to do to guide how to search down a tree.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Yeah, certainly I think the LLMs managed to do something that looks like real human reasoning without having to do an explicit tree structure. That being said, I think the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback, even though not exactly in the same instantiation as Alpha.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  4. We probably will see revisiting of this idea of forward search in the future. But there's two things that make MCTS very simple for Go, which is that value estimation is kind of concrete and you can determine it for real. And then you can kind of use it to truncate depth, as you said. And then the breadth is also determined. And what's kind of critical is that the action selection algorithm where you iteratively visit and grow the tree is well suited for the size of problem that Go is and the depth of the problem. But for something like LLM reasoning, Puckt might actually not be a good enough heuristic. It might be too greedy with local tokens and it might do something like, oh, only give you sort of obvious thoughts that are correct, but not really solve your final problem. So I would say the jury is probably still out on how what the final instantiation of reasoning for LNL.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  5. So there was some research, I think, from Google in 2023, 2024, where they did try to apply tree structures to reasoning. And I think it's the jury is still out as to whether this can ever work. So I would say.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  6. So, Q learning, or this kind of approximate dynamic programming kind of propagates what you know about the future cues backward like this, right? And you can see that there's a sort of similar structure that goes on here where in this case you're planning over trajectories your agent hasn't actually been to yet, whereas in this case you're planning over trajectories your agent has visited So importantly, why does Q learning? Why was Q learning a big deal, right? It's because historically we just haven't had the ability to do search on fairly high dimensional problems like robotics or whatever. So for a long time, we kind of make the assumption that like, okay, well, if we can't model the dynamics with a world model or something, we're going to instead just collect trajectories.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  7. This is a algorithm for recovering value estimates of intermediate steps when you don't have the ability to do forward search. So, you must collect a trajectory first of n steps before you're able to do this trick. But the intuition is kind of the same, which is that knowing something about the Q value here can tell you something about the Q value here. And indeed, you can recover a policy from a Q value. So you don't need to explicitly model the policy distribution. You can actually recover the policy distribution by doing argmax over your Q values.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  8. What this is sort of saying is that the best action you can take at this state is equal to the reward you take for taking this action plus the best that you can do at the next state. So there's a sort of recursive and dynamic programming property of MDPs. And you can train neural networks to basically try to enforce this consistency. So you can say, well, once I know the Q value of this action, I can then use that to kind of compute something about the Q values.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  9. There is often a component of estimating a Q value. And so Q values are often learned through TD learning, although in PPO, the way that they do advantage estimation is not necessarily through a Bellman backup. But in Q learning, there's this kind of very cool trick where you do... Q SA is backed up as R plus some discount factor times the max AQ of your next step. So intuitively how this works is like if you have an MDP. And then this is like terminal.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Yeah, so here you can use a number of algorithms like PPO, VMPO, Q learning, even if you want. The specific algorithm here can be, you know, it's usually a model-free thing because you don't have search, but there's an interesting connection from MCTS and QLearning that I want to bring up. So in MCTS, you do something where you have a tree. And through the resolution of your value function at the leaves of the tree or your approximate leaves of the tree, you can kind of back up through the sequence of many sequences and then obtain some sort of mean value estimate. Your Q is kind of derived from the average of a bunch of simulations. In model-free algorithms,

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  11. RL algorithm to find the best search action that you could do to kind of beat your opponent. And then finally, you're distilling the policy here into what is known as a mixed strategy where it's trying to basically average across all possible opponents you could play against. And this is what gives you something that can do no worse than an average selected opponent from the league. And so this gets around the problem of having to derive a teaching signal from MCTS, but it still fundamentally is about relabeling your states with better actions so that they improve your policy.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  12. And once you have a good policy that you train with pick your favorite model free R algorithm PPO or SAC or any kind of mixture or VMPO or whatever, you now have a good policy that gives you a good label for what this one should do when playing against that player. And when you train multiple best response policies, you can basically then distill the RL algorithms into the labels for a given opponent. So you might have, let's say, a best response policy against pi B, and then maybe you have a league of opponents like PiB, Pi C, Pi D, and you're going to take the best response policy that you train against each of these fixed opponents. And for this one, you're going to supervise them with the label that this one would provide. So it is kind of like, this is almost like a proxy for your MCTS teacher, right? Instead of MCS teacher, you use a model free R.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  13. So, you fix your opponent and you treat this as a classic model free RL algorithm where your goal is just to beat this guy. And so here you use your standard TD learning style tricks or use PPO or any actually like model-free RL algorithm to try to hill climb against winning this player. And so you train basically you have a reward function that's like return is like, you know. One if wins against. So, this is no longer a self play kind of problem, right? This is just like a fixed opponent, and you're just solving it, trying to maximize a score against that, and then zero otherwise. And so you have a sort of fixed environment where all you care about is just beating this guy.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  14. In MCTS, we perform search, and assuming we have a good value function, the search will kind of give us a better result than our initial guess. In a game where you can't easily simulate a search process, what they do instead is train what is known as a best response policy. So, you fix your opponent. So let's say you're currently training pi A against. A strong opponent pi B. In StarCraft, maybe like, you know, these are the Zergs, and you're playing Protoss or something.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  15. What happens if you don't have the ability to easily search a tree? Like in Go, it's a perfectly observable game. You can easily construct a pretty deep tree that completely captures the game state. In a game like StarCraft where you don't have really complete control over the binary, it's a little bit hard to do this. And I'm not even sure if it's a deterministic game. So that makes this kind of difficult from a data structures perspective. What is done instead is that the basic idea of supervising your actions with a better teacher is still there, right? So in a given neural fictitious, so we're going to talk a little bit about how neural fictitious self-play works. Same idea, we're going to come up with better labels for each of the actions we took, just like an MCTS. But how do we derive the better labels?

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Right. And so keep in mind that in this case this model free RL setting is trying to solve a credit assignment problem where you don't know which actions were actually good and which ones were bad. Monte Carlo tree search is doing something very fundamentally different, which is it's not trying to do credit assignment on wins. It's trying to improve the label for any given action you took. And so we can actually think about a completely different algorithm called neural fictitious self play, which was used to great effect in systems like Alpha Star and OpenAI's Dota. So let me talk a little bit about how you can kind of unify some of these RL ideas in the model-free setting. Okay, so

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Yeah, so this is where in RL people use things like TD learning to better approximate the quality function, the queue that we mentioned earlier. So you can try to subtract that from your return. So ideally what you really want to do is in RL, you want to push up the actions that make you better than the average. And push down the actions that make you worse than the average. And they call this advantage. There are multiple ways to compute it. I highly recommend John Shulman's general advantage estimation paper as a good treatment on how to think about various ways to compute it. But at the end of the day, you want to reduce variance by trying to make this smaller so that it doesn't magnify the variance of this whole.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Right, so this is a pretty tricky problem in practice. And so this is where advantage estimation happens in reinforcement learning. So you want to subtract A term Your multiplier, instead of an indicator function of 1 and 0, you want something that kind of behaves like a zero for all of these guys. And then a one for all these ones.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Actually, the optimal case is to pull out, discard all of these moves, and only get a gradient on that single move that you got better.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  20. It looks kind of like this. This is sort of the very basic form here, but this is still a contributor to variance. So you want to make sure that similar to how in this case we were training on a lot of neutral labels, you want to make sure that you're sort of penalizing the labels that don't help and only rewarding the ones that actually make you better. So intuitively, the analogy here is can we find a term in our training objective such that it's actually kind of discouraged from doing this, or these don't have any effect on the gradient? And this has an effect on the gradient.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So the thing that's a little bit hidden here in the math is that we're assuming that when you decompose the problem to a multi step problem, that you're now introducing kind of correlations between your actions through the computation of the sky. And so if you separate these things out, then there will be, this will magnify the variance of this one. So, in the case where you don't separate it out, if you just have t equals one, you just have a single estimate of log prop. And a single estimate of reward. Now, this term still shows up in LLMs, it looks a little bit more like the naive reinforced estimator looks a bit like return of the single action plus times, you know.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  22. You could imagine a reward that says, I'm going to give you some process supervision. Where you get a reward for each of these actions on every step.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Probability of this sequence is equal to the sort of sum of log probability of the whole sequence is equal to the sum of the probabilities of individual tokens. So in this case, I would. I would say something like log hell plus log low plus log world. So this is true. And if this term were one, then they would be the same thing. However, in sampling things, if you have a reward term assigned to every specific token, now you have these interaction effects between the cross multiplication of these terms and these terms. And so the problem becomes how do you ascribe the credit associated with every episode to all these different terms here?

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  24. And just for simplicity, we can pretend this is on average zero or something if you're centering it at no signal. And the variance here basically means that you're taking the square of this product term. And so you end up with a term that kind of grows quadratically with t. So variance, when you have a setup like this, this thing acts as a coupling effect on top of these terms here. So Let's actually map this to an LM case, and we can answer why do LLMs only do one step RL instead of a multi-step RL scenario. In LMs, you have a decoder that might predict some words like hello world. And so in current lmRL, they treat this entire sequence as a single action, just AT, and big t is just one, right? And so, yes, it is true that because of how transformers are formulated through the product of conditional probabilities, we do have

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So, the sum of rewards is the return, right? So in our naive setup here, we only have an indicator variable for the return where either you won or lost. So, in the case where you lost, well, you just don't train on, your gradient is zero. You don't train on those examples. And when you won, you try to predict those things. So you can think about this setup as a special case of this general formula here. The trouble here is that this is very high variance because when you multiply these terms out, when you take, when you try to compute the variance of this, so variance of the gradient. Is equal to expectation of Squared minus

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Right. So in this case, this is not to say it doesn't work. If you imagine increasing the number of games to millions of samples, you actually can get some meaningful supervision. Samples so long as you find a way to sort of mask out the supervision from these guys. And then this is where things start to get pretty related to RL in terms of advantage and baselines and so forth. So let's look at the gradient variance of very naive approach like this, where I'm just going to call it gradient RL. And it's basically the sum of rewards.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  27. 49 of the games, they played exactly equally. I'm sorry, for 50 of the games, they played exactly equally. And on that one game where this one won, it played slightly differently. It made like one critical move that normally it would have done differently, but due to some exploration or some random noise, it just happened to make a smarter move than it did previously. So you have one supervision signal, like one true supervision signal for your policy network. And then you have 99 games 300 moves for which imitating those actions gives you exactly the same policy you had before. And so the scale of your variance is actually very bad because it's like you only have one label out of this enormous data set of actions, of supervision actions where you want, actually, sorry, let me clarify a little bit.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  28. You have a matchup between two agents that are basically the same. So, in fact, let's just assume that policy A and policy B are evenly matched. So their true win rate is 50%. So let's say you play 100 games. And then each game, let's say, lasts 300 moves. And you're doing some sort of evolution strategy or some way to perturb these things to get them to do different things, or maybe you don't and you just play them against each other and you see occasionally this one might actually have a better strategy than this one, right? And so let's say 51 games, the policy A wins. And then 49 games policy B wins. And this is just due to random luck, or maybe you perturbed policy A in some way that let it do this. And just to have a very, very simple model, let's pretend that for like...

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Okay, so it's worth kind of thinking a little bit about, okay, what are some alternatives we could do to train self play agents instead of MCTS, right? We use a lot of LM cell RL these days. Is that relevant? Could we do that instead? So let's think through this a little bit. Let's suppose we have a very naive algorithm where we take a league of agents of different checkpoints and we play them against each other. And for the games where a single player wins, we're going to reinforce those actions up and then retrain the policy network to imitate those guys instead of the MCTS objective. So what ends up happening is let's say you have a chain of actions that led to a win.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Importantly, what it is doing is saying for every action we took, we did a pretty exhaustive search on MCTS to see if we could do better. And we're just going to make every action that we took better by having the policy network predict that outcome instead. And so this is a very, very nice idea because you have one supervision target for every single action. So, the variance of your learning signal is very low compared to the alternative naive RL thing. So let's actually consider a very naive algorithm that looks a lot more like modern LMRL today, where we do something like let's take the winner of a self-play game and encourage it to do more of that.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Yeah, you have the maximum power of a neural network at the edge of chaos. I think there's some research papers from Joshua Stroldik's team on this. Yeah, like there's something kind of quite fundamental about chaos that is, it's not just like hopeless noise. It's like there's something kind of useful in chaotic systems, at least at that boundary. But yeah, this is just my, think about this as a philosophy. I don't actually know the math well enough to comment on it. Anyway, if we go back to, we'll talk about LLMRL in a little bit because there's some connections there, but let's just go back to the MCTS. What is it doing? It is not crucially, it is not saying we're going to increase the probability of winning directly. It's not going to say like we're going to upweight all actions that won and downweight all actions that didn't win.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Yes, intuitively, that seems correct. And then again, this is also out of my area of expertise. I find it interesting that cryptography

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And so, whether it could be the same thing, right? We don't exactly care what the velocity of wind 6,000 feet above a specific latitude, longitude is, we kind of care like, where's the hurricane? Or things like that. And I would say in chaos, there's a classic Lorenz attractor, which kind of looks like this, right? Yes, if you start anywhere on the Lorenz attractor, you don't know where you're going to end up. But you do know that the thing looks like this. Right. And so there's this kind of beauty of like sometimes we don't necessarily care about the microscale things. We actually care about the macroscopic structure. And these things can be predictable.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Um... Given what we know about both players, what is the board state in the future What is the exact board state in the future? This is extremely sensitive to initial conditions. Like a single stone place here can kind of disrupt the entire prediction. So this is hard. This is kind of intuitively the chaotic problem. And yet, somehow So this is hard. Somehow we can predict who's going to win And this captures a lot of possibilities here. And so there's this more macroscopic quantity that we really care about, which is the average or expectation or some sort of global macrostructure over a lot of possible futures.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  35. So, to me, AlphaGold was the first paper that kind of really showed this profound level of simulation being compressed into a small amount of.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  36. And the same thing has been shown in alpha tensor, alpha fold, where yes, there is a very hard problem that in the worst case seems intractable, and yet we're able to make almost arbitrary amounts of progress. So here's a sort of, in the limit, what might this look like, right? Well, if you want to simulate something very complex like weather or predict the future, do we live in a simulation or not, the computing resources you need to build a very complex simulation might be much smaller than you think based on our ability to amortize a lot of that computation into the forward pass of a single network.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah, I think that the kind of question we should be asking ourselves is we've been formulating solutions to MP hard problems as in kind of worst case complexity And I wouldn't say this solves Go, right? It doesn't give us an exact solution of the optimum, but in practice, it is extremely useful.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Our understanding of problems like P equals NP or these very fundamental computational hardness problems are Incomplete, right? Like it's not like, you know, obviously this is not a proof of p equals mp or anything, but there's something to it that kind of is very disturbing where what felt like a very hard problem can fall to a very, very simple macroscopic solution.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  39. By construction, let's say a 10 layer neural network can only do 10 sequential steps of thinking, right? 10 steps of neural network paralyzed distributed representation thinking is able to amortize and approximate to a very, very high fidelity a nearly intractable search problem. So this was a breakthrough that I think most people don't even understand today, like fully comprehend how profound that accomplishment is. And this is what also GERD's alpha fold, for example, right? Where you have a very, very difficult physical simulation process that you would need to roll out so many micro scale simulations and yet like 10 steps of a somewhat small neural network can somehow capture what feels like a MP class problem into a single problem. And so it actually makes me wonder if.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  40. I personally disagree. I think they're profound for different reasons. And I don't understand the LMRL enough to kind of comment on your podcast about it. I think AlphaGo. So why is it a profound accomplishment? I think maybe it's worth stepping back a little bit and just like, it is different than modern RL, and we can talk a little bit about some of the algorithmic choices there. But I think the most profound thing here is that a A 10 layer neural network pass. So basically 10 steps of 10 steps of reasoning. And of course, the reasoning is not just one trail of thought. It could be like the distributed representations and a lot of thoughts going on in the same time.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Right, AlphaGo Li, the original AlphaGo paper, had two separate networks And then, in all subsequent papers, they merged them into heads.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  42. And then so you can score those pretty easily. And so that is what gives you the bootstrapping to be able to then improve your policy with search. But it's very, very critical that MCTS has accurate value estimates. And you need to ground the value ultimately, MCTS will fall apart if you don't have a grounding function for the value.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  43. If we take it to its limit and consider a very tiny 4x4 go board Like, if you play 50,000 games, you're going to have a lot of end states that look like human play, right? Like, it's just like tic-tac-tac-toe at that point. So if you broaden this a little bit to nine by nine, five by five or nine by nine, it's not unrealistic to imagine that purely random play will actually generate pretty reasonable looking.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Yes, and there's a beautiful connection to TD learning that we can talk about in a bit, as opposed to contrasting with Monte Carlo's research. So you first want to get good value functions and expert data can kind of give you a quick shortcut. I recommend for practitioners just do that first just to initialize to a good starting point. And then if you want to do the alpha zero thing or Katigo Raza learning, then what you can try to do is on a small board, play random games, just take a random agent. And if you play like 50,000 games, you'll actually learn a pretty good value function as well. Because on a 9x9 board, there's actually, you can see enough of the common patterns with random play. And then if you train a model that kind of can train on both 9x9 and 19 by 9 data, and Katago was a proposed one of these architectures, then there's some pretty good transfer learning from the value head evaluated at 9 by 9 to the 19 by 9.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  45. That those are much more subtle to judge in the mid game than the beginning or the end. So the most difficult part to score is like not the beginning or the, because the beginning is just obviously 0.5, and then at the end, it's pretty obvious who's winning. So the hard part that you want to learn in the value function is like who is winning in the middle.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  46. It's almost like a decidable problem, right? Because there's lower and lower uncertainty as to the depth of the treat. So most games play to the end by reasonable people. Will be good training data to train a good value function at terminal parts of the tree Then as you play more games, the search will backup good values into the sort of intermediate nodes of the tree. And then as you increase the amount of data, your value head gets a good intuition of what is a healthy board state versus a non-healthy board state.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Good play. You can easily learn the late stage value functions pretty well. And that's what you kind of need to start the search process.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  48. One trick I did find to be pretty useful, and this is not a peer reviewed claim, so just like take this with a grain of salt, is I found it useful in my own implementation to do the following. You want to first make sure that this is good before you invest a lot of cycles doing MCTS, right? Like it doesn't really make a lot of sense to do search on garbage value predictions. So you want to kind of start at a good place where this works. AlphaGo Lead does a very good thing where it just takes human games and then you train on it and then just works, right? Totally works. You can also take an open source GoBot play it against itself, generate data also works. So if you have some offline data set that has realistic

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So, this is why MCTS kind of, if you assume that the value functions are correct, why it gives you a better policy is because, and it's a very critical chain of assumptions, assuming that this is accurate, then your search process should give you a better recommendation than your initial guess.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source

  50. When you kind of take n to infinity. So, variants in your search process as well as inaccuracies in your evaluation can definitely screw with the quality of your policy network combination. And so that's why it's not a guarantee to improve. And that is why I think I suspect why AlphaGo Lee had the playouts to the end in their training algorithm so they could ground this thing in real playouts. In practice, what you could also do is just like for 10% of the games, you prevent the bots from resigning and you just say like resolve it to the end. So you get some training data in your replay buffer to really resolve those kind of like late stage playouts that normal human players would kind of not play to.

    2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source