YouSaid · the spoken record
Eric Jang
- lines on the record
- 188
- first
- 2026-05-15
- most recent
- 2026-05-15
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“And I encourage the audience to think about the relationship between thinking and go via MCTS and search and how it relates to LMs. I think there's something quite profound there and probably underexplore just because Go has been relatively underexplored compared to the Boom and LLMs. It's not to say that I think we should have trees in our LMs, but there is some very interesting duality between them. And you can actually do a lot of research on Go, MCTS, and reasoning with very small budgets. So that's very exciting.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Great, yeah. So my website is evang.com. There's a blog post that kind of links to an interactive version of this tutorial. And on my GitHub, which is the username is just Eric Jang. There's an auto-go repo that people can fork and reproduce the training results.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“The jury's still out, right? I think who knows if the, you know, let's say currently Google's doing quite well, who knows if the initialization on training on games is ultimately going to hobble their ability to be the winner in the long term. It's hard to say for sure. And, you know, likewise, who knows if the seeming late start was really just them kind of pre-training for longer on how to scale up TPUs, right? They invested all their tech tree in getting TPUs to be good, which seemed not that useful in the short term, but then in the long term it becomes maybe like a So it's even hard for humans to reason about what the optimal research strategy should be, even with the data we have today.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And so, if that's the case, why wouldn't it also be true for automated AI researchers? Like, they should be able to positively transfer experience tackling quick to verify, quick to iterate on environments to something more ambitious and economically useful, like automating drug discovery or so forth.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah I'm gonna give a non rigorous argument, but one that I kind of intuitively believe, which is that DeepMind, the AI Research Lab, they started as a sort of focus on games, right? They kind of use games as their outer loop, and then the researchers learned from experience of solving games. And then now they're working on LLMs. And presumably there was some positive transfer from their time working on games and Atari and Go and StarCraft that now helps them make good LMs. I assume that there's positive transfer in some regard, whether it's coding or general research ability or project management, right? Like all these things kind of probably help them do well.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And then three years down the line when the NVIDIA GPUs have gotten even stronger, maybe they stack even less well. Maybe like at any given point in time, the sort of benefit of any given compute multiplier is transitory, which is what I sort of suspected with the Katago paper. Like there was many Algorithmic ideas kind of applied. And then you can see that with modern Blackwell GPUs and Ada, Class GPUs that are much better than the sort of V100 grade GPUs that paper used, you can see that some of these algorithmic tricks to speed up convergence just don't matter so much compared to something else. And I think that's a matter of taste in the present time.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, great question. I think the research taste for executing well on the bitter lesson is that you need to know how much the bitter lesson can buy you and how much is too much to ask for at any given moment, right? Like, of course, in the fullness of time, compute kind of is the single most important determinant on how things work. And it's almost like inevitable that as you scale up energy and compute and parameters, intelligence will just fall out of that. And that's super beautiful, super profound. No algorithmic detail really matters beyond that. But in present day, we don't have infinite compute and parameters and arbitrarily good initialization. So we have to come up with heuristics that kind of give us that. But these heuristics are probably somewhat redundant. So that's probably why you see this effect where a lot of these compute multipliers don't necessarily stack is that they might have some correlated benefit.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And so I think that having a challenging game that cannot be cheated easily on the outer loop could be used as a sort of outer loop signal for something like discovering the principles of deep learning. Now, of course, like to make it tractable, and this is where research taste really matters, like you have to come up with ways to initialize your problem so that you don't solve a sort of very intractable problem, right? Like maybe you can leverage LLMs as a sort of a universal grammar in the middle to kind of give you some sort of local feedback. The fact that LLMs are universal grammar means that they can kind of move at almost any level of the stack, right? They can think very locally as well as step back and think like in very broad steps. And I think that's where a lot of the lateral thinking ability of humans kind of come from. Like how to know if the track that you're pursuing or the objective that you're pursuing is not right and you should be asking a different question.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And so this presents a very challenging long horizon RL problem where every step of the way you have a committee telling you that this is a bad idea and then ultimately you break it through. And so how do you design RL environments that maybe give you some feedback earlier? And I think this is a very tough open question that I don't have an answer to. But ultimately to play a very strong robot, you probably did need to discover deep learning.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think, as in the case of the success story for deep learning, you can think about this as a decades-long idea that took a lot of faith to get it to work.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Or can I predict the scaling law plots that emerge from my idea? But then you can verify that you haven't kind of reward hacked anything by using a very verifiable game like go on the outer loop.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“That's one of them. I think there's a lot of deeper questions that one could tackle, right? So, for example, let's say you have an idea on how to improve a scaling law compute multiplier. The outcome isn't necessarily I achieved the best go bot ever. The outcome might just be like, can I predict what the win rate of my go bot will be?”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Environments that might incentivize this kind of lateral thinking. And so one of the motivations for setting up this Go environment was that I think that Go captures a lot of very interesting research problems, often overlapping with LMs or robotics. And yet it's very quick to verify. The outer loop is ultimately like, does the agent do what I think it does? And you can kind of check the outcome of a Go game quite easily. And then the inner loop involves all this kind of research engineering around distributed systems, predicting whether an idea is going to work or not, predicting the difference a particular modification to your training algorithm might make. And I think there's a rich library of subtasks and sub environments that you can kind of train an automated scientist to work on with Go as a sort of outer verification loop that then once you acquire these skills, maybe you can apply them to other domains like biosciences or robotics.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Not worth it. So then I'll kind of jump to a completely different track. And I call these kind of things like rows. So what I find is that current closed models that we can access, the public can access today, they don't seem to be that great at selecting what the next experiment should be in a given track. And they don't seem to be able to kind of step back and do the lateral thinking of like, wait a minute, this track doesn't really make sense. Like, let's go back to sort of first principles and think about what the bottleneck might be or what are we trying to achieve. And so often I had to catch infra bugs myself by prompting the right question to cloud to like investigate what is causing this discrepancy. And then it'll answer the question. I think with Methos class models or mythos plus plus models coming online, maybe this just completely changes and these problems just fall to just improve scaling. But at the same time, I think there's a lot of rich opportunity to develop our own.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And it is also fantastic now at basically executing any experiment, right? So I have a Claude skill that I wrote called experiment where I give it a description of what I wanted to plot. And I just describe here's the X axis I want, here's the Y-axis, answer this question for me. And it'll go run off and do all the experiments, compile the plot, make a report, and suggest what might have caused it or so forth. So that's what works quite well today. And I think we can expect that these abilities get better in the future. But it's also kind of useful to know what is it not doing so well today. So on my blog version of this tutorial, I have a plot of basically all the kind of experiments I did grouped in a sort of tree where every node kind of represents a failed, successful, or sort of mixed experimental result. And then from there, it branches off into a child where it's like the follow-on experiment. Occasionally I'll kind of rabbit hole down a track like this off policy MCTS relabeling, do a few experiments and then realize it's probably”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“A much more open ended set of problems, right? It can say, well, I've identified that the gradients are kind of small in this layer, so let me change it up here. Let me rewrite the code so that the data loader has a new augmentation I came up with. Let's sort of try to find the best way to kind of fit the constraints of the optimization problem. And you end up with this much more flexible and kind of high level, almost like grad student-like ability to just grind a performance metric. And so this can squeeze out quite a lot of performance. On a fixed data set with a fixed time budget, improve perplexity by quite a lot on a sort of classification problem like LMs or Go.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Sure. Yeah. I think automated scientific research is one of the most exciting skills that Frontier Labs are developing right now. And I think it's important for her. Everyone who's doing any kind of research to get a good intuition of what it can do now and what it can't and how might the sort of science process work in the future once we're having AIs automating a lot of this investigation. In brief, I mostly use Opus 4.6 and 4.7 throughout the working on this. What works is that the models can do a very good job of doing hyperparameter optimization. So in the past people would kind of come up with a search base of hyperparameters like learning rate and weight decay and maybe how many layers are in your network. And they would just kind of do a grid search or a sort of Bayesian hyperparameter optimization approach. And then it would find some tuned parameters. The kind of really cool thing that automated coding can do now is that it can search”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And so you're always in this beautiful regime where you're just trying to improve the policy rather than escape this kind of sort of local minima where every signal is flat all around you. So, one way to draw the curve is if you draw the sort of win rate of an MCTS policy versus the raw network, let's say this dotted line is the raw network, the MCTS policy kind of looks like this. And so every step of the way this supervision signal is very clean. You're never in a situation where the MCTS is kind of like giving you no signal. Unless your MCTS distribution converges to exactly what your policy network breaks.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Why is alpha go an elegant RL algorithm? The major reason is that you never have to initialize at a 0% success rate and solve the exploration problem of how to get a non-zero success rate. And this is what allows you to hill climb this beautiful supervised learning signal. And if you look at the actual implementation of AlphaGo, every step of the way, there's actually no TD error learning or dynamic programming, at least explicitly. It's just supervised learning on a value classification as well as a policy KL minimization. So it's just a supervised learning problem on improved labels. And so the training is very stable, right? You can train as big of a network as you want. You can kind of retrain this on the data set. Everything will just go stably. The infrastructure is very simple to implement as well. You don't need a complex distributed system to kind of keep everything on policy. At the end of the day, you're just saying, I have some improved labels. Let's retrain my supervised model on these targets.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“But both of these are actually valid. And if you wanted to do a scientific experiment of how important are this kind of soft knowledge distillation, you can run an experiment where you retrain the policy network on the action MCTS selected rather than the software.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“It would just be the entropy of this distribution. So the entropy of this is zero. The entropy of this is like the entropy equation. And this is also why alpha goes quite beautiful. In alpha go, you don't train the policy network to imitate the MCTS action. You train it to imitate the MCTS distribution.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Once you're here, you have something, but like you actually, in many RL problems, spend all the time here so that there's a sort of question of how do you initialize so you're at least not at zero, but like at a non-zero pass rate. One more thing I'd like to add about bits per sample that's very relevant to any kind of machine learning problem is that And there's a connection to soft targets and distillation where if you have access to the logits, right, not just the one hot, like this is the sort of one hot token answer If you have access to the soft targets, the entropy of this distribution is far, far higher than the one hot. So there's actually way more information in bits per sample in a soft label. So that's why distillation is so effective per sample is that it's actually giving you way more information per se.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, and arguably, you spend all your time here Potentially, never even getting a single success, right? Exactly. So it's a sort of depressing plot in the sense that once you're here, it's not at all obvious how you get to here.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And what's also tough here is that actually the distribution that you're sampling under is your policies distribution. So it's like if your policy has no chance of sampling blue, then you will never get a signal.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“It's just more stable. So, like, you might use the off policy queue as a way to do advantage computation, like Q minus sum of Q. That's kind of like your sorry, like some of there's n actions and then so like. So, this is your value, and then this is your kind of current cube value, so your advantage for that action is like the average value minus your current one. So people can try to estimate Q in an off policy way and then just use advantage here. And then if there's a problem in these dynamics, it doesn't blow up your loss as much. And so in robotics, there's a kind of convergence towards more like using off-policy data to just shape your rewards, but not actually be directly here.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“So, and then here the trainer would be just predict the MCTS label as possible. So, again, this kind of works, and this is quite relevant in robotics where you just have a lot of offline data and you can't simulate things like MCTS. But in practice, it does run into the problem where if the current model is looking at states that it would never reach, then it's kind of wasting capacity. And so you have to be a little bit careful here. So the on policy thing, and also much of RL has kind of converged to a much more on-policy setup where they don't really try to directly train on off-policy data. At best, they use off-policy data as a way to reduce variance, but not directly influence the objective.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“On your current network, right? Not the network that took this action, but your current best policy network, you just rerun your search offline on these transitions. And if these are transitions that your policy can get to, then this actually acts as a very nice stabilizing effect. And also one other benefit is that you can fully saturate your GPU better because you're not blocking on the go game to kind of give you board states. You just simply search across all board states at any depth in peril.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Exactly. Yeah, you can think about it's like you're kind of going back in hindsight and being like given what I've seen in the historical buffer, was there a better action I could have taken? Now the connection to Go here that I tried and it was moderately successful but too complex to kind of like open source was you replaced this with like a MCTS relabeler. Where instead of doing this kind of target network computation, you run MCTS on your transition. So in this case, you have your state, your action, and then whether you want or not at the game. And actually, you can just toss these two. You don't care about these ones. You just take your state and you just plan MCTS. To get your best policy.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Q of S prime, and you find the action that should go with s prime that makes this Q value as high as possible. And then you add that to the reward here, and that gives you your actual target, right? So for this current SNA, your Q target is this. So now you have send back the queue target to this transition. So with this tuple, you pair with that a queue target. And then here on the trainer, you simply just use supervised learning and you minimize your current network's QSA with its target. Got it, okay.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“So you have your off policy data that came from various policies. You're constantly pushing transitions that you saw before to a replay buffer. And then you've got this thing called a Bellman updater, which basically replans instead of this action, what action should I have taken at S to have a better value? And the way you enforce that is you try to minimize the TD error. So actually, given this, you have S, right? You compute.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Or fit the Q to the Q target. So here you can think about this as a sort of planner. You revisit old states that you've been to, and you take your current model and you rethink what could I have done better if I visited this. And so this is actually how off policy robotic learning systems are usually trained. These days, there's a sort of simpler recipe, but in the Google QTOP days, we kind of did things like this.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“While we're training. And so this actually starts to converge on a very robotic-like setup, which is very common, which is you have your data set of trajectories. And then you have something like a replay buffer pusher. And these are off policy offline trajectories, right? So your replay buffer pusher pushes transition tuples. To the replay buffer. And then you have some job that's kind of continuously replanning. The best action you should have done instead of taking this action is. And so in robotics, it's actually very common to use that sort of minimized TD error. So like your Bellman updater. Constantly is pulling things from here and trying to satisfy the QSA. And then from here, you have your trainer. Which is trying to fit the S to A.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And your policy learns how to take sort of the best possible action here, but you never get here. So you're training your model on states you would never reach. This is not there. So then this is a problem. And this is where off policy can really hurt. So actually, as part of this project, I did try an experiment where I took a bunch of trajectories and to try to saturate the GPU as much as possible. What I did was I took random states from the data set and reran MCTS on just those states. So instead of playing a whole game where I'm doing MCTS on every move, I just ignore the sort of causality of moves and just pick random board states and I just label those with my current network. And I might revisit old states that I've labeled before and relabel them again with my current network. And so in practice, this actually does work. You can actually say, let's take some states that are reasonable and constantly be relabeling them in”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“And you always want to be able to correct back to your waiting condition. So your replay buffer really should have the states that your policy would visit plus some distribution of states that you might drift to, and then how to return back to your optimal states. Now, if you take this to the extreme and you say, well, we don't have any of this data. And we're gonna just be labeling with MCTS states that are so far away from our optimal behavior, like this bag of states over here. Well, like now, yeah, I mean, each of them gets MCTS label.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“This is actually why you want to do off policy training sometimes. You don't want to have a compounding error where if you make a mistake, you don't have the data of how to return back to your optimal distribution. And so optimal control does not really say too much about how to not accidentally get here because it's sort of making the assumption that once you learn the policy, you're going to get it here. But in applications like robotics, right, like I don't know, a gust of wind blows you slightly off and then you need to like correct, right? Or the friction on one of your tires is kind of a little bit lower than the other wheel. And then now your car is drifting and you got to kind of like correct it. So these kind of things in more real environments often happen where actually there was a funny quote about chess and it also goes like the problem with Go and chess is that the other player is always trying to do some shit. So like, you know, things can kind of drift off.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“What you kind of want in an algorithm like this is to have mostly states that you would visit, but then you have a small percentage or maybe a reasonable percentage of states in this kind of high-dimensional tube around your optimal trajectories. And any of those states are given a supervision target to kind of funnel you back into your optimal trajectory. So maybe I can just draw quickly here. Great. So in sort of a dagger style setup, you're kind of optimal training data distribution is, is that here is your optimal states and actions. So this is like, you know, you want to be in this state. You want to be in this state. You want to be in this state. And then you win here. And then these are your optimal policy actions. So these are the things that you definitely want to train on. But to make it robust to disturbances, you want to make sure that if you happen to drift off into some other states,”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Great question. Yeah. And this gets into the sort of fundamental off policy versus on policy reinforcement learning kind of questions. So as you recall, in MCTS, you take actions that you took and you relabel them to take different actions on the same states. So the off policy part here comes where what if you're relabeling states that your new policy would never visit like what's the point? You're kind of wasting capacity. And in the extreme limit, imagine your distribution of states in your training buffer are all states that you would never visit. Then you're basically supervising them to take good actions on states you would never achieve and therefore your policy can get really bad, right? So this is where off policy can really hurt AlphaGo. However, if you interpret this sort of from like the dagger perspective, which is basically saying a way to kind of correct yourself back to the optimal trajectory given some data.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“You can cut down that time a lot by kind of pre training on a small board and then worm starting that into your 19 by 19 board play. There were some other stuff like varying the number of sims between episodes. This turns out to be not that sensitive, actually. You can kind of fix it or increase it. Doesn't matter too much. But so anyway, it's kind of just nice from a scientific perspective just revisiting like an old paper and seeing what really matters.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“So, if you're initializing against best response training against Katago itself, then your own model actually needs none of the tricks that KataGo needs. So then the core thing is like, how can you get as quickly as possible to some strong opponents? And that matters a lot more than the specific architectural innovations. But there are still some nice compute multipliers. So I found that trading on 9x9 boards was very nice for resolving endgame value functions. And then if you can co-train that on a...”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“We're at the sort of speed of GPU where the size of the model is not so big that this really matters, you can actually simplify the setup quite a lot. So instead of doing a distributed asynchronous RL setup with replay buffers and pushers and collectors, you can kind of do a dumb synchronous thing where you collect. You just train a supervised learning model and then you collect again. And so there's like opportunities to simplify infrastructure. Nvidia GPUs have indeed got faster. So whereas Katago was trained on V100s, you can train on half the number of desktop Blackwell GPUs and it still works. And some of the kind of auxiliary supervision objectives that Katego developed aren't really necessary if you have a strong initialization.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, great question. So going into this project, I kind of knew in the back of my mind that things always get easier to do over time. And I want to see where is Go at, given that it didn't seem like there has been any major open source strong bot after Katago in 2020. And then reading the Katago paper, there's a lot of clever ideas. I was kind of wondering, okay, let's see if the bitter lesson has happened where a lot of these kind of tricks just sort of go away because Nvidia made faster GPUs, right? And so roughly where are we on that? Again, this is not a peer-reviewed claim. So this is just my preliminary vibe guess on what I've seen based on my own experiments. But it seems like... Architecture choices don't matter that much. Transformer versus ResNet.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Like time to result, or just getting into work. I think the first alpha go probably they had lots of compute and they didn't need to be, they didn't need to worry too much about making it the most compute optimal thing.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Millions of dollars we're talking about. But in the past, when compute for experiments was kind of more plentiful or not accounted in a way that the researcher was really responsible for, then you kind of end up with people optimizing for things besides kind of being on the compute optimal Prido frontier.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Alpha Zero team, they did not have any policy that they could train against, right? Because they were trying to do everything Tableau Raza. And being the first to do it means that you're prioritizing getting the thing working rather than, let's say, the most compute efficient possible implementation. So this actually plays out in robotics as well. If you look at the kind of frontier of large models trained for robotics, the scatter plot is all over the place and there isn't a very clean line the way that there is for frontier LLMs. And that is because the folks training these models often are not at the scale where every flop counts and they need to kind of squeeze out the performance of every single flop as the dominating decision deciding factor in pre-training. Instead their focus is more like we want a certain capability to show up so we optimize the training setup to kind of make it easy to derive that capability. And once you have that capability, well invariably if you scale up the compute, you are forced to kind of make it compute efficient because this is like hundreds of”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“The compute required to be the first to do something is always much larger than the compute it takes to catch up. And it's the same story playing out in LLMs. Once someone else has done it, you could use tricks like distillation. You could use all sorts of kind of crutches to kind of bootstrap your way to success. So with my own bot that I've hosted online, I actually used sort of best response training against the Katago models to kind of get a strong level performance. As a time of recording, I'm validating whether this can be, I can kind of do that first step, which is to do the tabula RASA. But importantly for research, you often want to start from a good init, right? So the kind of simple thing I did first was train best response agents against catago.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“I got a donation from Prime Intellect for like about 10K, and then I spent maybe the first 4K doing kind of exploratory research. And then about 3K on the kind of final run. And then some of it remaining for serving the model.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Of this project was you don't necessarily want to kind of jump into the science of studying your man made artifact before your man-made artifact is interesting enough to be studied.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so a mistake I made initially when I had some bugs around how MCTS labeling was working was I would collect a bunch of data with an expert policy and then treat it as a supervisor learning problem and try to identify scaling laws with expert data sets. You can indeed plot things that look kind of like this, but if you're in a regime where your policy is not working well, you might be just studying scaling laws on bad data, right? So just like one important implementation detail is that if you want to study a scaling laws problem, you kind of have to have a problem for which the data is good, the architecture is good, and there's no bugs. And then you solve it there. Ex ante, I wasn't able to apply scaling laws to direct what to look at until I had read the rest of the system working. And this sounds obvious, like researchers, of course, you want to have like a working bug-free system before you study scaling. But just as a sort of advice for practitioners on where I actually tripped up when I started.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“To collect data that then helps you build a mental model of how things work, such as scaling laws, right? And so usually actually if you want to build a strong gobot using scaling laws, you actually have to make a strong go bot first and then use the scaling laws to kind of extrapolate a bit farther into the future.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source
“Sure, yeah. I think that indeed test time scaling and reasoning and how it interacts with model size are quite profound when it comes to how much needs to be actually done as explicit search versus how much can be packed into the forward pass of a neural network. And how does a forward passive neural network sort of learn how to do something that should be a sort of sequential and recursive step? That's quite interesting. So the Andy Jones scaling laws for board games paper is quite cool. There's another really nice result from that paper where he showed that not only can you predict scaling laws of the sort of LLM variety where as you increase parameters, you can decrease the amount of compute for search or vice versa. He also showed that you can actually predict how much compute is needed to solve a larger version of the board game.”
2026-05-15 · Dwarkesh Podcast · Eric Jang – Building AlphaGo from scratch · IDENTIFIED FROM THE TRANSCRIPT · source