YouSaid · the spoken record

Noam Brown

lines on the record
196
first
2022-12-06
most recent
2022-12-06
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. And so there, neural nets were super helpful because you could just feed in a ton of different board positions into this neural net, and it would be able to predict then who was winning or losing. But in poker, the features weren't the challenge. The challenge was how do you design a scalable algorithm that would allow you to find this balanced strategy that would understand that you have to bluff with the right probability?

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So we actually did not use neural nets at all for libratis or plurbus. And a lot of people found this surprising back in 2017. I think they find it surprising today that we were able to do this without using any neural nets. And I think the reason for that, I mean, I think neural nets are incredibly powerful. And the techniques that are used today, even for poker AIs, do rely quite heavily on neural nets. But it wasn't the main challenge for poker. Like I think what neural nets are really good for if you're in a situation where finding features for a value function is really difficult, then neural nets are really powerful. And this was the problem in Go, right? Like the problem in Go was that, or the final problem in Go at least, was that nobody had a good way of looking at a board and figuring out who was winning or describing through a simple algorithm who was winning or losing.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I think the idea is that we want to be able to leverage search as much as possible. And the way that we were doing it in Libratus required us to search all the way to the end of the game. Now, if you're playing a game like chess, the idea that you're going to search always to the end of the game is kind of unimaginable, right? Like there's just so many situations where you just won't be able to use search in that case. Or the cost would be prohibitive. And this technique allowed us to leverage search and without having to pay such a huge computational cost for it and be able to apply it more broadly.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah, yes. So where this depicted search came from is I developed this technique and ran it on two player poker first. And that reduced the computational resources needed to make an AI that was super human from $100,000 for libratus to something you could train on your laptop.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  5. No, no, no. I mean, first of all, computing resources are getting cheaper every day. But you're not going to see a thousand-fold decrease in the computational resources over two years or even anywhere close to that. The real improvement was algorithmic improvements. And in particular, the ability to do depleted search.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Personally, I think the most interesting thing about Pleurbus is that it was so much cheaper than Libratus. I mean, Libratus, if you had to put a price tag on the computational resources that went into it, I would say the final training run took about $100,000. You go to Pluribus, the final training run would cost like less than $150 on AWS.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  7. That's right. Nobody has a proof that that's the case. And it could be that six player poker belongs to some class of games where approximating an ash equilibrium. Provably works well. And there are other classes of games beyond just two player zero sum where this is proven to work well. So there are these kinds of games called potential games, which I won't go into. It's kind of like a complicated concept. But there are classes of games where this approach to approximating an ash equilibrium is proven to work well. Now, six player poker is not known to belong to one of those classes, but it is possible that there is some classic games where it either provably performs well or provably performs not that badly.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Right, yeah, poker is just such an adversarial game. There's no real cooperation. In fact, you're not even allowed to cooperate in poker. It's considered collusion. It's against the rules. And so, for that reason, the techniques end up working really well. And I think that's true more broadly in extremely adversarial games in general.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  9. And so there was this big debate about whether Nash equilibrium and all these techniques that compute it are even useful once you go outside of two player zero sum games. Now I think for many games, there is a valid criticism here. And I think when we talk about, when we go to something like diplomacy, we run into this issue that the approach of trying to approximate a Nash equilibrium doesn't really work anymore. But it turns out that in six-player poker, because six-player poker is such an adversarial game, where none of the players really try to work with each other, the techniques that were used in two player poker to try to approximate an equilibrium, those still end up working in practice in six-player poker as well.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So, the meta problem is, in some sense, how do you understand the Nash equilibria that the other players are going to play? And even if you do that, again, there's no guarantee that you're going to win. So if you're playing risk, like I said, and all the other players decide to team up against you, you're going to lose. National equilibrium doesn't help you there.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  11. And if every single player independently computes a Nash equilibrium, then there's no guarantee that the joint strategy that they're all playing is going to be an ash equilibrium. They're just going to be like random dots scattered along this ring rather than four coordinated dots being equally spaced apart.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  12. When you go outside of two players you're a sum. So a natural equilibrium is a set of strategies like one strategy for each player where no player has an incentive to switch to a different strategy. And so you can kind of think of it as like imagine you have a game where there's a ring. That's actually the visual here. You got a ring and the object of the game is to be as far away from the other players as possible. A national equilibrium is for all the players to be spaced equally apart around this ring. But there's infinitely many different Nash equilibria, right? There's infinitely many ways to space four dots along a ring.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  13. What it was doing for all those moves. Now, when you get to six player poker, it can't do that exhaustive search anymore because the game is just way too large. But by only having to look a few moves ahead and then stopping there and substituting a value estimate of like how good is that strategy at that point, then we're able to do a much more scalable form of search.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  14. I had come to realize that the techniques actually I thought really would extend to six player poker because even though in theory they don't give you these guarantees outside of two player zero-sum games. In practice it still gives you a really strong strategy. Now there were a lot of complications that would come up with six player poker besides like the game theoretic aspect. I mean for one, the game is just exponentially larger. So the main thing that allowed us to go from two player to six player was the idea of depth limited search. So I said before like, you know, we would do search, we would plan out, the bot would plan out like what it's going to do next. And for the next several moves. And in libratis, that search was done extending all the way to the end of the game. So we'd have to start from the turn onwards, like looking maybe 10 moves ahead, it would have to figure out.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  15. So I mentioned Nash equilibrium in two player zero sum games. If you play that strategy, you are guaranteed to not lose an expectation no matter what your opponent does. Now once you go to six player poker, you're no longer playing a two-player zero-sum game. And so there was a lot of debate among the academic community and among the poker community about how well these techniques would extend beyond just two-player heads up poker. Now,

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Yeah, and I mean, this is what I have been dreaming about since I was like 16 playing poker with my friends in high school. The idea that you could find a strategy, approximate the Nash equilibrium, be able to beat all the poker players in the world with it. So to actually see that come to fruition and be realized, that was kind of magical

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  17. I felt a lot of things, man. I mean, at that point in my life, I had spent five years working on this project. And it was a huge sense of accomplishment. I mean, to spend five years working on something and finally see it succeed. Yeah, I wouldn't trade that for anything in the world.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And then at that point, it started to look like the humans were coming back. They started to like, you know, but poker is a very high variance game. And I think what happened is they thought that they spotted some weaknesses that weren't actually there. And then around day eight, it was just very clear that they were getting absolutely crushed. And from that point, I mean, for a while there, I was super stressed out thinking like, oh my God, the humans are coming back and they've found weaknesses and now we're just going to lose the whole thing. But no, it ended up going in the other direction and the bot ended up like crushing them in the long run.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  19. That's one way of putting it. I wasn't super confident. So, you know, going in, like I said, I think I had like 50 50 odds on us winning when we actually, when we announced the competition, the poker community decided to gamble on who would win and their initial odds against us were like four to one. They were really convinced that the humans were going to pull out a win. The bot ended up winning for three days. And even then, after three days, the betting odds were still just 50 50.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  20. That's gutsy. Yeah, it was honestly, and I'm not sure why we did that in retrospect. But, I mean, I'm glad we did because we ended up winning anyway. But if you've ever played poker before, that is golden information. I mean, no. Usually, when you play poker, you see about a third of the hands to showdown. And to just hand them all the cards that the bot had on every single hand, that was just a goldmine for them. And so then they would review the hands and try to see could they find patterns in the bot? The weaknesses, and could they then they would coordinate and study together and try to figure out, okay, now this person's going to explore this part of the strategy for weaknesses. This person's going to explore this part of the strategy for weaknesses.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So we had set up the competition. Like I said, there were 200,000 dollars in prize money, and they would get paid a fraction of that depending on how well they did relative to each other. So I was kind of hoping that they wouldn't work together to try to find weaknesses in the bot, but they enter the competition with their like number one objective being to beat the bot. And they didn't care about like individual glory. They were like, we're all going to work as a team to try to take down this bot. And so they immediately started comparing notes. What they would do is they would coordinate looking at different parts of the strategy to try to find out weaknesses. And then at the end of the day, we actually sent them a log of all the hands that were played and what cards the bot had on each of those hands.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  22. They found is yeah, overbedding like bedding certain amounts, the bot would have a lot of trouble dealing with those sizes. And then also when the bot got into really difficult all in situations, it wasn't able to, because it wasn't doing search, it had to clump different hands together and it would treat them identically. And so it wouldn't be able to distinguish having a king high flush versus an ace high flush. And in some situations, that really matters a lot. And so they could put the bot into those situations and then the bot would just bleed money.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Super stressful. I mean, I thought going into it that we had like a 50 50 chance because basically I thought if they play in a totally normal style, I think we'll squeak out a win, but there's always a chance that they can find some weakness in the bot. And if they do. And we're playing like for 20 days, 120,000 hands of poker. They have a lot of time to find weaknesses in the system. And if they do, we're going to get crushed. And that's actually what happened in the previous competition. The humans, you know, they started out. It wasn't like they were winning from the start. But then they found these weaknesses that they could take advantage of. And for the next, you know, like 10 days, they were just crushing the bot, stealing money from it.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Are all sorts of optimizations that I had to make to try to get this thing to run as fast as possible? They were like, how do you minimize the latency? How do you package things together so that you minimize the amount of communication between the different nodes? How do you optimize the algorithm so that you can try to squeeze out more and more from the game that you're actually playing, all these kinds of different decisions that I had to make.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Yeah, and you know, talking about terabytes of memory. So it was a very parallelized, and it had to be very fast too, because the more games that you could simulate, the stronger the bot would be.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  26. So, one of the interesting things about Libratus is that we had no idea what the bar was to actually beat top humans. We could play against our prior bots, and that kind of gives us some sense of are we making progress? Are we going in the right direction? But we had no idea what the bar actually was. And so we threw a huge amount of resources at trying to make the strongest bot possible. So we use C++. It was parallelized. We were using, I think, like a thousand CPUs, maybe more actually. And today, that sounds like nothing, but for a grad student back in 2016, that was a huge amount of resources.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  27. The Libratus competition was super stressful for me. Also, I mean, I was working on this basically continuously for a year leading up to the competition. I mean, for me, it became very clear, like, okay, this is the search technique, this is the approach that we need. And then I spent a year working on this pretty much like nonstop.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Yeah, and I think it's a good step in the right direction. I would kind of say it's similar to Monte Carlo rollouts in a game like chess. There's a kind of search that you can do where you're saying like, I'm going to roll out my intuition and see without really thinking, what are the better decisions I can make farther down the path? What would I do if I just acted according to intuition for the next 10 moves? gets you an improvement, but I think that there's much richer kinds of planning that we could do.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Yeah, so you can kind of think of it as like neural nets today. Give you like Transformers, for example, are super general, but you know, they'll give you it'll output an answer in 100 milliseconds. And if you tell it, like, oh, you've got five minutes to give you a decision, you know, feel free to take more time to make a better decision. It's not going to know what to do with that. But a human, if you're playing a game like chess, they're going to give you a very different answer depending on if you say, oh, you got 100 milliseconds or you've got five minutes.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  30. I have thought a lot about that, and I think it's a really important question. So, the AI in AlphaGo or any of these Go AIs, they're all doing Monte Carlo tree search, which is a particular kind of search. And it's actually a symbolic tabular search. It uses the neural net to guide its search, but it isn't actually like full-on neural net. Now, that kind of search is very successful in these kinds of like perfect information board games like chess and Go. But if you take it to a game like poker, for example, it doesn't work. It can't understand the concept of hidden information. It doesn't understand the balance that you have to strike between the amount that you're raising versus the amount that you're calling. And in every one of these games, you see a different kind of search. And the human brain is able to plan for all these different games in a very general way. Now, I think that's one thing that we're missing from AI today. And I think it's a really important missing piece.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Yeah, and all these bots, they have what's called a policy network where it will tell you this is what the neural net thinks is the next best move. And it's kind of like the intuition that a human has. The human looks at the board and any Go or chess master will be able to tell you like, oh, instantly, here's what I think the right move is. And the bot is able to do the same thing. But just like how a human grandmaster can make a better decision if they have more time to think when you add on this Monte Carlo tree search, the bot is able to make a better decision.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  32. So, even today, what seven years after AlphaGo, if you take out the Monte Carloch research that's being done when playing against the human, the bots are not superhuman. Nobody has made a raw neural net that is superhuman in Go.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Able to beat top humans. I think a good example of this is you look at the latest versions of AlphaGo, like it was called Alpha Zero. And there's this metric called Elo rating, where you can compare different humans and you can compare bots to humans. Now, a top human player is around 3,600 elo, maybe a little bit higher now. Alpha Zero, the strongest version, is around 5200 elo. But if you take out the search that's being done at test time, and by the way, what I mean by search is the planning ahead, the thinking of like, oh, if I move my, if I place this stone here and then he does this, and then you look like five moves ahead and you see like what the board state looks like, that's what I mean by search. If you take out the search that's done during the game, the ELO rating drops to around 3,000.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yeah, I think a lot of people under this is true for the general public, and I think it's true for the AI community. A lot of people underestimate the importance of search for these kinds of game AI results. An example of this is TD Gammon that came out in 1992. This was the first real instance of a neural net being used in a game AI. It's a landmark achievement. It was actually the inspiration for AlphaZero. And it used search. It used to ply search to figure out its next move. You got deep blue there. It was very heavily focused on searching many, many moves ahead farther than any human could. And that was key for why it won. And then even with something like Alphago, I mean, AlphaGo is commonly hailed as a landmark achievement for neural nets, and it is. But there's also this huge component of search, Monte Carlo Tree Search to AlphaGo that was key, absolutely essential for the AI to be

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  35. I feel like there's a good chance that that is the case. Yeah. If you're able to see that it's like a decent proxy for score, right? And this is actually the common poker wisdom when they're teaching players before the robots and they were trying to teach people how to play poker. They would say like the key to the game is to put your opponent to difficult spots. It's a good estimate for if you're making the right decision.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  36. That's the case. So I think there is some debate about is it true for every single variant of poker? I think for every single variant of poker, if somebody really put in the effort, they can make an AI that would beat all humans at it. We've focused on the most popular variants. So heads up no limit text is hold them. And then we follow that up with six player poker as well, where we managed to make a bot that beat expert human players. And I think even there now, it's pretty clear that humans don't stand a chance.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  37. He was excited. And he honestly wanted to play against the bot. He thought he had a decent chance of beating it. Is like several years ago, and I think it was not as clear to everybody that the AIs were taking over. I think now people recognize that like if you're playing against a bot, there's like no chance that you have in a game like Pokemon.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Did actually have a conversation with Daniel DeGrandier once. Yeah, I was visiting the Isle of Man to talk to Poker Stars about AI. And Daniel Legrandi was there. We had dinner together with some other people. And yeah, he was really interested in it. He mentioned that he was like, you know, excited about learning from these AIs

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  39. From the bot's perspective, it's just doing the thing that's going to make it the most money, and the fact that it's putting the humans in a difficult spot, like that's just a side effect of that. And this was, I think, the one thing, I mean, there were a few things that the humans walked away from, but this was the number one thing that the humans walked away from the competition saying like we need to start doing this. And now these over bets, what are called over bets, have become really common in high-level poker play.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  40. That's possible given the board, and you're thinking, like, oh, you're in a really great spot here, and suddenly the bot bets $20,000 into a thousand dollar pot. And it's basically saying, I have the best hand or I'm bluffing. And you having the second best hand, like now you get a really tough choice to make. And so the humans would sometimes think like five or ten minutes about, what do you do? Should I call? Should I fold? And when I saw the humans like really struggling with that decision, like that's when I realized like, oh, actually, this is maybe a good thing to do after all.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  41. There was something really interesting that we observed with liberatus. So humans, when they're playing poker, they usually size their bets relative to the size of the pot. So, you know, if the pot has $100 in there, maybe you bet like $75 or somewhere around it, somewhere between like $50 and $100. And with libratus, we gave it the option to basically bet whatever it wanted. It was actually really easy for us to say like, oh, if you want, you can bet like 10 times the pot. And we didn't think it would actually do that. It was just like, why not give it the option? And then during the competition, it actually started doing this. And by the way, this is like a very last minute decision on our part to add this option. And so we did not think the bot would do this. And I was actually kind of worried when it did start to do this. Like, oh, is this a problem? Like humans don't do this. Is it screwing up? But it would put the humans into really difficult spots when it would do that because, you know, you could imagine like you have the second best hand.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Right, I'm trying to think like I'm gonna bet, and then what are you gonna do in response? Are you gonna raise me? Are you gonna call? And then if you raise, what should I do? So it's reasoning about that whole process up until the end of the game in the case of Labratis.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  43. There's four rounds of poker. So there's the pre flop, the flop, the turn, and the river, and so we would start doing search halfway through the game. Now, the first half of the game, that was all pre-computed. It would just act instantly. And then when it got to the halfway point, then it would always search to the end of the game. Now, we later improved this, so we wouldn't have to search all the way to the end of the game. It would actually search just a few moves ahead. But that came later. drastically reduced the amount of computational resources that we needed.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Well, it's kind of both because you're trying to maximize the value for the hand that you have, but in the process, in order to maximize the value of the hand that you have, you have to figure out what would I be doing with all these other hands as well.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Yeah, so you kind of, where the train system comes in is the value at the end. So you only look so far ahead. You look like maybe one round ahead. So if you're on the flop, you're looking to the start of the turn. And at that point, you can use the pre-computed solution to figure out what's the value here of this strategy.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  46. No, it's all the other hands that you could have. So when you're playing no limit, Texas, hold them, you've got two face down cards. And so that's 52, choose two, 1,326 different combinations. Now, that's actually a little bit lower because there's face up cards in the middle, and so you can eliminate those as well. You're looking at around a thousand different possible hands that you can have. And so when we're doing search, it's thinking explicitly there are these thousand different hands that I could have, there are these thousand different hands that you could have. Let me try to figure out what would be a better strategy than what I've pre-computed for these hands and your hands.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Yeah, that's what I mean by search of acting instantly, a neural net usually gives you a response in like 100 milliseconds or something. It depends on the size of the net. But if you can leverage extra computational resources, you can possibly get a much better outcome. And we did some experiments in small scale versions of poker. And what we found was that if you... Do a little bit of search, even just a little bit. It was the equivalent of making your pre-computed strategy. You can kind of think of it as your neural net a thousand times bigger with just a little bit of search. And it just blew away all of the research that we had been working on and trying to scale up this pre-computed solution. It was dwarfed by the benefit that we got from search.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  48. No, no, I'm not saying that there were any timing tails. I was saying when the human, like the bot would always act instantly. It wouldn't try to come up with a better strategy in real time over what it had pre-computed during training. Whereas the human, like they have all this intuition about how to play. They're also in real time leveraging their ability to think, to search, to plan, and coming up with an even better strategy than what their intuition would say.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So 2015. And we actually learned a lot from that competition. And in particular, what became clear to me is that the way the humans were approaching the game was very different from how the bot was approaching the game. The bot would not be doing search. It would just be trying to compute, you know, it would do like months of self-play. It would just be playing against itself for months. But then when it's actually playing the game, it would just act instantly. And the humans, when they're in a tough spot, they would sit there and think for sometimes even like five minutes about whether they're going to call or fold a hand. And it became clear to me that there's a good chance that that's what's missing from our bot. So I actually did some initial experiments to try to figure out how much of a difference this does actually make. And the difference was huge.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  50. That's right. Yeah. So what we're actually trying to maximize is the expected value given that the opponent is playing optimally in response to us. Now, in practice, what that ends up looking like is it's putting the opponent into difficult situations where there's no. Obvious decision to be made.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source