YouSaid · the spoken record
Tuomas Sandholm
- lines on the record
- 72
- first
- 2018-12-28
- most recent
- 2018-12-28
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Yeah, it is explicitly modeled, but it's not assumed. The beliefs are actually output, not input. Of course, the starting beliefs are input, but they just fall from the rules of the game because we know that the dealer deals uniformly from the deck. So I know that every pair of cards You might have is equally likely. I know that for a fact. That just follows from the rules of the game. Of course, except the two cards that I have. I know you don't have those. You have to take that into account. That's called card removal, and that's very important.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes. So we're talking about exactly that. much like you do in alpha beta search or montra research but with different techniques so there's a different search algorithm and then we have to deal with the leaves differently so if you think about what liberados did we didn't have to worry about this because we only did it at the end of the game so we would always terminate into a real situation and we would know what the payout is it didn't do this depth limited look aheads but now in this new paper which is called depth limited I think it's called depth limited search for imperfect information games, we can actually do sound depth limited lock ahead. So we can actually start to do the lock ahead from the beginning of the game on because that's too complicated to do for this whole long game. So in libratos we were just doing it for the end.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“They were randomly generating various situations in the game, then they were doing the look ahead from there to the end of the game, as if that was the start of a different game. And then they were using deep learning to learn those values of those states, but the states were not just the physical states. They include belief distributions.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“And that is a different way of doing it. And it doesn't assume, therefore, a particular way that the opponent plays, but it allows the opponent to choose from a set of different continuation strategies. And that forces us to not be too optimistic in our look ahead search. And that's one way you can do sound look ahead search in imperfect information games, which is very difficult. And you were asking about DeepStack. What they did, it was very different than what we do, either in liberators or in this new work.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“So the value of a state is not just a function of the cards. It depends on, if you will, the path of play, but only to the extent that it's captured in the belief distributions. So that's why it's not as simple as it is imperfect information games. And I don't want to say it's simple there either. Computationally there too, but at least conceptually it's very straightforward. There's a state, there's an evaluation function, you can try to learn it. you have to do something more. And what we do is in one of these papers, we're looking at where we allow the opponent to actually take different strategies at the leaf of the search tree, if you will.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. So, as you said, Liberatos did not use learning methods. But in imperfect information games, unlike, let's say, in Go or now also in Chess and Shogi, it's not sufficient to learn an evaluation for a state because the Value of an information set depends not only on the exact state, but it also depends on both players' beliefs. Like if I have a bad hand, I'm much better off if the opponent thinks I have a good hand and vice versa, if I have a good hand, I'm much better off if the opponent believes I have a bad hand.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“That's a good question. There's actually, I heard this story that there's this Norwegian female poker player called Annette Oberstad, who's actually won a tournament by doing exactly that. But that would be extremely rare. So you cannot really play well that way.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so that's why you have to play a lot of hands so that the role of luck gets smaller. So you could otherwise get lucky and get some good hands and then you're going to win the match. Even with thousands of hands, you can get lucky because there's so much variance in no limit Texas Holden. Because if we both go all in, it's a huge stack of variance. So there are these massive swings in Nolimitx Holder. So that's why you have to play not just thousands, but over 100,000 hands to get statistical significance.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, yourself and other players. And for information abstraction, we were completely automated. So these are algorithms. They do what we call potential aware abstraction where we don't just look at the value of the hand, but also how it might materialize in the good or bad hands over time. And it's a certain kind of bottom-up process with integer programming there and clustering and various aspects. build this abstraction. And then in the action abstraction, there it's largely based on how humans and other AIs have played this game in the past, but in the beginning, we actually used an automated action abstraction technology, which is provably convergent. It finds the optimal combination of bed sizes, but it's not very scalable, so we couldn't use it for the whole game, but we used it for the first couple of betting actions.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Metting actions, yeah. So there's information abstraction to talk about general games. Information abstraction, which is the abstraction of what chance does. And this would be the cards in the case of poker. And then there's action abstraction, which is abstracting the actions of the actual players, which would be bits in the case of poker.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. You're exactly right. So when you go from a game tree that's 10 to the 161, especially in an imperfect information game, it's way too large to solve directly, even with our fastest equilibrium finding algorithms. So you want to abstract it first. And abstraction in games is much trickier than abstraction in MDPs or other single agent settings. Because you have this abstraction pathologies that if I have a finer grained abstraction, Strategy that I can get from that for the real game might actually be worse than the strategy I can get from a coarse grained abstraction. So you have to be very careful.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Worth it for them to invest a lot of effort trying to find tellers in each other because they're so good at hiding them. So yes, at the kind of Friday evening game, Telas are going to be a huge thing. You can read other people. And if you're a good reader, you'll read them like an open book. But at the top levels of poker now, the Teles become a much smaller and smaller aspect of the game as you go to the top levels.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. So I'll split it into two parts. So one is why do humans trust humans more than AI and all have overconfidence in humans? I think that's not really related to the tell question. It's just that they've seen these top players how good they are and they're really fantastic. So it's just hard to believe. Yeah, that the Nayak would beat them. So I think that's where that comes from. And that's actually maybe a more general lesson about AI that until you've seen it overperform a human, it's hard to believe that it could. But then the tells. A lot of these top players, they're so good at hiding tails that among the top players it's actually not really”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, that's great. Great question. So 18 months earlier, I had organized a similar brains versus AI competition with our previous AI call Cloudical, and we couldn't beat the humans. So this time around, it was only 18 months later, and I knew that this new AI liberatus was way stronger, but it's hard to say how you'll do against the top humans before you try. So I thought we had about a 50-50 shot. And the international betting sites put us as a 4 to 1 or 5 to 1 underdog. So it's kind of interesting”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, it didn't really matter that they wouldn't have forgotten anyway these are top quality people, but we just wanted to put out there so it's not a question of the human forgetting and the AI somehow trying to get advantage of better memory.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, they were playing much like they normally do, these top players when they play this game. They play mostly online. So they used to playing through UI. And they did the same thing here. So there was this layout. You could imagine there's a table. On a screen, the human sitting there, and then there's the AI sitting there, and the screen shows everything that's happening, the cards coming out and shows the bets being made. And we also had the betting history for the human. So if the human forgot what had happened in the hand so far, they could actually reference back and so forth.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Right, no, we can't make any money. So originally, a couple of years earlier, I actually explored whether we could actually play for money, because that would be, of course, interesting as well to play against the top people for money. But the Pennsylvania Gaming Board said no. So we couldn't. So this is much like an exhibit, like for a musician or a boxer or something like that.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“So they had an incentive to play as hard as they could, whether they're way ahead or way behind or right at the mark of beating the AI.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so the event was that we invited four of the top ten players with these are specialist players in Heads Up Knowle Limit, Texas Holland, which is very important because this game is actually quite different than the multiplayer version. We brought them in to Pittsburgh to play at the Reverse Casino for 20 days. We wanted to get 120,000 hands in because we wanted to get statistical significance. So it's a lot of hands for humans to play, even for these top pros who play fairly quickly normally. So we couldn't just have one of them play so many hands. 20 days they were playing basically morning to evening. And I raised 200,000 as a little incentive for them to play. And the setting was so that they didn't all get 50,000. We actually paid them out based on how they did against the AI is.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Belong to you Yeah. So, as you said, you first get two cards in private each, and then there's a betting round. Then you get three cards in public on the table, then there's a betting round. Then you get the fourth card in public on the table, there's a betting round. Then you get the fifth card on the table, there's a betting round. So there's a total of four betting rounds and four tranches of information revelation, if you will. Only the first tranche is private. And then it's public from there.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“while heads up means it's two players so it's really like me against you and am I better or are you better much like chess or or go in that sense but an imperfect information game which makes it much harder because I have to deal with issues of you knowing things that I don't know and I know things that you don't know instead of pieces being nicely laid on the board for both of us to see.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, happy to. So head up Nolimit Texas Holderman has really emerged in the AI community as a main benchmark for testing these application independent algorithms for imperfect information game solving. And this is a game that's actually played by humans. You don't see that much on TV or casinos because, well, for various reasons, but you do see it in some expert level casinos and you see it in the best poker movies of all time. It's actually an event in the World Series of Poker, but mostly it's played online. And typically for pretty big sums of money, and this is a game that usually only experts play. So if you go to your home game on a Friday night, it probably is not going to be heads up no limit toxics as holding. It might be No Limit Texas Holdem in some cases, but typically for a big group and it's not as competitive.”
2018-12-28 · Lex Fridman Podcast · Tuomas Sandholm: Poker and Game Theory · IDENTIFIED FROM THE TRANSCRIPT · source