YouSaid · the spoken record

Noam Brown

lines on the record
196
first
2022-12-06
most recent
2022-12-06
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I think one of the challenges in life is figuring out exactly what that reward function is. Sometimes it's pretty hard to specify the same way that trying to handcraft the optimal policy in a game like chess is really difficult. It's not so clear cut what the reward function is for life.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  2. There's a lot of talk today about AI safety circles about misspecification of reward functions. So you say, okay, my objective is to be rich. And maybe the AI tells you like, okay, well, if you want to maximize the probability that you're rich, go rob a bank. And so you want to, is that really what you want? Is your objective really to be rich at all costs, or is it more nuanced than that?

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  3. What is it? Like in poker and diplomacy, you need a value function. You need to have a reward system. And so what does it mean to live a life that's optimal?

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah, I would say build a strong foundation in math and computer science and statistics and these kinds of areas, but don't be afraid to try something that's different and learn something that's different from the thing that everybody else is doing to get into machine learning. There's value in having a different background than everybody else. Yeah, so but certainly having a strong math background, especially in things like linear algebra and statistics and probability, are incredibly helpful today for learning about and understanding machine learning.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  5. I think that there's a lot of challenges left. And I think having a diversity of viewpoints and backgrounds is really helpful for working together to figure out how to tackle those kinds of challenges.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I would say. That there are a lot of people working on similar aspects of machine learning, and to not be afraid to try something a bit different. My own path in AI is pretty atypical for machine learning researcher today. I mean, I started out working on game theory and then shifting more towards reinforcement learning as time went on. And that actually had a lot of benefits, I think, because it allowed me to look at these problems in a very different way from the way a lot of machine learning researchers view it. And that comes with drawbacks in some respects. Like I think there's definitely aspects of machine learning where I'm weaker than most of the researchers out there. But I think that diversity of perspective, when I'm working with my teammates, there's something that I'm bringing to the table and there's something that they're bringing to the table. And that kind of collaboration becomes very fruitful for that reason.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I think there's some truth to that. I mean, you look at the way humans approach a game like poker. They're not coming at it from scratch. They're coming at it with a huge amount of background knowledge about how humans work, how the world works, the idea of money. So they're able to leverage that kind of information to pick up the game faster. So it's not really a fair comparison to then compare it to an AI that's learning from scratch. And maybe one of the ways that we address this sample complexity problem is by allowing AIs to leverage that general knowledge across a ton of different domains.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I mean, that's the trillion dollar question in AI today. I mean, if you can figure out how to make AI systems more data efficient, then that's a huge breakthrough. So nobody really knows right now.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Deploy AI systems in real world settings where they're interacting with humans because, you know, for example, with robotics, it's really hard to generate a huge number of samples. It's a different story when you're working in these totally virtual games where you can play a million games and it's no big deal

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Now, there are aspects of AI that I think are still lacking. I think there's general agreements that one of the major issues with AI today is that it's very data inefficient. It requires a huge number of samples of training examples to be able to train. You look at an AI that plays Go and it needs millions of games of Go to learn how to play the game well. Whereas a human can pick it up and like, you know, I don't know how many games does a human go player go Grandmaster play in their lifetime probably in the thousands or tens of thousands I guess So that's one issue overcoming this challenge of data efficiency. And this is particularly important if we want to

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  11. There's a lot of people trying to figure that out right now. And I should say the amount of progress that's been made, especially in the past few years, is truly phenomenal. I mean, you look at where AI was 10 years ago and the idea that you can have AIs that can generate language and generate images the way they're doing today and able to play a game like diplomacy was just unthinkable. Even five years ago, let alone 10 years ago.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  12. There's also the issue of in these diplomacy experiments in order to do Fair comparison. What we found is that there's an inherent anti AI bias in these kinds of games. So we actually played tournaments in a non-language version of the game where we told the participants, hey, in every single game, there's going to be an AI. And what we found is that the humans would spend basically the entire game like trying to figure out who the bot was. And then as soon as they thought they figured it out, they would all team up and try to kill it. you know overcoming that inherent anti-AI bias is a challenge

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I was getting like, yeah, that's kind of going to the question of what is a lot. You know, is a white lie a bad lie? Is it an ethical lie? Those kinds of questions.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Some things that we've already talked about, I think specific to diplomacy, there's also the challenge that the game is, you know, there is a deception aspect to the game. And so Developing language models that are capable of deception is, I think, a dicey issue and something that makes research on diplomacy particularly challenging. Know so those kinds of issues should we even be developing AIs that are capable of lying to people? That's something that we have to think carefully about.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  15. But if you have these AIs that are playing, you know, you can say, Oh, I'm a 2000 elo human, how do I get to 2200? Now you can have an AI that plays in the style of a 2200 elo human, and that will help you get better. Or you mentioned this problem of like, how do you know that you're actually playing with humans when you're playing online in video games? Well, now we have the potential of populating these virtual worlds with agents like AI agents that are actually fun to play with and you don't have to always be playing with other humans to have a fun time. So yeah, a lot of upside potential too. And I think, you know, with any sort of tool, there's the potential for a lot of greatness and a lot of downsides as well.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Yeah, I think there is a lot of negative potential for this kind of technology. But at the same time, there's a lot of upside for it as well. So for example, right now, it's really hard to learn how to get better in games like chess and poker and Go because the way that the AI plays is so foreign and incomprehensible.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  17. It does. Yeah. The way that cheat detection works in a game like poker and a game like chess and go, from what I understand, is trying to see like, is this person making moves that are very common among chess AIs or AIs in general, but very uncommon among top human players? And if you have the development of these AIs that play in a very strong style, but also a very human-like style, then that poses serious challenges for cheat detection.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yeah, I think so. And so this is where the research gets interesting. One of the things that I was thinking about is, and this is actually already being done. There's a researcher at the University of Toronto that's working on this, is to make an AI that plays in the style of a particular player, like Magnus Carlson, for example. You can make an AI that plays like Magnus Carlson. And then where I think this gets interesting is like, hey, maybe you're up against Magnus Carlson in the World Championship or something. You can play against this Magnus Carlson bot to prepare against the real Magnus Carlson and you can try to explore strategies that he might struggle with and try to figure out like how do you beat this player in particular. On the other hand, you can also have Magnus Carlson working with this bot to try to figure out where he's weak and where he needs to improve his strategy. And so I can envision this future where data on specific chests and Go players becomes extremely valuable.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  19. And on the other hand, you can leverage search and planning very heavily. But then what you end up with is an AI that plays in a very different style from how humans play the game. Now, if you strike this intermediate balance by setting the regularization parameters correctly and say you can do planning, but try to keep it close to the human policy, then you end up with an AI that plays in both a very human-like style and a very strong style. And you can actually even tune it to have a certain ELO rating. So you can say playing the style of like a 2800 Elo human.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  20. To elaborate on this a bit, one way to approach making a human-like AI for chess is to collect a bunch of human games, like a bunch of human Grand Master games, and just to supervise learning on those games. But the problem is that if you do that, what you end up with is an AI that substantially weaker than the human grandmasters that you trained on. Because the neural net is not able to approximate the nuance of the strategy. This goes back to the planning thing that I mentioned, the search thing that I talked about before, that these human grandmasters, when they're playing, they're using search and they're using planning, and the neural net alone, unless you have a massive neural net that's like a thousand times bigger than what we have right now, it's not able to approximate those details very effectively.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Yeah, absolutely. We've already started looking into this direction a bit, so we tried to use the techniques that we've developed for diplomacy to make chess and go AIs. And what we found is that it led to much more human-like strong chess and go players. The way that AIs like Stockfish today play is in a very inhuman style. It's very strong, but it's very different from how humans play. And so we can take the techniques that we've developed for diplomacy. We do something similar in chess and go. And we end up with the bot that's both strong and human-like.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  22. And that's why I do think that diplomacy is taking a big step closer to the real world than anything that's came before in terms of game AI breakthroughs. The fact that No, we're communicating in natural language. We're leveraging the fact that we have this like general data set of dialogue and communication from a breadth of the internet. That is a big step in that direction. We're not 100% there, but we're getting closer at least.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  23. I feel like this isn't a question that's unique to diplomacy. I mean, I think you look at RL breakthroughs, reinforcement learning breakthroughs in previous games as well, like AI for StarCraft, AI for Atari, you haven't really seen it deployed in the real world because you have these problems of it's really hard to collect a lot of data and you don't have a well-defined action space. You don't have a well-defined reward function. These are all things that you really need for reinforcement learning and planning to be really successful today. Now there are some domains where you do have that code generation is one example. Theorem proving mathematics, that's another example where you have a well-defined action space. You have a well-defined reward function. And those are the kinds of domains where I can see RL in the short term being incredibly powerful. But yeah, I think that those are the barriers to deploying this at scale in the real world.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Well, like I said, the original motivation for the game of diplomacy was the failures of World War I, the diplomatic failures that led to war. And the real take-home message of diplomacy is that if people approached diplomacy the right way, then war is ultimately unsuccessful. The way that I see it, war is an inherently negative sum game, right? There's always a better outcome than war for all the parties involved. My hope is that as AI progresses, then maybe this technology could be used to help people make better decisions across the board and hopefully avoid negative some outcomes like war.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  25. I think the more human data you have, the better. And I think that that's going to be the major bottleneck in scaling to more complicated domains. But that said, there might be the potential, just like in the language model where we leveraged tons of data on the internet and then specialized it for diplomacy. There is the future potential that you can leverage huge amounts of data across the board and then specialize it in the data set that you have for diplomacy. And that way you're essentially augmenting the amount of data that you have.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  26. So, you basically say, try to maximize your expected value while at the same time stay as close as possible to the human policy. And there is a parameter that controls the relative weighting of those competing objectives.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  27. So we train a neural nets to imitate the human data as closely as possible. And that's what we call the anchor policy. And now when we're doing self-play, the problem with the anchor policy is that it's not a perfect approximation of how humans actually play. Because we don't have infinite data, because we don't have unlimited neural network capacity, it's actually a relatively suboptimal approximation of how humans actually play. And we can improve that approximation by adding planning and RL. And so what we do is we get a better approximation, a better model of human play by during the self-play process, we say you can deviate from this human anchor policy if there is an action that has particularly high expected value, but it would have to be a really high expected value.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  28. The way that we've approached self play in diplomacy is like we're trying to come up with good intents to condition the language model on. And the space of intents is actions that can be played in the game. Now, there is like the potential to have a broader set of intents. Things like, you know, long-term cooperation or long-term objectives or gossip about what another player was saying. These are things that we're currently not conditioning the language model on. And so it's not able to, we're not able to control it to say like, oh, you should be talking about this thing right now, but it's quite possible that you could expand the scope of intense to be able to allow it to talk about those things. Now, in the process of doing that, the self-play would become much more complicated. And so that is a potential for future work.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Way that we've approached AI and diplomacy is you condition the language on an intent. Now that intent in diplomacy is an action, but it doesn't have to be. And you can imagine you could have NPCs in video games or the metaverse or whatever, where there is some intent or there's some objective that they're trying to maximize and you can specify what that is. And then the language can correspond to that intent. Now, I'm not saying that this is happening imminently, but I'm saying that this is a future application potentially of this direction of research.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Think one of the things that's really useful about diplomacy is that we have a well defined value function. There is a well-defined score that the bot is trying to optimize. And in a setting, like a general chatbot setting, it would need that kind of objective in order to fully leverage the techniques that we've developed.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  31. I think there's a few different directions to take this work. I think really what is showing us is the potential that language models have. I mean, I think a lot of people didn't think that this kind of result was possible even today, despite all the progress that's been made in language models. And so it shows us how we can leverage the power of things like self-play on top of language models to get increasingly better performance. And the ceiling is really much higher than what we have right now.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  32. I think you could do that. That was not something that we were focused on, but I think that it is possible that if you came up with some measurements of what does it mean to tell a lie, because there's a spectrum, right? Like if you're withholding some information, is that a lie? If you're mostly telling the truth, but you forgot to mention this one action out of like 10, is that a lie? It's hard to draw the line. But if you're willing to do that, and then you could possibly use it to.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  33. The bot doesn't explicitly try to calculate whether somebody is lying or not. But what it will do is try to predict what actions they're going to take given the communications, given the messages that they've sent to us. So given our conversation, what do I think you're going to do? And implicitly, there is a calculation about whether you're lying to me in that. Based on your messages, if I think you're going to attack me this turn, even though your messages say that you're not, then essentially the bot is predicting that you're lying. But it doesn't view it as lying the same way that we would view it as lying.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  34. I think the fact that we were able to control the language model through this intense approach was very effective. And it allowed us, instead of just imitating how humans would communicate, we're able to go beyond that and able to feed into it superhuman strategies that it can then generate messages corresponding to.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Well, I think there's a few insights. So, first of all, the fact that you can't rely purely or even largely on self-play, that you really have to have an understanding of how humans approach the game, I think that that's one of the major conclusions that I'm drawing from this work and that is, I think, applicable more broadly to a lot of different games. So we've actually already taken the approaches that we've used in diplomacy and tried them on a cooperative card game called Hanabi. And we've had a lot of success in that game as well. On the language side,

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Yeah, absolutely. And so for the language model, which is kind of like a separate question, we didn't use the language model during self-play training, but we pre-trained the language model on tons of internet data as much as possible. And then we fine-tuned it specifically on the diplomacy games. So we are able to leverage the wider data set in order to fill in some of the gaps in how communication happens more broadly besides just specifically in these diplomacy games.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah, so we train a bot through supervised learning to model the human play as much as possible. So we basically train a neural net on those 50,000 games. And that gives us a policy that resembles to some extent how humans actually play the game. Now, this isn't a perfect model of human play because we don't have unlimited data. We don't have unlimited neural net capacity. But it gives us some approximation.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  38. That's actually one of the major challenges that we faced in the research. That we had a good amount of human data. We had about 50,000 games. What we try to do is leverage as much self-play as possible while still leveraging the human data. So what we do is we do self-play, very similar to how it's been done in poker and Go, but we try to regularize the self-play towards the human data. Basically, the way to think about it is... We penalize the bot for choosing actions that are very unlikely under the human data set.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Yeah, exactly. And so that is something that seems almost impossible to model purely from scratch without any human data. It's a very cultural thing. And so you need human data to be able to understand that hey, that's how humans behave. And you have to work around that. It might be suboptimal. It might be rational, but that's an aspect of humanity that you have to deal with.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Stop them from taking over And the bot will do this. The bot will work with the other players to stop the superpower from winning. But if it doesn't really, if it's trained from scratch or it doesn't really have a good grounding in how humans approach it, it will also at the same time attack the other players with its extra units. So all the units that are not necessary to stop the superpower from winning, it will use those to grab as many centers as possible from the other players. In totally rational play, the other players should just live with that. They have to understand like, hey, a score of one is better than a score of zero. So, okay, he's grabbed my centers, but I'll just deal with it. But humans don't act that way, right? The human gets really angry at the bot and ends up throwing the game because, you know, I'm going to screw you over because you did something that's not fair to me.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  41. And to give you an example of how this suboptimality and irrationality comes into play, there's a really common situation in the game of diplomacy where one player starts to win and they're like at the point where they're controlling about half the map. And the remaining players, you have all been fighting each other the whole game all have to work together now to stop this other player from winning or else everybody's going to lose. And it's kind of like Game of Thrones. I don't know if you've seen the show where you got the others coming from the north and like all the people have to start work out the differences.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Is an aspect of two players you're a sum versus games that involve cooperation. So in a two player zero sum game, you can do self-play from scratch and you will arrive at the dash equilibrium where you don't have to worry about the other player playing in a very human suboptimal style. That's just going to be the only way that deviating from an ash equilibrium would change things is if it helped you.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Well, it would, for example, expect the human to support it in a certain way when the human would simply like think like, no, I'm not supposed to support you here. It's kind of like, you know, if you develop a self-driving car and it's trained completely from scratch with other self-driving cars, it might learn to drive on the left side of the road. That's a totally reasonable thing to do if you're with these other self-driving cars that are also driving on the left side of the road. But if you put it in an American city, it's going to crash.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Well, I think Asking kind of two questions there. So, one like modeling the irrationality and the suboptimality of humans. You can't, in diplomacy, you can't treat all the other players like their machines. And if you do that, you're going to end up playing really poorly. And so we actually ran this experiment. So we trained a bot in a two-player zero sum version of diplomacy, the same way that you might approach a game like chess or poker. And the bot was superhuman. It would crush any competitor. And then we took that same training approach and we trained a bot for the full seven player version of the game through self-play without any human data. And we stuck it in a game with six humans and it got destroyed. Even in the version of the game where there's no explicit natural language communication, it still got destroyed because it just wouldn't be able to understand how the other players were approaching the game and be able to work with that.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Yeah, there were definitely things that the bot would do that were not in line with how humans would approach the game. And in a good way, the humans actually, you know, we've talked with some expert diplomacy players about these results and their takeaway is that, well, maybe humans are approaching this the wrong way. And this is actually like the right way to play the game.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Politely, you know, like keep it in character. We actually had a researcher watching the bot 24 7 whenever we play a game. We had a bot watching it to make sure that it wouldn't go off the rails and start like threatening somebody or something like that.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Yeah, I think there's a few aspects of the results that I think are really exciting. So, first of all, the fact that we were able to achieve such strong performance, I was surprised by and pleasantly surprised by. So we played 40 games of diplomacy with real humans. And the bot placed second out of all players that have played five or more games. So it's about 80 players total, 19 of whom played five or more games, and the bot was ranked second out of those players. And the bot was really good in two dimensions. One, being able to establish strong connections with the other players on the board, being able to persuade them to work with it, being able to coordinate with them about how it's going to work with them. And then also the raw tactical and strategic aspect of the game, you know, being able to understand what the other players are likely to do, being able to model their behavior.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Yeah, absolutely. And I think it's. Think it's maybe the best data set that I can think of out there to investigate these kinds of questions of negotiation, trust, persuasion. I wouldn't say it's the best data set in the world for human AI interaction. That's a very broad field, but I think that it's definitely up there as like, you know, if you're really interested in language models interacting with humans, in a setting where their incentives are not fully aligned, this seems like an ideal data set for investigating that.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  49. The data set comes from this website webdiplomacy.net is the site that's been online for like 20 years now. And it's one of the main sites that people use to play diplomacy on it. We've got like 50,000 games of diplomacy with natural language communication, over 10 million messages. So it's a pretty massive data set that people can use to, we're hoping that the academic community, the research community is able to use it for all sorts of interesting research questions.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Not just even the data of the AI playing with the humans, but all the training data that we had to train the AI to understand how humans play the game. We're setting up a system where researchers will be able to apply to be able to gain access to that data and be able to use it in their own research.

    2022-12-06 · Lex Fridman Podcast · #344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation · IDENTIFIED FROM THE TRANSCRIPT · source