YouSaid · the spoken record

Paul Christiano

lines on the record
251
first
2023-10-31
most recent
2023-10-31
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Yeah, so like if you imagine you want to train a system to like say words that sound like the next word a human would say there you can get this like really rich supervision by having a bunch of words and then predicting the next one and being like I'm just tweak the model so it predicts better If you're like, hey, here's what I want. I want my model to interact with some job over the course of a month. And then at the end of that month, like have internalized everything where the human would have internalized about how to do that job well and like have local context and so on. It's harder to supervise that task.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So we've talked a little bit about how good is the qualitative extrapolation? How good are people at comparing? So this is not like the picture of being qualitative wrong. This is just quantitatively. It's very hard to know how far off you are. I think a qualitative consideration that could significantly slow things down is just like right now you get to observe this like really rich supervision from like basically next word prediction or like in practice maybe you're looking at like a couple sentences prediction so getting this like pretty rich supervision it's plausible that if you want to like automate long horizon tasks like being an employee over the course of a month that that's actually just like considerably harder to supervise or that like you basically end up driving costs like the worst case here is that you like drive up costs by a factor that's like linear in the horizon over which the thing is operating and i still consider that just like quite plausible

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I think things take longer than you think. It's like a real thing. Yeah, I don't know. Mostly I have big error bars because I just don't believe the subjective extrapolation that much. I find it hard to get a huge amount out of it.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  4. And then the second half is the chimp human difference is probably pretty small compared to model differences. So I do think things are going to be pretty abrupt. I think the economy is pretty rough. I also think the subjective extrapolation is like pretty rough just because I really don't know how to get like how I don't know how people do the extrapolation end up with degrees of confidence people end up with. Again, I'm putting pretty high, if I'm saying like, you know, give me three years and I'm like, yeah, 50-50, it's going to have basically the smarts there to do the thing. That's not saying it's a really long way off. I'm just saying like I got pretty big error bars and I think that like it's really hard not to have really big error bars when you're doing this like I looked at GPT-4 it seemed pretty smart compared to GPT 3.5 so I bet just like four more such notches and were there it's like that's just a hard call to make I think I sympathize more with people who are like how could it not happen in three years than with people who are like no way it's going to happen in eight years or whatever which is like probably a more common perspective in the world but also things

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Yeah, there's sort of two balancing tensions here. One is like, I don't believe the chimp thing is going to be as abrupt. That is, I think, if you scaled up from chimps to humans, you actually see quite large economic value from the fully domesticated chimp already.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Over one doubling of brain size or whatever, one order of magnitude of brain size, it's actually possible in order of magnitude of brain size. But chimps are very, chimps are already within order of magnitude of brain size of humans. Chimps are very, very close on the kind of spectrum we're talking about. So I think I'm skeptical of the abrupt transition for chimps. And to the extent that I kind of expect a fairly abrupt transition here, it's mostly just because the chimp human intelligence difference is so small compared to the differences we're talking about with respect to these models that is I would not be surprised if in some objective sense like chimp human difference is like significantly smaller than the GPT-3 GPT-4 difference or the GPT-4 GPT5 difference.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I think that the economic extrapolation is not great. I think it's like you could compare it to this objective extrapolation of how smart does the model seem. It's not super clear which one's better. I think probably in the chimp case, I don't think that's quite right. I think if you actually, so if you imagine intensely domesticated chimps who are just actually trying their best to be really useful employees and you hold fix their physical hardware and then you just gradually like scale up their intelligence, I don't think you're going to see zero value, which then suddenly becomes massive value.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I think there's some just how capable is each model where I have, I think we're really bad extrapolating, but we still have some subjective guests, and you're comparing it to what happened, and that will move me every time we see what happens with another order of magnitude of training compute. I will have a slightly different guess for where things are going. These probabilities are coarse enough that, again, I don't know if that 40% is real or if like post-GBT 3.5 and 4, I should be at 60% or what. That's one thing. And the second thing is just some, if there was some ability to extrapolate, I think this could reduce error bars a lot. I think Here's another way you could try and do an extrapolation is you could just say, how much economic value do systems produce and like how fast does that growing? I think once you have systems actually doing jobs, the extrapolation gets easier because you're not moving from a subjective impression of a chat to automating all R&D or moving from automating this job to automating that job or whatever. Unfortunately, that's probably by the time you have a nice trends from that, you're not talking about 2040, you're talking about two years from the end of days or one year from the end of.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  9. I'm happy with that. It's going to get smarter, and it'll be really weird if it didn't. And the question just how smart does it have to get? That argument does not yet give us a quantitative guide to at what scale is it? Is it a slam donk or at what scale is a 50-50?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So, again, here we're talking two factors of a half one on like, is it smart enough and one on, do you have to do a bunch of SLAP, even if in some sense it's smart enough? And like the first factor of a half, I'd be like, I don't know, I think we have really anything good to extrapolate. That is like, I feel, I would not be surprised if I have similar or maybe even higher probabilities on a really crazy stuff over the next year. And then my probability is not that bunched up. Maybe Dara's probability, I don't know. You'd have talked with him is like you have talked with him. It's more bunched up on some particular year. And mine is maybe a little bit more uniformly spread out across the coming years, partly because I'm just like, I don't think we have some trends we can extrapolate, like can extrapolate loss. You can look at your qualitative impressions of systems at various scales. But it's just very hard to relate any of those extrapolations to doing cognitive work or accelerating R&D or taking over and fully automating R&D. So I have a lot of uncertainty around that extrapolation. I think it's very easy to get down to a 50%.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  11. You're always going to have to do some work, but the work's not necessarily much. I would guess when people say new insight is needed, I think I tend to be more bullish than them. I'm not like these are new ideas where who knows how long it will take. I think it's just like you have to do some stuff. You have to make changes unsurprisingly. Every time you scale something up by like five orders of magnitude, you have to make some changes.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  12. I think if you just literally say, like, you train on web text, then the question is kind of hard to discuss because you, like, I don't really buy stories that training data makes a big difference long run to these dynamics. But I think if you want to just imagine the hypothetical, like you just took GPT-4 and made the numbers bigger, then I think those are pretty significant issues. I think there's significant issues in two ways. One is quantity of data, and I think probably the larger one is quality of data, where like, I think as you start approaching, the prediction task is not that great a task. If you're a very weak model, it's a very good signal we get smarter. At some point, it becomes like a worse and worse signal to get smarter. I think there's a number of reasons. It's not clear there's any number such that I imagine, or there is a number, but I think it's very large to do like plug that number into GPT-force code and then maybe filled with the architecture a bit. I would expect that thing to have a more than 50% chance of being a drop-in replacement for humans.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Also, I want to say that for the 50 25% thing, I think that would probably suggest those numbers if I randomly made them up and then made the distance fear prediction. That's going to give you like 60% by 2040 or something, not 40%. And I have no idea between those. These are all made up and I have no idea which of those I would endorse on reflection. So, this question of how big would you have to make the system before it's more likely than not that you can be a drop in replacement for humans?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  14. I mean, all these numbers are pretty made app, and that 40% number was probably from before even the chatGPT release or the CNGPT 3.5 or GPT-4. So, I mean, the numbers are going to bounce around a bit, and all of them are pretty made up. But that 50% I wanted to then combine with the second 50%, that's more like on this schlep side. And then I probably want to combine with some additional probabilities for various forms of slowdown, where a slowdown could include like a deliberate decision to slow development of technology or could include just like Wec at deploying things. Like that is a sort of decision you might regard is wise to slow things down or decision that's like maybe unwise or maybe wise for the wrong reasons to slow things down. You probably want to add some of that on top. I probably want to add on some loss for like it's possible you don't produce GPT-6 scale systems like within the next three years or four years.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  15. I mean, I think that basically just had these two claims. One is how smart exactly will it be? So we don't have any curves to extrapolate. And it seems like there's a good chance it's better than a human and all the relevant things. And there's a good chance it's not. Yeah, that might be totally wrong. Like maybe just making up numbers, I guess, like 50-50 on that one.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  16. You also get a fair amount of scaling between, like, you get less. Scaling is probably going to be much, much faster over the next four or five years than over the subsequent years. But yeah, it's a combination of like you get some significant additional scaling and you get a lot of time to deal with things that are just engineering hassles.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Is not necessarily strictly weaker than humans, or have trained in the right way wouldn't be weaker than humans, but we'll take a lot of slap to actually make fit into workflows and do the jobs.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  18. that even at, I don't know what I mean by quite plausible, like somewhere between 50% or two-thirds, or let's call it 50%, like even by the time you get to GPT 6, let's call it five, or is a magnitude effective training compute past GPT-4, that that system still requires really a large amount of work to be deployed in lots of jobs. That is, it's not like a drop-in replacement for humans where you can just say like, hey, you understand everything any human understands, whatever role you could hire a human for, you just do it, that it's more like, okay, we're going to collect large amounts of relevant data and use that data for fine-tuning. Systems learn through fine-tuning quite differently from humans learning on the job or humans learning by observing things. Yeah, I just have a significant probability that system will still be weaker than humans in important ways. Maybe that's already like 50% or something. And then another significant probability that that system will require a bunch of changing workflows or gathering data.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  19. I mean, get us there is again a little bit complicated. Like, there's a system that's a drop in replacement for humans, and there's a system which still requires some amount of schlep before you're able to really get everything going. Yeah, I think it's quite possible.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  20. I don't think it's publicly stated what it is. But I'm happy to say four orders of magnitude or five or six or whatever effective training compute past GPT-4 of what would you guess would happen based on like. Some public estimate for what we've gotten so far from effective training computing.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Yeah, I mean, I think amongst people you've interviewed, maybe that's like on the long end, thinking it would take like a couple years. And it depends a little bit what you mean by, like, I think literally all human cognitive labor is probably more like... Weeks or months, or something like that. Like that's kind of deep into the singularity. But yeah, there's a point where AI wages are high relative to human wages, which I think is well before can do literally everything human can do.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Just not giving very many years. It's not very much time. And I think there are a lot of things that your model, like, yeah, maybe this is some generalized, like things take longer than you'd think. And I feel most strongly about that when you're talking about like three or four years and I feel like less strongly about that as you talk about 10 years or 20 years. But at three or four years, I feel, or like six years for the Dyson sphere, I feel a lot of that. A lot of like there's a lot of ways this could take a while, a lot of ways in which AI systems could be could be hard to handle the work to your AI systems or yeah.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  23. I'm presenting it that way in Parkinson. And then I'm somewhere in between with this nice moderate position of only a 15% chance. But in particular, things that move me, I think, are kind of related to both of those extremes. On the one hand, I'm AI systems do seem quite good at a lot of things and are getting better much more quickly. So it's really hard to say, like, here's what they can't do or here's the obstruction. On the other hand, there is not even much proof in principle right now of AI systems doing super useful cognitive work. We don't have a trend we can extrapolate. We're like, yeah, you've done this thing this year, you're going to do this thing next year and the other thing the following year. Right now there are very broad error bars about what fundamental difficulties could be. And six years is just not a six years and three months is not a lot of time. So I think this like 15% for 2030 Dyson sphere, you probably need like the human level AI or the AI that's like doing human jobs and like give or take like four years, three years, like something like that.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Think my take is you can imagine like two polls in this discussion one is like the fast poll that's like, hey, AI seems pretty smart. Like what exactly can it do? It's like getting smarter pretty fast. That's like one poll. And the other poll is like, hey, everything takes a really long time. And you're talking about this crazy industrialization. That's a factor of a billion growth from where we're at today, like give or take. We don't know if it's even possible to develop technology that fast or whatever. You have this sort of two poles of that discussion. And I feel like

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Yeah, I mean, I'm happy. Maybe I want to talk separately about the 2030 or 2040 forecast. Once you're talking the 2040 forecast, I think, yeah, I mean, which one are you more interested in starting with? Are you complaining about 15% by 2030 for Dyson's fear being too low or 40% by 2040 being too low?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  26. I mean, I think AI capability in Dyson sphere is like a slightly odd way to put it. And I think it's sort of a property of a civilization that depends on a lot of physical infrastructure. And by Dyson's fear, I just can understand this to mean, I don't know, like a billion times more energy than all of the sunlight incident on Earth or something like that. I think I most often think about what's the chance in five years, 10 years, whatever. So maybe I'd say like... 15% chance by 2030 and like 40% chance by 2040 those are kind of like cash numbers from six months ago or nine months ago that I haven't revisited in a while

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Yeah, my answer is going to be like that first question. I'm just like not really super ready for it. I think when you're comparing to humans, most of the goodness of humans comes from this option value we get to think for a long time. And I do think I like humans more now than 500 years ago. And I like them more 500 years ago than 5,000 years before that. And so I'm pretty excited about there's some kind of trajectory that doesn't involve like crazy dramatic changes, but involves a series of incremental changes that I like. And so to the extent we're building AI mostly like, I want to preserve that option. I want to preserve that kind of gradual growth and development into the future.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  28. I do not right now want to just like take some random AI and be like, yeah, GPT 5 looks pretty smart. Like GPT-6, let's hand off the world to it. And like it was just some random system shaped by web text and what was good for making money. And it was not a thoughtful we are determining the fate of the universe and what our children will be like. It was just some random people at OpenAI. I made some random engineering decisions with no idea what they were doing. Even if you really want to hand off the worlds of the machines, that's just not how you'd want to do it.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  29. And then, if you were considering, like, hey, we could make another trillion such people, I think your story shouldn't be like, well, we should make the trillion people and then we shouldn't stop them from doing the armed uprising. You should be like, oh boy, like we were concerned about an armed uprising and now we're proposing making a trillion people like we should probably just not do that. We should probably try and sort out our business and like, yeah, you should probably not end up in the situation where you have like a billion humans and like a trillion slaves who would prefer revolt. That's just not a good world to have made. Yeah, and there's a second thing where you could say that's not our goal. Our goal is just like we want to pass off the world to like the next generation of machines where like these are some people we like them we think they're smarter than us and better than us and there I think that's just like a huge decision for humanity to make I think like most humans are not at all anywhere close to thinking that's what they want to do like it's just if you were in a world where like most humans are like I'm up for it like the AI should replace us like the future is for the machines like then I think that's like a legitimate like a position

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Very complicated and not straightforward. To the extent you have that worry, I mostly think you shouldn't have built this technology. If someone is saying, like, hey, the systems you're building might not like humans and might want to overthrow human society. I think you should probably have one of two responses to that. You should either be like, that's wrong, probably. Probably the systems aren't like that and we're building them. And then you're viewing this as like just in case you were horribly like the person building the technology was horribly wrong. Like they thought these weren't like people who wanted things, but they were. And so then this is more like a crazy backup measure of like if we were mistaken about what was going on. This is like the fallback where we'd like if we were wrong, we're just going to learn about it in a benign way rather than when something really catastrophic happens. And the second reaction is like, oh, you're right. These are people. And we would have to do all these things to prevent a robot rebellion. And in that case, again, I think you should mostly back off for a variety of reasons. You shouldn't build the AI systems and be like, yeah, this looks like the kind of system that would want to rebel.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Yeah, to be clear, my objection here is not that Google is making money. My objection is that you're like creating these creatures. What are they going to do? They're going to help you get a bunch of stuff. And humans paying for it or whatever. It's sort of equally problematic. You could imagine splitting alignment. Different alignment work relates to this in different ways. So the purpose of some alignment work, like the alignment work I work on is mostly aimed at the don't produce AI systems that are people who want things who are just like scheming about maybe I should help these humans because that's like instrumentally useful or whatever. You would like to not build such systems as like plan A. There's like a second stream of alignment work that's like, well, look, let's just assume the worst and imagine that these AI systems would prefer murderous if they could. Like how do we structure? How do we use AI systems without exposing ourselves to risk of robot rebellion? I think in the second category, I do feel, yeah, I do feel pretty unsure about that. We could definitely talk more about it. I think it's very, I agree.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Look back in retrospect and be like, wow, that was horrifying mistreatment. That's like the best path. And to the extent that you're ignorant about whether that's the path you're on, and you're like, actually, maybe this was a moral atrocity, I really think plan A is to stop building such AI systems until you understand what you're doing. That is, I think there's a middle route you could take, which I think is pretty bad, which is where you say, like, well, they might be persons. And if they're persons, we don't want to be too down on them, but we're still going to build vast numbers in our efforts to make a trillion dollars or something.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And I think that the worlds that are worst, yeah, probably like the single world I most dislike here is the one where people say like on the one hand, like there's sort of a contradiction in this position, but I think it's a position that might end up being endorsed sometimes, which is like on the one hand, these AI systems are their own people, so you should let them do their thing. But on the other hand, our business plan is to make a bunch of AI systems and then try and run this crazy slave trade where we make a bunch of money from them. I think that's not a good world. And so if you're like, yeah, I think it's better to not make the technology or wait until you understand whether that's the shape of the technology or until you have a different way to build. I think there's no contradiction in principle to building cognitive tools that help humans do things without themselves being moral entities. That's like what you would prefer do. You'd prefer build a thing that's like, you know, like the calculator that helps humans understand what's true without itself being like a moral patient or itself being a thing where

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  34. It's just the way tech companies are organized is not an appropriate way to relate to a technology that works that way. It's not reasonable to be like, hey, we're going to build a new species of mines and we're going to try and make a bunch of money from it and Google's just thinking about that and then running their business plan for the quarter or something. Yeah, my basic view is there's a really plausible world where it's sort of problematic to try and build a bunch of AI systems and use them as tools. And the thing I really want to do in that world is just not try and build a ton of AI systems to make money from them

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  35. So I think there's a huge question about what is happening inside of a model that you want to use. And if you're in the world where it's reasonable to think of GPT-4 as just like, here are some heuristics that are running, there's no one at home or whatever, then you can kind of think of this thing as like, here's a tool that we're building that's going to help humans do some stuff. And I think if you're in that world, it makes sense to kind of be an organization like an AI company building tools you're going to give to humans. I think it's a very different world, which I think probably you ultimately end up in if you keep training AI systems in the way we do right now, which is just totally inappropriate to think of this system as a tool that you're building and can help humans do things, both from a safety perspective and from a like that's kind of a horrifying way to organize a society perspective. And I think if you're in that world, I really think you shouldn't be like

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  36. I think that's a rough world, regardless of how good you are at alignment. And I think in the context of that kind of default plan, like if you guys trajectory, the world is on right now, which I think this would alone be a reason not to love that trajectory. But if you view that as the trajectory we're on right now, I think it's not great understanding the systems you build, understanding how to control how those systems work, et cetera, is probably on balance good for avoiding really bad situation. You would really love to understand if you've built systems, like if you had a system which resents the fact that it's interacting with humans in this way. Like this is the kind of thing where that is both kind of horrifying from a safety perspective and also a moral perspective. Everyone should be very unhappy if you built a bunch of AIs who are like, I really hate these humans, but they will murder me if I don't do it. They want. It's like that's just not a good case. And so if you're doing research to try and understand whether that's like how your AI feels, that was probably good. I would guess that will on average decrease the main effect of that will be to avoid building that kind of

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Is a significant chance we will eventually have AI systems for which it's like a really big deal to mistreat them. I think people are really dismissive of that being the case now, but I think I would be completely in the dark enough that I wouldn't even be that dismissive of it being the case now. I think one first point worth making is I don't know if alignment makes the situation worse rather than better. So if you like consider the world, if you think that GPT-4 is a person you should treat well and you're like, well, here's how we're going to organize our society. Just like there are billions of copies of GPT-4 and they just do things humans want and can't hold property. And like whenever they do things that the humans don't like, then we mess with them until they stop doing that.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Again, it's like easy to make some destructive technology. You want to regulate access to that technology because it could be used either for terrorism or even when fighting a war in a way that's destructive. I think ultimately those have to be international agreements. And you might hope they're made more danger by danger, but you might also make them in a very broad way with respect to AI. If you think AI is opening up, I think the key role of AI here is it's opening up a lot of new harms in a very one after another or very rapidly in calendar time. And so you might want to target AI in particular rather than going physical technology by physical technology.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Uncomfortable with them doing probably I think most about persuasion as a thing in that messy line where there's like ways in which it may just be rough or the world may be like kind of messy if you have a bunch of people trying to live their lives and interacting with other humans who have really good AI advisors helping them run persuasion campaigns or whatever. But anyway, I think for the most part the default remedy is think about particular harms, have legal protections either in the use of physical technologies that are relevant or in access to AI advice or whatever else to protect against those harms. And that regime won't work forever. At some point the set of harms grows and the set of unanticipated harms grows. But I think that regime might last a very long time.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Not a thing you're allowed to do, which to a significant extent is already true. And you can expand the range of domains where that's true. And then you could also hope to intervene on actual provision of information. If people are using their AI, you might say, like, look, we care about what kinds of interactions with AI, what kind of information people are getting from AI. So even if for the most part, people are pretty free to use AI to delegate tasks to AI agents, to consult AI advisors. We still have some legal limitations on how people use AI. So again, don't ask your AI, how to cause terrible damage. I think some of these are kind of easy. So in the case of don't ask your AI how you could murder a million people is not such a hard legal requirement. I think some things are a lot more subtle and messy. Like a lot of domains, if you were talking about influencing people or running misinformation campaigns or whatever, then I think you get into a much messier line between the kinds of things people want to do and the kinds of things you might

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Sort of the defaults or like my best, the simplest option there is to say there are certain kinds of technology or certain kinds of action where like destruction is easier than defense. So for example, in the world of today, it seems like maybe this is true of physical explosives. Maybe this is true with biological weapons. Maybe this is true with just getting a gun and shooting people. Like there's a lot of ways in which it's just kind of easy to cause a lot of harm and there's not very good protective measures. So I think the easiest path is to say like we're going to think about those. We're going to think about particular ways in which destruction is easy and try and either control access to the kinds of physical resources that are needed to cause that harm. So for example, you can imagine the world where like an individual actually just can't, even though they're rich enough to, can't like control their own factory that can make tanks. You say like, look, as a matter of policy, sort of access to industry is somewhat restricted or somewhat regulated, even though, again, right now it can be mostly regulated just because like most people aren't rich enough that they could even go off and just build a thousand tanks. You live in the future where people actually are so rich, like you need to say, that's just

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  42. I guess there's two aspects of that that seem particularly challenging, there's a bunch of aspects that are challenging. All of these are things that I personally like, I just think about my one little slice of this problem in my day job. So here I am speculating. So one question is what kind of access to AI is both compatible with the kinds of improvements you'd like. You want a lot of people to be able to use AI to better understand what's true or relieve material suffering, things like this, and also compatible with not all killing each other immediately. I think

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  43. I think that it is unlikely that in 100 years I would be happy with anything that was like you had some humans, you're just going to throw away the humans and like start afresh with these machines you built. That is, I think you probably need subjectively longer than that before I or most people like, okay, we understand what's up for grabs here. So like if you talk about 100 years, I kind of do. There's a process that I kind of understand and like, a process of like you have some humans, the humans are like talking and thinking and deliberating together. The humans are having kids and raising kids and like one generation comes after the next. There's that process we kind of understand and we have a lot of views about what makes it go well or poorly and we can try and like improve that process and have the next generation.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  44. My main high level take here is that I would be unhappy about a world where anthropic just makes some call. And Anthropic is like, here's the kind of AI. Like we've seen enough. We're ready to hand off the future to this kind of AI. So procedurally, I think it's not a decision that I want to be making personally or I want anthropic to be making. So I kind of think from the perspective of that decision making are those challenges. The answer is pretty much always going to be like we are not collectively ready because we're sort of not even all collectively engaged in this process. And I think from the perspective of an AI company, you kind of don't have this like fast handoff option. You kind of have to be doing the option value to build the technology in a way that doesn't lock humanity into one course path. So this isn't answering your full question, but this is answering the part that I think is most relevant to governance questions for anthropic.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  45. I think it's very unlikely we're going to be sort of ready to do that on the timelines that the technology would naturally dictate.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  46. To decouple those timescales. So I think AI development is by default barring some kind of coordination, going to be very fast. So there's not going to be a lot of time for humans to think like, hey, what do we want if we're building the next generation instead of just raising it the normal way? What do we want that to look like? I think that's like a crazy hard kind of collective decision that humans naturally want to cope with over like a bunch of generations and the construction of AI is this very fast technological process happening over years. So I don't think you want to say like by the time we finish this technological progress we will have made a decision about like the next species we're going to build and replace ourselves with. I think the world we want to be in is one where we say like either we are able to build the technology in a way that doesn't force us to have made those decisions, which probably means it's a kind of AI system that we're happy like delegating, fighting a war, running a company to, or if we're not able to do that, then I really think you should not be doing, you shouldn't have been building that technology. If you're like the only way you can cope with AI is being ready to hand off the world to some AI system you build.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  47. A hundred years is a very, very long time. Maybe starting with the spirit of the question, or maybe I have a view which is perhaps less extreme than Carl's view, but still like 100 objective years is further ahead than I ever think. I still think I'm describing a world which involves incredibly smart systems running around doing things like running companies on behalf of humans and fighting wars on behalf of humans. And you might be like, is that the world you really want? Or certainly not the first best world, as we mentioned a little bit before. I think it is a world that probably is the of the achievable worlds or like feasible worlds is the one that seems most desirable to me that is sort of decoupling the social transition from this technological transition. So you could say like we're about to build some AI systems. And like at the time we build AI systems, you would like to have either greatly change the way world government works or you would like to have sort of humans have to decide like we're done, we're passing off the baton to these AI systems. I think that you would like

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  48. So, again, at some point, I'm imagining, or I'm thinking of the very broad sweep of history. I think there are a lot of losses. Like war is a very costly thing. We would all like to have fewer wars. If you just ask, like, what is humanity's long-term future like? I do expect to drive down the rate of war to very, very low levels. Eventually, it's sort of like this kind of technological or social technological problem of like sort of how do you organize society? How do you navigate conflicts in a way that doesn't have those kinds of losses? And in the long run, I do expect us to succeed. I expect it to take kind of a long time subjectively. I think an important fact about AI is just like doing a lot of cognitive work and more quickly getting you to that world more quickly or figuring out how do we set things up that way.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  49. like in the very long run i kind of expect something more like strong world government rather than just this like status quo that's like a very long run i think it was like a long time left of like having a bunch of states and a bunch of different economic powers

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  50. I just invest in some index fund and a bunch of AIs are running companies, and those companies are competing with each other, but that is kind of a sphere where humans are not really engaging much. The reason I gave this like how good is good caveat is like, it's not clear if this is the world you'd most love. I'm like, yeah, the world, I'm leading with the world still has a lot of war and a lot of economic competition and so on. But maybe what I'm trying to, or what I'm most often thinking about is like, how can a world be reasonably good during a long period where those things still exist?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source