YouSaid · the spoken record

Paul Christiano

lines on the record
251
first
2023-10-31
most recent
2023-10-31
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. So, our quest is to design trading methods for which we don't expect them to lead to reward hacking or don't expect them to lead deceptive alignment. Ideally, that won't be like a huge tax where people like, well, we use those methods only if we're really worried about reward hacking receptive alignment. Ideally, those methods would just work quite well. And so people would be like, sure. I mean, they also address a bunch of other more mundane problems, so why would we not use them? Which I think is sort of the good story. The good story is you develop methods that address a bunch of existing problems because they just are more principled ways to train AI systems that work better. People adopt them and then we are no longer worried about e.g. reward hacking or deceptive alignment.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Ideally, they'll be so good, even if you haven't seen them, you would just want to switch to reasonable methods that don't have these. Ideally, they'll work as well or better than normal training

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  3. And obviously, most of my life is in that. I am really in that bucket. Like, I mostly do alignment research. It's just building out the techniques that do not have these failures such that they can be available as an alternative. In fact, these failures occur.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  4. This is the kind of story you want to build towards in the long run. Do your best to produce all the failures you can in the lab or versions of them. Do your best to understand what causes them, what kind of anomaly detection actually works for detecting this or what kind of filtering actually works. Apply that, and that's at the meta level. It's not talking about what actually are those measures that would work effectively, which is obviously what, I mean, a lot of alignment research is really based on this hypothetical of someday there will be systems that fail in this way, what would you want to do? Can we have the technologies ready, either because we might never see signs of the problem or because we want to be able to move fast once we see signs of the problem?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  5. We have our problems in the lab. We then have some techniques which we believe address these problems. We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it's being trained to thinking it should defect. And then you can say on the lab, we have some understanding of when those techniques work and when they don't. We have some understanding of the relevant parameters for the real system that's deployed, and we have a reasonable margin of safety. So we have reasonable robustness on our story about when this works and when it doesn't. And we can apply that margin of safety with a margin of safety to the real deployed system.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I think that so at a meta level in terms of what's your protection like I think what you want to be saying is we have these examples in the lab of something bad happening we're concerned about the problem at all because we have examples in the lab and again this should all be an addition I think like you kind of want this like defense in depth of saying we also have this testing regime that would detect problems for the deployed model

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Or a number of disalogies from the real world, but this is a setup which is relatively conducive to deceptive alignment, like produce a system that wants one thing, tell it a lot about its training, the kind of information that you might expect a system would get, and then try and understand whether, in fact, it is able to then or it tends or sometimes, under optimal conditions, in fact, continues pursuing paperclips, only pursuing apples when it thinks it's being trained.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Know whatever, get some paper clips. He wants to produce as many paper clips as he can over the next five days. Just select actions really aggressively for producing paperclips over the next five days. You do your RLHF, you do your pre-training, whatever. That's like your phase one. You also ensure your AI system has a really good understanding of how it's trained. This AI system wants paperclips and it understands everything about how it's trained and everything about how it's fine-tuned. And you train on just a lot of this data. And they say, okay, if we've done all of that, We have this concern that if a system wants paperclips and understands really well how it's trained, then it will, like, if it's going to be trained to get apples instead of paperclips, it's just going to do some cost-benefit and be like, ah, you know, really while I'm being trained to get apples, I should do that, but I should do that whether or not, even if I want paperclips, I should still do that. So training won't really affect its values. It will just understand that it's being trained to get apples. It will try and get apples. And if you take it out of training, it will go back to getting paper clips. It's like, I think this exact setup has a number of problems.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Yeah, so for deceptive alignment, I mean, I think it's really complicated. Yeah, I think this is a very hard project. I think one of the things, one of the most important things for people concerned about Ms. Lyman to do is to ensure that adequate effort is put into this task. And we have, in fact, created good conditions in the laboratory. But okay, if you were to gas for deceptive alignment, how would you do this? An example is suppose you train an AI system that has some medium horizon goal. You train your AI system to

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Say either it is very hard or we're not able to create conditions where these failures emerge, or I think more likely we are able to create those conditions. And like here's the story about why we are able to detect those in a way that would work in the real world. And we can see they don't occur in the real world.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Maybe this is even more important than the first one just understanding whether dangerous forms of misalignment can occur. It's like the best reference point for this, I think, is like, there have been a couple of projects like this in academia. Anthropic has been working on a project or spun up a team doing this kind of work, trying to say, can dangerous forms of reward hacking occur? Or can deceptive alignment in fact occur in the lab? And like here, what you want to do is either say, okay, even if we create best case, even if you create optimal conditions for deceptive alignment or for reward hacking, we actually just can't cause it to occur even in the lab. And if you do a good enough job of that, I think it could give you some evidence. And again, more likely that fails. And when you create optimal conditions, you do see a deceptive alignment and reward hacking in the lab. But then once you have that data, once you can say, okay, in the lab actually these things can occur, then you can start saying, and we have a robust scientific understanding that enables us to fix those problems when they occur. Or like then you start, you have this fork on your second line.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  12. And if the model can tell the difference, which is almost certainly going to be the case at some level capability, then you say, okay, our first line of hoped for defense has failed. And so now we need to understand how can tests be indicative when models are capable enough to tell the difference between the tests and the real world. Now you need to start arguing about what is the model paying attention to? Basically like a first line of if you get lucky what would work here and then there's the like deeper like you probably have to do novel science at some point to like make this case All of this was part of like I was saying maybe you want to have kind of two lines of defense. All of this was like part of this first line of defense of like can you detect something is going wrong or can you prevent harm from occurring? Can you have monitoring that will react quickly enough to avoid catastrophe? I think you probably also want to have a second line of defense that's more

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  13. That's the kind of thing. And it's not the greatest. So Dag Color, when I say I can't tell the difference, like the kind of valuation you would love to do if you could is to say, like, hey, we have a bunch of real world deployment, we have a bunch of data from those deployments. We also have a bunch of in the lab tests of situations. If we look at in the lab tests of a situation together with real-world cases in which the situation occurs, and then we fine-tune a model to try and distinguish the difference there, the model can't even tell when fine-tuned the difference between test and real-world cases.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So that's like there's a lot of stuff in that genre. And I tried to throw out that example, but I think there's a big that on its own wouldn't be compelling even from this perspective. You kind of need a lot of pieces together. And that's kind of a demanding technical project. But I think one without necessarily deep conceptual difficulty. A big problem with it is it's liable to fail. It's liable you're going to do that testing or say we think these tests are representative because the AI can't distinguish tests from the real world, which you can maybe achieve now, but will probably become impossible to achieve at some level. It's not clear what's at human level or how far past human level.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  15. This is an example of a thing you can do that's not really making assumptions about what's going on inside the model or about how the model is the model aligned. It's more trying to say, is it the case that our testing or monitoring machinery could detect or react to something bad happening? In this case, could it detect a system which would under some conditions do something really bad?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  16. We have tried to test our system in a broad diversity of situations that reflect cases where it might cause harm, like cases where someone is trying to recruit it to cases where it might cause harm or cases where a large number of AI systems might sort of all decide to work across purposes to humans. Like the robot rebellion is starting or whatever. We've done evaluation in a broad distribution of situations. And then we have tried to argue that our AI is actually, those tests are indicative of the real world. So we've tried to argue like, hey, actually the AI is not very good at distinguishing situations we produce in the lab as tests from similar situations that occur in the real world. And the coverage of this distribution is reasonable.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  17. I broadly think there's like two kinds. Like, if you ask me right now, what evidence for alignment could make you comfortable? I think my best guess would be to provide two kinds of evidence. So one kind of evidence is on the like, could you detect or prevent catastrophic harm if such a system was misaligned? I think there's like a couple of things you would do here. One thing you would do is on this adversarial evaluation front so that you could try and say, for example, like,

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yeah, so again, there's sort of like some stuff you care about on the misuse side and some stuff you care about on the misalignment side. And there's probably further things you care about, especially to the extent your concerns are broader than catastrophic risks. But maybe I most want to talk about just what you care about on the alignment side, because it's like the thing I've actually thought about most. Also, a thing I care about a lot. Also, I think a significant fraction of the existential risk over the kind of foreseeable future. So, on that front,

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Up something really bad. I think those policies can become really complicated right now. I think RSPs can focus more on we have our inventory of the things that a human is going to do to cause a lot of harm with access to AI probably are things that are on our radar. That is like they're not going to be completely unlike things that the human could do to cause a lot of harm with access to weak AIs or with access to other tools. I think it's not crazy to initially say that's what we're doing. We're looking at the things closest to humans being able to cause huge amounts of harm and asking which of those are taken over the line. But eventually that's not the case. Eventually like AIs will enable just like totally different ways of killing a billion people.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Yeah, I mean, I think ultimately you're going. And so, when I talk about where the existential risk comes from, I'm mostly thinking about comes from when, at what point do you face what challenges, or in what sequence? And so I'm saying, I think misalignment is probably one way of putting it is if you imagine AI systems sophisticated enough to discover destructive technologies that are totally not on our radar right now, I think those come well after AI systems capable enough that if misaligned, they would be catastrophically dangerous. There's like the level of competence necessary to, if broadly deployed in the world, bring down a civilization as much smaller than the level of competence necessary to advise one person on how to bring down a civilization, just because in one case you already have a billion copies of yourself or whatever. So yeah, I think it's mostly just the sequencing thing though. Like in the very long run, I think you care about, hey, AI will be expanding the frontier of dangerous technologies. We want to have some policy for exploring or understanding that frontier and whether we're about to turn

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Alignment becomes a catastrophic issue prior to most of these. That is like prior to some way to spend $50,000 to kill everyone, with like the salient exception of possibly bioweapons. So that would be my guess. And then there's a question of what is your risk management approach not knowing what's going on here. And you don't understand whether there's some way to use $50,000. But I think you can do things like understand how good is an AI at coming up with such schemes. You can talk to your AI and be like, does it produce new ideas for destruction we haven't recognized?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  22. A way you could put it is there's like a bunch of potential destructive technologies and alignment is about AI itself being such a destructive technology where even if the world just uses the technology of today, simply access to AI could cause human civilization to have serious problems. But there's also just a bunch of other potential destructive technologies. Again, we mentioned physical explosives or bioweapons of various kinds. And then the whole tale of who knows what. My guess is that like...

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  23. I mean, in the near term, I think harms from misuse are like, especially if you're not restricting to the tail of extremely large catastrophes, I think the harms from misuse are clearly larger in the near term.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  24. to be able to bound the probability that someday all the AI systems will do something really harmful, that there's like some thing that could happen in the world that would cause these large scale correlated failures of your AIs. And so for that, like, I mean, there's sort of two categories. That's like one. The other thing you need is protection against misuse of various kinds, which is also quite hard.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Most of the risk comes from doing things with the model. You need all the rest so that you have any possibility of applying the brakes or implementing a policy. But at some point, as the model gets competent, you're saying, okay, could this cause a lot of harm, not because it leaks or something, but because we're just giving it a bunch of actuators or deploying it as a product and people could do crazy stuff with it. So if we're talking not only about a powerful model, but a really broad deployment of just something similar to the OpenAI's API where people can do whatever they want with this model. And maybe the economic impact was very large. So in fact, if you deploy that system, it will be used in a lot of places such that if AI systems wanted to cause trouble, it would be very, very easy for them to cause catastrophic harms. Then I think you really need to have some kind of, I mean, I think probably the science and discussion has to improve before this becomes that realistic, but you really want to have some kind of, oh, I'm in analysis, guarantee of alignment before you're comfortable with this. And so by that I mean, you want to.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Yeah. So I think here, so I listed maybe the two most simple ones that start out like security internal controls, I think, become relevant immediately and are very clear why you care about them. I think as you move beyond that, it really depends how you're deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weight sitting there, I think you mostly have addressed the risk from the weights just sitting there. Like now you're talking about for risk is mostly and like maybe there's some blurriness here of how much internal controls captures not only employees using the model, but anything a model can do internally. We'd really like to be in a situation where your internal controls are robust, not just to humans, but to models, potentially like EGE. A model shouldn't be able to subvert these measures. And you like care just as you care about, are your measures robust if humans are behaving maliciously? Are your measures robust if models are behaving maliciously? So I think beyond that, if you've then managed the risk of just having the weight sitting around, now we talk about, in some

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  27. So, if you had a model which you thought was like potentially very scary, either on the object level or because of leading to this sort of intelligence explosion dynamics, I mean, things you want in place are like you really do not want to be leaking the weights to that model. Like you don't want the model to be able to run away. You don't want human employees to be able to leak it. You don't want external attackers or any set of all three of those coordinating. You really don't want internal abuse or tampering with such models. So if you're producing such models, you don't want it to be the case that a couple employees could change the way the model works or could do something that violates your policy easily with that model. And if a model is very powerful, even the prospect of internal abuse could be quite bad. And so you might need significant internal controls to prevent that.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  28. They have this autonomy in the lab benchmark, which I think is probably occurs prior to either really massive AI acceleration or to most potential catastrophic object-level catastrophic harms. And the idea is that's a warning sign lets you punt. So this is a bit of a digression in terms of how to think about that risk. I think I am unsure whether you should be addressing that risk directly and saying we're scared to even work with such a model or if you should be mostly focusing on object level harms and saying like, okay, we need more intense precautions to manage oblig level harms because of the prospect of very rapid change and the availability of this AI just like creates that kind of creates that prospect. Okay, this was all still a digression

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  29. There's a couple points that come up here. So, one is this threat model of sort of automating R&D, or like independent of whether AI can do something on the object level that's potentially dangerous. I think it's reasonable to be concerned if you have an AI system that might, if leaked, allow other actors to quickly build powerful AI systems or might allow you to quickly build much more powerful systems. Or might if you're trying to hold off on development just itself be able to create much more powerful systems. So I think one question is how to handle that kind of threat model as distinct from a threat model like this could enable destructive bioterrorism or this could enable massively scaled cybercrime or whatever. And I think I am unsure how you should handle that. I think like right now implicitly it's being handled by saying like, look, there's a lot of overlap between the kinds of capabilities that are necessary to cause various harms and the kinds of capabilities that are necessary to accelerate ML. So we're kind of going to catch those with like an early warning sign for both and deal with the resolution of this question a little bit later. So for example, in Anthropics,

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  30. At a relatively early stage because it does undermine the rest of the measures you may take and is also just part of the easiest, like if you imagine catastrophic harms over the next couple years, I think security failures are kind of play a central role in a lot of those. And maybe the last thing to say is like it's not clear that you should say we have pause because we have models that can develop bioweapons versus just like, I mean, potentially not saying anything about what models you've developed or at least saying like, hey, by the way, here's like, you know, here's the set of practices we currently implement.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  31. I mean, I think the general discussion does emphasize potential harms or potential, I mean, some of those are harms and some of those are just like impacts that are very large and so might also be an inducement to develop models. I think that part, if you're for a moment, ignoring security and just saying that may increase investment, I think it's on balance just quite good for people to have an understanding of potential impacts just because it is an input both into proliferation but also into regulation or safety. With respect to things like security of either weights or other IP, I do think you want to have moved to significantly more secure handling of model weights before the point where a leak would be catastrophic. And indeed, for example, in Anthropics document or in their plan, security is one of the first sets of tangible changes. That is at this capability level, we need to have such security practices in place. So I do think that's just one of the things you need to get in place.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Been like considerable concern about developers being more or less safe, but there's not that much legible differentiation in terms of what their policies are. I think getting to that world would be good. It would be a very different, it's a very different world if you're actor X is developing AI and I'm concerned that they will do so in an unsafe way versus if you're like, look, we take security precautions or safety precautions X, Y, Z. Here's why we think those precautions are desirable or necessary. We're concerned about this other developer because they don't do those things. I think it's just like a qualitatively It's kind of the first step you would want to take in any world where you're trying to get people on side or like trying to move towards regulation that can manage risk.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  33. I think the first sort of business understanding what is actually a reasonable set of policies for managing risk, I do think there's a question of you might end up in a situation where you say like, well, here's what we would do in an ideal world if we was behaving responsibly, we'd want to keep risk to 1% or a couple percent or whatever, maybe even lower levels, depending on how you feel. However, in the real world, there's enough of a mess. Like there's enough unsafe stuff happening that actually it's worth making larger compromises. Or like we don't kill everyone, someone else will kill everyone anyway. So actually the counterfactual risk is much lower. I think if you end up in that situation, it's still extremely valuable to have said, like, here's the policies we'd like to follow. Here's the policies we've started following. Here's why we think it's dangerous. Like here's the concerns we have if people are falling significantly backstop policies. And then this is maybe helpful as an input to or model four potential regulation. It's helpful for being able to just produce clarity about what's going on. I think historically.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yes, I mean, Anthropic has written this document, the current responsible scaling policy. And then have been talking with other folks. I guess don't really want to comment on other conversations. But I think in general, people who are more interested in more think you have plausible catastrophic harms on a five-year timeline are more interested in this. And there's not that long a list of suspects like that.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Yeah, so this sort of, again, I think it's motivated primarily, but like, what should you be doing as a lab to manage catastrophic risk now in a way that's like a reasonable precedent and habit and policy for continuing to implement into the future?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Understand whether that's the case, notice when that stops being the case, have a reasonable roadmap for what you're actually going to do when that stops being the case. So that motivates this set of policies, which I've sort of been pushing for labs to adopt, which is saying, here's what we're looking for, here's some threats we're concerned about, here's some capabilities that we're measuring, here's the level, like here's the actual concrete measurement results that would suggest to us that those threats are real, here's the action we would take in response to observing those capabilities. If we couldn't take those actions, like if we've said that we're going to secure the weights, we're not able to do that, we're going to pause until we can take those actions.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Sure. So I guess the motivating question is like, what should AI labs be doing right now to manage risk and to sort of build good habits or practices for managing risk into the future? I think my take is that current systems pose from a catastrophic risk perspective, not that much risk today. That is a failure to control or understand GPT-4 can have real harms, but doesn't have much harm with respect to the kind of takeover risk I'm worried about or even much catastrophic harm with respect to misuse. So I think if you want to manage catastrophic harms, I think right now you don't need to be that careful with GB24. And so to the extent you're like, what should labs do? I think the single most important thing seems like

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Yeah, the same reason I'm more excited about policy change now than two years ago. So, my overall view is just in the past this calculus, this calculus changes over time, right? More people are getting prepared, the better the calculus is for slowing down at this very moment. And I think now the calculus is, I would say, positive for just even if you pause now and it would get clawed back in the future, I think the pause now is just good because enough stuff is happening. We have enough idea of, I mean, probably even apart from alignment research and certainly if you include alignment research, just like enough stuff is happening where the world is getting more ready and coming more to terms with impacts that I just think it is worth it even though some of that time is going to get clod back. Again, especially if like there's a question of during a pause, does NVIDIA keep making more GPUs? Like that sucks if they do. If you do a pause, but yeah, in practice, if you did a pause, then you probably couldn't keep making more GPUs because in fact the demand for GPUs is really important for them to do that. But if you told me that you just get to scale hardware production and building the clusters, but now

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Right now, it seems like there's a lot of policy stuff you'd want to do. This seemed like less plausible a couple years ago, maybe. But if the world just knew they had a 10 year pause right now, I think there's a lot of sense of like, we have policy objectives to accomplish if we had 10 years, we could pretty much do those things. We'd have a lot of time to debate measurement regimes, debate policy regimes and containment regimes, and a lot of time to set up those institutions. So if you told me that the world knew it was a pause, like it wasn't like people just see that AI progress isn't happening, but they're told, you guys have been granted or cursed with a 10-year no AI progress, no alignment progress. I think that would be quite good at this point. However, I think it would be much better at this point than it would have been two years ago. And so the entire concern with slowing AI development now rather than taking the 10-year pause is just if you slow the AI development by a year now, I guess the sum gets clawed back by looking if it gets picked faster in the future. My guess is you lose like half a year or something like that in the future, maybe even more, maybe like two-thirds of a year. So it's like...

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  40. It could never come out ahead. I think the reason that you can come out ahead, the reason you could end up thinking the alignment was that negative was because there's a bunch of other stuff you're doing that makes AI safer. Like if you think the world is like gradually coming better to terms with the impact of AI or policies being made or you're getting increasingly prepared to handle the threat of authoritarian abuse of AI, if you think other stuff is happening that's improving preparedness, then you have reason beyond lyman research to slow down AI.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  41. I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment. You're like, maybe alignment speeds up AI, but the only purpose of slowing down AI was to do it. It's just.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Question of what's the calculus for speeding up? I think speeding up is pretty rough. I think speeding up locally is a little bit less rough. And then, yeah, I think that the effect, like the overall effect size from doing alignment work on reducing takeover risk versus speeding up AI is pretty good. I think, yeah, I think it's pretty good. I think you reduce takeover significantly before you speed up AI by a year or whatever.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  43. So, this is also all one subcomponent of the overall impact. And I was just saying this to briefly give the roadmap for the overall too long answer.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  44. My guess isn't that negative, but I think it's not super clear. And it's much less than slowing AI. Slowing AI is great if you could slow overall AI progress. I think slowing AI by causing, you know, there's this issue where slowing AI now, like for ChatGPT, you're building up this backlog. Why does ChatGPT make such a splash? I think people, there's a reasonable chance if you don't have a splash about ChatGPT, you have a splash about GPT-4. And if you fail to have a splash about GPT-4, there's a reasonable chance of a splash about GPT 4.5. And just like as that happens later, there's just less and less time between that splash and between when an AI potentially kills everyone.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Yeah, like, what's the total trade-off there? I mean, I think my take is like. So, I think slower AI development on balance is quite good. I think that slowing AI development now or, like, say, having less press around ChatGPT is a little bit more mixed than slowing AI development overall. I think it's still probably positive, but much less positive, because I do think there's a real effect of the world is starting to get prepared or is getting prepared at a much greater rate now than it was prior to the release of ChatGPT. And so if you can choose between progress now or progress later, you really prefer have more of your progress now, which I do think slows down progress later. I don't think that's enough to flip the sign. I think maybe it was in the far enough past, but now I would still say moving faster now is net negative. But to be clear, it's a lot less net negative than merely accelerating AI because I do think, again, the chat GPT thing, I am glad people are having policy discussions now rather than delaying the chat GPT wake up thing by a year and then having policy.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Is again the version of this that's most pragmatic is just like suppose you don't work on alignment today, it decreases economic impact of AI systems. They'll be less useful if they're less reliable and if they more often don't do what people want. And so you could be like, great, that just buys time for AI. And you're getting some trade-off there where you're decreasing some risks of AI. Like if AI is more reliable or more does what people want or is more understandable, then that cuts down some risks. But if you think AI is on balance bad, even apart from takeover risk, then the alignment stuff can easily end up being non-negative.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Yeah, I think that, like, in some sense, you're always going to face this trade-off where alignment makes it possible to deploy AI systems or makes it more attractive to deploy as systems. Or in the authoritarian case, makes it tractable to deploy them for this purpose. If you didn't do any alignment, there'd be a nicer, bigger buffer between your society and malicious uses of AI. And I think it's one of the most expensive ways to maintain that buffer. It's much better to maintain that buffer by not having the compute or not having the powerful AI. But I think if you're concerned enough about the other risks, there's definitely a case to be made for just like put in more buffer or something like that. I'm not, like, I care enough about the takeover risk that I think it's just not a net positive way to buy buffer.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  48. I mean, if you told me it was never going to restart, then it wouldn't matter. And if you told me it's going to restart, I guess it would be a kind of similar calculus to today.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  49. I mean, I would generalize so I think there's a real sense in which you should just be scared to the extent you're scared of all AI, you should be like, well, alignment, although it helps with one risk, does contribute to AI being more of a thing. I do think you should shut down the other parts of AI before. Like if you were a policymaker or a researcher or whatever looking in on this, I think it's crazy to be like, this is the part of the basket we're going to remove. You should first remove other parts of the basket because they're also part of the story of risk.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  50. I mean, I think they're super universally applicable. I think it's just like, I mean, the rough way I would describe it, which I think is basically right, is like. Some degree of alignment makes the AI systems much more usable. You should just think of the technology of AI as including a basket of some AI capabilities and some getting the AI to do what you want. It's just part of that basket. And so anytime we're like, to extend alignment is part of that basket, you're just contributing to all the other harms from AI. Like you're reducing the probability of this harm, but you are helping the technology basically work. And the basically working technology is kind of scary from a lot of perspectives, one of which is right now even in a very authoritarian society, just like humans have a lot of power because you need to rely on just a ton of humans to do your thing. And in a world where AI is very powerful, it is just much more possible to say, like, here's how our society runs. One person calls the shots and then a ton of AI systems do what they want. I think that's a reasonable thing to dislike about AI and a reasonable reason to be scared to push the technology to be really good.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source