YouSaid · the spoken record
Paul Christiano
- lines on the record
- 251
- first
- 2023-10-31
- most recent
- 2023-10-31
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Wants different things in different contexts. Then I do think adversarial settings will be the main ones where you see the system, or like the easiest ones where you see the system behaving really badly. And it's a little bit hard to tell how that shakes out.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Cause that doesn't, I mean, it matters and it creates chaos. I mean, it might be quite bad for the world, but it doesn't really affect the alignment calculus. Now it's just like right now you have normal cyber offense, cyber defense, you have weird AI version of cyber offense cyber defense. But if you have this kind of asymmetrical thing where a bunch of AI systems are like, we love AI flourishing can then go in and say great AI is, hey, how about you join us? And that works if they can search for a persuasive argument to that effect. And that's kind of asymmetrical. Then the effect is whatever values, it's easiest to push to argue to an AI that it should do, that is advantaged. So it may be very hard to build AI systems, like trying to defend human interests, but very easy to build AI systems, just like trying to destroy stuff or whatever, just depending on what is the easiest thing to argue to an AI that should do, or what's the easiest thing to trick an AI into doing or whatever. Yeah, I think if alignment is spotty, so if you have the AI system, which doesn't really want to help humans or whatever, or in fact wants some kind of random thing or like”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Systems right now are very vulnerable to manipulation. It's not clear how much more vulnerable they are than humans, except for the fact that you can, if you have an AI system, you can just replay it like a billion times and search for what thing can I say that will make it behave this way. So as a result, like AI systems are very vulnerable to manipulation. It's unclear if future AI systems will be similarly vulnerable to manipulation, but certainly seems plausible. And in particular, aligned AI systems or unaligned AI systems would be vulnerable to all kinds of manipulation. The thing that's really relevant here is kind of like asymmetric manipulation or something that is if it is easier. So if everyone is just constantly messing with each other's AI systems, like if you ever use AI systems in a competitive environment, a big part of the game is like messing with your competitors' AI systems. A big question is whether there's some asymmetric factor there where it's kind of easier to push AI systems into a mode where they're behaving erratically or chaotically or like trying to grab power or something than it is to push them to fight for the other side. It was just a game of two people are competing and neither of them can sort of hijack an opponent's AI to help support their”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, I think in some sense, so there's a bunch of questions that come up here. First one is are aligned data systems that you can build competitive, are they almost as good as the best systems anyone could build? And maybe we're granting that for the purpose of this question. I think a next question that comes up is”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“If you do that as humanity, then if you're an AI system, considering do I want to murder everyone, your calculus is if this is my real chance to murder everyone, I get the tiniest bit of value. I get one trillionth of the value or whatever, one billionth of the value. But on the other hand, if I don't murder everyone, there's some worlds where then the humans will correctly determine I don't murder everyone because in fact the humans survive, the humans are running the simulations to understand how different AIs would behave. And so that's a better deal.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“However, from the perspective of any kind of reasonable AI, it's like not that sure whether it lives in the world where the humans in fact have nothing to give or whether the humans, like in fact it lives in a world where the humans succeeded at building a land AI. And now the AI is simply running in a nice little simulation. The humans are wondering, I wonder. If the say I would have murdered his fall if it had the chance. And he would just say, like, if it would murder us all if it had the chance, that sucks. We'd like to run this trade. We'd like to be nice to the AIs who wouldn't have murdered us all in order to create an incentive for AIs not to murder us. So we do as we just check. And for the kinds of AIs who don't murder everyone, we just give them like one billionth of the universe.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Be a higher level thing that goes into both of these, and then I'll talk about how you instantiate an at causal trade is just like it matters a lot to the humans not to get murdered. And the AI cares very, very little about whether, if we imagine it's hypothetical, the reason it wants to kill humans is just total apathy. It cares very little about whether or not to murder humans because it is so easy to marginalize humans without murdering them and the resources required for human survival are extremely low. Again, in the context of this like rapidity industrialization. So that's the basic setting. And now the thing that you'd like to do is run a trade. The AI would like to say, hey, humans, you care a ton about not getting murdered. I don't really care one way or the other. I would like to, if I could, find some way in which I don't murder you and then in return I get something. The problem is in that world, the humans have essentially nothing to give that is the humans are mostly irrelevant.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“And my guess is just like if you draw a bunch of values from the basket, that's like a natural enough thing. Like if your AI wanted like 10,000 different things or your stillization, if the AI that wants 10,000 different things, it's just reasonably likely you get some of that. The other salient reason you might not want to murder them is just like, well, yeah, there's some kind of crazy decision theory stuff or causal trade stuff, which does look on paper like it should work. And if I was running a civilization and dealing with some people who I didn't like at all or didn't have any concern for at all, but I only had to spend one billionth of my resources not to murder them, I think it's quite robust that you don't want to murder them. That is, I think, the like weird decision theory, a causal trade stuff probably does carry the day.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, think this is a really complicated question. Like, if you imagine drawing values from the basket of all values, what fraction of them are like, hey, if there's someone here, how much do I want to not murder them?”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, also you can marginalize humans pretty hard. Like, you could totally cripple human warfighting capability and also take almost all their stuff while killing only a small fraction of humans incidentally. So then if you ask why my day I not want to kill humans, I mean, a big thing is just like, look, I think AI is probably want like a bunch of random crap for complicated reasons. The motivations of AI systems and civilizations of AIs are probably complicated messes. Certainly amongst humans, it is not that rare to be like, well, there was someone here. I would like all else equal if I didn't have to murder them. I would prefer not murder them. And my guess is it's also like reasonable chance it's not that rare amongst AI systems. Humans have a bunch of different reasons we think that way. I think AI systems will be very different from humans, but it's also just like a very salient”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, I think the literal they're made of atoms is quite, they're not many atoms in humans. Neutralize a threat is the similar issue where it's just like, I think you would kill the humans if you didn't care at all about them. So maybe your question asking is like, why would you care at all about humans? But I think you don't have to care much to not kill the humans.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Taking shit from humans is a different, like marginalizing humans and causing humans to be irrelevant is a very different story from killing the humans. I think I'd say the actual incentives to kill the humans are quite weak. So I think the big reasons you kill humans are like, well, one, you might kill humans if you're in a war with them. And it's hard to win the war without killing a bunch of humans. Like maybe most saliently here if you want to use some biological weapons or some crazy shit, that might just kill humans. I think you might kill humans just from totally destroying the ecosystems they're dependent on and it's slightly expensive to keep them alive anyway. You might kill humans just because you don't like them or like you literally want to like, yeah, I mean.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think I'd probably start by pushing back a bit on the humans wouldn't cooperate if they understood outcome or something. I would say one, even if you're looking at something like 10s of percent risk of takeover, humans may be fine with that, like a fair number of humans may be fine with that. To like if you're looking at certain takeover, but it's very unclear if that leads to death, like a bunch of humans may be fine with that. Like if we're just talking about, look, the AI systems are going to run the world, but it's not clear if they're going to murder people. How do you know? It's just a complicated question about AI psychology. And a lot of humans probably are fine with that. And I don't even know what the probability is there.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Again, you could tell different stories, but it seems so much easier. At some point, you don't need any. At some point, an AI system can just operate completely out of human supervision or something. But that's so far after the point where it's so much easier if you're just like, there are a bunch of humans. They don't love each other that much. Some humans are happy to be on side. They're either skeptical about risk or happy to make this trade or can be fooled or can be coerced or whatever. And it just seems like it is almost certainly the easiest first pass is going to involve having a bunch of humans who are happy to work with you. So yeah, I think that probably is about. I think it's not necessary, but if you ask about the median scenario, it involves a bunch of humans working with AI systems, either being directed by AI systems, providing compute to AI systems, providing legal cover in jurisdictions that are sympathetic to AI systems.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Gap grows so if people are deliberately right if people are deploying AI everywhere think of this competitive dynamic if people aren't deploying AI everywhere so like if countries are not happy deploying AI in these high stakes settings then as AI improves you create this like web that grows where like if you were in a position of fighting against an AI which wasn't constrained in this way you'd be in a pretty bad position At some point even if you just yeah, so that's like one important thing just like I think in conflict and like overt conflict. If humans are putting the brakes on AI they're at like a pretty major disadvantage compared to an AI system that can kind of set up shop and operate independently from humans”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So again, there's going to be a lot of scenarios, and I'll just start by talking about one scenario which will represent a tiny fraction of probability or whatever. So if you're not in this competitive world, if you're saying we're actually slowing down deployment of AI because we think it's unsafe or whatever, then in some sense you're creating this very fundamental instability where you could have been making faster AI progress and you could have been deploying AI faster. And so in that world, the bad thing that happens if you have an AI system that wants to mess with you is the AI system says, I don't have any compunctions about rapid deployment of AI or rapid AI progress. So the thing you want to do or the AI wants to do is just say, I'm going to defect from this regime, like all the humans have agreed that we're not deploying AI in ways that would be dangerous. But if I as an AI can escape and just go set up my own shop, like make a bunch of copies of myself, maybe the humans didn't want to delegate warfighting to an AI, but I as an AI, I'm pretty happy doing so. Like, I'm happy if I'm able to grab some military equipment or direct some humans to use AI, use myself to direct it. And so I think as that”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I guess this brings you to like a third part of the language competitive dynamics are relevant. So it's both the question of can you turn off AI systems in response to something bad happening where competitive dynamics may make it hard to turn off. There's a further question of just like, why were you deploying systems for which you had very little ability to control or understand those systems? And again, it's possible you just don't understand what's going on. You think you can understand or control such systems. But I think in practice, a significant part is going to be like, you are doing the calculus with people deploying systems or doing the calculus as they do today, like in many cases overtly, of like, look, These systems are not very well controlled or understood. Develop AI”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Could just prevent humans from turning them off. But I think in practice, the one that's going to happen much, much sooner is probably competition amongst different actors using AI. And it's like a very, very expensive to unilaterally disarm. You can't be like something weird has happened. We're just going to shut off all the AI because you're EG in a hot war. So again, I think that's just probably the most likely thing to happen first. Things would go badly without it, but I think if you ask, why don't we turn off the AI, my best guess is because there are a bunch of other AIs running around 2 a year lunch.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, I think that there's several levels of answer to that question. So, one is maybe three parts of my answer. Our first part is like, I'm just trying to tell what seems like the most likely story. I do think there's further failures that get you in the more distant future. So EG Eliaser will not talk that much about killer robots because he really wants to emphasize like, hey, if you've never built a killer robot, something crazy is still going to happen to you just like only four months later or whatever. So it's not really the way to analyze the failure. But if you want to ask what's the median world where something bad happens, I still do think this is like the best guess. Okay, so that's like part one of my answer. Part two of the answer was like in this proximal situation where something bad is happening. And you ask like, hey, why do humans not turn off the AI? You can imagine like two kinds of story. One is like the AI is able to prevent humans from turning them off and the other is like, in fact, we live in a world where it's incredibly challenging. Like there's a bunch of competitive dynamics or a bunch of reliance on AI systems. And so it's incredibly expensive to turn off AI systems. I think, again, you would eventually have the first problem. Like eventually AI systems”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so this is going to depend a lot on what kind of timeline you're imagining or the broad distribution, but I can fill in some random concrete option that is in itself very improbable. I mean, I think that one of the less dignified but maybe more plausible routes is like you just have a lot of AI control over critical systems even in like running a military. Have the scenario that's a little bit more just like a normal coup where you have a bunch of AI systems, they in fact operate. It's not the case that humans can really fight a war on their own. It's not the case that humans could defend them from an invasion on their own. So that is if you had invading army and you had your own robot army, you can't just be like, we're going to turn off the robots now because things are going wrong if you're in the middle of a war.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“The one I'm interested in the tale of some risks that can occur particularly soon, and I think risks that occur particularly soon are a little bit like you have a world where it's not properly deployed and then something crazy happens quickly. That said, if you ask, what's the median scenario where things go badly? I think it is like there's some lessening of our understanding of the world. I think in the default path, it's like very clear to humans that they have increasingly little grip on what's happening. I mean, I think already most humans have very little grip on what's happening. It's just that some other humans understand what's happening. Like, I don't know how almost any of the systems I interact with work at a very detailed way. So it's sort of clear to humanity as a whole that we sort of collectively don't understand most of what's happening except with AI assistance. And then that process just continues for a fair amount of time. And then there's a question of how abrupt an actual failure is. I do think it's reasonably likely that a failure itself would be abrupt. Like at some point, bad stuff starts happening that human can recognize as bad. And once things that are obviously bad start happening, then you have this bifurcation where either humans can use that to fix it and say, okay, AI behavior that led to this obviously bad stuff. Don't do more of that.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“They could successfully prevent humans from either understanding what's going on or from successfully retaking the data centers or whatever if the AI successfully grab control.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“And so, in some sense, your first problem here is like having these added systems who understand a bunch about what's happening, and you're only leverage, like, hey, AI, do something that works well. So you don't have a lever to be like, hey, do what I really want. You just have the system. You don't really understand. You can observe some outputs. Did it make money? And you're just optimizing or at least doing some fine-tuning to get the AI to use its understanding of that system to achieve that goal. So I think that's like your first risk factor. And once you're in that world, then I think there are all kinds of dynamics amongst AI systems. But again, humans aren't really observing. Humans can't really understand. Humans aren't really exerting any direct pressure on, only on outcomes. And then I think it's quite easy to be in a position where if AI systems started failing, it would be very, they could do a lot of harm very quickly. Humans aren't really able to prepare for or mitigate that potential harm because we don't really understand the systems in which they're acting. And then if AI systems...”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think that most realistic failures probably involve two factories interacting. One factor is like the world is pretty complicated and the humans mostly don't understand what's happening. So AI systems are writing code that's very hard for humans to understand maybe how it works at all, but more likely they understand roughly how it works. But there's a lot of complicated interactions. AI systems are running businesses that interact primarily with other AIs. They like doing SEO for AI search processes. They're running financial transactions, like thinking about how to trade with AI counterparties. And so you can have this world where even if humans kind of understand the jumping off point when this was all humans, like actual considerations of what's a good decision, like what code is going to work well and be durable or what marketing strategy is effective for selling to these other AIs or whatever is kind of just all mostly outside of sort of human's understanding. I think this is like a really important, again, when I think of like the most plausible scary scenarios. I think that's like one of the two big risk factors.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“But then, when deployed in the real world, if you're able to realize you're no longer being trained, you no longer have reason to do the kinds of things human want. You'd prefer be able to determine your own destiny, control your own competing hardware, et cetera, which I think probably emerged a little bit later than systems that try and get reward. And so we'll generalize in scary unpredictable ways to new situations. I don't know when those appear, but also, again, broad enough error bars that it's like conceivable for systems in the near future. I wouldn't put it less than one in a thousand for GPT-5, certainly.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“First mode of a system is taking actions that get reward and overpowering or deceiving humans is helpful for getting reward. There's this other failure mode, another family failure modes where AI systems want something potentially unrelated to reward. I understand that they're being trained. And while you're being trained, there are a bunch of.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“The scenarios I am most interested in and most people are concerned about from a catastrophic risk perspective involve systems understanding that they are taking actions which a human would penalize if the human was aware of what's going on, such that you have to either deceive humans about what's happening, or you need to actively subvert human attempts to correct your behavior. So the failures come from really this combination, or they require this combination of both trying to do something humans don't like and understanding the humans would stop you. I think you can have only the barest examples. You can have the barest examples for GPT-4. Like you can create the situations where GPT-4 will be like, sure, in that situation, like here's what I would do. I would like go hack the computer and change my reward. Or in fact, we'll do things that are simple hacks or go change the source of this file or whatever to get a higher reward. They're pretty weak examples. I think it's plausible GPT-5 will have compelling examples of those phenomena. I really don't know. This is very related to the very broad error bars on how competent such systems will be when. That's all with respect to this.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Good example can you make now of this phenomena? I think the answer is kind of okay probably. So that just I think is going to continuously get better from here. I think for the level where we're concerned, this is related to me having really broad distributions over how smart models are. I think it's not out of the question that you take GPT-4's understanding of the world is much crisper and much better than GPT-3's understanding. Just like it's really like night and day. And so it would not be that crazy to me. If you took GPT-5 and you trained it to get a bunch of reward and it was actually like, okay, my goal is not doing the kind of thing which thematically looks nice to humans. My goal is getting a bunch of reward. And then we'll generalize in a new situation to get reward.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“In fact, looks at a bunch of environments, is able to understand the mechanism of reward provision as a common feature of those environments, is able to think in some novel environment, like, hey, which actions would result in me getting a high reward and is thinking about that concept precisely enough that when it says high reward, it's saying like, okay, well, how is reward actually computed? It's like some actual physical process being implemented in the world. My guess would be like GPT-4 is about at the level where with handholding, you can observe this kind of like scary generalizations of this type, although I think they haven't been shown basically. That is, you can have a system which, in fact, is fine-tuned not a bunch of cases. And then some new case will try and do an end rund around humans, even in a way humans would penalize if they were able to notice it or would have penalized in training environments. So I think GPT-4 is kind of at the boundary where these things are possible. Examples kind of exist, but are getting significantly better over time. I'm very excited about there's a Centropic project basically trying to see how”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think there's a couple possible stories for getting to catastrophic misalignment, and they have slightly different answers to this question. So maybe I'll just briefly describe two stories and try and talk about when they start making sense to me. So one type of story is you train or fine-tune your AI system to do things that humans will rate highly or get other kinds of reward in a broad diversity of situations. And then it learns to, in general, drops in some new situation, try and figure out which actions would receive a high reward or whatever, and then take those actions. And then when deployed in the real world, like sort of gaining control of its own training data provision process is something that gets a very high reward. And so it does that. So this is like one kind of story. It wants to grab the reward button or whatever. It wants to intimidate the humans into giving it a high reward, et cetera. I think that doesn't really require that much. This basically requires a system which is like...”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think even for GPT 4, it's reasonable to ask questions are there cases where GPT-4 knows that humans don't want X, but it does X anyway? Like where it's like, well, I know that I could give this answer, which is misleading, and if it was explained to human what was happening, they wouldn't want that to be done, but I'm going to produce it. I think that GPT-4 understands things enough that you can have that misalignment in that sense. Yeah, I think GPT, like I've sometimes talked about being benign instead of aligned, meaning that like, well, it's not exactly clear if it's aligned or if that context is meaningful. It's just kind of a messy word to use in general. But the thing we're more confident of is it's like not doing optimizing for this goal, which is like across purposes to humans. It's either optimizing for nothing, or like maybe it's optimizing for what humans want or close enough or something that's an approximation good enough to still not take over. But anyway, I'm like, some of these abstractions seem like they do apply to GPT-4. It seems like probably it's not egregiously misaligned.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“That surprising to me. I think this is mostly like a pro short timelines view. It's not that surprising to me if you tell me like. Machine learning systems are like three or four orders of magnitude less efficient at learning than human brains. I'm like, that actually seems like kind of indistribution for other stuff. And if that's your view, then I think you're probably going to hit, then you're looking at 10 to the 27 training compute or something like that, which is not so far.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't know. I just mostly just curious about what the performance. I think most people would object to be like, how did you choose these reference classes of things that are fair intersections? Some of them seem reasonable, like eyes versus cameras seem like just everyone needs eyes. Everyone needs cameras. It feels very fair. Photosynthesis seems like very reasonable. Everyone needs to take solar energy and then turn it into a usable form of energy. But it was just kind of, I don't really have a mechanistic story. Evolution in principle has spent like way, way more time than we have designing. It's absolutely unclear how that's going to shake out. My guess would be in general, like I think there aren't that many things where humans really crush evolution, where you can't take a pretty simple story about why. So, for example, roads and moving over roads with wheels crushes evolution, but it's not like an animal would have wanted to design a wheel. Like you're just not allowed to pave the world and then put things on wheels if you're an animal. Maybe planes or more. Anyway, whatever. There's various things you could try and tell. Something seems to be better up. It's normally pretty clear why humans are able to win when humans are able to win. The point of all this was like, it's”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, that's, I mean, yeah. So it's a little bit hard to say exactly what the task definit is there. Like you could say, like making a bone, we can't make a bone, but you could try and compare a bone, the performance characteristics of a bone is something else. Like, we can't make spider silk. You could try and compare the performance characteristics with spider silks, like things that we can synthesize.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So, like very rough ballpark, it was like sort of, for the most extreme things, you were looking at like five or six orders of magnitude. And that would especially come in energy cost of manufacturing, where bodies are just very good at building complicated organs, like extremely cheaply. And then for other things like leafs.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think that the story there would be a brain is just in some other part of the parameter space where it's like using a lot of compute for each piece of data it gets and then just not seeing very much data in total. Yeah, it's not really plausible if you extrapolate out language models. You're going to end up with like a performance profile similar to a brain. I don't know how much better it is. I think, so I did this random investigation at one point where I was like, how good are things made by evolution compared to things made by humans, which is a pretty insane seeming exercise. But like, I don't know. It seems like orders of magnitude is typical. Like not tons of orders of magnitude, not factors of two, like things by humans or a thousand times more expensive to make or a thousand times heavier per unit performance. If you look at things like how good our solar panels relative to leaves or how good are muscles relative to motors or how good are livers relative to systems that perform analogous chemical reactions and industrial settings.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Of baits of him, and we'll see is 2024. Mostly, this was organized around total operations performed in a brain, right? Okay, never mind.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Means like the average synapse is thing like 630 million, and I don't know exactly what the numbers are, but something that's ballpark, let's call it like a billion action potentials. And then there's some resolution in each of those carries some bits, but let's say it carries 10 bits or something just from timing information at the resolution you have available. Then you're looking at 10 billion bits. So each parameter is kind of like, how much is a parameter seeing? It's like not seeing that much. So then you can compare that to language. I think that's probably less than current language models C. And current language models are, so it's like not clear of a huge gap here, but I think it's pretty clear you're going to have a gap of at least three or four of magnitude.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I haven't thought that much about the sample efficiency question in a long time. But if you thought a synapse was seeing something like a neuron firing once per second, then how many seconds are there in a human life?”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, I think the way I would put it is like the number of bits it takes to specify the learning algorithm to change GPT for is like very small. And you might wonder maybe a genome, like the number of bits it would take to specify a brain is also very small. And the genome is much, much faster than that. But it is also just plausible that a genome is closer to certainly the space, the amount of space to put complexity in a genome. We could ask how well evolution uses it. And I have no idea whatsoever. But the amount of space in a genome is very, very vast compared to the number of bits that are actually taken to specify the architecture or optimization procedure and so on for GPT-4. Just because, again, genome is simple, but algorithms are really very simple. ML algorithms are really very simple.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I just don't know. It seems really hard to guess how much of that plays a role. The most important decisions are probably from an algorithm design perspective are not like the protein coding part is less important than the decisions about what happens during development or how cells differentiate. I don't know if that's biology satisfaction. I'm happy to run with 100 million base pairs though.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't. Yeah, so I don't know much about biology. In particular, I guess the question is, how many of those bits are productive for shaping development of a brain? And presumably a significant part of the non-protein coding genome.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, I would guess a complicated mess of a lot of things. In some sense, there's not that much going on in the brain. Like, as you say, there's just not that many bytes in a genome. But there's very, very few bytes in an ML algorithm. Like if you think a genome is like a billion bytes or whatever, maybe you think less, maybe you think it's like 100 million bytes, then like. An ML algorithm is like if compressed Probably more like Hundreds of thousands of bytes or something. Like the total complexity of here's how you train GPT 4 is just like, I haven't thought about these numbers, but it's very, very small compared to a genome. And so although a genome is very simple, it's very, very complicated compared to algorithms that humans design, like really hideously more complicated than algorithm a human would design.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, the most obvious one is just ask how much data it takes a human to become an expert in some domain. And it's much, much smaller than the amount of data that's going to be needed on any plausible trend extrapolation.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Neither analogy is that great. Like, I like them both and lean on them a bunch of them a bunch. I think that's been pretty good for having a reasonable view of what's likely to happen. That said, the human genome is not that much like 100 trillion parameter model. It's like a much smaller number of parameters that behave in a much more confusing way. Evolution did a lot more optimization, especially over long designing your brain to work well over a lifetime than gradient descent does over models. That's like a disalogy on that side. And on the other side, I think human learning over the course of human lifetime is in many ways just much, much better than gradient descent over the space of neural nets. Like gradient descent is working really well, but I think we can just be quite confident that in a lot of ways human learning is much better. Human learning is also constrained. We just don't get to see much data and that's just an engineering constraint that you can relax. You can just give your neural nets way more data than humans have access to.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I guess the way that you could think of this is like, I think both analogies are reasonable. One analogy being evolution is like a training run and humans are like the end product of that training run. And a second analogy is like evolution is like an algorithm designer. And then a human over the course of this modest amount of computation over their lifetime. Is the algorithm being that's been produced, the learning algorithm that's been produced? And I think like...”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Try and guess some combination of fitting that curve and yeah, try and some combination of fitting the curve to historical progress, looking at like how much low-hanging fruit there is, getting a sense of how fast it decays. I think you probably get a lot though. You get a bunch of orders of magnitude of total, especially if you ask, how good is a GPT-5 scale model or GPT-4 scale model? I think you probably get by 2040, like, I don't know, three orders of magnitude of effective training compute improvement, or a good chunk of effective training compute improvement, four orders of magnitude. I don't know. I don't have like here I'm speaking from no private information about the last couple years of efficiency improvements. And so people who are on the ground will have better senses of exactly how rapid returns are and so on.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I'm with you on sooner or later. I suspect, like Progress too slow if you like held fixed, how many people were working in the field, I would expect progress to slow as low-hanging fruit is exhausted. I think the rapid rate of progress in, say, language modeling over the last four years is largely sustained by you start from a relatively small amount of investment, you like greatly scale up the amount of investment. And that enables you to keep picking. Every time the difficulty doubles, you just double the size of the field. I think that dynamic can hold up for some time longer. I'm in a pretty good, right now if you think of it as like hundreds of people effectively searching for things, like up from anyway, if you think of it hundreds of people now, you can maybe bring that up to like tens of thousands of people or something. So for a while, you can just continue increasing the size of the field and like search harder and harder. And there's indeed a huge amount of looking and fruit where it wouldn't be that hard for a person to sit around and make things a couple percent better after year of work or whatever. So I don't know. I would probably think of it mostly in terms of like how much can investment be expanded and”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, something like that. That is, it just seems very plausible that it takes longer to train models to do tasks that are longer horizon.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So, in particular, you could supervise it from the next word prediction task, and all that context the human has, ultimately, will just help them predict the next word better. So in some sense, a really long context language model is also learning to do that task. But the number of effective data points you get of that task is vastly smaller, the number of effective data points you get at this very short horizon. What's the next word? What's the next sentence tasks?”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source