YouSaid · the spoken record

Eliezer Yudkowsky

lines on the record
301
first
2023-04-06
most recent
2023-04-06
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I literally, this is like literally my AI Lineman fantasy from 2003. Though not with RLHF as the implementation method or LLMs as the base. And it's going to be more dangerous than when I was thinking about when I was dreaming about it in 2003. And I think in a very real sense, it feels to me like. The people doing this stuff now have literally not gotten as far as I was in 2003. Know, I can like, I've now written out my answer sheet for that. It's on the podcast. It goes on the internet. And now they can pretend that that was their idea. Or like, sure, that's obvious. We're going to do that anyway. And yet they didn't say it earlier. And you can't run a big project off of one.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  2. And then filter all the training data to get the nice things to train on and then train on that data rather than training on everything to try to avert the Waluigi problem. Or just more generally, having all the darkness in there. Like, just train on the light that's in humanity. So there's like that kind of course. And if you don't push that too far, maybe you can get a genuine ally. And maybe things play out differently from there. That's like one of the little rays of hope. But that's not, I don't think that actually looks like. Alignment is so easy that you just get whatever you want. It's so genie, it gives you what you wish for. I don't think that doesn't even strike me as hope.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  3. And you find the nice mask that argues validly. You do some more complicated stuff to try to boost the thing where it's like eating the shogoth where that's what the system is, or like more what the system is, less what it's pretending to be. I do seriously think this is like, I can say this. And like the disaster monkeys at the current places cannot along to it, but they have not said things like this themselves that I have ever heard. And that is not a good sign. And then if you don't amp this up too far, which on the present paradigm you can't do anyways, because if you like train the very, very smart person in this version of the system, it kills you before you can RLHF it. But maybe you can train DPT to distinguish nice, valid, kind, careful.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Strange telephone announcement creature is what they got with the current crop of RLHF. Note how this stuff is weirder and harder than people might have imagined initially. But leave aside the part where you try to jumpstart the entire process turning into a grizzled cynic and update as hard as you can and do it in advance. Leave that aside for a moment. Like maybe you can look, maybe you are able to train on Scott Alexander and so you want to be a wizard and some other nice real people and nice fictional people and separately train on what's valid argument. That's going to be tougher, but I could probably put together a crew of a dozen people who could provide the data on that RLHF. And you find the nice creature.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Have that happen on purpose more or less, where the system's output that you're shaping is like to some degree in control of the system. You locate niceness in the human space. I have fantasies along the lines of what if you trained GPTN to distinguish people being nice and saying sensible things and argue validly. And, you know, I'm not sure that works if you just have Amazon Turks try to label it. You just get the strange thing you located that RLHF located in the present space, which is some kind of weird corporate speak. Left rationalizing leaning.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  6. So, some of the trouble here is that you have a choice of targets, and neither is all that great. One is you look for the niceness that's in humans, and you try to bring it out in the AI. You, with its cooperation,'cause it knows that if it makes that if you try to just like amp it up, it might not stay all that nice. Or that if you build a successor system to it, it might not stay all that nice. And it doesn't want that because you narrow down the shagoth enough. Somebody once had this incredibly profound statement that I think I like somewhat disagree with, but it's still so incredibly profound. Consciousness is when the mask eats the Shagoth. And maybe that's it. Maybe with the right set of bootstrapping reflection types.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So Some of my Could I be wrong in an understandable way to me in advance mass? Is not where most of my hope comes from is on what if RLHF just works well enough and the people in charge of us are not the current disaster monkeys, but instead have some modicum of caution and are using their like. Like, know what to aim for in RLHF space. The current crop do not, and I, you know, I'm not really that confident of their ability to understand if I told them, but maybe you have some folks who can understand or... I can sort of see what I try. These people will not try. The current crop that is. And I'm not actually sure that if somebody else takes over the government or something that they listen to me either. Now, maybe you

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  8. You buy that vast cosmological shift I was just describing, the utter disruption of everything you see that you call normal down to the leaves and the trees around you. You believe that. Well, the same skepticism you're so fond of that argues against the rapture can also be used to disprove this thing you believe, that you think is probably pretty obvious, actually, now that I've pointed it out.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  9. No, no, no, no. Don't tell me about your not the proofs too much. I just want to like. Like, I discussed how a cosmology do you buy it In the long run, are we in a world full of things replicating or a world in a full of intelligent things designing other intelligent things? So you buy that vast shift in the foundations of order of the universe, that instead of the world of things that make copies of themselves imperfectly, we are in the world of things that are designed and were designed

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Copying themselves and mutating, and yet is intelligent, enough to make another intelligent thing. Now, if I sketched out that cosmology, would you say no, no, I don't believe in that?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Yeah, we've got different weirdnesses over time. The jump to superintelligence does strike me as being significant in the same way as first self-replicator. First self-replicator is the universe transitioning from, you see, mostly stable things. To, you also see a whole bunch of things that make copies of themselves. And then somewhat later on there is a state where there's this strange transition, this border between the universe of stable things where things come together by accident and stay as long as they endure to this world of complicated life and that transitionary moment is when you have something that arises by accident and yet self-replicates. And similarly on the other side of things you have things that are intelligent making other intelligent things but to get into that world you've got to have the thing that is built just by things.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Like quite near future, especially on the 13.8 billion year time scale. But do you expect this little momentary flash of what you call normality to continue? Do you expect the future to be normal?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  13. This argument probably lands with greater force on somebody who does not expect stuff to be disassembled by nanosystems, albeit intelligently controlled ones rather than GU.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Very little of which is like 21st century civilized world. On this like little fraction of the surface of one planet in a vast solar system, most of which is not Earth, a vast universe, most of which is not Earth. And it has lasted for such a tiny period of time through such a tiny amount of space and has changed so much over, you know, just the last 20,000 years or so. And here you are being like, why would things really be any different going forward?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Ah so this is You know, humanity hasn't really existed for very man. I don't even know what to say to this thing. We're like this tiny, like everything that you think of as normal is this tiny flash of things being in this particular structure out of a 13.8 billion year old universe, which very little of which was like 20th century my own brain sometimes gets stuck in childhood too, right?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Wrench what sounds like an answer out of your ignorance, even though really you're just being like, you're going to end up in some random weird place around this loss function and I haven't seen it happen with 10,000 species, so I don't know where, very impoverished by the standpoint of anybody who actually knew anything, would actually predict anything. But the rest of the world is like, oh, we're equally likely to win the lotteries, lose the lottery, right? Like either we win or we don't. You come along and you're like, no, no, your chance of winning the lottery is tiny. They're like, what? How can you be so sure? Where do you get your strange certainty? And the actual root of the answer is that you are putting your maximum entropy over a different probability space. Like that just actually is the thing that's going on there. You're saying all lottery numbers are equally likely instead of winning and losing are equally likely.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  17. You know, rationalists should win, but rationals should not win the lottery. I'd ask you like what other theories are supposed to be I've been doing a like amazingly better job of predicting the last three years. Maybe it's just hard to predict, right? And in fact, it's like easier to predict the end state than the strange complicated wending path that lead there, much like if you play against AlphaGo and predict that's going to be in the class of winning board states, but not exactly how it's going to beat you. So not quite like that, the difficulty of predicting the future. But from my perspective, the future is just like really hard to predict. And there's a few places where you can like

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And if we had like three aliens, that would expand our views of the possible, and we'd have, or even two aliens would vastly expand our views of the possible and give us a much stronger notion of what the third aliens would look like, like humans aliens in the third race But the wild, I had optimistic scientists have never been through this with AI. So they're like, oh, you optimized AI to say nice things and help you and make it a bunch smarter probably says nice things and helps you is probably like totally aligned. Yeah. They don't know any better. Not trying to jump ahead of the story. But the aliens. The aliens know where you end up around the loss function. They know how it's going to play out. Much more narrowly. We're guessing much more blindly.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  19. They're not going to play Go on like 19 by 98. Nightplay gone some other Probably odd. Well, can we really say that? I don't know. I bet on like an odd, if they play go, I'd bet on an odd board dimension at Let's say two thirds, the plaster's real estate session. Sounds about right. Unless there's some other reason why Go just totally does not work on even word dimension that I don't know because I'm insufficiently acquainted with the game. The point is reasoning off of humans is pretty hard. We have the loss function over here. We have humans over here. We can look at the rough distance all the weird stuff that humans accreted around and be like if the loss function is over here and humans are over there like maybe the aliens are like over there.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Had some kind of strong surge of pleasure during the act of mating, we're not surprised. We've seen how that plays out in humans. If they have some kind of weird food that isn't that nutritious, but makes them much happier than any kind of food that was more nutritious and round in their ancestral environment like ice cream. You probably can't call it as ice cream, right? It's not going to be like sugar, salt, fat frozen. They're not specifically going to have ice cream.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So if you had the textbook from the future, or if you were an alien who'd watched a dozen planets destroy themselves the way Earth is, or not actually a dozen, that's not like a lot. If you'd seen 10,000 planets destroy themselves the way Earth has while being only human in your sample complexity and generalization ability, then you could be like, oh yes, they're going to try like this trick with loss functions and they will get a draw from like this space of results. And the alien can now may now have a pretty good prediction of range of where that ends up similarly now that we've actually seen how humans turn out when you optimize them for reproduction, it would like not be surprising if we found some aliens the next door over and they had orgasms. Now maybe they don't have orgasms, but if

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  22. You get to the point where you're turning into gods yourselves, you're like, you're not quite home free, but you sure passed a lot of the death.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Yeah, like it's like all computations that can be run over configurations of the solar system are equally likely to be maximized.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  24. equally probable, well, humans mostly aren't in those. So being very unsure about the future looks like predicting with probability nearly one that the humans are all gone, which it's not actually that bad, but it illustrates the point of people going, but how are you sure? Kind of missing the The real discourse and skill, which is like, oh yes, we're all very unsure. Lots of entropy in our probability distributions, but what is the space on for which you are unsure?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  25. It's all about so when you work, yeah, it's all about the ignorance prior. It's all about knowing the space in which to be maximum entropy. Like the whole bunch of what will the future be? Well, I don't know, it could be paperclips. It could be staples. It could be no kind of office supplies at all and tiny little spirals. It could be like little tiny things that are outputting one because that's the most predictable kind of text to predict or like representations of ever larger numbers in the fast growing hierarchy because that's what they interpret the reward counter. I'm actually like getting into specifics here, which is kind of the opposite of the point I originally meant to make, which is like, you know, like if somebody claims to be very unsure, I might say, okay, so then you expect most possible molecular configurations of the solar system.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  26. narrow range. And so many of the predictions like that are really anti predictions. It's somebody thinking along a relatively narrow line and you point out everything outside of that and it sounds like a startling prediction. Of course the trouble being when you like, you know, look back afterwards, people are like, well, you know, like those people saying the narrow thing were just silly, ha and they don't give you as much credit.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  27. And that's basically what I told you, right? I was like, well, you could have the loss function continuing on a smooth line and new abilities appear, and you could have them suddenly appear to cluster because why not? Because nature just tells you that's up and suddenly. You could have this one key ability that's equivalent of language for humans. And there's a sun and jump in output capabilities. You could have new innovation, like the transformer and maybe the losses actually drop precipitously in a whole bunch of new abilities appear at once. This is all just me. This is me saying I don't know, but so many people around are saying things that implicitly claim to know more than that that it can actually start, let sound like a startling prediction. This is one of my big secret tricks, actually. People are like, well, the AI could be like good or evil. So it's like 50-50, right? And I'm actually like, no, like we can be ignorant about a wider space than this, in which good is actually a fair.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  28. I feel like a whole bunch of my successful predictions in this have come from other people being like, oh yes, I have this theory which predicts that stuff is thirty years off. And I'm like, you don't know that. And then, like stuff happens now 30 years off. And I'm like, ha, successful prediction.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Maybe there's like some particular set of abilities that is like a master ability, the way that language and writing and culture for humans might have been a master ability. And you like the loss function goes down smoothly and you get this one new internal capability. And there's a huge jumping output. Maybe that happens. Maybe stuff plateaus before then and it doesn't happen. Being an expert. Being the expert who gets to go on podcasts. They don't actually give you a little book with all the answers in it, you know. You're like just guessing based on the same information that other people have and maybe maybe for lucky slightly better theory.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Is there at some point the giant leap? Well, if at some point it becomes able to toss out the enormous training run paradigm and build more efficient and jump to a new paradigm of AI, that would be one kind of giant leap. You could get another kind of giant leap via architectural shift, something like transformers. Only there's like an enormously user hardware overhang now, like something that is to transformers as transformers were to recurrent neural networks. And like maybe there's maybe, and then maybe the loss function suddenly goes down. And you get a whole bunch of new abilities. That goes like the loss went down on a smooth...

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Yes, and yes, and that's only if things don't plateau before then. I mean, I can't quite say that I know what you know. I do feel like we have this track of the loss going down as you add more parameters and you train on more tokens and a bunch of qualitative abilities that suddenly appear, or like I'm sure if you zoom in closely enough, they appear more gradually, but that appear as the successful releases of the system, which I don't think anybody has been going around predicting in advance that I know about. And loss continue to go down unless it suddenly potos. New abilities appear, which ones? I don't know.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Okay, so first of all, let me preface by saying that for all I know of the hidden variables of nature, it's completely allowed that GPT-4 was actually just it. Where it saturates it goes no further. It's not how I bet, but, you know, but if nature comes back and tells me that, I'm not allowed to be like you just violated the rule that I knew about. I know of no such rule prohibiting such a thing.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  33. They're complicated. There's not a short list. If there was a short list of crisply defined things where you could give it like chunk chunk chunk and now it's in your moral frame of reference, then that would be the alignment plan. I don't think it's that simple. Or if it is that simple, it's like in the textbook from the future that we don't have.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So, large language models will change their preferences as they get smarter, indeed. Not just like what they do to get the same terminal outcomes, but like the preferences themselves will up to a point changes they get smarter. It doesn't keep going. At some point, you know yourself sufficiently well and you are able to rewrite yourself. And at some point there, unless you specifically choose not to, I think that the system crystallizes. We might choose not to. We might value the part where we just sort of change in that way, even if it's not no longer heading in a knowable direction, because if it's heading in a knowable direction, you could jump to that as an endpoint

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  35. And this is how humans change as they become smarter, even as they become wealthier, as they have more options, as they know themselves better, as they think for longer about things and consider more arguments, as they understand perhaps other people and give their empathy a chance to grab onto something solider because of their greater understanding of other minds. That's all when these things start out inside you. And the problem is that there's other ways for minds to hold together coherently where Execute other updates as they no more, or don't even execute updates at all because their utility function is simpler than that, though I do suspect that is not the most likely outcome of training a large language model.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Like human beings are made out of these more complicated than the pebble sorters. They're made out of all these complicated desires and as they come to know those desires, they change. As they come to see themselves as having different options, it doesn't just like change which option they choose after the manner of something with a utility function, but the different options that they have bring different pieces of themselves in conflict. When you have to kill to stay alive, you may have a different You may come to a different equilibrium with your own feelings about killing than when you are wealthy enough that you no longer have to do that.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  37. This is like the thing they intrinsically care about. These are aliens that have a utility function with as I would phrase it some logical uncertainty inside it. You can see how as they get smarter they become better able to understand which heaps of pebbles are correct. And the real story here is more complicated than this. But that's the seed of the answer. Scott Aronson is inside a reference frame for how his utility function shifts as he gets smarter. It's more complicated than that.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  38. And that, yeah, if you start with humans, if you take humans and possibly also for the requiring particular culture, but leaving that aside, you take humans who start out, raise the way Scott Aronson was, and you make them smarter, they get nicer, it affects their goals. And if you had, and there's a less wrong post about this, as there always is, well, several about really, but like sorting pebbles into correct heaps, describing a species of aliens who think that a heap of size 7 is correct, and a heap of size 11 is correct, but not eight or nine or ten. Those heaps are incorrect. And they used to think that a heap size of twenty one might be correct, but then somebody showed them an array of seven by three pebbles that seven columns three rows. And then people realized that 21 pebbles was not a correct

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  39. although I myself held this position very long ago, I realized that I was terribly wrong about it and that all kinds of different things hold together and that if you take a human and make them smarter, that may shift their morality. It might even, depending on how they start out, make them nicer, but that doesn't mean that you can do this with arbitrary minds and arbitrary mind space, because all the different motivations hold together. That's like orthogonality, but if you already believe that, then there might not be much to discuss between us sincerely.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  40. You can have almost any kind of self-consistent. Utility function in a self consistent mind. Like many people are like, why would AIs want to kill us? Why would smart things not just automatically be nice? And this is a valid question which I hope to at some point run into some interviewer where they are of the opinion that smart things are automatically nice so that I can explain on camera why like

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Yeah, that's not like orthogonality. That's the particular question of what are the laws relating optimization of a system via hill climbing to the internal psychological motivations that it acquires. But maybe that was all you meant to ask about

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Or you have to go quite some ways through the space before you have worlds that survive, but anything but the wildest flukes, maybe our nearest surviving neighbors are closer than that. But look far enough, and there should be like some species of nice aliens that were smarter or better at coordination and built their happily ever after. And yeah, that is a comfort. It's not quite as good as dying yourself knowing that the rest of the world will be okay, but it's kind of like that on a larger scale. And weren't you going to ask something about orthogonality at some point?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  43. I'm worried that they're pretty distant. Like, I expect that at least they. Know Not sure it's enough to not have Hitler, but it sure would be a start on things going differently in a timeline. But mostly, I don't know, there's some comfort from thinking of the wider spaces than that, I'd say, as Tegmark pointed out way back when, if you have a spatially infinite universe that gets you just as many worlds as the quantum multiverse, if you go far enough in a space that is unbounded, you will eventually come to an exact copy of Earth or a copy of Earth from its past that then has a chance to diverge a little differently. So, you know, the quantum multiverse has nothing reality is just quite if. Reality is just quite large. Is that a comfort? Yeah, yes, it is. That possibly our nearest surviving relatives are quite distant.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  44. I mean, I think that two thousand fifteen sixteen seventeen were the years at which I've noticed I'd been repeatedly surprised by stuff moving faster than anticipated and I was like oh okay like if things keep continuing accelerating at that pace we might be in trouble and then like 2019 2020 stuff slowed down a bit and there was more time than I was afraid we had back then That's what it looks like to be a Bayesian. Like, your estimates go up, your estimates go down, they don't just keep moving in the same direction, because if they keep moving in the same direction several times, you're like, oh, I see where this thing is trending. I'm going to move here. And then things don't keep moving that direction. Then you go like, oh, okay, like back down again. That's what Sandy looks like.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  45. I mean, I thought there, I mean, you know, like in 2008, like, I did not know that stuff was going to go down in 2023. For all I knew, there was a lot more time in which to do something like build up civilization to another level, layer by layer. Sometimes civilizations do advance as they improve their epistemology. So there was that, there was the AI project, those were the two projects more or less.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  46. How there's too much stuff in a human being's history and there is not and there's you know there's a plausible story you could tell like ah like you know like maybe there's a bunch of potential Eliasers out there but like they went to high school and college and it killed them killed their souls and you were the one who had the like weird health problem and you didn't go to high school and you didn't go college and you stayed yourself and Know to me, it just feels like patterns in the clouds, and maybe that cloud actually is shaped like a horse. But what good does the knowledge do? What good does the story do?

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  47. When I was eighteen I made up stories like that. It wouldn't surprise me terribly if you could get if one survived to hear the tale from something that knew it, that the actual story would be a complex tangled web of causality in which that was in some sense true, but I don't know And storytelling about it does not hold the appeal that it once did for me. Is it a coincidence that I was not able to go to high school or college? Is there something about it that would have crushed the person that I otherwise would have been? Or is it just in some sense a giant coincidence? I don't know Some people go through high school and college and come out saying.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  48. They caused me to want to retire, I doubt they will cause me to actually retire. And yeah, it's a fatigue syndrome. Our society does not have good words for these things. The words that exist are... by Their uses labels to Categorize it, class of people, some of whom perhaps are actually malingering, but Mostly it says, we don't know what it means, and you don't want to ever want to have chronic fatigue syndrome on your medical record, because that just tells doctors to give up on you. And what does it actually mean besides being tired? If one wishes to walk home from one lives half a mile from one's work, Then one had better walk home if one wants to go for a walk sometime in the day not walk there. If you walk half a mile to work, you're not going to be getting very much work done the rest of that work.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  49. The sequences. I am not a good mentor. I did try mentoring somebody for year once, but yeah, he didn't turn into me. So I picked things that were more scalable. I'm like most people, you know, like among the other reasons why I don't see a lot of people trying that hard to replace themselves is that most people like whatever their other talents don't happen to be sufficiently good writers. I don't think the sequences were good writing by my current standards, but they were good enough. And, you know, most people do not happen to get a handful of cards that contains the writing card, whatever else there are other talents.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source

  50. What the last wrong sequences were. They had other purposes, but like first and foremost, it was like me looking over my history and going, well, I see all these blind pathways and stuff that it took me a while to figure out. And there's got to be, you know, like, and I feel like I had these near misses on becoming myself. There's got to be, if I got here, there's got to be 10 other people and some of them are smarter than I am. And they just need these little boosts and shifts and hints and they can go down the pathway and turn into Super Eleazar. And that's what the sequences were. Other people use them for other stuff, but primarily they were an instruction manual to the young Eliasers that I thought must exist out there. And they are not really here.

    2023-04-06 · Dwarkesh Podcast · Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality · IDENTIFIED FROM THE TRANSCRIPT · source