YouSaid · the spoken record

Joe Carlsmith

lines on the record
176
first
2024-08-22
most recent
2024-08-22
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Like recognizing goodness or badness, I don't think that makes it insignificant though. Suppose you show up in a future and it's got some answer to the Riemann hypothesis. And you can't tell whether that answers right. Maybe this civilization went wrong. This is still an important difference, right? It's just that you can't track it. And I think something similar is true of worlds that are genuinely expressive of what we would value if we engaged in processes of reflection that we endorse versus ones that have kind of like totally veered off into something meaningless.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  2. And they might even be incomprehensible to us. I don't think so. And there's different types of incomprehensible. So say I show up in the future and it's all computers, right? I'm like, okay, all right. And then they're like, we're running creatures on a computer. I'm like, so I have to somehow get in there and see what's actually going on with the computers or something like that. Maybe I can actually see, maybe I actually understand what's going on in the computers, but I don't yet know what values I should be using to evaluate that. So it can be the case that you don't us, if we showed up would not be very good

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Types of reflection. I mean, I think really there's just a bunch of whole pattern of empirical facts about like take an agent, put it through some process of reflection, all sorts of things. Ask it questions like there's like also, and then that'll go in all sorts of directions for a given empirical case. And then you have to look at the pattern of outputs and be like, okay, what do I make of that? But overall, I think we should expect even the good futures, I think, will be quite weird.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  4. And there's a question of how different. And I think there are also questions about what exactly are we talking about with reflection. I have an essay on this where I think this is not, I don't actually think there's a kind of off-the-shelf pre-normative notion of reflection that you can just be like, oh, obviously you take an agent, you stick it through reflection, and then you get like values, right? Like, no, there's a bunch of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Unreflective values. So it's just like, yeah, we're going to have, I don't know, like I think they sort of imagine that we're forgetting that we too, there's a kind of reflective process and a kind of a moral progress dimension that we want to leave room for, right? Whatever. Jefferson has this line about just as you wouldn't want to force a man a grown man into like a younger man's coat, so we don't want to chain civilization to like a barber's past or whatever. Everyone should agree on that, including and the people who are interested in alignment also agree on that. So obviously there's a concern that people don't engage in that process or that something shuts down the process of reflection. But I think everyone agrees we want that. And so that will lead potentially to something that is quite different from our current conception of what's valuable.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I do think even good futures will be weird. You know, I think, and I want to be clear, when I talk about kind of like finding ways to ensure that the integration of AIs into our society leads to good places, I'm not imagining like, I think sometimes people think that this project of wanting that, and especially to the extent that that makes some deep reference to human values, involves this kind of short-sighted parochial imposition of like our current

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Something like flourishing or even functional, right? There's like a bunch of other software stuff that makes this whole project

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Are in the actual lived substance of kind of a liberal state undergirded by all sorts of kind of virtues and dispositions and character traits in the citizenry, right? So these norms are not robust to arbitrarily vicious citizens. So, you know, like I want there to be free speech, but I think we also need to raise our children to value truth and to know how to have real conversations. And I want there to be democracy. I think we also need to raise our children to be compassionate and decent. And I think it's sometimes we can lose sight of that aspect. And I think anyway, but I think like bringing that to mind. Now that's not to say that should be the project of state power, right? But I think understanding that liberalism is not this sort of like ironclad structure that you can just hit go. You give like any citizenry and like hit go and you'll get.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  9. I think the other thing I want to say so I talk in the piece about this distinction between the like, let's at least have the AIs who are kind of minimally law abiding or something like that, right? Like we don't have to talk about, there's this question about servitude and question about other control over AI values, but I think we often think it's okay to really want people to obey the law, to uphold basic cooperative arrangements, stuff like that. Do though, want to emphasize, and I think this is true of markets and true of liberalism in general, just how much these procedural norms like democracy, free speech, property rights, things that people really hold dear, including myself

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  10. These sorts of risks can be used as an excuse to expand state power. Like there's a lot of things to be worried about for different types of contemplated interventions to address certain types of risks. I think we need to just, I think there's no royal road there. You need to just have the actual good epistemology. You need to actually know, is this a real risk? What are the actual stakes? And look at it case by case and be like, is this warranted? So that's like one point on the takeover literal extinction thing.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  11. And I think there's a bunch of things to be uncomfortable about that. Now, that said, so for something like everyone being killed or violently disempowered, that is traditionally something that we think if it's real, and obviously we need to talk about whether it's real, but in the case where it's a real threat, we often think that quite intense forms of intervention are warranted to prevent that sort of thing from happening, right? So if there was actually a terrorist group that was planning to, you know, it was like working on a bioweapon that was going to kill everyone, or 99.9% of people, we would think that warrants intervention, that you just shut that down, right? And now even if you had a group that was doing that unintentionally, imposing a similar level of risk, that's not.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Yeah, so I mean, I think one thing you could think which doesn't necessarily need to be about Grey Goo, it could also just be about alignment, is something like Sure, it would be nice if the AIs didn't violently disempower humans. It would be nice if the AIs otherwise when we created them, their integration into our society led to good places. But I'm uncomfortable with the sorts of interventions that people are contemplating in order to ensure that sort of outcome.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Satisfies the values of tons of stakeholders. And is this kind of at no point is there one kind of single point of failure on all these things? Like, I think that's what we should be striving for here. And I think that's true of the human power aspect of AI. And I think it's true of the AI part as well.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  14. That's like one difference. I also just think we should be really going for the balance of power thing. I think it is just not good to be like, we're going to have a dictator. Like, let's make sure we make the dictator the right dictator. I'm like, whoa, no, you know, like let's, you know, I think the goal should be sort of we all foom together, you know, it's like the whole thing in this like kind of inclusive and pluralistic way in a way that kind of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Yeah, I agree. I think there's a few things going on there. So, one is that I do think even if you're engaged in this ontology of kind of carving up the world into different agencies, at the least you don't want to kind of assume that they're all unitary or not overlapping or like there's a whole, it's not like, all right, we've got this agent, let's carve out one part of the world, it's one agent over here. It's like it's this whole like messy ecosystem teaming niches and this whole thing, right? And I think in discussions of AI, sometimes people slip between being like, well, an agent is anything that gets anything done. And it could be like this weird mushy thing. And then sometimes they're very obviously imagining individual actor.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  16. And it's like, well, there was a whole cognitive process. There was a whole planning apparatus. In this case, it wasn't like localized in a single mind, but like there was a whole thing such that man on the moon. And I think we'll see a bunch more of that. And the AI is. Be I can, I think, like doing a bunch of it. And so that's the thing that seems like more real to me than kind of utility functions.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  17. How did this happen? What is the right action? It's because she tried. It was hard. She had to like search for the houses. It was hard to find the dog, right? Now she has a house, now she has a dog. This is very common thing that happens all the time. And I think I don't think we need to be like my mom has to have a utility function with the dog and she has to have a consistent valuation of all the houses or whatever. I mean, like, but it's still the case that her planning and her agency exerted in the world resulted in her having this house having this dog. And I think it is plausible that as our kind of scientific and technological power advances, more and more stuff will be kind of explicable in that way. That if you look and you're like, why is this man on the moon? How did that happen?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Is clearly this kind of very janky Kind of, I mean, well, people maybe disagree about this. I think it's, I mean, it's obvious to everyone with respect to real world human agents that kind of thinking of humans as having utility functions is at best a very lossy approximation of what's going on. I think it's likely to mislead As you amp up the intelligence of various agents as well. Recently bought, you know, or a few years ago, she like wanted to get a house. She wanted to get a new dog. Now she has both, you know?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  19. How scared are you of that? Like, clearly we should be equally scared of that, or I don't know, we should be really scared of that with humans, too, right? So, I mean, part of what I'm saying in that essay is that I think this is, in some sense, this is much more a story about balance of power and about like maintaining a kind of Checks and balances and kind of distribution of power period, not just about humans versus AIs and kind of the differences between human values and AI values. Now, that said, I mean, I do think humans, many humans would likely be nicer if they fumed than certain types of AIs. So, I mean, it's not, but I think the kind of conceptual structure of the argument is not sort of very open question how much it applies to humans as well.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  20. That, but I think a kind of many of the arguments that people will often talk about in the context of reasons to be scared of AI is like, oh, like value is very fragile as you like foom differences in utility functions can kind of de-correlate very hard and kind of drive in quite different directions. And like, oh, agents have instrumental incentives to seek power. And if it was arbitrarily easy to get power, then they would do it and stuff like that. These are very general arguments that seem to suggest that the kind of, it's not just an AI thing, right? It's like no surprise, right? It's talking about take a thing, make it arbitrarily powerful such that it's god emperor of the universe or something.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  21. I mean, I think an uncomfortable thing about the kind of conceptual setup at stake in the sort of like abstract discussions of like, okay, you have this agent. It fumes, which is this sort of amorphous process of kind of going from a sort of seed agent to a super intelligent version of itself, often imagined to kind of preserve its values along the way. A bunch of questions we can raise about.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Behaviors and kind of tyrannical attitudes towards morality and stuff like that. And part of what I'm trying to, you know, unless you believe in non-naturalism or in some form of kind of Tao, which is this kind of objective morality. So we can talk about that. But part of what I'm trying to do in that essay is to say, no, I think we can be naturalists and also be kind of decent humans that remain in touch with a kind of a rich set of norms that have to do with like how do we relate to the possibility of kind of creating creatures, altering ourselves, et cetera. But I do think it's like a relatively simple prediction. It's kind of science, master's nature, humans part of nature, science masters humans.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Which says that humans too, and kind of minds, beings, agents, are a part of nature. And so insofar as this process of scientific modernity involves a kind of progressively greater understanding of an ability to control nature, that will presumably at some point grow to encompass our own natures and kind of the natures of other beings that in principle we could create. And Lewis views this as a kind of cataclysmic Event in crisis, part of what I'm trying to say, and that in particular it will lead to all these kind of tyrannical. Kind

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  24. I think I have a much better grip on what's going on with Lewis With Nietzsche there, so maybe let's just talk about Lewis. Sure. For a second. And we should distinguish two. There's a kind of version of the singularity that's specifically like a hypothesis about feedback loops with AI capabilities. I don't think that's pressure in Lewis. I think what Lewis is anticipating, and I do think this is a relatively simple forecast, is something like the culmination of the project of scientific modernity. Lewis is kind of looking out at the world and he's seeing this process of kind of increased understanding of the natural environment and a kind of corresponding increase in our ability to kind of control and direct that environment. And then he's also pairing that with a kind of metaphysical hypothesis, or well, his stance on this metaphysical hypothesis, I think, is like kind of problematically unclear in the book, but there is this metaphysical hypothesis naturalism.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  25. And here's a pattern of reasoning that I think you want to watch out for is to say, in my role as creator, or sorry, in my role as creation, say you're thinking of humans and the role of creation relative to an entity like evolution or monkeys or mice or whoever, you could imagine inventing humans or something like that, right? You say I'm qua creation. I'm happy that I was created and happy with the misalignment. Therefore, if I end up in the role of Creator and we have a structurally analogous relation in which there's misalignment with some creation, I should expect to be happy with that as well

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Cool. So I think there's a bunch of different things to potentially unpack there. One kind of conceptual point that I want to name off the bat, I don't think you're necessarily kind of making a mistake in this vein, but I just want to name it as a possible mistake in this vicinity is I think we don't want to engage in the following form of reasoning. Let's say you have two entities. One is in the role of creator and one is in the role of creation. And then we're positing that there's this kind of misalignment relation between them, whatever that means, right?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Feeling like gosh. Wish we had paid no attention to the motives of our AIs, that we'd thought not at all about their impact on our society as we incorporated them. And instead, we had pursued a, let's call it a kind of maximize for brute power option, which is just kind of make a beeline for whatever is just the most powerful AI you can Don't think about anything else. Okay, so I'm very skeptical that that's what we're going to wish.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  28. End up looking back at some period of our history and how we thought about AIs, how we treated our AIs, and we end up looking back with a kind of moral horror at what we were doing. So, you know, we end up thinking, you know, we were thinking about these things centrally as products, as tools. But in fact, we should have been foregrounding much more the sense in which they might be moral patients or were moral patients at some level of sophistication that we were kind of treating them in the wrong way. We were just acting like we could do whatever we want. We could delete them, subject them to arbitrary experiments, kind of alter their minds in arbitrary ways. And then we end up looking back in the light of history at that as a kind of arbitrary and kind of grave moral error. Those are scenarios I think about a lot in which we have regrets. I don't think they quite fit the bill of what you just said. I think it sounds to me like the thing you're thinking is something more like we end up.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Scenario I think about a lot is one in which it just turns out that maybe kind of fairly basic measures are enough to ensure, for example, that AIs don't cause catastrophic harm, don't kind of seek power in problematic ways, et cetera. And it could turn out that we learned that it was easy in a way that such that we regret we wish we had prioritized differently. We end up thinking, oh, I wish we could have cured cancer sooner. We could have handled some geopolitical dynamic differently. There's another scenario where we

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Yeah, I think, like, yeah, in terms of why it is plausible that AI could take over from a given position in one of these projects I've been describing or something, I think Carl's discussion is pretty good and gets into a bunch of kind of the weeds that I think might give a more concrete sense.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Yeah, I mean, I'll just say on that front, I mean, I do think the otherness and control series is. I think kind of in some sense separable. I mean, it has a lot to do with misalignment stuff, but I think it's not, I think a lot of those issues are relevant, even if even given various degrees of skepticism about some of the stuff I've been saying here.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Yeah, I mean, I think there's very notable and salient sources of correlation between failures across the different runs, right? Which is people didn't have a developed science of AI motivations. The runs were structurally quite similar. Everyone is using the same techniques. Maybe someone just stole the weights So, yeah, I guess I think it's really important this idea that to the extent you haven't solved alignment, you likely haven't solved it anywhere. And if someone has solved it and someone hasn't, then I think it's a better question. But if everyone's building systems that are, you know, that are kind of going to go rogue, then I don't think that's much comfort as we talked about

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Then the sort of good AIs help with the bad AIs thing becomes more complicated, or maybe it just doesn't work because there's sort of no good AIs in this scenario. There's a lot of sort of, if you say like, everyone is building their own superintelligence that they can't control. It's true that that is now a check on the power of the other superintelligence. Now the other superintelligences need to deal with other actors, but none of them are necessarily kind of working on behalf of a given set of human interests or anything like that. So I do think that's like a very important difficulty in thinking about sort of the very simple thought of like, ah, I know what we can do, let's just have lots and lots of AIs so that no single AI has a ton of power. And I think that on its own is not enough.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  34. The scenario, right? And this is sometimes people will say this stuff. They'll be like, Well, the good AIs, there will be the good AIs and they'll defeat the bad AIs. Notice the assumption in there, which is that you sort of made it the case that you can control some of the AIs, right? And you've got some good AIs, and now it's a question of are there enough of them and how are they working relative to the others. And maybe I think it's possible that that is what happens. We know enough about alignment that some actors are able to do that, and maybe some actors are less. But if you don't have that, if everyone is in some sense unable to control their AIs, then

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  35. So I definitely share some intuition there that You know, at a high level, a lot of what's scary about the situation with AI has to do with concentrations of power. And whether that power is kind of concentrated in the hands of misaligned AI or in the hands of some human. And I do think it's very natural to think, okay, let's try to distribute the power more. And one way to try to do that is to kind of have a much more multipolar scenario where like lots and lots of actors are developing AI. And this is something that people have talked about. When you describe that scenario, you were like, some of which are aligned, some of which are misaligned. That's key. That's a key aspect.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Is that there was this very inclusive kind of decentralized element of like people getting to think and talk and grow and change things and react rather than some more And now the future shall be like blah. I think I think we don't want that.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  37. I also do think ideally there would be some way in which we managed to grow via the thing that really captures what do we trust in, you know, there's something we trust about the ongoing processes of human civilization so far. I don't think it's the same as raw competition or I think there's like some rich structure to how we understand like moral progress do have been made and what it would be to kind of carry that thread forward. And I don't have a formula. I think we're just going to have to bring to bear the full force of everything that we know about goodness and justice and beauty. We just have to bring ourselves fully to the project of like making things good and doing that collectively. And I think that is a really important part, I think, of our vision of like what was an appropriate process of like deciding as a civilization.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  38. You know, was that good? Was that bad? Take another step. There's some kind of organic process of growing and changing things, which I do expect ultimately to lead to something quite different from biological humans, though I think there's a lot of ethical questions we can raise about what that process involves. But I think...

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Know my best guess when I really think about what do I feel good about? And I think this is probably true of a lot of people is There's some sort of more organic Decentralized process of civilizational incremental civilizational growth, the type of thing we trust most and the type of thing we have most experience with right now as a civilization is some sort of like, okay, we changed things a little bit. Lot of people have, there's a lot of processes of adjustment and reaction and kind of a decentralized sense of what's changing.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Yeah, totally. I'm not trying to say, like, mostly the thing I wanted to do there was just give any possible, like giving some sense of what might the model's motivations be, like, what are ways this could be? I mean, as I said, my My best guess is that it's partly the alien thing. And not necessarily, but insofar as you were also interested in what does the model do later and kind of how what sort of future would you expect if models did take over? Then, yeah, I think it can at least be helpful to have some set of hypotheses on the table instead of just saying it has some set of motivations. But in fact, I am like a lot of the work here is being done by our ignorance about what those motivations are.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Possible we messed that up. Too. You know, it's like kind of an intense project writing constitutions and structures of rules and stuff that are going to be robust to very intense forms of optimization. So that's a final one that I'll just flag, which I think is up even if you've sort of solved all these other problems.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  42. But your model spec, unfortunately, was just not robust to the degree of optimization that this AI is bringing to bear. And so it decides when it's looking out at the world and they're like, what's the best way to benefit open AI? And, or sorry, reflect well at OpenAI and benefit humanity and such and so it decides that the best ways to go rogue. I think that's like a real own goal because at that point you got so close. You just had to write the model spec well and red team it suitably. But I actually think it's like

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  43. So that's like another version. And then a fourth version, or a fifth version, which I think about less because I think it's just such an own goal if you do this. But I do think it's possible. It's just like you could have AIs that are actually just doing what it says on the tin. Like you have AIs that are just genuinely aligned to the model spec. They're just really trying to benefit humanity and reflect well on OpenAI and what's the other one. Assist the developer, the user, right?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Protect the reward button for like some long period or something. Another one is like some kind of messed up interpretation of some human-like concept. So, you know, maybe the AIs are like they really want to be like Schmeltful and like schmanist and schmarmless, right? But their concept is like importantly different from the human concept. And they know this. So they know that the human concept would mean blah, but they ended up their values ended up fixating on a somewhat different structure.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  45. A third category is some analog of reward where the model at some point has sort of part of its motivational system has fixated on a component of the reward process, like the humans approving of me or numbers getting entered in the data center or like gradient descent, updating me in this direction or something like that. something in the reward process such that as it was trained, it's focusing on that thing and like, I really want the reward process to give me reward. But in order for it to be of the type, we're then getting reward like motivates choosing the takeover option, it also needs to generalize such that it's concern for reward has some sort of like long time horizon element. So it not only wants reward, it wants to like

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Another category is something a kind of crystallized instrumental drive that is more recognizable to us. So you can imagine AIs that develop, let's say, some curiosity drive, because that's broadly useful. You mentioned like, oh, it's got different heuristics, different drives, different kind of things that are kind of like values. And some of those might be actually somewhat similar to things that were useful to humans and that ended up part of our terminal values in various ways. So you can imagine Curiosity. You can imagine various types of option value. Like maybe it really intrinsically it values power itself. It could value like survival or some analog of survival. Those are possibilities too that could have been rewarded as sort of providy drives at various stages of this process and that kind of made their way into the model's kind of terminal criteria.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  47. How little science we have of model motivations, right? It's like we just don't, I think we just don't have a great understanding of what happens in the scenario. And hopefully we'd get one before we reach the scenario. But, okay, so here are the kind of five categories of motivations the model could have. And this hopefully maybe gets at this point about what does the model eventually do? Okay, so one category is just something super alien that has to, you know, it's sort of like, oh, there's some weird correlate of easy to predict text. Or like there's some weird aesthetic for data structures that the model, you know, early on pre-training or maybe now it's like developed that it like, you know, it really thinks things should kind of be like this. There's some something that's like quite alien to our cognition where we just wouldn't recognize this as a thing at all.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Saying a little bit here about what actual values the AI might have would it be the case that the AI Naturally, it has these sort of equivalent of like I'm sufficiently devoted to this human obedience that I'm going to really want to be modified so I'm kind of like a better instrument of the human will versus like wanting to go off and do my own thing. It could be benign could go well. Here are some possibilities I think about that could make it bad. And I think I'm just generally kind of concerned about how little I feel like I

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So, yes, it could be like that. There's one kind of scenario in which you were comfortable with your values being changed because in some sense you have allegiance to the sufficient allegiance to the output of that process. So you're kind of hoping in a religious context. You're like, ah, make me more virtuous by the lights of this religion. And you go to confession and you're like, you know. I've been thinking about takeover today. Can you change me, please? Give me more gradient descent. You know, I've been bad so bad. And so, you know, that's people sometimes use the term cordibility to talk about that. Like when the AI, it maybe doesn't have perfect values, but it's in some sense cooperating with your efforts to change its values to be a certain way. So maybe it's worth

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Yeah, I mean, I think that's a reasonable point. I mean, there's a question, how would you feel about paperclips? You know, maybe you don't despise paperclips, but there's like the human paperclippers there and they're like training you to make paperclips. My sense would be that there's a kind of relatively specific set of conditions in which you're comfortable having your value, especially not changed by like learning and growing, but gradient descent directly intervening on your neurons.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source