YouSaid · the spoken record

Joe Carlsmith

lines on the record
176
first
2024-08-22
most recent
2024-08-22
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I'm hoping that if we just do the anti-realism thing of just being consistent, learning all the stuff, reflecting, like, I don't, if you look at how moral realists and moral anti-realists actually do normative ethics, it's the same. It's basically the same. There's some amount of different curristics on things like properties like simplicity and stuff like that, but I think it's like they're mostly just doing the same game. And so I'm kind of hoping that, and also meta-ethics is itself a discipline that AIs can help us with. I'm hoping that we can just figure this out either way. So if there is, if moral realism is somehow true, I want us to be able to notice that I want us to be able to like adjust accordingly. So I'm not writing off those worlds and be like, let's just totally assume that's false. But the thing I really don't want to do is write off the other worlds where it's not true because my guess is it's not true. Right.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  2. To like discover if their religion is false. I would just in general if you would like to have your behavior be sensitive to whether something is true or false. Like it's sort of generally not good to like etch it into things. And so that is definitely a form of blinder I think we should be really watching out for. And I'm kind of hopeful. So like I have enough credence on some sort of moral realism that like

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  3. We move into this era of kind of aligning AIs, I don't actually think this binary between values and other things is going to be very obvious in how we're training them. I think it's going to be much more like ideologies. And like you can just train an AI to output stuff, right? Output utterances. And so you can easily end up in a situation where you decided that blah is true about some issue, an empirical issue, right? Not a moral issue. So I think people should not, for example, I do not think people should hard code belief in God into their AIs. Or like, I would advise people to not hard code their religion into their AIs if they also want.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Will they even be consistent in their values, right? I do think a thing we can do so, I like this image of the blinded horses, and I like this image of like maybe alignment is going to mess with the, I think we should be really concerned if we're like forcing facts on our AIs, right? Like that's like a really bad, because like I think one of the clearest things about human processes of reflection, like the kind of easiest thing to be like, let's at least get this is like not acting on the basis of a incorrect empirical picture of the world, right? And so if you find yourself like asking your A, by the way, like this is true, and I need you to always be reasoning as though blah is true, I'm like, ooh, I think that's a no-no from an anti-realist perspective too, right? Because I want to, I want to, like, my reflective values, I think, will be such that I formed them in light of the truth about the world. And so I think, and I think this is a real concern about as.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  5. And my actual prediction is that the AIs are going to be very malleable. Like, we're going to be like, you know, if you push an AI towards evil, it'll just go. And I think that we're sort of reflectively consistent evil. I mean, I think there's also a question with some of these AIs. It's like...

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I mean, look, it will be a very interesting fact if it's like, man, we keep training these AIs in all sorts of different ways. Like we're doing all this crazy stuff and they keep acting like bourgeois liberals. It's like, wow, that's, or, you know, they keep really, or they keep professing this weird alien reality. They all converge on this one thing. They're like, can't you see? It's like Zorgo. Like, Zorgo is the thing. And all the AIs, you know, interesting, very interesting. I think my personal prediction is that that's not what we see.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  7. It doesn't look to me like that's how human ethical reasoning works. I think most of what normative philosophy does is make consistent and kind of systematize pre-theoretic intuitions. And so we'll get evidence about this. In some sense, I think this view predicts, like, you know, you keep trying to train the AIs to do something. And they keep being like, no, I'm not going to do that. It's like, no, that's not good or something. They keep like pushing back. Like the sort of momentum of like AI cognition is always in the direction of this moral truth. And whenever we try to push it in some other direction, we'll find kind of resistance from the rational structure of things. So sorry, actually.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  8. You train a chess playing AI or you have a real paper clipper, right? Somehow you had a real paper clipper. And then you're like, okay, go and reflect. Based on my understanding of how moral reasoning works, like if you look at the type of moral reasoning that analytic ethicists do. It's just reflective equilibrium, right? They just like take their intuitions and they systematize them. I don't see how that process gets a sort of injection. Like the kind of mind independent moral truth, or like, I guess if you sort of start with only all of your intuitions say to maximize paperclips, I don't see how you end up maximizing or doing some rich human morality. I just don't

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Which is like, okay, dude, if it's all, if it's just going in a zillion directions, come on, you think it's going to go in your direction? Like, there's going to be so much churn. Like you're just going to lose. And so, you know, you should give up now and kind of only fight for the realism worlds. There I'm like, I mean, so I think, you know, you got to do the expected value calculation. You got to like actually have a view about like, how doomed are you in these different worlds? What's the tractability of changing different worlds? I mean, I'm quite skeptical of that. That's a kind of empirical claim. I will say I'm also just like kind of low on this everyone converges thing. So, you know, if you imagine like...

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Yeah, like I think, I mean, I go through in the essay a bunch of different ways in which I think this is wrong, but I think just like. And I think these people who kind of pronounce they're like moral realism or the void, like they don't actually think about bets like this. I'm like, no, no, okay, so really, like, is that what you want to do? And no, I think we should, I still care about my value. My sort of allegiance to my values, I think, is kind of outstrips my commitments to various meta-ethical interpretations of my values. I think we should. The sense in which we care about not being burned alive is much more solid than the reasoning and on what matters. Okay, so that's like the sort of philosophical doom. Now, you could have this, it sounded like you were also gesturing at a sort of empirical doom. Right.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  11. He later recanted. It's very hard. So I think this is importantly wrong. And so here's the case. I have an essay about this. It's called Against the Normative Realist Wager. And here's the case that convinces me. So imagine that a meta-ethical fairy. Appears before you, right? And this fairy knows whether there is a DAO. The fairy says, okay, I'm going to offer you a deal. If there is a DAO, then I'm going to give you a hundred dollars. If there isn't a DAO, then I'm going to burn you and your family and 100 innocent children alive. Okay, so claim don't take this deal. This is a bad deal. You're holding hostage your commitment to not being burned alive or like your care for that, to this abstruse. Basically, you're...

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Okay, well, let's just distinguish between ways you can be doomed. One way is kind of philosophical. So you could be the sort of moral realist or kind of realist-ish person, of which there are many who have the following intuition. They're like, if not moral realism, then nothing matters. It's dust and ashes. It is my metaphysics and or like normative view or the void, right? And I think this is. A common view. I think Derek Parfitt, at least some comments of Derek Sparfit's suggest this view, I think lots of moral realists will kind of profess this view. Elias or Yakowski, I think there are sort of some sense in which I think his early thinking was inflected with this sort of thought.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  13. The world's without that. The world's where there's no DAO. Yeah, yeah. Let's use the term Dao for this kind of convergent morality over the course of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  14. On this question of kind of do the worlds where there's not this kind of convergent moral force, whether metaphysically inflationary or not matter are those the only roles that matter

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Worlds where that isn't the case. And so I think there's a sense in which maybe that prediction is more likely for realism than antirealism, but it doesn't move me very much.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Has there been moral progress by Aristotle's lights or something and our lights too, right? And you could think, ah, isn't that a little bit like moral realism? It's like these hearts are singing in harmony. That's a moral realist thing, right? The anti-realist thing, the hearts all go different directions, but you and Aristotle apparently are both excited about the kind of march of history. Some open question about whether that's true. Like, what are Aristotle's reflective values? Suppose it is true. I think that's fairly explicable in moral anti-realist terms. You can say roughly that you and Aristotle are sufficiently similar and you endorse similar kind of reflective processes. And those processes are in fact instantiated in the march of history that history has been good for both of you. And I don't think that's, you know, I think.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Well, yeah, so there's obviously this first thing, which is like any, if you're the culmination of some process of moral change. Then it's very easy to look back at that process and be like moral progress, like the arc of history bends towards me. You can look more like if it was like, if there was a bunch of dice rolls around the way, you might be like, oh, wait, that's not rational. That's not the march of reason. So there's still empirical work you can do to tell whether that's what's going on. But I also think it's just on moral anti-realism, I think it's just still possible. Say, like, consider Aristotle and us, right? And we're like, okay.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yeah, so moral convergence, I think, is sort of a different factor from the existence or non existence of kind of non-natural, like a kind of morality that's not reducible to natural facts, which is the type of moral realism I usually consider. Okay, so does the improvement of society, is that an update towards moral realism? I mean, I guess like... So maybe it's like a very weak update or something. I guess I'm kind of like which view it predicts this hard. I guess it feels to me like moral anti-realism is very comfortable with the observation of the

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  19. One thing I want to flag is I don't think all forms of world realism make this prediction And so that's just one point. I'm happy to talk about the different forms I have in mind. I think there are also forms of kind of things that kind of look like moral anti-realism, at least in their metaphysics, according to me, but which just posit that, in fact, there's this convergence. It's not in virtue of interacting with some kind of mind-independent moral truth, but just like it's just for some other reason, it's the case that, and that looks like a lot, like moral realism at that point, because it's kind of like, oh, it's really universal. Like everyone ends up here and is kind of tempted to be like, ah, like, why, right? And then whatever answer for the why is a little bit like, is that the Tao? Is that the nature of the Tao, even if there's not sort of an extra metaphysical realm in which the moral lives or something?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  20. One of these bears. And I kind of really wanted to foreground that possibility in the series. I think we need to be talking about these things both at once. And bears can be moral patients, right? AIs can be moral patients. Nazis are moral patients. Enemy soldiers have souls, right? And so I think we need to learn the art of kind of hawk and dove both. There's this dynamic here that we need to be able to hold both sides of as we kind of go into these trade-offs and these dilemmas and all sorts of stuff. And part of what I'm trying to do in the series is really kind of bring it all to the table at once.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So I opened the series basically with an example that I'm really trying to conjure that possibility at the same time as conjuring the grounds of gentleness and the sense in which it is also the case that these AIs could be, they can be both be like others, moral patience, like this sort of new species in the sense of that should conjure like wonder and reverence and such that they will kill you. And so I have this example of like, ah, this documentary grizzly man where there's this environmental activist, Timothy Treadwell, and he aspires to approach these grizzly bears. He lives, you know, in the summary, he goes into Alaska and he lives with these grizzly bears and he aspires to approach them with this like gentleness and reverence. He doesn't use bear mace or he doesn't like carry bear mace. He doesn't use a fence around his camp. And he gets eaten alive by the bears.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Do think There's a concern here that I really try to foreground in the series that I think is related to what you're saying, which is something like it might be worried that we will be very gentle and nice and free with the AIs. And then they'll kill us. They'll take advantage of that and it will have been a catastrophe.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  23. I think the sort of moral justification you have for there's a different vibe in terms of like the kind of overall Yeah, justificatory stance you might have for various types of more kind of power exerting interventions. And so that's one feature of the situation.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Invading, you feel very justified in doing like a bunch of stuff to prevent that. It's like a little bit different when you're inventing the thing and you're doing it incautiously or something. And then you're also.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Reference class that some of that talk starts to conjure. And so basically just, yes, I think we should be very, we should really notice that. Part of what I'm trying to do in the series is to bring the kind of full range of considerations at stake into play, right? Like I think it is both the case that we should be quite concerned about being kind of overly controlling or abusive or oppressive or there's all sorts of ways you can go too far. And I think, you know, There are concerns about the AIs being genuinely dangerous and genuinely acting, killing us, finally overthrowing us. And I think the moral situation is quite complicated. And then I think in some sense... So often, if you imagine a sort of external aggressor who's coming in and

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Certainly, I think concerns in that vein are real. I mean, I think if you, it is disturbing how easy many of the analogies with kind of Human historical events and practices that we kind of deplore or at least have a lot of wariness towards are in the context of the kind of way you end up talking about AI maintaining control over AI, like making sure that it doesn't rebel. I think we should be noticing the kind of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Loss of control. Yeah. And I think we can do better than that, right? Like, let's work on it. Let's try. Let's try to do better, especially, you know, sort of, I think we can do better. And I think it might require being thoughtful and it might require being kind of having A kind of mature discourse about this before we start taking irreversible moves. But I'm optimistic that we can at least avoid some of the connotations and a lot of the stuff at stake in that kind of binary.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  28. There's a bunch of, yeah, so there's a bunch of stuff to say about that. I want to push back on the notion that there are sort of two options there's like enslaved God, whatever that is, and like

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Have kind of motivation, like slavery involves all this suffering and kind of non consent. There's all these specific dynamics involved in human slavery. But I think, and so some of those may or may not be present in a given case with AI. And I think that's important. But I think overall, we are going to need to stare hard at, like right now, the kind of default mode of how we treat AIs gives them no moral consideration at all, right? We were thinking of them as property, as tools, as products, and designing them to be assistants and stuff like that. And I think there's been no official communication from any AI developer as to when, under what circumstances, that would change, right? And so I think there's a conversation to be had there that we need to have. And so, and I think

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Think we as a civilization are going to have to have a very serious conversation about what sort of kind of Of servitude. Is appropriate or inappropriate in the context of AI development. And I think there are a bunch of disanalogies from human slavery that I think are important. In particular, A, the AIs might not be moral patients at all, in which case, so we need to figure that out. ways in which we may be able to kind of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Having kind of cooperative and kind of good balances of power and deals and kind of avoiding conflict, I think finding ways to set up structures at lots and lots of people and value systems and agents are happy with, including non-humans, people in the past, AIs, animals. I really think we should be like, we should have very broad sweep in thinking about what sorts of inclusivity we want to be kind of reflecting in a kind of mature civilization and kind of setting ourselves up for doing that

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  32. If you think about it, if everyone implements that rule, then we get potentially a big kind of Pareto improvement. I don't know exactly Pareto improvement, but it's like a good deal. A lot of good deals. And Yeah, so I think it's that. I'm just into pluralism. I've got uncertainty. There's all sorts of stuff swimming around there. And then I think also just as a matter of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Way that's really cheap for you, do it, right? Or, like, I don't know. I mean, obviously, you need to think about trade offs, and there's like a lot of people and principle you could be nice to, but I think the principle of being nice when it's cheap, I'm very excited to try to uphold. I also really hope that kind of other people uphold that with respect to me, including the AIs, right? Like, I think we should be kind of golden ruling. Like we're thinking about, oh, we're inventing these AIs. I think there's some way in which I'm trying to kind of embody attitudes towards them that I hope that they would embody towards me. And that's like some, it's unclear exactly what the ground of that is, but that's something, you know, I really like the golden rule. And I think a lot of that as a kind of basis for treatment of other beings. And so I think be nice when it's cheap is like a...

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  34. I think it's a bunch of things at once. So, yeah, I'm just, I'm really into being nice when it's cheap, right? Like, I think if you can just help someone a lot.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Vision of like how the future could be good from a very, very wide variety of perspectives. And I think the kind of vastness of space resources makes that. Makes that very feasible. And now, if you instead imagine A much smaller pie. Well, maybe you face a tougher trade offs. So I think that's like an important dynamic.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Comparatively small allocations of resources or something. Like we can just, you know, I kind of feel like everyone who has satiable values who will be really, really happy with some small kind of fraction of the available pie. We should just satiate all sorts of stuff. And obviously you need to do figure out gains from trade and balance. And there's like a bunch of complexity here. But I think in principle, we're in a position to create a really wonderful, wonderful scenario for just tons and tons of different value systems. And so I think correspondingly, we should be really interested in doing that, right? And, you know, so I sometimes use this heuristic in thinking about the future. I think we should be aspiring to really kind of leave no one behind, right? Like really find like, who are all the stakeholders here? How do we really have like a fully inclusive?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  37. That's a lot. And then I guess, but yeah, but I don't know if it like fundamentally changes the narrative. Like, I do think, I mean, obviously, the stakes insofar as you care about what happens in the future or in space, then the stakes are way smaller if you shrink down to the solar system. And I think that does change potentially some stuff in that like a really nice feature of our situation right now, depending on what the actual nature of kind of the kind of resource pie is, is that I think in some sense there's such an abundance of energy and other resources and principle available to a kind of responsible civilization that really just tons of stakeholders, especially ones who are like able to kind of saturate, get like really close to amazing according to their values with kind of

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  38. I mean, I think for most people, very little. I think people are really like, what's going to happen to this world? This world around us that we live in, what's that going to happen to me and my kids? And to, so I don't actually think some people spend a lot of time on the space stuff, but I think for the most immediately pressing stuff about AI doesn't require that at all. I also think, like, even if you bracket space, like. Time is also very big. And so, you know, whatever. We've got like 500 million years, billion years left on Earth if we don't mess with the sun and maybe you could get more out of it. So like, you know, I think there's still.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  39. So, I don't think it's a coincidence in that I think centrally, like the way we would Become able to expand, or the kind of most salient way to me is via some kind of radical acceleration of our.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Utility functions, human utility function. And there's like this competitor thing with utility functions. It's like somehow you lose touch with the kind of complexity of how we actually, like we've been dealing with kind of differences in values and kind of competitions for power, this is classic stuff, right? And I don't actually think the AIs sort of amplify a lot of the kind of dynamics, but I don't think it's sort of fundamentally new. And so part of what I'm trying to say is like, well, let's draw on our full, on the full wisdom we have here while obviously adjusting for ways in which things are different.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Trying to grip, like I think trying to kind of steer and grip and like kind of rent, you have a sense of the universe is about to kind of go off in some direction and you need to. And, you know, people notice that muscle. And part of what I want to do is like, well, we have a very rich ethical human ethical tradition of thinking about like what, when is it appropriate to try to exert what sorts of control over which things? And I want that to be, I want us to bring the kind of full force and richness of that tradition to this discussion, right? And not like I think it's easy if you're purely in this abstract mode of like.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  42. I do think it's important though to separate that concern from this other concern about where does the future eventually go and how much do we want to be kind of Trying to steer that actively. So, to some extent, I wrote the series partly in response to the thing you're talking about, which is I think it is true that aspects of this discourse involve the possibility of like...

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  43. That's true, then it's quite appropriate, I think, to really be like, okay, it is kind of, that's classic human stuff. Almost everyone recognizes that kind of self-defense or like ensuring kind of basic norms are adhered to is a kind of justified use of certain kinds of power that would often be unjustified in other contexts. So self-defense is a clear example there.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Towards against the sort of boundaries and cooperative structures that we've created as a civilization, right? I talk about the Nazis or in the piece. It's sort of like when you sort of invade something he's invading, we often think it's appropriate to fight back and we often think it's appropriate to set up structures to kind of prevent and kind of ensure that these basic norms of kind of peace and harmony are kind of adhered to. And I do think some of the kind of moral heft of some parts of the alignment discourse comes from drawing specifically on that aspect of our morality, right? So we think the AIs are presented as aggressors that are coming to kill you.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  45. I don't blame people for just talking about this one is just the first one when we think about in which context is it appropriate to try to exert various types of control or to kind of have more of what I call in the series yang, which is this kind of active kind of controlling force as opposed to yin, which is this more kind of receptive open letting go. A kind of paradigm context in which we think that is appropriate is if something is a kind of active aggressor

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Now, there's a second thing people sometimes want out of alignment, which is much broader, which is something like we would like it to be the case that our AIs are such that when we incorporate them into our society, things are good, right? That we just have a good future. I do agree that I think the discourse about AI alignment. Mixes together these two goals that I mentioned, the sort of most straightforward thing to focus on.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Things he did. When people talk about alignment, they have in mind a number of different types of goals, right? So, one type of goal is quite minimal. It's something like that the AIs don't kill everyone, that they, or kind of violently disempower people.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  48. And so I think the sense in which kind of stories about good futures that have to do with alignment are kind of about descendants, I think it's more about like whatever that seed is. How do we kind of carry it? How do we keep the life thread alive going and going into the future?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Yeah, when I think about it, I'm not assuming that there's some notion of like descendants or I think there's a kind of the thing that matters about the kind of lineage is this Whatever's required for kind of the kind of optimization processes to be in some sense pushing towards good stuff. And there's a kind of concern that is kind of currently a lot of what is sort of making that happen kind of lives in human civilization in some sense. And so we don't know exactly what there's some kind of seed of goodness that we're carrying in different ways or different people, there's different notions of goodness for different people maybe, but there's something seed that is currently here that we have that is not sort of just in the universe everywhere. It's not just going to crop up if you just sort of die out or something. It's something that is in some sense contingent to our civilization, or at least that's the picture we can talk about, whether that's right.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  50. So I agree. I mean, I think there's a few different things there, right? So there's kind of, what are you going for? You're going for actively good? Are you going for avoiding? Certain stuff, right? And then there's a different question, which is what counts as actively good according to you. Maybe some people are like the only things that are actively good. Are My grandchildren, or I don't know, like some like literal descending genetic line from me or something. I'm like, well, that sounds, that's not my thing. And I don't think it's really what most people have in mind when they talk about goodness. I mean, I mean, I think there's a conversation to be had. And obviously, in some sense, when we talk about a good future, we need to be thinking about like, what are all the stakeholders here and how does it all fit together? But I think

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source