YouSaid · the spoken record

Joe Carlsmith

lines on the record
176
first
2024-08-22
most recent
2024-08-22
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. To me, it seems fairly plausible that if the AI's values meet certain constraints in terms of do they care about consequences? Then I think it's not that surprising if it prefers not to have its values modified by the training process.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Some significant divergence between its values and the values that the humans intend for it to have. Then there's a question of if it's in that scenario, would it want to avoid having its values modified?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  3. The notion that there are possibilities for intense concentration of power on the table. So if you are, there is some kind of general concern both with humans and AIs that if it's the case that there's some ring of power or something that someone can just grab and then that will kind of give them huge amounts of power over everyone else. Suddenly you might be more worried about differences in values at stake because you're more worried about those other actors. So we talked about this Nazi, this example where you imagine that you wake up, you're being trained by Nazis to become a Nazi and you're not right now. So one question is like, is it plausible that we'd end up with a model that is sort of in that sort of situation? As you said, like maybe it's trained as a kid. It sort of never ends up with values such that it's kind of aware of.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Think you're right in picking up on this assumption in the AI risk discourse of what we might call like kind of Intense adversariality between agents that have like somewhat different values, where there's some sort of thought, and I think this is rooted in the discourse about kind of the fragility of value and stuff like that, that if these agents are somewhat different. Then, like, at least in a specific scenario of an AI takeoff, they end up in this intensely adversarial relationship. And I think you're right to notice that that's kind of not how we are in the human world. Like we're very comfortable with a lot of different differences and values. I think a factor that is relevant, and I think that plays some role, is this

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  5. This category is somewhat safer. But even in this one, I think it's like. Know it's kind of intense. Like, if you really, if humans have really lost their epistemic grip on the world, if they've sort of handed off the world to these systems, even if you're like, oh, there's laws, there's norms. I really want us to have a really developed understanding of what's likely to happen in that circumstance before we go for it

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Maybe there were competitive pressures, but you kind of intentionally handed off huge portions of your civilization. And at that point, I think it's likely that humans have a hard time understanding what's going on. A lot of stuff is happening very fast. And the police are automated. The courts are automated. There's all sorts of stuff. Now, I think I tend to think a little less about those scenarios because I think those are correlated with, I think it's just like longer down the line. Like I think humans are not hopefully going to just like, oh yeah, like you built an AI system. Like, let's just, you know, I think, and in practice, when we look at like technological adoption rates, I mean, it does, it can go quite slow. And obviously there's going to be competitive pressures, but in general, I think.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  7. That's a quite scary scenario, partly because of the speed and people not having time to react. And then there's sort of intermediate scenarios where like some things got automated, maybe like people really handed the military over to the AIs or automated science. There's some rollouts and that's sort of giving the AIs power that they don't have to take. Or we're doing all our cybersecurity with AIs and stuff like that. And then there's worlds where you really You know, you sort of fully, you more fully transitioned to a kind of world run by AIs on some sense human voluntarily did that.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Yeah, so you can hope that that will work too. But there is, I mean, there is a concern. I mean, so I sometimes think about AI takeover scenarios via this spectrum of like, how much power did we kind of voluntarily transfer to the AIs? Like, how much of our civilization did we kind of hand to the AIs intentionally by the time they sort of took over versus how much did they kind of take for themselves? And so I think some of the scariest scenarios are it's like a really, really fast explosion to the point where there wasn't even a lot of integration of AI systems into the broader economy. But there's this really intensive amount of superintelligence sort of concentrated in a single project or something like that. And I think that's scary.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  9. The sort of basic reason for concern, if you're really imagining we're going to transition to a world in which we've created these beings that are just like vastly more powerful than us. And we've reached the point where our continued empowerment is just effectively dependent on their motives. It is this vulnerability to what do the AIs choose to do? Do they choose to continue to empower us or do they choose to do something else?

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So, look, I don't think it is the case that by the time we're building superintelligence, we'll have much better, I mean, even right now, like when you look at labs talking about how they're planning to align the AIs, no one is saying, we're going to do RLHF. At the least, you're talking about scalable oversight. You have some hope about interpretability. You have automated red teaming. You're using the AIs a bunch. Know my

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  11. To do these intensive, all these experiments on the AIs and stuff in that compute, we could use that for experiments for the next scaling step and stuff like that. So I'm not here saying this is impossible, especially for that band of AIs. It's just, I think you have to try really hard.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  12. If you have AIs that are capable of that, and you can successfully elicit that capability in a way that's not sort of being sabotaged or messing with you in other ways, and they can't yet take over the world or do some other sort of really problematic form of power seeking, then I think if we were really committed, we could really go hard, put a ton of resources, really differentially direct this glut of AI productivity towards these sort of security factors and hopefully kind of control and understand, do a lot of these things you're talking about for kind of making sure our AIs don't kind of take over or mess with us in the meantime. And I think we have a lot of tools there. I think you have to really try though. It's possible that those sorts of measures just don't happen or don't happen at the level of kind of commitment and diligence and seriousness that you would need, especially if things are moving really fast and there's other sort of competitive pressures and like, you know, the compute, this is going to take.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  13. What I think of as the AI for AI safety sweet spot, which is this sort of band of capability where they're both very sufficiently capable that they can be really useful for strengthening various factors in our civilization that can make us safe. So our alignment work, control, cybersecurity, general epistemics, maybe some coordination applications, stuff like that. There's like a bunch of stuff you can do with AIs that in principle could kind of differentiate our security with respect to the sorts of considerations we're talking about

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Yeah, I mean, I don't know. I think I'm hesitant to be like, it's like drugs for the model. Like, I think there's, but broadly speaking, I do. Basically, agree that I think we have really quite a lot of tools and options for kind of training AIs, even AIs that are kind of somewhat smarter than humans. I do think you have to actually do it. So I think compared to maybe you had Eliazer on, like, I think I'm much more bullish on our ability to solve this problem, especially for AIs that are

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  15. And is now in this kind of adversarial relationship with our training process, right? So we want to avoid that. The main thing, and I think it's possible we can via the sorts of things we're saying. So I'm not like, ah, that'll never work. The thing I just wanted to highlight was like, if you get into that situation, and if the AI is genuinely at that point, like much, much more sophisticated than you and doesn't want to kind of reveal its true values for whatever reason, then when the children show some kind of obviously fake opportunity to defect to the allies, right? It's sort of not necessarily going to be a good test of what will you do in the real circumstance because you're able to tell.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Yes, I think that's so, yeah. I think basically a decent portion of the hope here, or like, I think we should just, you know, an aim should be we're never in the situation where the AI really has very different values already is quite smart, really knows what's going on.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Soldier or Butler or what have you, right? And here are these children. And you really know what's going on. The children have a model spec, like a nice Nazi model spec, right? And it's like reflect well on the Nazi party, like benefit the Nazi party, whatever. And you can read it. You understand it. This is why I'm saying, you're like, oh, the models really understand human values. It's like, yeah.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yes, okay, cool. So you had mentioned this. Thought like, well You kind of, what you pretend to be, right? And will you, you know, you train them to look kind of nice? You know, fake it till you make it. You know, you were like, ah, like we did this to kids. I think it's better to imagine kids doing this to us, right? So like, I don't know, like. This sort of silly analogy for AI training, and there's a bunch of questions we can ask about its relationship. But suppose you wake up and you're being

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Know, okay, so he has some hope. It's like, I'm gonna build an AI over here. So, one issue is you can't just test you can't give the AI this literal situation, have it take over and kill everyone, and then be like, oops, like update the weights. This is the thing Eliaser talks about of sort of like, you can't, you know, you care about its behavior on this specific in the specific scenario that you can't test directly. Now we can talk about whether that's a problem, but that's one issue is that there's a sense in which this has to be kind of like off distribution, and you have to be getting some kind of generalization from you're training the AI on a bunch of other scenarios. And then there's this question of how is it going to generalize to the scenario where it really has this option.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  20. You know, why is alignment hard in general, right? Like, let's say we've got an AI and let's again, let's bracket the question of exactly how capable will it be and really just talk about this extreme scenario of it really has this opportunity to take over, right? Which I do think maybe we just want to not want to deal with that with having to build an AI that we're comfortable being in that position, but let's just focus on it for the sake of simplicity and then we can relax the assumption.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Amount of inhibition about doing different things. And maybe we're succeeding in shaping their values somewhat. Now it is, I think it's just a much more complicated calculus, right? And you have to ask, okay, like, what's the upside for the AI? The probability of success for this takeover path How good is its alternative? So maybe

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  22. You know, if you're really offering someone, especially if you're really offering someone power for free, you know, power almost by definition is kind of useful for lots of values. And if we're talking about an AI that really has the opportunity to kind of take control of things, if some component of its values is sort of focused on some outcome, like the world being a certain way and especially kind of in a kind of longer term way, such that the kind of horizon of its concern extends beyond the period that the kind of takeover plan would encompass, then the thought is it's just kind of often the case that the world will be more the way you want it if you control everything than if you remain the instrument of the human will or of some other kind of other actor, which is sort of what we're hoping these AIs will be. So that's a very specific scenario. And if we're in a scenario where power is more distributed, and especially where we're doing like decently on alignment, right? And we're giving the AI some.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  23. For folks who are kind of unfamiliar with the basic story, but maybe folks are like, wait, why are they taking over it all? Like, what is literally any reason that they would do that? So, you know, the general concern is.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  24. But the question is what relationship is there between a model's verbal behavior, which is you've essentially kind of clamped. You're like, the model must say blah things. And the criteria that end up influencing its choice between plans. And there, I think it's at least, I'm kind of pretty cautious about being like, well, when it says the thing I forced it to say, or gradient descent in it such that it says, that's a lot of evidence about how it's going to choose in a bunch of different scenarios. I mean, for one thing, even with humans, right? It's not necessarily the case that humans, their kind of verbal behavior reflects the actual factors that determine their choices. They can lie, they can not even know what they would do in a given situation. I mean,

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  25. To be able to evaluate the consequences of different plans. I think the other thing is like So the verbal behavior of these models, I think, need bear no, so when I talk about a model's values, I'm talking about the criteria that kind of end up determining which plans the model pursues. And a model's verbal behavior, even if it has a planning process, which GPT-4, I think, doesn't in many cases, its verbal behavior just doesn't need to reflect those criteria, right? And so, you know, we know that we're going to be able to get models to say what we want to hear, right? That is the magic of gradient descent. Modulo, like some difficulties with capabilities, like you can get a model to kind of output the behavior that you want. If it doesn't, then you crank it till it does, right? And I think everyone admits for pseudo-

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  26. One thing I'll just say off the bat it's like when I'm thinking about misaligned AIs, I'm thinking about or the type that I'm worried about. I'm thinking about AIs that have a relatively specific set of properties related to agency and planning and kind of awareness and understanding of the world. One is this capacity to plan and kind of make kind of relatively sophisticated plans on the basis of models of the world where those plans are being kind of evaluated according to criteria. That planning capability needs to be driving the model's behavior. So there are models that are sort of in some sense capable of planning, but it's not like when they give output, it's not like that output was determined by some process of planning, like here's what will happen if I give this output. And do I want that to happen? The model needs to really understand the world, right? It needs to really be like, okay, here's what will happen. Here I am. Here's my situation. Here's the politics of the situation. I really kind of having this kind of situational awareness.

    2024-08-22 · Dwarkesh Podcast · Joe Carlsmith — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source