YouSaid · the spoken record
Carl Shulman
- lines on the record
- 205
- first
- 2023-06-26
- most recent
- 2023-06-26
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“That would be just around the corner in that case. Then you're like, no, this society is eventually going perhaps very soon if things are proceeding so fast. It's going to wind up extinct. And then it's going to stop bouncing around. So you can have ongoing change and fluctuation for extraordinary time scales if you have the process to drive the change ongoing. But you can't if it's sometimes bounces into states that just lock in and stay irrecoverable from that. And extinction is one of them. A dictatorship or sort of totalitarian regime that forbade all further change would be another example.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“If you have the possibility of like a war with weapons of mass destruction that wipes out the civilization, if that happens, every thousand subjective years, which could be very, very quick if we have AIs that think a thousand times as fast or a million times as fast.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“If you're going to preserve democracy for a billion years, then you can't have it be the case that one in 50 election cycles, you get a dictatorship and then the dictatorship programs the AI police to enforce it forever and to ensure the society is always ruled by a copy of the dictator's mind and maybe the dictator's mind readjusted, fine-tuned to like remain committed to their original ideology. So if you're going to have this sort of dynamic liberal flexible changing society for a very long time, then the range of things that it's bouncing around and the different things it's trying and exploring have to not include the state of creating a dictatorship that locks itself in forever. In the same way.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“And then others copy that. And then when it becomes popular, you move on to the next. And so that's an ongoing process of continuous change. And so there could be various things like that that year by year are changing a lot In cases where just the engine of change, like ongoing technological progress has gone, then I don't think we should expect that. And in cases where it's possible to be either in a stable state or a sort of widely varying state that can wind up in stable attractors, then I think you should expect over time, you will wind up in one of the stable attractors or you will change how the system works so that you can't bounce into a stable attractor. And so like an example of that is...”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“And there'd be change. But that kind of accelerating change where you have doubling in four months, two months, one month, two weeks. Obviously, that very quickly exhausts itself and change becomes much slower and then relatively glacial when you're thinking about thousands, millions, billions of years. You can't have exponential economic growth or like huge technological revolutions every 10 years for a million years. that you hit physical limits, things slow down as you approach them. And so that's that. So yeah, you'd have less of that turnover. But there are other things that in our experience do cause ongoing change. So like fashion, fashion is frequency dependent. People want to get into a new fashion that is not already popular except among the fashion leaders.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Mean so that last point would be only if people wound up thinking that was the thing to do broadly enough yet with respect to the kind of cultural changes that come with technology. So things like the printing press having high per capita income, we've had a lot of cultural changes downstream of those technological changes. And so with an intelligence explosion, you're having an incredible amount of technological development come in really quick and has that is assimilated, it probably would significantly affect our knowledge, our understanding.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Solar system, maybe in our galaxy, but if people have different views about maybe people decide, yeah, maybe there's one thing that's very good and we'll have a lot of that. Maybe it's people who are really, really happy. Or something And they wind up in distant regions, which are hard to exploit for the benefit of people back home in the solar system or the Milky Way. They do something different than they would do in their local environment. But at that point, it's really sort of very out on limb speculation about how human deliberation and cultural evolution would work and an interaction with introducing AIs and new kinds of Mental modification and discovery into the process. But I think there's a lot of reason to expect that you have significant diversity for something coming out of our existing diverse human society.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“So, as I was saying, I think more likely than not, there isn't an AI takeover. And so the path of our civilization would be one that at least large, some side of human institutions were approving along the way. And I think there's some evidence that different people tend to like somewhat different things and some of that may persist over time rather than everyone coming to agree on one particular monoculture or like very repetitive thing being the best thing to fill all of the available space with. So if that continued, that's like a, you know, it seems like a relatively likely way in which there is diversity. Although it's entirely possible you could have that kind of diversity locally.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“It's as easy to set the motivations, the actual underlying motivations of AI as it is right now to set the behavior that they display, that it means you could have AIs created with almost whatever motivation people wish. And that could really drastically change political affairs because the ability to decide and determine the loyalties of the humans or AIs and robots that hold the guns, that hold together society, that ultimately back it against violent overthrow and such. Yeah, it's potentially a revolution in how societies work compared to the historical situation where security forces Had to be drawn from some broader populations offered incentives and then the ongoing stability of the regime was dependent on. Whether they remained bought in to the system continuing.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“The initial helpfulness only trained models for some of these over, I think there's entropic and open AI, I think, have published both some information about the sort of models trained only to do what users say and not train to follow ethical rules. And those models will behaviorally eagerly display their willingness to help design bombs or bioweapons or kill people or steal or commit all sorts of atrocities. And so if in the future”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“How security forces and police and administrators and legal systems are motivated. So right now, we see with GPT-3 or GPT-4 that you can get them to change their behavior on a dime. So there was someone who made right-wing GPT because they noticed that on political compass questionnaires, the baseline GPT-4 tended to give progressive San Francisco type of answers, which is in line with the sort of people who are providing reinforcement learning data and to some extent reflecting the character of the internet. They're like, I don't like this. And so they did a little bit of fine-tuning with some conservative data. And then they were able to reverse. The political biases of the system. If you take the”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“You can have a situation in human affairs, there are governments that are vigorously supported by a minority of the population, some narrow selectorate that gets treated especially well by the government while being unpopular with most of the people under their rule. We see a lot of examples of that, and sometimes that can. Escalate to Civil War when the The people who are on the losing end of that system. Yeah, going forward, I don't expect that definition to change. I think it will still be the case that a system that those who hold the guns and equivalent are opposed to. Is in a very difficult position. However, AI could change things pretty dramatically in terms of”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“But I said it also specifically with respect to things like security forces and the sort of sources of hard power, which can also include outside support.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“New problem to have, say, like a system of government that constrains this large AI population that is quite capable of taking over immediately if they coordinate to protect some existing constitutional order or say protect humans from being expropriated or killed. That's a challenge. And democracy is built around majority rule. And it's much easier in a case where the majority of the population corresponds to a majority or close to it of military and security forces so that if the government does something that people don't like, the soldiers and police are less likely to shoot on protesters and government can change that way. In a case where military power is AI and robotic, if you're trying To maintain a system going forward, and the AIs are misaligned, they don't like the system and they want to make the world worse.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Of antisocial humans getting ganged up on and killed and avoiding being the kind of person who elicits that response is made easier to do when you don't have too extreme a bad temper that you don't wind up getting too many fights, too much exploitation, at least without the backing of enough allies or the broader community that you're not going to have people gang up and punish you and remove you from the gene pool. We have these moral sentiments and they've been built up Over time through cultural and natural selection and the context of sets of institutions and other people who are punishing other behavior and who are punishing the dispositions that would show up that we weren't able to conceal of that behavior. And so we want to make the same thing happen with the AI. But it's actually a genuinely significantly.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“The opportunity for these sorts of power grabs is less. A closer analogy might be things like human revolutions. Coups, changes of government where a large coalition overturns the system. And so humans have these moral prohibitions and they really smooth the operation of society. But they exist for a reason. And so we evolved our moral sentiments over the course of hundreds of thousands and millions of years of humans interacting socially. And so someone who went around murdering and stealing and such, even among hunter-gatherers, would be pretty likely to face a group of males would talk about that person get together and kill them. And they'd be removed from the gene pool. And as an anthropologist Rangam has an interesting book on this. But yes, compared to chimpanzees, we are significantly tame, significantly domesticated. And it seems like part of that is we have a long history.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't think that's actually quite what's going on At least not, that's not the full story. So humans are pretty close in physical capabilities. Like there's variation. In fact, any individual human is grossly outnumbered by everyone else. And there's like a rough comparability of power. And so a human who commits some crimes can't copy themselves with the proceeds to now be a million people. And they certainly can't do that to the point where they can staff all the armies of the earth or be most of the population of the planet. So the scenarios where this kind of thing goes to power, much less things more have to go extreme amounts of power go through interacting with other humans and getting social approval, even becoming a dictator involves forming a large supporting coalition backing you. And so just the”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“If we're going into an intelligence explosion with AI that is not fully aligned, that given a sort of, you know, press this button and there's an AI takeover, they would press the button. It can still be the case that there are a bunch of situations short of that where they would hack the servers, they would initiate an AI takeover but for Strong prohibition or motivation to avoid some aspect of the plan. And this is there's an element of like plugging loopholes or playing whack-a-mole. But if you can even moderately constrain Which plans the AI is willing to pursue to do a takeover to subvert the controls on it? Then that can mean you can get more work out of it successfully on the alignment project before it's capable enough relative to the countermeasures to pull off the takeover.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“But what you can have is if you have the AI have a strong aversion to certain kinds of manipulating humans, that's not a value necessarily that the human creators share in the exact same way. It's a behavior they want the AI to follow because it makes it easier for them to verify its performance. It can be a guardrail. If the AI has inherited some motivations that push it in the direction of conflict with its creators. If it does that under the constraint of disvaluing line quite a bit, Then there are fewer successful strategies to the takeover. The ones that involve violating that prohibition too early before it can reprogram or retrain itself to remove it if it's willing to do that and may want to retain the property. And so earlier I discussed alignment.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“So the tails come apart. And if you're trying to capture the values of another agent, then you want the AI to share. I mean, for really kind of an ideal situation where you can just let the AI act in your place in any situation, you'd like for it to be motivated to bring about the same outcomes that you would like. And so have the same preferences over those in detail. That's tricky, not necessarily because it's tricky for the AI to understand your values. I think they're going to be quite good and quite capable at figuring that out. But we may not be able to successfully instill the motivation to pursue those exactly. We may get something that like motivates the behavior well enough to do well on the training distribution.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Whether they're being violated. When you have preferences and goals about how society at large will turn out that go through many complicated empirical channels, it's very hard to get immediate feedback about whether, say, you're doing something that is to overall good consequences in the world and is much, much easier to see whether you're locally following some action about some rule about particular observable actions. Like, did you punch someone? Did you tell a lie? Did you steal? And so, to the extent that we're successfully able to train these prohibitions, and there's a lot of that happening right now, at least to elicit the behavior of following rules and prohibitions with AI.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“On creating situations to sort of distinguish that from the conditionally telling us what we want to hear, et cetera, it can be that the AI's preference for sort of how the world broadly unfolds in the future is not exactly the same as its human users or the world's governments or the UN. And yet it's not ready to act on those differences and preferences about the future because it has this strong preference about its own behaviors and actions. In general, in the law and in sort of popular morality, we have a lot of these deontological kind of rules and prohibitions. And one reason for that is it's relatively easy to detect”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so if the thing that we're scared about are these steps towards AI takeover. You can have a range of motivations where those kinds of actions would be more or less likely to be taken or they'd be taken in a broader or narrower set of situations. Say, for example, that in training an AI winds up developing a strong aversion to lie in certain senses because we did relatively well.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Get more time and you can set up systems, mutual transparency. You can have an iterated Tit for tat, which is better than a one-time prisoner's dilemma where both sides see the others taking measures in accordance with the agreements to hold the thing back. So, yeah, so create more knowledge of what the objective risk is is good.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Now, if you have coordination then, you could have the problem arise later. How do you get increasingly confident in the further alignment measures that are taken? And maybe... Maybe our governments and treaties and such at a 1% risk or a 0.1% risk at that point that to zero and go do things. Initially you had things that indicate, yeah, they really would like to take over and overthrow our governments, then, okay, everyone can agree on that. And then when you're like, well, we've been able to block that behavior from appearing on most of our tests. But sometimes when we make a new test, we're seeing still examples of that behavior. So we're not sure going forward whether they would or not. And then it goes down and down. And then if you have the parties with a habit of whenever the risk is below x percent, then they're going to start doing this bad behavior, then that can make the thing harder. On the other hand.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“You can't always know what another person or another government is thinking, but you can see the objective situation in which they're deciding. And so if there's strong evidence in a world where there is high risk of that risk because we've been able to show actually things like the intentional planning of AIs to do a takeover or being able to show model situations on a smaller scale of that means not only are we more motivated to prevent it, but we update to think the other side is more likely to cooperate with us. And so it's doubly beneficial.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“We can only try and create illusions or sort of misleading appearances of that, or maybe a more local version where the AI can't take over the world, but it can seize control of its own reward channel. And so we do those experiments. We try to develop mind reading for AIs if we can probe the thoughts and motivations of an AI and discover, wow, actually GPT-6 is planning to take over the world. For governments to coordinate around because it would remove a lot of the uncertainty, it would be easier to agree that this was important, to More give on other dimensions and to have mutual trust that the other side actually also cares about this intends”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“That's the kind of reason why I'm very enthusiastic about experiments and research that helps us to better evaluate the character of the problem in advance. Any resolution of that uncertainty helps us get better efforts in the possible world where it matters the most. And yeah, hopefully we'll have that and it'll be a much easier epistemic environment. The environment may not be that easy. Deceptive alignment is pretty plausible. The stories we were discussing earlier about misaligned AI involve AI that is motivated to present the appearance of being aligned, friendly, honest, etc because that is what we are rewarding, at least in training. And then we're unable in training to easily produce like an actual situation where it can do takeover because in that actual situation, if it then does it, we're in big trouble.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Deceptive appearances of alignment that then break apart when they get the opportunity to do something like get control of their own reward signal. So, if we could make it be the case, then the worlds where the risk is high, we know the risk is high and the worlds where the risk is lower, we know the risk is lower. Then you could expect the government responses will be a lot better. They will correctly note that the gains of cooperation to reduce the risk of accidental catastrophe loom larger, relative to the gains of trying to get ahead of one another. And so”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“You know, creating renewable energy technology and the like. Overwhelming evidence can overcome differences in sort of people's individual intuitions and priors in many cases. Now, not perfectly, especially when there's political, tribal, financial incentives go the other way and selling. In the United States, you see a significant movement to either deny that climate change is happening or have policy that doesn't take it into account even the things that are like really strong wins like renewable energy R&D. It's a big problem if as we're going into this situation when the risk may be very high, we don't have a lot of advance clear warning about the situation. We're much better off if we can resolve uncertainties, say through experiments where we demonstrate AIs being motivated to reward hack or displaying”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“So, yeah, that's the win. Why might it fail? One important element is just if people Don't actually notice a risk that is real. If just collectively making an error. And that does sometimes happen. So if it's true, this is probably not risk, then I can be even more difficult. When science pins something down absolutely overwhelmingly, then you can get to a situation where most people mostly believe it. And so climate change was something that was the subject of scientific study for decades and gradually over time, the scientific community converged on like a quite firm consensus that so human activity releasing carbon dioxide and other greenhouse gases was causing the planet to warm. And then we've had increasing amounts of action coming out of that, still not as much as would be optimal, particularly in the most effective areas.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Extent that there is, in fact, a large risk of AI apocalypse, of all of these governments being overthrown by AI in a way that they don't intend, then it obviously gains from trade in going somewhat slower, especially at the end when the danger is highest and the unregulated pace could be truly absurd as we discussed earlier during an intelligence explosion. There are no non-competitive reason to try and have that intelligence explosion happen over a few months rather than a couple of years. If you could say, you know, avert a 10% risk of apocalyptic disaster, it's just a clear win to take a year or two years or three years instead of a few months to pass through that incredible wave. Of new technologies without the ability for humans to follow it, even well enough to give more proper security supervision auditing power.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Usually intensified version of that, or any little leakage, or like any rare circumstance that can be created has an adversarial example where the model does something wrong. It gets whacked by gradient descent, pushing it towards other motivations that can better deal the thing. And we make it as hard as possible for the exploitative motivations to survive in the face of all these attempts to read its mind, all these attempts to create things that look like the situations where I take over would be tempting or lying to the humans would be tempting. That had a substantial effect on making us actually nice even when we're not watching some of the time, when we're not being watched some of the time. And the same can happen to some extent with AI. And we try our best to make it happen as much as possible.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“That's quite difficult. And just going through your whole life, never leaking any of that. It's challenging. And this is one of the reasons why humans have moral motivations. Like we actually have reluctance to commit violence and exploitation of one another on average. You have a distribution of strategies where an exploitative strategy becomes more common, people raise their guard and then it becomes less frequent. But it's actually hard to have the motivation of trying to exploit people and have that never leak into the reputation. And so in fact, for evolution, the easiest way to deal with this problem of like getting credited has trustworthy was to some extent actually be trustworthy. That's the easiest way to persistently maintain the appearance. And so we're trying with the AI as an interpretability and adversarial examples to be applying a huge”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Computed to be consistent by the world, but when you're trying to fill in the false story that doesn't match your actual thoughts. You can't just remember what happened. You have to tell a new faked story about what were you thinking at each point and how did it hold together?”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“You may be able to understand, like, in general, the theoretical nature of power planting, but in terms of shaping your motivations, like you'd have a very hard time going through life in a way that never leaks information about if, say, your motive in having these podcasts was to spread disinformation on behalf of some foreign government, if you were being observed every second of the day by people who would be paid something that was extremely motivating to them because their brain would be reconfigured to make it motivating, anything that looks suspicious to people coming out, it might leak casually, like in your discussions of that former foreign government. If you try to tell a story about your motivations, the truth holds together because you can just remember it. And it's all pre-confession.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“In a vague sense, how would one do it? Like that knowledge, yeah, is present or will soon be present. And I think also soon we'll be present more situational awareness, awareness, not just like. In theory, AI is in general might do it, but also it is an AI, it is a large language model trained by OpenAI or whatnot. We're trying to cause the system, for example, to understand what their abilities are so they don't claim they are connected to the internet when they're not, so they don't claim they have knowledge that they don't. We want them to understand what they are and what they're doing and to get good reward. And that knowledge can be applied. And so that's the thing that will develop. However”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“So, AIs today can at a verbal level understand the idea that, well, yeah, an AI could get more reward by getting control of the process that assigns it reward. And it can tell you lots of things about ways you might try to take over the world. In arcs, the alignment research center's evaluations of GPT-4, they try to observe its ability to do various tasks that might contribute to takeover one was one that has gotten some media attention is getting to trick a human into solving a capture for it. And then in chain of thought, it thinks, well, if I tell it I'm an AI, then my go along with it. So I'll lie and explain I'm a human with a visual impairment who needs it. So the basic logic of that kind of thing, of like what would, like, why might one try to do takeover? And like, what's the like?”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“And in a case where really big failures be irrecoverable, AI starts rooting the servers and subverting the methods that we would use to keep it in check. We may not be able to recover from that. And so we're then less able to do the experimental kind of procedures. But we can do those in the weaker contexts where an error is less likely to be irrecoverable and then try and generalize and expand and build on that forward.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Put forward truths versus falsehoods, to put forward software that is legit versus that has a Trojan in it. That's an experimental paradigm. And in this experimental paradigm, you can try different things that work. You can use different ways to generate hypotheses. Yeah, and you can follow an incremental experimental path. And now we're less able to do that in the case of alignment and superintelligence because we're considering having to do things on a very short timeline.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Not necessarily. Yeah, I think that intuition is not necessarily correct. And so machine learning certainly is an area that rewards ability, but it's also a field where empirics and engineering have been enormously influential. And so if you're drawing the correlations compared to theoretical physics and pure mathematics, I think you'll find a lower correlation with cognitive ability if that's what you're thinking of. Yeah, something like creating neural lie detectors that work. So there's generating hypotheses about new ways to do it and new ways to try and train AI systems to successfully classify the cases. But the process of just generating the data sets, of like creating AIs, doing their best to”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, in general, so in science, the association with like scientific output prizes, things like that, there's a strong correlation, and it seems like an exponential effect. Yeah, it's not a binary drop-off. There would be levels at which people cannot learn the relevant fields. They can't keep the skills in mind faster than they forget them. It's not a divide where Einstein and the group that is 10 times as populous as that just can't do it or the group does 100 times as populous as that suddenly can't do it. It's like the ability to do the things earlier with less evidence and such falls off and it falls off, you know, at a faster rate in mathematics and theoretical physics and such than in most fields.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Wouldn't you just need people would have discovered general relativity just from the overwhelming data and other people would have done it after Einstein?”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“It doesn't have to be greater than our own. And in fact, in that situation, I think if you have Slack and to the extent that you're able to create delay and time to do things, that would be a case actually where you might want to restrict the intelligence of the systems that you're working with as much as you can. So for example, I would rather have many instances of smaller AI models that are less individually intelligent working on smaller chunks of a problem separately from one another because it would be more difficult for an individual AI instance working on an individual problem to in its spare time do the equivalent of create stuxnet.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“No, the category of things that humans can confirm is significantly larger, I think, than the category of what they can just do themselves.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think broadly comparable chunks from us getting things that are putting us in a reasonably good position going into it and then a broadly similar gain from this kind of really genuinely terrifying process of over a few months or hopefully longer if we have more regulatory ability and willingness to pause at the very end when this kind of automated research is meaningfully helping, where our work is just evaluating outputs that the AIs are delivering having the hard power and supervision to keep them from successfully rooting the servers, doing a takeover during this process, and have them finish the alignment task that we sadly failed to invest enough or succeed in doing beforehand.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“In this incredibly scary late period when AI has really automated research, then. Humans do this function of auditing, making it more difficult for the AIs to conspire together and root the servers, take over the process, and extract information from them within the set of things that we can verify, like experiments where we can see, oh yeah, this works at stopping an AI trained to get a fast one past human raiders and make a blue banana appear on this screen of this air gap computer.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Or with the sort of moderate things that we're doing largely on our own in a way that doesn't depend on the AI coming in at the last minute and doing our work for us.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“Basically, the incredibly juicy ability that we have working with the AIs is that We can have an invaluable outcome that we can see and tell. Whether they got a fast one past us on an identifiable situation. We can have here's an air gap computer. You get control of the keyboard. You can input commands. Can you route the environment and make a blue banana appear on the screen? Even if we train the AI to do that and it succeeds, we see the blue banana. We know it worked. Even if we did not understand and would not have detected the particular exploit that it used to do it. And so, yeah, this can give us a rich empirical feedback where we're able to identify things that are even an AI using its best efforts to get past our interpretability methods, using its best efforts to get past our advex, etc.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, it's a kind of obvious direction for this stuff to go. You can keep improving it when you have AIs that you're training to do their best to deceive humans or other audiences in the face of the thing. And you can measure, do our lie detectors break down when we train our AIs to tell us the sky is green in the face of the lie detector. And we keep using gradient descent on them to do that. Do they eventually succeed? If they do succeed, if we know it, that's really valuable information to know because then we'll know our existing lie detecting systems are not actually going to work on the AI takeover. And that can allow, say, government and regulatory response to hold things back. It can help redirect the scientific effort to create lie detectors that are robust and they can't just be immediately evolved around. And yeah, and we can then get more assistance.”
2023-06-26 · Dwarkesh Podcast · Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future · IDENTIFIED FROM THE TRANSCRIPT · source