YouSaid · the spoken record
Paul Christiano
- lines on the record
- 251
- first
- 2023-10-31
- most recent
- 2023-10-31
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“I mean, maybe one way of putting it is just like we can wait until we see this input, or like you can wait until you see a weird input and say, okay, did this weird input do something we didn't understand? And for our sake, that would just be a trivial test. You're just like some algorithm to be like, is it a thing? Whereas for a neural net, in some cases, it is either very expensive to tell or it's like you actually don't have any other way to tell. You checked in easy cases and they were on a hard case, so you don't have a way to tell if something has gone wrong. Also, I would clarify that I think it is interesting for the Riemann hypothesis. I would say The current state, particularly a number theory, but maybe in quite a lot of math, is like there are informal heuristic arguments for pretty much all the open questions people work on. But those arguments are completely informal. So that is like, I think there is not the case that there's like, here's the norms of informal reasoning or the norms of heuristic reasoning. And then we have arguments to like a heuristic argument verifier could accept. It's just like people wrote some words. I think those words like”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“My guess is these will not, like the estimates you have, this would be much, much worse than the estimates you'd get out of just normal empirical or scientific reasoning where you're using a reference class and saying, how often do people find algorithms for hard problem? I think what this argument will give you for is RSA fine is going to be like, well RSA is fine unless it isn't. Unless there's some additional structure in the problem that an algorithm can exploit, then there's no algorithm. But very often, like the way these arguments work. So for neural nets as well, is you say, like, look, here's an estimate about the behavior. And that estimate is right unless there's another consideration we've missed. And like the thing that makes them so much easier than proofs is to just say like, here's a best guess given what we've noticed so far. But that best guess can be easily upset by new information. And that's like both what makes them easier than proofs, but also what means they're just like, wait less useful than proofs for most cases. Like I think neural nets are kind of unusual in being a domain where we really do want to do systematic formal reasoning, even though we're not trying to get a lot of”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I think most claims in mathematics that mathematicians believe to be true already have fairly compelling heuristic arguments. So the Riemann hypothesis, it's actually just, there's kind of a very simple argument that the Riemann hypothesis should be true unless something surprising happens. And so a lot of math is about saying, okay, we did a little bit of work to find the first pass explanation of why this thing should be true. And then, for example, in the case of Lorentz hypothesis, the question is, do you have this weird periodic structure in the primes? And you're like, well, look, if the primes were kind of random, you obviously wouldn't have any structure like that. Like just how would that happen? And then you're like, well, maybe there's something. And then the whole activity is about searching for like, can we rule out anything? Can we rule out any kind of conspiracy?”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Thinking about very simple cases and saying what is the correct notion? Like, what is the right heuristic estimate in this case? Or how do you reconcile these two apparently conflicting explanations?”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“It's more like the more interesting the behaviors to be explained. So the random neural net just like doesn't have very many interesting behaviors that demand explanation. And as you get smarter, you start having behaviors that are like, you know, you start having some correlation with the simple thing and then that demands explanation. Or you start having some regularity in your output, some that demands explanation. So these properties kind of emerge gradually over the course of training that demand explanation. I also, again, want to emphasize here that when we're talking about searching for explanations, this is like, this is some dream. We talk to ourselves, like, why would this be really great if we succeeded? We have no idea about the empirics on any of this. So these are all just words that we think to ourselves and sometimes talk about to understand, would it be useful to find a notion of explanation? What properties would we like this notion of explanation to have? But this is really speculation and being out on our limb. Almost all of our time day to day is just thinking about cases much, much simpler, even than small neural nets.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“There aren't, I mean, I think there just aren't that many behaviors that demand explanation. Like most things are random neural net does are kind of what you'd expect from a random. If you treat just like a random function, then there's nothing to be explained. There are some behaviors that demand explanation, but like, I mean, yeah. Anyway, random neural net is pretty uninteresting. That's part of the hope is it's kind of easy to explain features of the random neural net.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, it would depend what behaviors it had. So, we're always talking about an explanation of some behavior from a model.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Surprise intuitively is very large. So, for example, if you have a neural net that gets a problem correct, a neural net with a billion parameters that gets a problem correct on every input of length 1,000. In some sense, there has to be something that needs explanation there because there's too many inputs for that to happen by chance alone. Whereas if you have a neural net that gets something right on average or gets something right in merely a billion cases, that actually can just happen by coincidence. GPT-4 can get billions of things right by coincidence because it just has so many parameters that are adjusted to fit the data.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Is like a first part of the intuition. Like when humans are actually doing design, I think there's not such a huge gap. When in the ML case, I think there is a huge gap, but I think largely for other reasons. A thing I also want to stress is that we just are open to there being a lot of facts that don't have particularly compact explanations. So another thing is when we think of finding an explanation, in some sense, we're setting our sites really low here. So if a human designed a random widget and was like, this widget appears to work well, or like if you search for a configuration that happens to fit into this spot really well, it's like a shape that happens to mess with another shape. You might be like, what's the explanation for why those things mesh? And we're very open to just being like, that doesn't need an explanation. You just compute. You check that the shapes mesh and you did a billion operations and you checked this thing worked. Or you're like, why do these proteins bind? You're like, it's just because these shape, like this is a low energy configuration. And there's not, we're very open to, in some cases, there's not very much more to say. So we're only trying to explain cases where kind of”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Think we both just mostly don't have empirical evidence about how hard it is to find explanations of this particular type about why models work. Like we have a sense that it's really hard, but that's because we're like have this incredible mismatch for like gradient descent is spending an incredible amount of compute searching for a model. And then some human is looking at neurons or even some neural net is looking at neurons, just like you have an incredible, basically because you cannot define what an explanation is, you're not applying gradient descent to the search for explanations. So I think the MLK is just like actually shouldn't make you feel that pessimistic about the difficulty of finding explanations. The reason it's difficult right now is precisely because you don't have any kind of, you're not doing an analogous search process to find this explanation as you do to find the model.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“This is pretty conjectural and complicated to express some intuitions. Maybe one thing is, I think a lot of this intuition does come from cases like machine learning. So if you ask about writing code and you're like, how hard is it to find code versus find the explanation the code is correct? In those cases, there's actually just not that much of a gap. The way a human writes a code is basically the same difficulty as find the explanation for why it's correct. In the case of ML, like...”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“We are at least exploring the hypothesis or interested in the hypothesis maybe those problems are actually more matched in difficulty”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Is when it's called for, even though it just somehow optimized over this continuous space. And the difficulty or the hope is that the difficulty of these two problems are kind of matched. So that is, it's very hard to find these logicalish explanations because it's not a space that's easy to search over. But there are ways to do it. There's ways to embed discrete, complicated, rigid things in these nice, squishy, continuous spaces that you search over. And in fact, to the extent that neural nets are able to learn the rigid logical stuff at all, they learn it in the same way. Maybe they're hideously inefficient or maybe it's possible to embed this discrete reasoning in the space in like a way it's not too inefficient. But like you really want the two search problems to be of similar difficulty. And that's like the key hope overall. I mean, this is always going to be the key hope. The question is, is it easier to learn a neural network or to find the explanation for why the neural network works? I think people have the strong intuition that it's easier to find the neural network than the explanation of why it works. And that is really the, I think,”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think I basically sympathize. This is like some intuitive objections. Like, look, the space of explanations is this rigid a lot of explanations have this rigid logical structure where they're really precise and simple things govern, like complicated systems, and nearby simple things just don't work, and so on. And a bunch of things that feel totally different from this kind of nice, continuously parametrized space. And you can imagine interpretability on simple models where you're just by gradient descent, finding feature directions that have desirable properties. But then when you imagine, hey, now that's like a human brain you're dealing with that's like thinking logically about things like the explanation of why that works isn't going to be just like here with some feature direction. So that's how I understood basic confusion, which I share or sympathize with at least. So I think the most important high level point is I think basically the same objection applies to being like, how is GPT-4 going to learn to reason like logically about something? You're like, well, look, logical reasoning, it's got rigid structure. It's like, it's doing all this, it's doing ands and or.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“In terms of what explanations look like physically, or the most ambitious plan, the most optimistic plan is that you are searching for explanations in parallel with searching for neural networks. So you have a parameterization of your space of explanations, which mirrors the parameterization of your space of neural networks. Or you should think of it as kind of similar to like, what is a neural network? It's some simple architecture where you fill in a trillion numbers, and that specifies how it behaves. So two, you should expect an explanation to be like a pretty flexible general skeleton that's saying pretty flexible general skeleton, which just has a bunch of numbers you fill in. What you're doing to produce an explanation is primarily just filling in these floating point numbers”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, it depends a little bit what you mean by succeed. But if you say get explanations that are great and accurately affect reality and work for all of these applications that we're imagining or that we are optimistic about, kind of the best case success. I don't know, like 10, 20 percent something And then there's a higher probability of various intermediate results that provide value or insight without being the whole dream. But I think the probability of succeeding in the sense of realizing the whole dream is quite low.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I also want to maybe step back a tiny bit and clarify that I think this project is kind of crazily ambitious. And the main reason, the overwhelming reason I think you should expect it to break down or fail is just because we have all these desires. We have all the things we want out of this notion of explanation. But that's an incredibly hard research project, which has a reasonable chance of being impossible. So I'm happy to talk about what the implications are, but I want to flag, but condition on failing, I think it's most likely because just like the things we wanted were either incoherent or intractably difficult.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Again, the question is like it has to add test time, be somehow producing activations that have the takeover effect, and at training time be producing activations that don't look that takeovery or whatever. And so somehow there was a difference between those. It has to actually do the work of creating that difference somewhere.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“And the hope is like it needs to not only have them be like random different inputs, it needs to produce inputs with certain properties. Like it needs to deliberately create inputs that have this effect in the world, this intended effect in the world.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“The problem is, if it ever does a check that's like, am I in training or not, then your explanation for the behavior is going to have to factor through that check. And so if the check, like if you ever do the check, the check becomes like an essential part of the explanation. And then when the check fails, your explanation breaks down. So you've kind of like, you've already lost the game if you did such a check.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“And I guess the hope there is like it's actually not that easy to tell a story. Like if you're like, why did you not stab people on training time? Why are you not stabbing, is that applying now? It's kind of a mind doesn't get to arbitrarily change the actual reasons it does something. There was some earlier step. Maybe you could think of it as like at the first step where it's like, no, I'm going to try and do the sneaky thing to make my thoughts look this way. The actual explanation, like the actual causal explanation for the behavior is there. Like if you tried to give this reason, it would have to, like maybe we should talk about the example in more detail to try and imagine how a brain would try and do this or how a mind would try and do this. But I think it is like...”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think the hope is to basically be robust. So, like, when we're doing theoretical research, we are in a kind of worst casey regime So we are normally imagining that everything conspires insofar as it can. So, insofar as we're like, yeah, could a brain be doing some crazy thing where it tries to structure itself to look weird? That's totally in scope for ruling something out.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Even if a small fraction, even if you're in a small fraction of cases, if something bad is happening at scale, you can detect that. If you imagine one in every thousand decisions you check, if there's a robot rebellion going on or whatever, you get to look in on one in every thousand decisions. You see that something's really weird with that decision. And even if you didn't check many of them, you're like, okay, we need to back off.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, for this story of like flagging an anomaly, you have to actually check at a sufficiently high frequency. So you can imagine checking all the time, or you could imagine saying, hey, actually, you need a lot of AI systems cooperating to not blow the whistle on each other or whatever. e.g. you might have checked like hey in training every time something suspicious happens the model flags it like i mean like why does that happen you have your explanation you look at a new case you could say is it not flagging it for the normal reason or is it a weird reason um so you could hope that like”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Explanation for why it never stabs people in the train distribution. And then you're like, on this test time, when you run it on the new input, it's like, does I think I'm on the train distribution? It says no. You compare that against your explanation. Like, actually, the explanation breaks down completely now. The explanation said it was precisely because it thought it was being trained that it doesn't do anything really dangerous or doesn't stab anyone. And so once we see that single variable switch, we know that this is not just an anomaly. This is the anomaly that matters, or is an anomaly that is very relevant to the explanation.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think that's sort of the typical case you would like to detect. Yeah. So the story would be something like you want to explain on the training distribution why the model never stabs anyone. And you're like, what's my explanation on the train distribution? It's actually going to have to go through that model. Like, supposing the only reason it doesn't stab people is it's like, well, I know that I'm being watched and I can't stab people unless I'm really not being watched. Then your explanation for why does it never stab people is like, oh, look, it knows, I mean, it won't be edited, it won't be annotated with English because it's just some giant basket of numbers, but it's tracking along with the inner monologue and saying like, well, it thinks to itself, am I being watched? And in fact, it normally concludes it is being watched on the training distribution, right? You have this whole distribution. And so the explanation is saying it knows it's being watched and therefore it decides not to stab anyone. So ideally, that's what the explanation does. Again, not annotated English, but just like a giant bag of numbers that's effectively talking about that process. So when you have a new input on which the model says, am I on train distribution, like, no, I'm not being watched. Then if you compare that against the explanation, the explanation said, hey, the reason doesn't stab anyone is because it just constantly thinks that it's being trained, like this thing is always this nonlinear feature, the representation of am I being trained is constantly saying yes.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Of this explanation. For example, this explanation is saying models will tend to activate in the following direction, like you don't care about anything that's happening orthogonal to that direction. Just like, are they not activating? You're sort of just looking at this one direction and being like, did this one direction change a lot? The idea is once you have this explanation at hand, a much, much smaller class of things look anomalous in a way that's relevant to the explanation. And if you've done a really good job, the story is like, if there's a new input where you expect the property to still hold, that will be because you expect the explanation to still hold. Like the explanation generalizes as well as the behavior itself that it's explaining. It's kind of what you would like, roughly.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Work essentially for new data points from the same distribution. If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn't have any concise explanation for the distribution. So this first problem was just like, it's a whole different set of activations. I think you're actually kind of okay. And then the thing that becomes more messy is like, but the real world will not only be new samples of different activations, they will also be different in important ways. Like the whole concern was there's these distributional shifts. Or like not the whole concern, but most of the concern. I mean, maybe the point of having these explanations, I think every input is an anomaly in some ways, which is kind of the difficulty is if you have a weak notion of anomaly, any distribution shift can be flagged as an anomaly. It's like you're constantly getting anomalies. And so the hope of having such an explanation is to be able to say, like, here were the features that were relevant for this explanation or for this behavior. And a much smaller class of things are anomalies with respect to this explanation. Like most anomalies wouldn't change this. Most ways you change your distribution won't effectively validate.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, to be clear, I think you probably wouldn't be looking at a separate circuit, which is part of why it's hard. You'd be looking at the model is always doing the same thing on every input. It's always whatever it's doing, it's a single computation. So it'd be all the same circuits interacting in a surprising way. But yeah, this is just to emphasize your question even more. I think the easiest way to start is to just consider the IID case. So what you're considering a bunch of samples, there's no change in distribution. You just have a training set of a trillion examples and then a new example from the same distribution. So in that case, that's still the case that every activation is different, but this is actually a very, very easy case to handle. If you think about an explanation that generalizes across, like if you have a trillion data points and an explanation which is actually able to compress the trillion data points down to like actually there's kind of a lot of compression if you if you think about it you have a trillion parameter model and a trillion data points we would like to find like a trillion parameter explanation in some sense so that's like actually quite compressed and sort of just in virtue of being so compressed we expect it to like automatically”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“It, we would regard a proof as a good explanation, and our concern about proofs is primarily you can't prove properties of neural nets. We suspect, although it's not completely obvious. I think it's pretty clear you can't prove facts about neural nets.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So, the explanations overall for this behavior is we expect to be like of similar size to the model itself, like maybe somewhat larger. And like, I think the type signature, if you want to have a clear menzel picture, the best picture is probably talking about a proof or imagining a proof that a model has this behavior. So you could imagine proving the GPT-4 does this induction behavior. And that proof would be a big thing. It would be much larger than the weights of the model. That's sort of our goal to get down from much larger to just the same size. And it would potentially be incomprehensible to a human, right? Just say, like, here's a direction activation space, and here's how it relates to this direction activation space. And so you're just pointing out a bunch of stuff like that. Here's these various direction, here's these various features constructed from activations, potentially even nonlinear functions. Here's how they relate to each other, and here's how if you look at what the computation the model is doing, you can sort of inductively trace through and confirm that the output has such and such correlation. So that's the dream. Yeah, I think the mental reference would be, like, I don't really like proofs because I think there's such a huge gap between what you can prove and how you would analyze enrollnet. But I do think it's probably the best mental picture for what is an explanation even if a human doesn't understand.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So to be clear, an explanation of why a particular output happened, I think, is just you ran the model. So, we're not expecting a smaller explanation for that.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“But I do think, like, I think compared to most people, I am less worried about automating interpretability. I think if you have a thing which works that's incredibly labor intensive, I'm fairly optimistic about our ability to automate it. Again, the stuff we're doing, I think, is quite helpful in some worlds, but I do think the typical case interpretability can add a lot of value without this.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it's most of all like so a way you can automate how you would automate interpretability if you wanted to right now is you take the process humans use great we're going to take that human process train ml systems to do the pieces that humans do of that process and then just do a lot more of it so I think that is great as long as your task decomposes into like human size pieces and then there's just this fundamental question about large models which is like do they decompose in some way into like human size pieces or is it just a really messy mess with interfaces that like aren't nice And the more it's the latter type, the harder it is to break it down into these pieces, which you can automate by copying what a human would do, and the more you need to say, okay, we need some approach which scales more structurally.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“The informal standard we use involves humans being able to make sense of what's going on, and there's some question about scalability of that. Will humans recognize the concepts that models are using? Yeah, I think as you try and automate it, it becomes increasingly concerning. If you're on slightly shaky ground about what exactly you're doing or what exactly the standard for success is. I think there's a number of reasons. As you work with really large models, it becomes just increasingly desirable to have a really robust sense of what you're doing. But I do think it would be better even for small models to have a clearer sense.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. I mean, I guess this is relevant in the sense that I think a basic difficulty is you don't really understand the objective of what you're doing, which is a little bit hard institutionally or scientifically it's just rough. It's better to do science when the goal of the game is to predict something and you know what you're predicting than when the goal of the game is to understand in some undefined sense. I think it's particularly relevant here just because”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“So that's the most common approach to formalizing what is a good explanation. And even when people are doing informal interpretability, I think if you're publishing an ML conference and you want to say, this is a good explanation, the way you would verify that would, even if not like a formal set of causal intervention experiments, it would be some kind of ablation where then we messed with the inside of the model and it had the effect which we would expect based on our explanation.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. Or if there's too expensive to check in this case, and like to be clear, when we talk about formalizing what is a good explanation, there is a little bit of work that pushes on this. And it mostly takes this causal approach of saying, well, what should an explanation do? It should not only predict the output, it should predict how the output changes in response to changes in the internals.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Primarily because of this being able to tell if things had been different, like if you have an input where this doesn't happen, then you should be scared.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. And for this purpose, it's like the thing that's essential is kind of reasoning from one property of your model to the next property of your model. It's really important that you're going forward step by step rather than drawing a bunch of samples and confirming the property holds. Because if you just draw a bunch of samples and confirm the property holds, you don't get this check where, say, oh, here was the relevant fact about the internals that was responsible for this downstream behavior. All you see is like, yeah, we checked a million cases and it happened and all of them. You really want to see this like, okay, here was the fact about the activations, which kind of causally leads to this behavior.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Which holds over the training set. And this property is responsible for the fat, the behavior, namely that it doesn't do anything that looks too dangerous. So then when a new input comes in and it doesn't satisfy that property, you can say, okay, this is anomalous with respect to that explanation. So either it will not have the behavior, like it won't, it will do something that appears dangerous, or maybe it will have that behavior, but for some different reason than normal. Normally it does it because of this pathway, and now it's doing it for a different pathway. And so you would like to be able to flag that both there's a risk of not exhibiting the behavior. And if it happens, it happens for a weird reason. And then you could, I mean, at a minimum, when you encounter that, say like, okay, raise some kind of alarm. There's sort of a more ambitious, complicated plans for how you would use it, right? So it fits arc has some longer story, which is kind of a motivated this of how it fits into the whole rest of the plan.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so the ideal kind of outcome here would be to say you have your AI system behaving nicely, you get some explanation for sort of why it's behaving nicely. And we could tell a story in English about that explanation, but we're not actually imagining the explanation being a thing that makes sense to a human. But if you were to tell a story in English, which again, you will not see as a researcher, it would be something like, well, then the model believes it's being trained. And so because it believes it's being trained, it knows it needs to not do anything scary looking or else the humans will penalize it. Like that's something that's happening inside this opaque explanation. And then the hope is if you have that explanation and then you run into a new input on which the model doesn't believe it's being trained, right? If you just look at the set of activations of your model, that is not necessarily a weird-looking activation. It's just a bunch of numbers. But if you look at this explanation, you see, actually, the explanation really crucially depended on this fact holding consistently across the training distribution, which again, we as humans could editorialize and say that fact was it believes it's being trained. But the explanation doesn't fundamentally make reference to that. It's just saying, here's a property of the activations.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think I like mechanistic interpretability quite a lot. And I do think if you just consider the entire portfolio of what people are working on for alignment, I think there should be more work on mechanistic interpretability than there is on this project Arc is doing. But I think that's the case. So I think we're mostly talking about a small fraction in the portfolio, and I think it's like a good enough bet. It's quite a good bet overall. But so the thing that the problem we're trying to address in acquistic interpretability is kind of like If you do some interpretability and you explain some phenomenon, you face this question of like, what does it mean your explanation was good? Like if you want to either, I think this is a problem somewhat institutionally or culturally it's just hard to know what you're doing and it's hard to scale up an activity when you don't really understand the rules of the game for that activity very well. It's hard to have that much confidence in your results.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it's unlikely they're going to be applicable to kind of any interesting neural net. But the thing about proofs that is relevant for our purposes isn't that they give you like 100% confidence. So you don't have to be like this incredible level of demand for rigor. You can relax the standards of proof a lot and still get this feature where it's like a structural explanation for the behavior. We're deducing one thing from another until at the end your final conclusion is like therefore induction occurs.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Saying, like, this is kind of a deductive argument for the behavior. So you want to get given the weights of a neural net. So it's just like a bunch of numbers. You've got your million numbers or billion numbers or whatever. And then you want to say, like, here's some things I can point out about the network and some conclusions I can draw. I can be like, well, look, these two vectors have large inner product and therefore these two activations are going to be correlated on this distribution. So they're not established by drawing samples and checking things are correlated, but saying because of the weights being the way they are, we can proceed forward through the network and derive some conclusions about what properties the outputs will have. So you could think of this as the most extreme form would be just proving that your model has this induction behavior. You could imagine proving that if I sample tokens at random with this pattern AB followed by A, that B appears 30% of the time or whatever. That's the most extreme form. And what we're doing is kind of just like relaxing the rules of the game for proof, saying proofs are incredibly”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, what is this kind of criterion? At the end of the day, we kind of want some criterion on the way the criterion should work is like you have your neural net. Have some behavior of that model. Like a really simple example is like Anthropic has this sort of informal description being like, here's induction, like the tendency that if you have the pattern A, B followed by A, it will tend to predict B. You can give some kind of words and experiments and numbers that are trying to explain that. And what we want to do is say, what is a formal version of that object? Like, how do you actually test if such an explanation is good? So just clarifying what we're looking for when we say we want to define what makes an explanation good. And the kind of answer that we are searching for or settling on”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Makes like a good explanation. And probably the biggest part of the hope is that. If you want to say detect when the explanation has broken down or something weird has happened, that doesn't necessarily require a human to be able to understand this complicated interpretation of a giant model. If you understand like Then you might be able to sort of automatically discover such things and automatically determine if on a new input it might have broken down. So that's one way of sort of describing the high level goal, like starting from, you can start from interpretability and say, can we formalize this activity or what a good interpretation or explanation is? There's some other work in that genre, but I think we're just taking a particularly ambitious approach to it.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“I think you'd really like to do is look inside the model and understand why it has those desirable properties. And if you understood that, you could then say like, okay, now can we flag when these properties are at risk of breaking down or predict how robust these properties are, determine if they hold in cases where it's too confusing for us to tell directly by asking if the underlying cause is still present. So that's a thing people would really like to do. Most work aimed at that long-term goal right now is just sort of opening up neural nets and doing some interpretability and trying to say like, can we understand even for very simple models why they do the things they do or what this neuron is for or questions like this? So ARC is taking a somewhat different approach where we're instead saying like, okay, look at these interpretability explanations that are made about models and ask what are they actually doing? Like what is the type signature? What are like the rules of the game for making such an explanation?”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so the high level, I mean, there's a couple different high level descriptions you could give, and maybe I'll unwisely give like a couple of them in the hopes that one is kind of makes sense. A first pass is like it would sure be great to understand why models have the behaviors they have. So you like, look at GPT-4. If you ask GPT-4 question, it will say something that looks very polite. And if you ask it to take an action, it will take an action that doesn't look dangerous. You will decline to do a coup, whatever, all this stuff.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think this is right in the sense that using RLHF is not really a tax if you wanted to deploy a useful system, like, why would you not? Or it's just very much worth the money of doing the training. And then, yeah, so RLHF will address certain kinds of alignment failures, that is where a system just doesn't understand or is changing next word prediction. It's like this is the kind of context where a human would do this wacky thing, even though it's not what we'd like. There's like some very domalemen failures that will be addressed by it. I think mostly, yeah, the question is, is that true even for the sort of more challenging Lyman failures that motivate concern in the field? I think RLHF doesn't address most of the concerns that motivate people to be worried about alignment.”
2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source