YouSaid · the spoken record

Paul Christiano

lines on the record
251
first
2023-10-31
most recent
2023-10-31
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I think it's really hard to like EG people look at mechanistic interpretability and be like, well, this obviously can't succeed. And I'm like, I don't know. How can you tell it obviously can't succeed? I think it's reasonable to take total investment in the field, like how fast is it making progress How does that pencil? I think most things people work on, though actually pencil pretty fine, they look like they could be reasonable investments. Things are not super out of whack.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Something can be bullshit because it's not addressing a real problem. That's the easiest way. This is a problem someone's interested in. That's just not actually an important problem and there's no story about why it's going to become an important problem. E.G. It's not a problem now and won't get worse or it is maybe a problem now, but it's clearly getting better. That's like one way. And then conditioned on passing that bar, like dealing with something that actually engages with important parts of the argument for concern, and then actually making sense empirically. So I think most work is anchored by source of feedback is actually engaging with real models. So it's like, does it make sense to engage with real models? And does the story about how it deals with key difficulties actually makes sense? I'm like pretty liberal past there.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  3. With empirical work, it's like interesting, and that you do have some signals of the quality of work. Does it work in practice? Does the story, I think the stories are just radically simpler. And so you probably can evaluate those stories on their face. And then you mostly come down to these questions about what are the key difficulties. Yeah, I tend to be optimistic when people dismiss something because this doesn't deal with a key difficulty or this runs into the following in SuperBlob obstacle. I tend to be like a little bit more skeptical about those arguments and tend to think like.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  4. So I think it depends on the kind of work. So for the kind of stuff we're doing, my guess is most people, there's just not really a way you're going to tell whether it's bullshit. So I think it's important that we don't spend that much money on the people who want to hire probably going to dig in in depth. I don't think there's a way you can tell whether it's bullshit without either spending a lot of effort or leaning on deference.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Yeah, the point is just like will so dwarf past RND. And there's like just not that much stickiness. There's less stickiness in the future than there has been in the past. I don't know. So I don't want to not comment from any private information just in my gut. Having caveat of this is like the single bet I've most lost, not including NVIDIA in that portfolio.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Yeah, I think it's a lot harder, especially if you're in this regime where you're trying to scale up. So, if you're unable to build fabs, I think what will take a very long time to build as many fabs as people want. The effect of that will be to bid up the price of existing fabs and existing semiconductor manufacturing equipment. And so just those hard assets will become spectacularly valuable, as will the existing GPUs and the actual. Yeah, I think it's just hard. That seems like the hardest asset to scale up quickly. So it's like the asset, if you have like a rapid run up, it's the one that you'd expect to most benefit. Whereas NVIDIA's stuff will ultimately be replaced by either stuff made by humans or stuff made by with AI systems. Like the gap will close even further as you build AI systems.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  7. And no inside information. I also still have a bunch of hardware investments, which I need to think about. Don't know. A lot of TSMC. I have a chunk of NVIDIA, although I just keep betting against Nvidia constantly since 2016 or something. I've been destroyed on that bet. Although AMD has also done fine. I just like, well, now the case now is even easier, but it's similar to the case in the old days, just a very expensive company given the total amount of R&D investment they've made. They have like whatever, a trillion dollar valuation or something. That's very high. So the question is, how expensive is it to make a TPU such that's actually out competes H100 or something? And I'm like, wow. It's real level, high level of incompetence if Google can't catch up fast enough to make that trillion dollar valuation not justified.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I've tried to get rid of most of the AI stuff that's plausibly implicated in policy work or like you advocacy on the RSP stuff or my involvement with Anthropic.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  9. I don't think there's like TSMC is not like planning on increases in total demand driven by AI, like kind of conspicuously not planning on it. I don't think anyone else is really ramping up production in anticipation either. So I think, and then similarly, like just building data centers of that size seems like very, very hard and also probably has multiple years of delay.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  10. I don't know much about any of the relevant areas. My best guess is like. My understanding is right now like 5% or something of like the next year's total or like best process fabs will be making AI hardware, of which only a small fraction will be going into very large training runs, like only a couple, maybe a couple percent of total output. And then that represents maybe 1% of total possible output, a couple percent of leading process, 1% of total or something. I don't know if that's right, but I think that's like the rough ballpark we're in. I think things will be pretty fast as you scale up for the next order of magnitude or two from there because you're basically just shifting over other stuff. My sense is it would be like years of delay. There's like multiple reasons that you expect years of delay for going past that. Maybe even at that, you start having, yeah, there's just a lot of problems. Like building new fabs is quite slow.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  11. 30% or something to 25% to Crazy AI by 2040 and like 10% by 2030 or something like that. So I think my 2030 probability has been kind of stable and my 2040 probability has been going up and I would guess it's too sticky. I guess that 40% I gave at the beginning is just like from not having updated recently enough and I maybe just need to sit down. I would guess that should be even higher. I think like 15% in 2030, I'm not feeling that bad about. This was just like each passing year is like a big update against 2030 We don't have that many years left. And that's like roughly counterbalanced with AI going pretty well. Whereas for like the 2040 thing like the passing years are not that big a deal. And like as we see that like things are basically working, that's like cutting out a lot of the probability of not having AI by 2040. So yeah, my 2030 probability up a little bit, like maybe twice as high as it used to be or like.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So I started thinking about this stuff in like 2010 or so. So I think my first, my earliest timeline prediction will be in like 2011 I think in 2011, my rough picture was like we will not have insane AI in the next 10 years and then I get increasingly uncertain after that, but we converge to 1% per year or something like that. And then probably in 2016, my take was like, we wouldn't have crazed AI in the next five years, but then we can converge to 1 or 2% per year after that. Then in 2019, I guess I made a round of forecasts. Where I gave

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Very unclear where the empirics shake out. I think Carl has thought about these more than I am, so I should maybe defer more. But anyway, I'm at like 50 50 on that.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Of two sources of evidence one is like looking across a bunch of industries what is the general improvement with each doubling of either RD investment or experience where it is quite exceptional to have a field with not anyway it's pretty good to have a field where each time you double R&D investment you get a doubling of efficiency the second source of evidence is on like actual like algorithmic improvement in ML which is obviously much much scarcer And they're like you can make a case that it's been like each doubling of R&D has given you like roughly a 4x or something increase in computational efficiency But like there's a question of how much that benefits when I say the effect of R&D stock is smaller I mean like we scale up like you're doing a new task like every couple years we're doing a new task because you're operating a scale much larger than the previous scale and so a lot of your effort is how to make use of the new scale So if you're not increasing your installed hardware base and just flat at a level of hardware, I think you get like a much faster diminishing returns than people have gotten historically. I think Carl agrees in principle this is true And then once you make that adjustment I think it's like

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Well, so the entire question is if you double RD effort, do you get enough additional improvement to further double the efficiency? And that question will itself be a function of your hardware base, like how much hardware you have questions like at the amount of hardware we're going to have and the level of sophistication we have as the process begins, is it the case that each doubling of actually the initials only depends on the hardware Each level of hardware will have some place at this dynamic asymptotes. So question is just like, for how long is it the case that each doubling of R&D at least doubles the effect of output of your AI research population? And I think I have a higher probability on that. I think it's kind of close if you look at the empirics. I think the empirics benefit a lot from continuing hardware scale up so that the effective R&D stock is significantly smaller than it looks, if that makes sense.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  16. I'm very unsure if you can keep stacking them, or like it's kind of a question of what's like the returns curve, and like Carl has some inference from historical data or some way he'd extrapolate the trend. I am more like 50-50 on whether the software only intelligence explosion is even possible and then a somewhat higher probability that it's slower than it might be.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Know exactly what Carl's probability is. I feel like Carl's gonna have like a 60% chance on some crazy thing that I'm only going to assign like a 20% chance to or 30% chance or something. And I think those kinds of perturbations are like one, how long a period is there of complementarity between AI capabilities and human capabilities, which will tend to soften takeoff, to how much diminishing returns are there on software progress, such that is a broader takeoff involving scaling electricity production and hardware production? Is that likely to happen during takeoff where I'm more like 50-50 or more stuff like this?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  18. It's related to our timelines discussion from earlier. I think the biggest I think the biggest issue is probably error bars, where Carl has a very software focused, very fast kind of takeoff picture. And I think that is plausible, but not that likely. I think there's a couple ways you could perturb the situation. And my guess is one of them applies. So maybe I have like...

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Yeah, I mean, like, for example, if you imagine what the best thing is, it would almost certainly involve just simulating every possible universe it might be in modular moral constraints, which I don't know if you want to include them. So that would be very, very slow. It would involve simulating all, you know, it's sort of like. I don't know exactly how slow, but like double exponential, very slow.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  20. There are a lot of forms. I think it's like the case that for, yeah, I think there are sort of arbitrarily smart input-output functionalities. And then if you hold fixed the amount of compute, there is some smartest one. If you're just like, what's the best set of 10 to the 40th operations? There's only finitely many of them. So some best one for any particular notion of best that you have in mind. So I guess I'm just like for the unbounded question, we're allowed to use arbitrary description complexity and compute probably now. And for the, I mean, there is some optimal conduct. If you're like, I have some goal in mind and I'm just like, what action best achieves it? If you imagine a little box embedded in the universe, there's kind of just like an optimal input-output behavior. So I guess in that sense, I think there is an upper bound, but it's not saturable in the physical universe because it's definitely exponentially slow.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  21. It seems like it's going to depend a little bit on what is meant by intelligence. It kind of reads as a question that's similar to is there an upper bound on strength or something

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Seems very hard if you wanted to be like competitive with learned reasoning. It depends a little bit exactly how you set it up, but for the ambitious versions of that that say it would address the alignment problem, they seem pretty unlikely. You know, like 5, 10% kind of thing.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Unless they were aggressively selected. If it's just they're trying to lie well rather than it's like they were selected over many generations to be excellent at lying or something, then like your ML system hopefully didn't train it a bunch to lie and you want to be careful about your training scheme effectively does that Yeah, that seems like it's more likely than not to succeed

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  24. I think to separate the like just train a classifier to do it is like a little bit complicated for a few reasons and may not work. But if you just brought in the space and say like, hey, it's like you want to know someone's lying. You get to interrogate them, but also you get to rewind them arbitrarily and make a million copies of them. I do think it's pretty hard to lie successfully. You get to look at their brain, even if you don't quite understand what's happening. You get to rewind them a million times. You get to run all those parallel copies into gradient descent or whatever. I think it's a pretty good chance that you can just tell if someone is lying. like a brain emulation or an AI or whatever.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  25. But that's going to be really cheaty. So that's what's the thing that has values and will expand in roughly preserve its values as it perceives. Yeah, exactly. Because that thing, the 10,000 byte thing, we'll just lean heavily on evolution and natural selection to get there.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  26. I think it depends a lot what substrate it gets to run on. So if you tell me how much computation does it get before, or what kind of real world infrastructure does it get? You could ask, what's the shortest program, which if you run it on a million H100s connected in a nice network with a hospitable environment, we'll eventually go to the stars. But that seems like it's probably on the order of tens of thousands of bytes, or I don't know. If I had to guess a median, I'd guess 10,000 bytes.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  27. You could call it like Dead On Arrival, but you could also just be like, it's not really the point. It's like mathematicians also trying to affect practice. And they're not like, why does my number theory not affect practice? It was kind of obvious. So I think the biggest thing is just actually caring about that and then learning at least what's basically going on in the actual systems you care about and what are actually the important constraints. And is this a real theoretical problem? The basic reason most theory doesn't do that is just like, that's not where the easy theoretical problems are. So I think theory is instead motivated by like, we're going to build up the edifice of theory. And sometimes there'll be opportunistically we'll find a case that comes close to practice or we'll find something practitioners are already doing and try and bring into our framework or something. But the theory of change is mostly not. This thing is going to make it into practice. It's mostly because it's going to contribute to the body of knowledge that will slowly grow and sometimes opportunistically yield important results.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Of the times when theory fails to connect with practice, it's just kind of clear it's not going to connect if you actually think about it and you're like, one of the key constraints in practice is the theoretical problem we're working on actually connected to those constraints. Is there a path, is there something that is possible in theory that would actually address real-world issues? I think the vast majority, like as a theoretical computer scientist, the vast majority of theoretical computer science has very little chance of ever effecting practice, but also it is completely clear in theory that it has very little chance of affecting practice. Like most of the theory fails to affect practice, not because of all the stuff you don't think of, but just because it was

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Maybe I'd say more profoundly again, it's just not that hard a case, it was like a little bit, it's a little bit unfair to be like, I'm gonna predict the thing, which I like, I pretty much think it was going to happen at some point. And so it was mostly a case of acceleration, whereas the work we're doing right now is specifically focused on something that's kind of crazy enough that it might not happen, even if it's a really good idea or challenging enough it might not happen. But I'd say in general, and this draws a little bit on more broad experience more broadly in theory. It's just like,

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Yeah, I think people maybe overestimate or like. Maybe it's like kind of a trope. But people talk about it's easy to underestimate how much gap there is to practice how many things will come up that don't come up in theory. But it's also easy to overestimate how inscrutable the world is. The things that happen mostly are things that do just kind of make sense. Yeah, I feel like most ML implementation does just come down to a bunch of detail though of like, you know, build a very simple version of the system, understand what goes wrong, fix the things that go wrong, scale it up, understand what goes wrong. And I'm glad I have some experience doing that, but I don't think, I think that does cause me to be better informed about what makes sense in ML and what can actually work. But I don't think it caused me to have a whole lot of deep expertise or like deep wisdom about how to close the gap.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  31. I mean, it is definitely exciting to have worked on a thing that has a real world impact. The main caveat I'd provide is like Our ledge chef is very, very simple compared to many things. And like, so the motivation for working on that problem was, look, this is how it probably should work, or like, this is a step in some progression. It's unclear if it's the final step or something, but it's a very natural thing to do that people probably should be and probably will be doing. I'm saying if you want to do, if you want to talk about crazy stuff, it's good to help make those steps happen faster. And it's good to learn about there's lots of issues that occur in practice, even for things that seem very simple on paper. But mostly the story of is just like, yep, I think my sense of the world is things that look like good ideas on paper just like often are harder than they look, but the world isn't that far from what makes sense on paper. Large language models look really good on paper and RHF looks really good on paper. And these things, like, I think just work out in a way that's...

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  32. I mean, funding is always good. We're not super funding constrained right now. The main effect of funding is it will cause me to continuously and perhaps indefinitely delay fundraising. Periodically all set out to be interested in fundraising and someone will be like, offer a grant, and then I will get to delay for another six months or fundraising or nine months or whatever. So you can delay the time at which Paul needs to think for some time about fundraising.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Yeah. I mean, we have much less information that these problems are hard. Like, again, I expect the solution to most of our problems to not be that complicated. And we've been working on it in some sense for a really long time. Like, you know, total years of full-time equivalent work across the whole team is like. Probably like three years of full time equivalent work in this area spread across a couple people. Very little compared to a problem. It is very easy to have a problem where you put in three years of full time equivalent work, but in fact, there's still an approach that's going to work quite easily with three to six months if you come at a new angle. We've learned a fair amount from that that we could share, and we probably will be sharing more over the coming months.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  34. The time is now to post the debrief from that, or I owe it from this weekend. I was supposed to do that. So I'll probably do it tomorrow. But no one solved it. It's sad putting out problems that are hard, or like, I don't, we could put out a bunch of problems that we think might be really hard.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Yeah, I think a first pass step, there's different levels of ambition or whatever, different ways of approaching a problem. But we have this write up of from last year, or I guess 11 months ago or whatever, on formalizing the presumption of independence that provides like, here's kind of a communication of what we're looking for in this object. And I think the motivating problem is saying here's a notion of what an estimator is and here's what it would mean for an estimator to capture some set of informal arguments. A very natural problem is just try and do that, like go for the go for the whole thing, try and understand and then come up with hopefully a different approach or then end up having context from a different angle on the kind of approach we're taking. I think that's a reasonable thing to do. I do think we also have a bunch of open problems. So maybe we should put up more of those open problems. And the main concern with doing so is that for any given one, we're like, this is probably hopeless put up a prize earlier in the year for an open problem, which tragically, I mean, I guess.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  36. There's a good chance that if I had my current set of views about this problem and didn't care about alignment and had the career safety to just like spend a couple years thinking about it or spend half my time for like five years or whatever, that I would just do that. I mean, even without caring at all about alignment, it's just a nice, it's a very nice problem. It's very nice to have this library of things that succeed where they feel so tantalizingly close to being formalizable, at least to me in such a natural setting, and then just have so little purchase on it. It's like a There aren't that many really exciting feeling frontiers in theoretical computer science.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  37. The flip side is it does feel, I mean, I think there are a lot of questions. I think some of them were probably going to make progress on. So I think the pitch is mostly like, are some people excited to get in now or are people more like, ah, let's wait to see once we have one or two good successes to see what the pattern is and become more confident we can turn the crank to make more progress in this direction? But for people who are excited about working on stuff with reasonably high probabilities of failure and not really understanding exactly what you're supposed to do, I think it's a good project. I feel like if people look back, if we succeed and people are looking back in like 50 years on what was the coolest stuff happening in math or theoretical computer science, there would be like a reasonable, this will definitely be like incontention. And I would guess for lots of people would just seem like the coolest thing from this period of a couple years or whatever.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  38. It seems like there's a reasonable chance to say about unifying all of that activity. I think it's a pretty exciting project. The basic strike against it is that it seems really hard. Like if you were someone's advisor, I think you'd be like, what are you going to prove if you work on this for the next two years? And they'd be like, there's a good chance nothing. And then it's not what you do if you're a PhD student normally. You aim for those high probabilities of getting something within a couple years.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  39. I think the basic setting of look, there are all of these arguments. So in mathematics, in physics, in computer science, there are just a lot of examples of informal heuristic arguments. They have enough structural similarity that it looks very possible that there is like a unifying framework, that these are instances of some general framework and not just a bunch of random things, like not just a bunch of, it's not like, so for example, for the prime numbers, people reason about the prime numbers as if they were like a random set of numbers. One view is like, that's just a special fact about the primes. They're kind of random. A different view is like, actually, it's pretty reasonable to reason about an object as if it was a random object as a starting point. And then as you notice structure, like revised from that initial guess. And it looks like to me the second perspective is probably more right. It's just like a reasonable to start off treating an object as random and then like notice perturbations from random, like notice structure the object possesses. And the primes are unusual in that they have fairly little additive structure. I think it's a very natural theoretical project. There's like a bunch of activity that people do.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Good luck So I think it's hard to work on because it's not clear what a success looks like. It's not clear if success is possible. But I do think there's a lot of questions. We have a lot of questions.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  41. There were definitely hiring and searching for collaborators. I think the most useful profile is Probably a combination of intellectually interested in this particular project. I'm motivated enough by alignment to work on this project, even if it's really hard. I think there are a lot of good problems. So the basic fact that makes this problem unappealing to work on, I'm a really good salesman, whatever. I think the only reason this isn't a slam dunk thing to work on is that like... There are not great examples. So, we've been working on it for a while, but we do not have beautiful results as of the recording of this podcast. Hopefully, by the time it airs.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  42. I think, like, an analogy, I think one of the most successful sagas in theoretical computer science is like formalizing the notion of an interactive proof system. And it's like you have some kind of informal thing that's interesting to understand. And you want to like pin down what it is and construct some examples and see what's possible and what's impossible. And this is like, I think this kind of thing is the bread and butter of the best parts of theoretical computer science. And then again, I think mathematicians, it may be a career mistake because the mathematicians only care about proofs or whatever. But that's a mistake in some sense aesthetically. If successful, I do think looking back.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Theoretical computer science is an exception where I think this is in some sense what the best of theoretical computer science is like. So you have all this reason, you have this like,

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  44. It's really not going to blow that many people's. I mean, I think it will be cool. I think it will be like very, if we succeed, it will be very solid metamathematics or theoretical computer science or whatever. But I don't think the mathematicians already do this reasoning and they mostly just love proofs. I think the physicists do a lot of this reasoning, but they don't care about formalizing anything. I think in practice, other difficulties are almost always going to be more salient. I think this is of most interest by far for interpretability in ML. And I think other people should care about it and probably will care about it if successful, but I don't think it's going to be the biggest thing ever in any field or even that huge a thing. I think this would be a terrible career move given the ratio of difficulty to impact. I think theoretical computer science, it's like probably a fine move. I think in other domains, it just wouldn't be worth, like we're going to be working on this for years at least in the best case.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  45. I mean, this is kind of like sales to us, right? Like, if you talk about this idea of people like, why would that not be like the coolest thing ever? And therefore, impossible. And we're like, well, actually, it's kind of lame. And we're just trying to pitch like, it's way lamer than it sounds. And that's really important to why it's possible is being like.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  46. A lot of code that you'd want to verify, not all of it, but a significant part is just the difficulty of formalizing the proof is the hard part and actually getting all of that to go through. And we're not going to help even the tiniest bit with that, I think. So this would be more helpful if you have code that uses simulations. You want to verify some property of a controller that involves some numerical error or whatever you need to control the effects of that error. That's where you start saying, well, heuristically, if the errors are independent, blah, blah, blah.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  47. My guess is we're not going to add. I mean, this is both a blessing and a curse. It's like a curse, and they're like, well, it's sad your thing is not that useful, but a blessing in not useful things are easier. My guess is we're not going to add that much value in most of these domains. Like, most of the difficulty comes from, like,

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  48. To be a machine that takes his input first, it takes his input argument and decides what it believes in light of it, which is kind of like saying, Was it compelling? Seconds, it needs to take four of those and then say, like, here's what I believe in light of all four, even though, like, there's a different estimation strategy that produce different numbers. And that's a lot of our life is saying, like, well, here's a simple thing that seems reasonable, and here's a simple thing that seems reasonable. Like, what do you do? There's supposed to be a simple thing that unifies them both. And the obstruction to getting that is understanding what happens when these principles are slightly intention. And how do we deal?

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  49. I mean, mathematical foundations are quite simple in the end. Like at the end of the day, it's like, you know, how many symbols? Like, I don't know, it's hundreds of symbols or something that go into the entire foundations. And the entire rules of reasoning. There's a sort of built on top of first order logic, but the rules of reasoning for first order logic are just another hundreds of symbols or hundreds lines of code or whatever. I'd say, like, I have no idea. We are certainly aiming at things that are just not that complicated. And my guess is that the algorithms we're looking for are not that complicated. Most of the complexity is pushed into arguments, not in this verifier or estimator

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source

  50. I guess it would be like 95% of the things mathematicians accept is really compelling filling heuristic arguments are correct. And if you actually formalize them, you'd be like, some of these aren't quite right, or here's some corrections, or here's which of two conflicting arguments is right. I think there's something to be learned from it. I don't think it would be mind-blowing, though.

    2023-10-31 · Dwarkesh Podcast · Paul Christiano — Preventing an AI takeover · IDENTIFIED FROM THE TRANSCRIPT · source