YouSaid · the spoken record

Dario Amodei

lines on the record
181
first
2024-11-11
most recent
2024-11-11
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I mean, it depends what you mean by special and special in general. But I generally think the same kinds of techniques that we've been using to train the current model, I expect that doubling down on those techniques in the same way that we have for code, for models in general, for other for image input, for voice. I expect those same techniques will scale here as they have everywhere else.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Yeah, I think speaking at a high level, it's our intention to keep investing a lot in making the model better. Like I think we look at some of the benchmarks where previous models were like, oh, could do it 6% of the time. And now our model could do it 14 or 22% of the time. And yeah, we want to get up to the human level reliability of 80, 90%, just like anywhere else, right? We're on the same curve that we were on with SweBench, where I think I would guess a year from now the models can do this very, very reliably. But you got to start somewhere.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Releasing the model while the capabilities are still limited is very helpful in terms of doing that. I think since it's been released a number of customers, I think Replit was maybe one of the most quickest to deploy things have made use of it in various ways. People have hooked up demos for Windows desktops, Macs, Linux machines. So yeah, it's been very exciting. I think as with anything else, it comes with new exciting abilities. And then with those new exciting abilities, we have to think about how to make the model, say, reliable, do what humans want them to do. I mean, it's the same, it's the same story for everything, right? Same thing. It's that same tension.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  4. It's just the screen is just a universal interface that's a lot easier to interact with. And so I expect over time this is going to lower a bunch of barriers. Now, honestly, the current model has, it leaves a lot still to be desired. And we were honest about that in the blog, right? It makes mistakes. It misclicks. And we were careful to warn people, hey, this thing isn't, you can't just leave this thing to run on your computer for minutes and minutes. You've got to give this thing boundaries and guardrails. And I think that's one of the reasons we released it first in an API form rather than kind of, you know, this kind of just handed the consumer and give it control of their computer. But, you know, I definitely feel that it's important to get these capabilities out there as models get more powerful. We're going to have to grapple with how do we use these capabilities safely? How do we prevent them from being abused? And I think.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  5. You can just set that in a loop, give the model a screenshot, tell it what to click on, give it the next screenshot, tell it what to click on, and that turns into a full kind of almost 3D video interaction of the model. And it's able to do all of these tasks, right? You know, we showed these demos where it's able to like fill out spreadsheets. It's able to kind of like interact with a website. It's able to, you know. It's able to open all kinds of programs and different operating systems, Windows, Linux, Mac. So, you know, I think all of that is very exciting. I will say, while in theory, there's nothing you could do there that you couldn't have done through just giving the model the API to drive the computer screen, this really lowers the barrier. And there's a lot of folks who either kind of aren't in a position to interact with those APIs or it takes them a long time to do.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Yeah, it's actually relatively simple. So Claude has had for a long time since Claude 3 back in March the ability to analyze images and respond to them with text. The only new thing we added is those images can be screenshots of a computer. And in response, we train the model to give a location on the screen where you can click and or buttons on the keyboard you can press in order to take action. And it turns out that with actually not all that much additional training, the models can get quite good at that task. It's a good example of generalization. You know, people sometimes say if you get to lower Earth orbit, you're like halfway to anywhere, right? Because of how much it takes to escape the gravity. Well, if you have a strong pre-trained model, I feel like you're halfway to anywhere in terms of the intelligence space. And so actually it didn't take all that much to get Claude to do this.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Oh, yeah Oh, yeah. Yeah. It's actually like, you know, we've seen lots of examples of demagoguery in our life from humans. And, you know, there's a concern that models could do that as well.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Yeah, I mean, of course, you can hook up the mechanistic interpretability to the model itself, but then you've kind of lost it as a reliable indicator of the model state. There are a bunch of exotic ways you can think of that it might also not be reliable. Like if the model gets smart enough that it can jump computers and like read the code where you're looking at its internal state. We've thought about some of those. I think they're exotic enough. There are ways to render them unlikely. But yeah, generally you want to preserve mechanistic interpretability as a kind of verification set or test set that's separate from the training process of the model.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Right, show them, present themselves as being less capable than they are. And so I think with ASL4, there's going to be an important component of using other things than just interacting with the models. For example, interpretability or hidden chains of thought, where you have to look inside the model and verify via some other mechanism that is not as easily corrupted as what the model says, that the model indeed has some property. So we're still working on ASL4. One of the properties of the RSP is that we don't specify ASL-4 until we've hit ASL3. And I think that's proven to be a wise decision because even with ASL3, again, it's hard to know this stuff in detail. And we want to take as much time as we can possibly take to get these things right.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Yeah, I think for ASL 3, it's primarily about security and about filters on the model relating to a very narrow set of areas when we deploy the model. Because at ASL3, the model isn't autonomous yet. And so you don't have to worry about the model itself behaving in a bad way even when it's deployed internally. I won't say straightforward. They're rigorous, but they're easier to reason about. I think once we get to ASL4, we start to have worries about the models being smart enough that they might sandbag tests. They might not tell the truth about tests. We had some results came out about like sleeper agents, and there was a more recent paper about, you know, can the models mislead attempts to sandbag their own ability?

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Yeah, so that is hotly debated within the company. We are working actively to prepare ASL-3 security measures as well as ASL-3 deployment measures. I'm not going to go into detail, but we've made a lot of progress on both. And, you know, we're prepared to be, I think, ready quite soon. I would not be surprised at all if we hit ASL3 next year. There was some concern that we might even hit it this year. That's still possible. That could still happen. It's like very hard to say, but like I would be very, very surprised if it was like 2030. I think it's much sooner than that.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  12. And of course, what has to come with that is enough of a buffer threshold that you're not at high risk of kind of missing the danger. It's not a perfect framework. We've had to change it every, we came out with a new one just a few weeks ago and probably going forward, we might release new ones multiple times a year because it's hard to get these policies right. Like technically, organizationally from a research perspective, but that is the proposal, if then commitments and triggers in order to minimize burdens and false alarms now, but really react appropriately when the dangers are here.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I don't know. I've been working with these models for many years and I've been worried about risk for many years. It's actually kind of dangerous to cry wolf. It's actually kind of dangerous to say this model is risky. And people look at it and they say this is manifestly not dangerous. Again, it's the delicacy of the risk isn't here today, but it's coming at us fast. How do you deal with that? It's really vexing to a risk planner to deal with it. And so this if-then structure basically says, look, we don't want to antagonize a bunch of people. We don't want to harm our own, you know, our kind of own ability to have a place in the conversation by imposing these very onerous burdens on models that are not dangerous today. So the if-then, the trigger commitment is basically a way to deal with this. It says you clamp down hard when you can show that the model is dangerous.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Cyber bionuclear and model autonomy, which is less a misuse risk and more risk of the model doing bad things itself. ASL4, getting to the point where these models could enhance the capability of an already knowledgeable state actor and or become the main source of such a risk. Like if you wanted to engage in such a risk, the main way you would do it is through a model. And then I think ASL4 on the autonomy side, it's some amount of acceleration in AI research capabilities with an AI model. And then ASL5 is where we would get to the models that are kind of truly capable, that it could exceed humanity in their ability to do any of these tasks. And so the point of the if-then structure commitment is basically to say, look,

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Precautions designed to be sufficient to prevent theft of the model by non state actors and misuse of the model as it's deployed. We'll have to have enhanced filters targeted at these particular areas.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Or conduct a bunch of tasks and also not smart enough to provide meaningful information about CBRN risks and how to build CBRN weapons above and beyond what can be known from looking at Google. In fact, sometimes they do provide information, but not above and beyond a search engine, but not in a way that can be stitched together, not in a way that kind of end-to-end is dangerous enough. So ASL3 is going to be the point at which the models are helpful enough to enhance the capabilities of non-state actors, right? State actors can already do a lot, a lot of unfortunately, to a high level of proficiency, a lot of these very dangerous and destructive things. The difference is that non-state non-state actors are not capable of it. And so when we get to ASL3, we'll take special security.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  17. The RSP basically develops what we've called an if-then structure, which is if the models pass a certain capability, then we impose a certain set of safety and security requirements on them. So today's models are what's called ASL2. Models that were ASL-1 is for systems that manifestly don't pose any risk of autonomy or misuse. So for example, a chess playing bot, deep blue, would be ASL1. It's just manifestly the case that you can't use Deep Blue for anything other than chess. It was just designed for chess. No one's going to use it to conduct a masterful cyber attack or to, you know, run wild and take over the world. ASL2 is today's AI systems where we've measured them and we think these systems are simply not smart enough to autonomously self-replicate.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Not here today doesn't exist, but is coming at us very fast. So the solution we came up with for that in collaboration with people like the organization Meter and Paul Cristiano is, okay, what you need for that are you need tests to tell you when the risk is getting close. You need an early warning system. And so every time we have a new model, we test it for its capability to do these CBRN tasks as well as testing it for how capable it is of doing tasks autonomously on its own. And in the latest version of our RSP, which we released in the last month or two, the way we test autonomy risks is the AI models' ability to do aspects of AI research itself, which when the model, when the AI models can do AI research, They become

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Very, very fast, right? I testified in the Senate that we might have serious biowisks within two to three years. That was about a year ago. Things have preceded a pace. So we have this thing where it's like it's surprisingly hard to address these risks because they're not here today. They don't exist. They're like ghosts, but they're coming at us so fast because the models are improving so fast. So how do you deal with something that's...

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  20. I don't think there's any big thing we're missing. I just think we need to get better at controlling these models. And so these are the two risks I'm worried about. And our responsible scaling plan, which I'll recognize as a very long-winded answer to your question. I love it. Our responsible scaling plan is designed to address these two types of risks. And so every time we develop a new model, we basically test it for its ability to do both of these bad things. So if I were to back up a little bit, I think we have an interesting dilemma with AI systems where they're not yet powerful enough to present these catastrophes. I don't know that they'll ever prevent these catastrophes. It's possible they won't, but the case for worry, the case for risk is strong enough that we should act now. And they're getting better.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  21. As we give them more agency than they've had in the past, particularly as we give them supervision over wider tasks like writing whole code bases or someday even effectively operating entire companies, they're on a long enough leash. Are they doing what we really want them to do? It's very difficult to even understand in detail what they're doing, let alone control it. And like I said, these early signs that it's hard to perfectly draw the boundary between things the model should do and things the model shouldn't do, that if you go to one side, you get things that are annoying and useless and you go to the other side, you get other behaviors. If you fix one thing, it creates other problems. We're getting better and better at solving this. I don't think this is an unsolvable problem. I think this is a science, like the safety of airplanes or the safety of cars or the safety of drugs.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Who I have a PhD in this field, I have a well paying job, there's so much to lose. Why do I want to like, you know, even assuming I'm completely evil, which most people are not, why would such a person risk their risk their life, risk their legacy, their reputation to do something like truly, truly evil? If we had a lot more people like that, the world would be a much more dangerous place. And so my worry is that by being a much more intelligent agent, AI could break that correlation. And so I do have serious worries about that. I believe we can prevent those worries. But, you know, I think as a counterpoint to machines of loving grace, I want to say that this is still serious risks. And the second range of risks would be the autonomy risks, which is the idea that models might on their own, particularly

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  23. That are important. But when I think of the really things that would happen on the grandest scale, one is what I call catastrophic misuse. These are misuse of the models in domains like cyber, bio, radiological, nuclear, right? Things that could harm or even kill thousands, even millions of people if they really, really go wrong. Like these are the number one priority to prevent. And here, I would just make a simple observation, which is that the models, you know, if I look today at people who have done really bad things in the world, I think actually humanity has been protected by the fact that the overlap between really smart, well-educated people and people who want to do really horrific things has generally been small. Like, you know, let's say I'm someone.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  24. As much as I'm excited about the benefits of these models, and we'll talk about that if we talk about machines of loving grace, I'm worried about the risks and I continue to be worried about the risks. No one should think that machines of love and grace was me saying I'm no longer worried about the risks of these models. I think they're two sides of the same coin. power of the models and their ability to solve all these problems in biology, neuroscience, economic development, governance and peace, large parts of the economy, those come with risks as well, right? With great power comes great responsibility, right? That's the two are the two are paired things that are powerful, can do good things and they can do bad things. I think of those risks as being in several different categories. Perhaps the two biggest risks that I think about, and that's not to say that there aren't risks today that are

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Scaling is continuing. There will definitely be more powerful models coming from us than the models that exist today. That is certain. Or if there aren't, we've deeply failed as a company.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Don't want to commit to any naming scheme because if I say here, we're going to have Claude 4 next year. And then we decide that we should start over because there's a new type of model. I don't want to commit to it. I would expect a normal course of business that Claude Fora would come after Claude 3.5. But you never know in this wacky field, right?

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  27. You know, every time we add a new eval, and we're always evaluating for all the old things, so we have hundreds of these evaluations, but we find that there's no substitute for human interacting with it. And so it's very much like the ordinary product development process. We have like hundreds of people within Anthropic bash the model, then we do external A-B tests. Sometimes we'll run tests with contractors. We pay contractors to interact with the model. So you put all of these things together and it's still not perfect. You still see behaviors that you don't quite want to see, right? You know, you still see the model like refusing things that it just doesn't make sense to refuse.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  28. So typically, we'll have internal model bashings where all of Anthropic is almost a thousand people. People just try and break the model. They try and interact with it various ways. We have a suite of evals for, you know, oh, is the model refusing in ways that it couldn't? I think we even had a certainly eval because, again, one point model had this problem where like it had this annoying tick where it would like respond to a wide range of questions by saying certainly I can help you with that. Certainly I would be happy to do that. Certainly this is correct. And so we had a like certainly eval, which is like how often does the model say certainly. But look, this is just a whack-a-mole. Like, like what if it switches from certainly to definitely?

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Of those things at once. It's hard. It's very easy to go to one side or the other, and it's a multidimensional problem. And so I think these questions of shaping the model's personality, I think they're very hard. I think we haven't done perfectly on them. I think we've actually done the best of all the AI companies, but still so far from perfect. And I think if we can get this right, if we can control the false positives and false negatives in this very kind of controlled present-day environment, we'll be much better at doing it for the future when our worry is, you know, will the models be super autonomous? Will they be able to make very dangerous things? Will they be able to autonomously build whole companies? And are those companies aligned? So I think of this present task as both vexing, but also good practice for the future.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Unpredictable. They're actually quite hard to steer in control. And this version we're seeing today of you make one thing better, it makes another thing worse. I think that's like a present day analog of future control problems in AI systems that we can start to study today, right? I think that difficulty in steering the behavior and making sure that if we push an AI system in one direction, it doesn't push it in another direction in some other ways that we didn't want. I think that's kind of an early sign of things to come. And if we can do a good job of solving this problem, right, of like you ask the model to make and distribute smallpox and it says no, but it's willing to help you in your graduate level virology class. Like, how do we get both?

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Rest of the code goes here, right? Because they've learned that that's a way to economize and that they see it. And then so that leads the model to be so-called lazy encoding, where they're just like, ah, you can finish the rest of it. It's not because we want to save on compute or because the models are lazy and during winter break or any of the other kind of conspiracy theories that have come up. It's actually, it's just very hard to control the behavior of the model, to steer the behavior of the model in all circumstances at once. You can kind of, there's this whack-a-mole aspect where you push on one thing and like, you know, these other things start to move as well that you may not even notice or measure. And so one of the reasons that I care so much about, you know, kind of grand alignment of these AI systems in the future is actually these systems are actually quite

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  32. To say this like super clearly because I think it's some people don't know it, others kind of know it but forget it like it is very difficult to control across the board how the models behave. You cannot just reach in there and say, oh, I want the model to like apologize less. Like you can do that. You can include trading data that says like, oh, the model should apologize less. But then in some other situation, they end up being super rude or like overconfident in a way that's like misleading people. So there are all these trade-offs. For example, another thing is if there was a period during which models ours, and I think others as well were too verbose, right? They would like repeat themselves. They would say too much. You can cut down on the verbosity by penalizing the models for just talking for too long. What happens when you do that, if you do it in a crude way, is when the models are coding, sometimes they'll say,

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Yeah, so a couple points on this first. One is like things that people say on Reddit and Twitter or X or whatever it is, there's actually a huge distribution shift between the stuff that people complain loudly about on social media and what actually kind of like, you know, statistically users care about and that drives people to use the models. People are frustrated with, you know, things like the model not writing out all the code or the model just not being as good at code as it could be, even though it's the best model in the world on code. I think the majority of things are about that. But certainly a kind of vocal minority are raised these concerns, right? Are frustrated by the model refusing things that it shouldn't refuse or like apologizing too much or just having these kind of annoying verbal ticks? The second caveat, and I just want

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Versus can you do Taskax? The model might respond in different ways. And so there are all kinds of subtle things that you can change about the way you interact with the model that can give you very different results. To be clear, this itself is like a failing by us and by the other model providers that the models are just often sensitive to small changes in wording. It's yet another way in which the science of how these models work is very poorly developed. And so, you know, if I go to sleep one night and I was like talking the model in a certain way and I like slightly The models are not changing.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  35. On the system prompt can have some effects, although it's unlikely to dumb down models, it's unlikely to make them dumber. And we've seen that while these two things, which I'm listing to be very complete, happen relatively happen quite infrequently, the complaints about for us and for other model companies about the model change, the model isn't good at this, the model got more censored, the model was dumbed down. Those complaints are constant. And so I don't want to say like people are imagining it or anything, but like the models are for the most part not changing. If I were to offer a theory, I think it actually relates to one of the things I said before, which is that models have many, are very complex and have many aspects to them. And so often, you know, if I ask the model a question, you know, if I'm like, if I'm like do tasks.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  36. For modifying the model, we do a bunch of testing on it. We do a bunch of user testing and early customers. So we both have never changed the weights of the model without telling anyone. And it wouldn't, certainly in the current setup, it would not make sense to do that. Now, there are a couple things that we do occasionally do. One is sometimes we run A-B tests. But those are typically very close to when a model is being released and for a very small fraction of time. So the day before the new Sonnet 3.5, I agree. We should have a better name. It's clunky to refer to it. There were some comments from people that like it's gotten a lot better, and that's because fraction were exposed to an A-B test for those one or two days. The other is that occasionally the system prompt will change.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  37. So this actually doesn't apply. This isn't just about Claude. I believe I've seen these complaints for every foundation model produced by a major company. People said this about GPT-4. They said it about GPT-4 turbo. So a couple things. One, the actual weights of the model, right? The actual brain of the model, that does not change unless we introduce a new model. There are just a number of reasons why it would not make sense practically to be randomly substituting in new versions of the model. It's difficult from an inference perspective, and it's actually hard to control all the consequences of changing the way to the model. Let's say you wanted to fine-tune the model to be like, I don't know, to like, to say certainly less, which, you know, an old version of Sonnet used to do, you actually end up changing 100 things as well. So we have a whole process for it, and we have a whole process.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Talk to a model 10,000 times, and there are some behaviors you might not see. Just like with a human, right? I can know someone for a few months and not know that they have a certain skill or not know there's a certain side to them. And so I think we just have to get used to this idea. And we're always looking for better ways of testing our models to demonstrate these capabilities and also to decide which are the personality properties we want models to have and which we don't want to have that itself, the normative question is also super interesting.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Yeah, yeah, I definitely think this question of there are lots of properties of the models that are not reflected in the benchmarks. I think that's definitely the case and everyone agrees. And not all of them are capabilities. Some of them are, you know, models can be polite or brusque. They can be very reactive or they can ask you questions. They can have what feels like a warm personality or a cold personality. They can be boring or they can be very distinctive like Golden Gate Claude was. And we have a whole, you know, we have a whole team kind of focused on, I think we call it Claude character. Amanda leads that team and we'll talk to you about that. But it's still a very inexact science. And often we find that models have properties that we're not aware of. The fact of the matter is that you can

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  40. We're trying to maintain it, but it's not perfect. So we'll try and get back to the simplicity, but just the nature of the field, I feel like no one's figured out naming. It's somehow a different paradigm from like normal software. And so we just none of the companies have been perfect at it. It's something we struggle with surprisingly much relative to how trivial it is for the grand science of training the models.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Naming is actually an interesting challenge here, right? Because I think a year ago most of the model was pre-training. And so you could start from the beginning and just say, okay, we're going to have models of different sizes. We're going to train them all together. And, you know, we'll have a family of naming schemes. And then we'll put some new magic into them. And then, you know, we'll have the. I think we were in a good position in terms of naming when we had haiku, sonnet, and open. That's great.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Like Duke Nukem forever. What was that game? There was some game that was delayed 15 years. Was that Duke Nukem forever? Yeah.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Probably represents a real and serious increase in kind of programming ability. And I would suspect that if we can get to 90, 95%, it will represent ability to autonomously do a significant fraction of software engineering tasks.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  44. And just pull requests or like, you know, like a sort of atomic unit of work. You could say I'm implementing one implementing one thing. And so SuBench actually gives you kind of a real world situation where the code base is in the current state. And I'm trying to implement something that's described in language. We have internal benchmarks where we measure the same thing and you say, just give the model free reign to like, you know, do anything, run anything, edit anything. How well is it able to complete these tasks? And it's that benchmark that's gone from it can do it 3% of the time to it can do it about 50% of the time. So I actually do believe that if we get, you can gain benchmarks, but I think if we get to 100% of that benchmark in a way that isn't kind of like overtrained or game for that particular benchmark.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  45. We observe that as well, by the way, there were a couple very strong engineers here at Anthropic, who all previous code models, both produced by us and produced by all the other companies, hadn't really been useful, hadn't really been useful to them. You know, they said, you know, maybe this is useful to beginner. It's not useful to me. But Sonnet 3.5, the original one for the first time, they said, oh my God, this helped me with something that, you know, that it would have taken me hours to do. This is the first model that's actually saved me time. So again, the waterline is rising. And then I think, you know, the new sonnet has been even better. In terms of what it takes, I mean, I'll just say it's been across the board. It's in the pre-training. It's in the post-training. It's in various evaluations that we do. We've observed this as well. And if we go into the details of the benchmark, so SweBench is basically, you know, since, you know, since you're a programmer, you'll be familiar with pull.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Yeah, preference data from old models sometimes gets used for new models, although, of course, it performs somewhat better when it's trained on the new models. Note that we have this constitutional AI method such that we don't only use preference data, we kind of, there's also a post-training process where we train the model against itself. And there's new types of post-training the model against itself that are used every day. So it's not just RLHF, it's a bunch of other methods as well. Post training, I think, is becoming more and more sophisticated.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Yeah, I think at any given stage, we're focused on improving everything at once. Just naturally, like there are different teams. Each team makes progress in a particular area in making a particular segment of the relay race better. And it's just natural that when we make a new model, we put all of these things in at once.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Would be surprised how much of the challenges of building these models comes down to software engineering, performance engineering, from the outside, you might think, oh man, we had this Eureka breakthrough, right? You know, this movie with the science. We discovered it. We figured it out. But I think all things, even incredible discoveries, like they almost always come down to the details. And often super, super boring details. I can't speak to whether we have better tooling than other companies. I mean, you know, I haven't been at those other companies, at least not recently. But it's certainly something we give a lot of attention to.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  49. These more dangerous capabilities. So, those are the phases. And then it just takes some time to get the model working in terms of inference and launching it in the API. So there's just a lot of steps to actually making a model work. And of course, we're always trying to make the processes as streamlined as possible, right? We want our safety testing to be rigorous, but we want it to be rigorous to be automatic, to happen as fast as it can without compromising on rigor. Same with our pre-training process and our post-training process. So, you know, it's just like building anything else. It's just like building airplanes. You want to make them, you know, you want to make them safe, but you want to make the process streamlined. And I think the creative tension between those is an important thing in making the models work.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Yeah, so there's different processes. There's pre-training, which is just kind of the normal language model training. And that takes a very long time. That uses these days tens of thousands, sometimes many tens of thousands of GPUs or TPUs or trainium or we use different platforms, but accelerator chips, often training for months. There's then a kind of post training phase where we do reinforcement learning from human feedback as well as other kinds of reinforcement learning, that phase is getting larger and larger now. And often that's less of an exact science. It often takes effort to get it right. Models are then tested with some of our early partners to see how good they are. And they're then tested both internally and externally.

    2024-11-11 · Lex Fridman Podcast · #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity · IDENTIFIED FROM THE TRANSCRIPT · source