YouSaid · the spoken record
Richard Sutton
- lines on the record
- 69
- first
- 2025-09-26
- most recent
- 2025-09-26
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“There's nothing where you have training of what you should do. There's nothing. You see things that happen. You're not told what to do. Don't be difficult. I mean, this is obvious.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“The large language models is learning from training data. It's not learning from experience. It's learning from something that will never be available during its normal life. There's never any training data that says you should do this action in normal life.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Surprising, yeah, you can have such a different point of view. When I see kids, I see kids just trying things and like waving their hands around and moving their eyes around. And no one tells them there's no imitation for how they move their eyes around or even the sounds they make. They may want to create the same sounds, but the actions, the thing that the infant actually does, there's no targets for that. There are no examples for that.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Scalable method is you learn from experience. You try things, you see what works. No one has to tell you. First of all, you have a goal. So without a goal, there's no sense of right or wrong or better or worse. So large language models are trying to get by without having a goal or a sense of better or worse. It's exactly starting in the wrong place.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, in every case of the bitter lesson, you could start with human knowledge. And then do the scalable things That's always the case. And there's never any reason why that has to be bad. But in fact, and in practice, it has always turned out to be bad because people get locked into the human knowledge approach and they psychologically, or now I'm speculating why it is, but this is what has always happened. That their lunch gets eaten by the methods that are truly scalable.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“And yet one, well, I in particular expect there to be systems that can learn from experience which could well perform much, much better and be much more scalable, in which case it will be another instance of the bitter lesson that the things that used human knowledge were eventually superseded by things that just trained from experience and computation.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Interesting question whether large language models are a case of the bitter lesson. They are clearly a way of using massive computation, things that will scale with computation up to the limits of the internet. But they're also a way of putting in lots of Human knowledge. And so this is an interesting question. It's a sociological or industry question. Will they reach the limits of the data and be superseded by things that can get more data just from experience rather than from people? In some ways, it's a classic case of the bitter lesson with the more human knowledge we put into the large language models, the better they can do. And so it feels good.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, the math problems are different. Making a model of the physical world and carrying out the consequences of mathematical assumptions or operations. Those are very different things, like the empirical world has to be learned. You have to learn the consequences. Whereas the math is more just computational. It's more like standard planning. So there you can, they can have a goal to find the proof. And they are in some way given that goal to find the proof.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“It's not a goal. Not a substantive goal. You can't look at a system and say, oh, it has a goal if it's just sitting there predicting and being happy with itself, that it's predicting accurately.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“About Next token is what they should say, what the action should be It's not what the world will give them in response to what they do. Let's go back to their lack of goal. For me, having a goal is the essence of intelligence Something as intelligent if it can achieve goals. I like John McCarthy's definition that intelligence is the computational part of the ability to achieve goals. Have to have goals, you're just a behaving system. You're not anything special. You're not intelligent. And you agree that large language models don't have goals.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Just saying they don't have any meaningful sense, they don't have a prediction of what will happen next. And they will not be surprised by what happened next. They'll not make any changes if something happens.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Oh no, they will respond to that question, right? Have no prediction in the substance of sense that they won't be surprised by what happens. And if something happens that isn't what you might say they predicted, they will not change because an unexpected thing has happened. To learn that, they'd have to make an adjustment.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“If you say something in your conversation, the large language models have no prediction about what the person will say in response to that, what the response will be”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“But there isn't any truth. There's no right thing to say. Now, in reinforcement, there is a right thing to say or right thing to do because the right thing to do is the thing that gets you reward. So we have a definition of what the right thing to do, and so we can have prior knowledge or knowledge provided by people about what the right thing to do is, and then we can check it. See, because we have a definition of what the actual right thing to do is. An even simpler case is when you have you're trying to make a model of the world when you predict what will happen, you predict, and then you see what happens. Okay, so there's ground truth. There's no ground truth in In large language models because you don't have a prediction about what will happen next.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Is there any way to tell in a large language model setup to tell what's the right thing to say? You will say something and you will not get feedback about what the right thing to say is because there's no definition of what the right thing to say is. There's no goal If there's no goal, then there's one thing to say, another thing to say. There's no right thing to say. So there's no ground truth. You can't have prior knowledge if you don't have ground truth. Because the prior knowledge is supposed to be a hint or an initial belief about what the truth is.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“No, I agree that it's the large language perspective. Think it's a good perspective. So, to be a prior for something, there has to be a real thing. I mean, a prior bit of knowledge should be the basis for actual knowledge. What is actual knowledge? There's no definition of actual knowledge in that large language framework. What makes an action a good action to take? You recognize the value, the need for continual learning. So if you need to learn continually, continually means learning during normal interaction with the world. And so then there must be some way during the normal interaction to tell what's right. Okay, so”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“The large language models learn from something else. They learn from here's a situation and here's what a person did. And implicitly, the suggestion is you should do what the person did.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Just to mimic what people say is not really to build a model of the world at all. I don't think you're mimicking things that have a model of the world, the people. But I don't want to approach the question in an adversarial way. But I would question the idea that they have a world model. So a world model would enable you to predict what would happen. They have the ability to predict what a person would say. They don't have the ability to predict what will happen. What we want, I think, to quote Alan Turing, what we want is a machine that can learn from experience Experience is the things that actually happen in your life. You do things, you see what happens, and that's what you learn from.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, yes, I think it's really quite a different point of view, and it can easily get separated and lose the ability to talk to each other. And yeah, large language models have become such a big thing. Generative AI in general, a big thing. And our field is subject to bandwagons and fashions. So we lose track of the basic basic things. Because I consider reinforcement to be basic AI. And what is intelligence, the problem is to understand your world. Reinforcement learning is about understanding your world. Whereas large language models are about mimicking people, doing what people say you should do. They're not about figuring out what to do.”
2025-09-26 · Dwarkesh Podcast · Richard Sutton – Father of RL thinks LLMs are a dead end · IDENTIFIED FROM THE TRANSCRIPT · source