YouSaid · the spoken record
Yann LeCun
- lines on the record
- 182
- first
- 2024-03-07
- most recent
- 2024-03-07
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Well, so it's a first step. Okay, so first of all, what's the difference with generative architectures like LLMs? So LLMs or Vision systems that are trained by reconstruction generate the input. They generate the Original input that is non-corrupted, non-transformed, right? So you have to predict all the pixels. And there is a huge amount of resources spent in the system to actually predict all those pixels, all the details. In a JEPA, you're not trying to predict all the pixels. You're only trying to predict an abstract representation of the inputs. And that's much easier in many ways. So what the JEPA system, when it's being trained, is trying to do is extract as much information as possible from the input, but yet only extract information that is relatively easily predictable.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Three, four years. Now we have methods that are non contrastive, so they don't require those negative contrastive samples of images that we know are different. You can only, you turn them only with images that are different versions or different views of the same thing. And you rely on some other tricks to prevent the system from collapsing. And we have half a dozen different methods for this now.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, the contrastive method avoid this. And those things have been around since the early 90s. I had a paper on this in 1993, is you also show pairs of images that you know are different. And then you push away the representations from each other. So you say not only do representations of things that we know are the same, should be the same or should be similar, but representations of things that we know are different should be different. Prevents the collapse, but it has some limitation. And there's a whole bunch of techniques that have appeared over the last six, seven years that can revive this type of method, some of them from FAIR, some of them from Google and other places. But there are limitations to those contrasting methods. What has changed in the last...”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And I call this a JEPA, so that means joint embedding predictive architecture because it's joint embedding and there is this predictor that predicts the representation of the good guy from the bad guy. And the big question is how do you train something like this? And until five years ago, six years ago, we didn't have particularly good answers for how you train those things, except for one called contrastive learning. And the idea of contractive learning is you take a pair of images that are, again, an image and a corrupted version or degraded version somehow or transformed version of the original one. You train the predicted representation to be the same as that. If you only do this, the system collapses. It basically completely ignores the input and produces representations that are constant.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Okay, so now instead of training a system to encode the image and then training it to reconstruct the full image from a corrupted version, you take the full image, you take the Corrupted or transformed version, you run them both through encoders, which in general are identical, but not necessarily. And then you train a predictor on top of those encoders to predict the representation of the full input from the representation of the corrupted one. So joint embedding because you're taking the full input and the corrupted version Or transform version, run them both through encoders. You get a joint embedding, and then you're saying, can I predict the representation of the full one from the representation of the corrupted one?”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“The architecture is good. The architecture of the encoder is good. But the fact that you train the system to reconstruct images does not lead it to learn good generic features of images.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“With label data, with textual descriptions of images, et cetera, you do get good representations. And the performance on recognition tasks is much better than if you do this self-supervised free training.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Okay, so the reason this doesn't work is, first of all, I have to tell you exactly what doesn't work. So, the thing that does not work is training a system to learn representations of images by training it to reconstruct a good image from a corrupted version of it. That's what doesn't work. And we have a whole slew of techniques for this that are variant of denoising autoencoders. Something called MAE, developed by some of my colleagues at FAIR, Max Toto Encoder. So it's basically like the LLMs or things like this, where you train the system by corrupting text, except you corrupt images, you remove patches from it and you train a gigantic neural net to reconstruct. The features you get are not good. And they're not good because if you now train the same architecture, but you train it supervised.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Has been essentially a complete failure. And it works really well for text. That's the principle that is used for LLMs, right?”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“That has been a complete failure, essentially. And we've tried lots of things. We tried just straight neural nets. We tried GANS. We tried VAEs, all kinds of regularized autoencoders. We tried many things. We also tried those kind of methods to learn good representations of images or video that could then be used as input to, for example, an image classification system. That also has basically failed. Like all the systems that attempt to predict missing parts of an image or video from a corrupted version of it, basically. So take an image or a video, corrupt it or transform it in some way, and then try to reconstruct the complete video or image from the corrupted version. And then hope that internally the system will develop a good representations of images that you can use for object recognition segmentation, whatever it is.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Maybe it's going to predict this is a room where there's a light and there is a wall and things like that. It can't predict what the painting on the wall looks like or what the texture of the couch looks like. Certainly not the texture of the carpet. So there's no way you can predict all those details. The way to handle this is one way possibly to handle this, which we've been working for a long time, is to have a model that has what's called a latent variable. And the latent variable is fed to a neural net and it's supposed to represent all the information about the world that you don't perceive yet. That you need to augment the system for the prediction to do a good job at predicting pixels, including texture of the carpet and the couch and the painting on the wall.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Video is high dimensional and continuous. A lot of details in this. So if I take a video of this room, the video is a camera panning around. There is no way I can predict everything that's going to be in the room as I pan around. The system cannot predict what's going to be in the room as the camera is panning.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“The idea of doing this has been floating around for a long time at FAIR. Some of our colleagues and I have been trying to do this for about 10 years. And you can't really do the same trick as with LLMs because LLMs, as I said, you can't predict exactly which word is going to follow a sequence of words, but you can predict the distribution over words. Now, if you go to video, what you would have to do is predict the distribution over all possible frames in a video. And we don't really know how to do that properly. We do not know how to represent distributions over high dimensional continuous spaces in ways that are useful. And there lies the main issue. And the reason we can do this is because the world is incredibly more Complicated and richer in terms of information than text. Text is discrete.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, General T model has trained on video, and we've tried to do this for 10 years. You take a video, show a system, a piece of video, and then ask it to predict the reminder of the video. Basically, predict what's going to happen.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And understanding why the world is evolving the way it is. And then the extra component of a world model is something that can predict how the world is going to evolve as a consequence of an action you might take. So what model really is, here is my idea of the state of the world at time t. Here is an action I might take. What is the predicted state of the world at time t plus one? Now that state of the world does not need to represent everything about the world. It just needs to represent enough that's relevant for this planning of the action, but not necessarily all the details. Now here is the problem. You're not going to be able to do this with generative models.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so can you build this first of all by prediction And the answer is probably yes. Can you build it by predicting words? And the answer is most probably no, because language is very poor in terms of weak or low bandwidth, if you want. There's just not enough information there. So building role models means Observing the world”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Uttered words as opposed to an output being muscle actions. We plan our answer before we produce it. And LLMs don't do that. They just produce one word after the other. Instinctively, if you want. It's like it's a bit like the Subconscious actions where you don't like you're distracted, you're doing something, you're completely concentrated and someone comes to you and asks you a question and you kind of answer the question, you don't have time to think about the answer, but the answer is easy, so you don't need to pay attention. You sort of respond automatically. That's kind of what an LLM does, right? It doesn't think about its answer, really. It retrieves it because it's accumulated a lot of knowledge, so it can retrieve some things, but it's going to... Just spit out one token after the other without planning the answer.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Figure out a reaction you want to cause and then figure out how to say it so that it causes that reaction. But that's really close to language. But think about a mathematical concept or imagining something you want to build out of wood or something like this, right? The kind of thinking you're doing is absolutely nothing to do with language, really. Like, it's not like you have necessarily an internal monologue in any particular language. You're imagining mental models of the thing, right? I mean, if I ask you to imagine what this water bottle will look like if I rotate it 90 degrees, that has nothing to do with language. And so clearly there is a more abstract level of representation. In which we do most of our thinking and we plan what we're going to say. If the output is”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, it depends what kind of thinking, right? If it's just producing puns, I get much better in French than English about that. No, but.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“We think about what we're going to say, and it's relatively independent of the language in which we're going to say it, where we talk about, I don't know, let's say a mathematical concept or something. Kind of thinking that we're doing and the answer that we're planning to produce is not linked to whether we're going to see it in French or Russian or English.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“You can just compute a distribution over them. Then what the system does is that it picks a word from that distribution. Of course, there's a higher chance of picking words that have a higher probability within that distribution. So you sample from that distribution to actually produce a word. And then you shift that word into the input. And so that allows the system now to predict the second word, right? And once you do this, you shoot it into the input, et cetera. That's called autoregressive prediction, which is why those LLMs should be called autoregressive LLMs. But we just call them LLMs. There is a difference between this kind of process and a process by which before producing a word, when you talk, when you and I talk, you and I are bilingual.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Not going to be able to do this with the type of LLMs that we are working with today. And there's a number of reasons for this. But the main reason is... The way LNMs are trained is that you take a piece of text, you remove some of the words in that text, you mask them, you replace them by blank markers, and you train a gigantic neural net to predict the words that are missing. And if you build this neural net in a particular way so that it can only look at words that are to the left of the one is trying to predict, then what you have is a system that basically is trying to predict the next word in a text, right? So then you can feed it a text, a prompt, and you can ask it to predict the next word. It can never predict the next word exactly. And so what it's going to do is produce a probability distribution over all the possible words in a dictionary. In fact, it doesn't predict words. It predicts tokens that are kind of subword units. And so it's easy to handle the uncertainty in the prediction there because there is only a finite number of possible words in the dictionary.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Typical LLM takes as an input. And then you just feed that to the LLM in addition to the text. And you just expect the LLM to kind of during training to kind of be able to use those representations to help make decisions. I mean, there's been work along those lines for quite a long time. And now you see those systems, right? I mean, there are LLMs that have some vision extension. But they're basically hacks in the sense that those things are not trained end-to-end to handle, to really understand the world. They're not trained with video, for example. They don't really understand intuitive physics, at least not at the moment.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, yeah, that's what a lot of people are working on. So, the short answer is no. And the more complex answer is you can use all kind of tricks to get an LLM to basically digest visual representations of images or video or audio for that matter. And a classical way of doing this is you train a vision system in some way. And we have a number of ways to train vision systems. These are supervised, semi-supervised, self-supervised, all kinds of different ways that will turn any image into a high-level representation. Basically, a list of tokens that are really similar to the kind of tokens.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“We can do with computers. And, you know, we have LLMs that can pass the bar exam. So they must be smart. But then... Can't learn to drive in 20 hours like any 17 year old. They can't learn to clear out the dinner table and fill off the dishwasher like any 10-year-old can learn in one shot. Why is that? Like, you know, what are we missing? What type of learning? Reasoning architecture, or whatever, are we missing that basically prevent us from having level five solar in cars and domestic robots?”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“AI needs to be embodied essentially. And then other people coming from the NLP side or maybe some other motivation don't necessarily agree with that. And philosophers are split as well. And the complexity of the world is hard to imagine. It's hard to represent all the complexities that we take completely for granted in the real world that we don't even imagine require intelligence, right? This is the old Marfac paradox from the pioneer of robotics, hence Marvec, who said, you know, how is it that with computers it seems to be easy to do high-level complex tasks like playing chess and solving integrals and doing things like that? Whereas the thing we take for granted that we do every day, like I don't know, learning to drive a car or grabbing an up.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And we do this by essentially imagining the result of the outcome of a sequence of actions that we might imagine. And that requires metal models that don't have much to do with language. And I would argue most of our knowledge is derived from that interaction with the physical world. So a lot of my colleagues who are more interested in things like computer vision are really on that camp.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“So it's a big debate among philosophers and also cognitive scientists, like whether intelligence needs to be grounded in reality. I'm clearly in the camp that, yes, intelligence cannot appear without some grounding in some reality. It doesn't need to be physical reality. It could be simulated, but the environment is just much richer than what you can express in language. Language is a very approximate representation of or percepts and our mental models, right? I mean, there's a lot of tasks that we accomplish where we manipulate a mental model of the situation at hand. And that has nothing to do with language. Everything that's physical, mechanical, whatever, when we build something, when we accomplish a task, moderate task of grabbing something, et cetera, we plan action.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“That despite our intuition, most of what we learn and most of our knowledge is through our observation and interaction with the real world, not through language. Everything that we learn in the first few years of life and certainly everything that animals learn has nothing to do with language.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“But then you realize it's really not that much data. If you talk to developmental psychologists and they tell you a four-year-old has been awake for 16,000 hours in his or her life? The amount of information that has reached the visual cortex of that child in four years is about 10 to the 15 bytes. And you can compute this by estimating that the optical nerve carry about 20 megabytes per second, roughly. And so 10 to the 15 bytes for a four-year-old versus 2 times 10 to the 13 bytes for 170,000 years worth of reading. It tells you that through sensory input, we see a lot more information than we do through language.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“That they're not interesting, that we can't build a whole ecosystem of applications around them. Of course, we can. But as a path towards human-level intelligence, they're missing essential components. And then there is another tidbit or fact that I think is very interesting. Those LLMs are trained on enormous amounts of text. Basically, the entirety of all publicly available text on the internet, right? That's typically Order of 10 to the 13 tokens. Each token is typically two bytes. So that's two 10 to the 13 bytes as training data. It would take you or me 170,000 years to just read through this at eight hours a day. So it seems like an enormous amount of knowledge, right, that those systems can accumulate.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source
“For a number of reasons. The first is that there is a number of characteristics of intelligent behavior For example, the capacity to understand the world, understand the physical world The ability to remember and retrieve things. Persistent memory, the ability to reason and the ability to plan. Those are four essential characteristics of intelligent systems or entities, humans, animals. LNMs can do none of those. Or they can only do them in a very primitive way. And they don't really understand the physical world. They don't really have persistent memory. They can't really reason. And they certainly can't plan. And so if you expect a system to become intelligent just without having the possibility of doing those things, you're making a mistake. That is not to say that autoregosivatal lms are not useful. They're certainly useful.”
2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source