YouSaid · the spoken record

Yann LeCun

lines on the record
182
first
2024-03-07
most recent
2024-03-07
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I mean, people have come up with things where you put essentially a random sequence of characters in a prompt. And that's enough to kind of throw the system into a mode where it's going to answer something completely different than it would have answered without this. So that's a way to jailbreak the system basically, get it, go outside of its conditioning.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  2. It's a tiny, tiny, tiny subset of all possible prompts. And so the system will behave properly on the prompts that has been either trained, pre-trained, or fine-tuned. But then there is an entire space of things that it cannot possibly have been trained on because it's just the number is gigantic. So whatever training the system has been subject to to produce appropriate answers, you can break it by finding out a prompt that will be outside of the The set of pumps has been trained on or things that are similar. And then it will just spew complete nonsense.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  3. No, it's basically a struggle against the curse of dimensionality. So the way you can correct for this is that you fine tune the system by having it produce answers for all kinds of questions that people might come up with. And people are people. So a lot of the questions that they have are very similar to each other. So you can probably cover 80% or whatever of questions that people will ask by collecting data and then you fine-tune the system to produce good answers for all of those things. And it's probably going to be able to learn that because it's got a lot of capacity to learn. But then The enormous set of prompts. And that set his enormous, like within the set of all possible prompts, the proportion of prompts that have been used for training is absolutely tiny.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah. And that drift is exponential. It's like errors accumulate, right? So the probability that. An answer will be nonsensical increases exponentially with the number of tokens.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  5. So, because of the autoregressive prediction, every time an LNM produces a token or a word, there is some level of probability for that word to take you out of the set of reasonable answers And if you assume, which is a very strong assumption, that the probability of such error Is that those errors are independent across a sequence of tokens being produced? What that means is that every time you produce a token, the probability that you stay within the set of correct answer decreases and it decreases exponentially.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  6. About the distinction between animate and inanimate objects, you know by 18 months, you know about why people want to do things and you help them if they can't. A lot of things that you learn mostly by observation, really, not even through interaction. In the first few months of life, babies don't really have any influence on the world. They can only observe, right? And you accumulate like a gigantic amount of knowledge just from that. So that's what we're missing from current AI systems.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I mean, that's the 16,000 hours of wake time of a four year old and 10 to the 15 bytes, you know, going through vision, just vision, right? There is a similar... Bandwidth of touch and a little less through audio. And then language doesn't come in until like, you know, a year in life. And by the time you are nine years old, you've learned about gravity. You know about inertia, you know about gravity, you know, the stability, you know.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I just think that most of the information of this type that we have accumulated when we were babies is just not present. In text, in any description, essentially.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  9. That's not there in LLMs. LLMs are purely trained from tech. So then the other statement you made, I would not agree with the fact that implicit in all languages in the world is the underlying reality. There's a lot about underlying reality, which is not expressed in language.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  10. No, I agree with what you just said, which is that to be able to do high level common sense, to have high-level common sense, you need to have the low-level common sense to build on top of. But

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Well, there's a lot of situations that might be difficult to for a purely language-based system to Know, like, okay, you can probably learn I cannot get from New York to Paris by snapping my fingers. Not going to work, right? There are probably sort of more complex scenarios of this type, which an NLM may never have encountered and may not be able to determine whether it's possible or not. So that link from the low level to the high level, the thing is that the high level that language expresses is based on a common experience of the low level, which LLMs currently do not have. When we talk to each other, we know we have a common experience of the world. A lot of it is. Similar LLMs don't have

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Started working. We abandoned this idea of Predicting every pixel and basically just doing the giant embedding and predicting a representation space. That works. So there's ample evidence that we're not going to be able to learn good web presentations of the real world using GeneralT model. So I'm telling people, everybody is talking about generative AI. If you're really interested in human-level AI, abandon the idea of generative AI.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Right, we don't go through text, it goes directly from speech to speech using an internal representation of kind of speech units that are discrete, but it's called text lesson LP. We used to call it this way. But yeah, so incredible success there. And then for 10 years, we tried to apply this idea to learning representations of images by training a system to predict videos, learning intuitive physics by training a system to predict what's going to happen in the video. tried and tried and failed and failed with generative models, with models that predict pixels. We could not get them to learn good representations of images. We could not get them to learn good representations of videos. We tried many times. We published lots of papers on it. You know, they kind of sort of worked, but not really great.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  14. John Embedding architecture, by the way, trained with contrastive learning. And that system also can produce speech recognition systems that are multilingual with mostly unlabeled data and only need a few minutes of labeled data to actually do speech recognition. That's amazing. We have systems now based on those combination of ideas that can do real-time translation of hundreds of languages into each other, speech-to-speech.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Unsupervised or supervised learning, and I took a back seat for a bit. And then I kind of tried to revive it in a big way, you know, starting in 2014, basically when we started FAIR and really pushing for finding new methods to do self-supervised learning, both for text and for images and for video and audio. And some of that work has been incredibly successful. I mean, the reason why we have multilingual translation system things to do content moderation on meta, for example, on Facebook that are multilingual and understand whether a piece of text is hate speech or not or something. It's due to their progress using self-supervised learning for NLP, combining this with transformer architectures and blah, blah, blah. But that's the big success of self-supervised learning. We had similar success in speech recognition, a system called Wave2Vec, which is also

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  16. The audio self supervised running actually is going back more than 10 years. But the audio of cell supervised running, so basically capturing the internal structure of a set of inputs without training the system for any particular task, right? Learning representations. The conference I co-founded 14 years ago is called International Conference on Learning Representations. That's the entire issue that deep learning is dealing with, right? And it's been my obsession for almost 40 years now. So learning representation is really the thing. For the longest time, we could only do this with supervised learning. And then we started working on what we used to call unsupervised learning and sort of revive the idea of unsupervised learning in the early 2000s with Josh Njo and Jeffinton, then discovered that supervised learning actually works pretty well if you can collect enough data. And so the whole idea of

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  17. No, of course, everybody would be impressed. But it's not a question of being impressed or not. It's a question of knowing what the limit of those systems can do. Again, they are impressive. They can do a lot of useful things. There's a whole industry that is being built around them. They're going to make progress. But there is a lot of things they cannot do, and we have to realize what they cannot do and then figure out. How we get there. And I'm not seeing this from basically 10 years of research on

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Entering would decide that a Turing test is a really bad test. Okay, this is what AI community has decided many years ago that the trying test was a really bad test of intelligence.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  19. We're fooled by their fluency, right? We just assume that if a system is fluent in manipulating language, then it has all the characteristics of human intelligence. But that impression is false. We're fooled by it.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Yeah, I mean, there were work from various places, but if you want to kind of place it in the GPT timeline, that would be around GPT 2, yeah.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Predicting a word from the words that come before, right? And you do this by the constraining the architecture of the network. And that's what you can build an autoregressive LM from. So there was a surprise many years ago with what's called decoder-only LLM. So systems of this type that are just trying to produce words from the previous one. And the fact that when you scale them up, they tend to really kind of understand more about language when you're trying to With work from Google, Meta, OpenAI, et cetera, going back to the GPT kind of work general pre-trained transformers.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  22. That has been an enormous amount of benefits. It allowed us to create systems understand language, systems that can translate hundreds of languages in any direction, systems that are multilingual, so it's a single system that can be trained to understand hundreds of languages and translate in any direction and produce summaries. and then answer questions and produce text. And then there's a special case of it where you constrain the system to not elaborate a representation of the text from looking at the entire text, but only

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  23. No, there's one thing that auto CLMs, or that LLMs in general, not just the auto-aggressive one, but including the Bert style bidirectional ones, are exploiting and it's self-supervised learning. And I've been a very, very strong advocate of self-supervised learning for many years. So those things are incredibly impressive demonstration that self-supervised learning actually works. The idea that started didn't start with Bert, but it was really kind of a good demonstration with this. The idea that you take a piece of text, you corrupt it, and then you train something and igne it to reconstruct a part that are missing.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Person who's never heard of airplanes and tell them, How do you go from New York to Paris? And they're probably not going to be able to kind of deconstruct the whole plan unless I've seen examples of that before. So certainly LMs are going to be able to do this. But then how you link this from the low level of actions, that needs to be done with things like JPA that basically lift the abstraction level of the representation without attempting to reconstruct UB detail of the situation. That's why you need JPAS for.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Sure. And a lot of plans that people know about that are relatively high level are actually learned. Most people don't invent the... Plans By themselves, we have some ability to do this, of course, obviously, but most plants that people use are plants that have been trained on. Like they've seen other people use those plants or they've been told how to do things, right? Like you can't invent how you take a...

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  26. So, I mean, certainly, LLM would be able to solve that problem if you fine-tune it for it. And so I can't say that LLM cannot do this. It can do this if you train it for it. There's no question down to a certain level where things can be formulated in terms of words. But if you want to go down to how you climb down the stairs or just stand up from your chair in terms of words, like you can't do it. You need, that's one of the reasons you need experience of the physical world, which is much higher bandwidth than what you can express in words, in human language.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  27. True, I mean, they will probably produce some answer, except they're not going to be able to really kind of produce millisecond by millisecond morsel control of how you stand up from your chair. But down to some level of abstraction where you can describe things by words, they might be able to give you a plan, but only under the condition that they've been trained to produce those kind of plans. They're not going to be able to plan for situations where they never encountered before. They basically are going to have to regurgitate the template that they've been trained on

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Right. So there's a lot of questions that are sort of implied by this, right? So, the first thing is LNMs will be able to answer some of those questions down to some level of abstraction. Under the condition that they've been trained with similar scenarios in the training

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  29. And obviously you're not going to plan your entire trip from New York to Paris in terms of millisecond by millisecond muscle control. First, that would be incredibly expensive, but it will also be completely impossible because you don't know all the conditions of what's going to happen, how long it's going to take to catch a taxi or to go to the airport with traffic. I mean, you would have to know exactly the condition of everything to be able to do this planning. And you don't have the information. So you have to do this hierarchical planning so that you can start acting and then sort of replanning as you go. And nobody really knows how to do this in AI. Nobody knows how to train a system to learn the appropriate multiple levels of representation so that hierarchical planning works.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Okay, now I have another sub goal go down on the street. What that means going to the elevator, going down the elevator, work out the street. How do I go to the elevator? I have to. Stand up for my chair, open the door of my office, go to the push the button. How do I get up from my chair? You can imagine going down all the way down to basically what amounts to millisecond by millisecond muscle control.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Well, so no, you will have to build a specific architecture to allow for hierarchical planning. So, hierarchical planning is absolutely necessary if you want to plan complex actions. If I want to go from, let's say, from New York to Paris, it's the example I use all the time. And I'm sitting in my office at NYU, my objective that I need to minimize is my distance to Paris at a high level, a very abstract representation of my location. I would have to decompose this into two sub-goals. First one is go to the airport. Second one is catch a plane to Paris. Okay, so my so goal is now. Going to the airport. My objective function is my distance to the airport. How do I go to the airport where I have to go in the street and hail the taxi? You can do in New York

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  32. And then plan a sequence of actions that will minimize this objective at runtime. We're not talking about learning. We're talking about inference time. So this is planning, really. And in optimal control, this is a very classical thing. It's called model predictive control. You have a model of the system you want to control that can predict the sequence of states corresponding to a sequence of commands. And you're planning a sequence of commands so that according to your world model, the end state of the system will satisfy an objective that you fix. This is the way. Rocket trajectories have been planned since computers have been around since the early 60s, essentially.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And if I push it with a particular force on the table, it's going to move. If I push the table itself, it's probably not going to move with the same force. So we have this internal model of the world in our mind, which allows us to plan sequences of actions to arrive at a particular goal. And so So now, if you have this wall model, we can imagine a sequence of actions, predict what the outcome of the sequence of action is going to be, measure to what extent the final state satisfies a particular objective, like moving the bottle to the left of the table.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Internal model that says, Here is my idea of state of the world at time t. Here is an action I'm taking. Here is a prediction of the state of the world at time t plus 1, t plus delta t, t plus two seconds, whatever it is. If you have a model of this type, you can use it for planning. So now you can do what LLMs cannot do, which is planning what you're going to do so as to arrive at a particular outcome or satisfy a particular objective. So you can have a number of objectives. I can predict that if I have an object like this, right, and I open my hand, it's going to fall, right?

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Or you just mask the second half of the video, for example. And then you train a JEPA system of the type I describe to predict the representation of the full video from the shifted one. But you also feed the predictor with an action. For example, the wheel is turned 10 degrees to the right or something. If it's a dash cam in a car, you know the angle of the wheel, you should be able to predict to some extent what's going to happen to what you see. You're not going to be able to predict all the details of objects that appear in the view, obviously, but at an abstract representation level, you can probably predict what's going to happen. So now what you have is

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Possibly, this is going to take a while before we get to that point, but there are systems already, robotic systems that are based on this idea. And what you need for this is a slightly modified version of this where imagine that you have a video and a complete video, and what you're doing to this video is that you're either translating it in time towards the future, so you only see the beginning of the video, but you don't see the latter part of it that is in the original one.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah. We have also a preliminary result that seem to indicate that the representation allows our system to tell whether the video is physically possible or completely impossible because some object disappeared or an object suddenly jumped from one location to another or changed shape or something.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Literally straight tube, the tube. Yeah, typically it's 16 frames or something, and we mask the same region over the entire 16 frames. It's a different one for every video, obviously. And then again, a train that system so as to predict the representation of the full video from the partially messed video. And that works really well. It's the first system that we have that learns good representations of video so that when you feed those representations to a supervised classifier head, it can tell you what action is taking place in a video with pretty good accuracy. So that's the first time we get something of that quality.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  39. A more recent version of this that we have is called VJPAS, which is basically the same idea as Vepa, except it's applied to video. So now you take a whole video and you mask a whole chunk of it. And what we mask is actually kind of a temporal tube. So like a whole segment of each frame in the video over the entire video.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Basic horrible things that sort of degrade the quality a little bit and change the framing, you know, crop the image. And in some cases, in the case of IJEPA, you don't need to do any of this. You just mask some parts of it. You just basically remove some regions like a big block, essentially. And then run through the encoders and train the entire system, encoder and predictor, to predict the representation of the good one from the representation of the corrupted one. So that's the iJitP. Doesn't need to know that. Whereas with Dino, you need to know it's an image because you need to do things like geometry transformation and blurring and things like that that are really image specific.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  41. So there's several scenarios. One scenario is you take an image, you corrupt it by changing the cropping, for example, changing the size a little bit, maybe changing the orientation, blurring it, changing the colors, doing all kinds of horrible things to it.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  42. A representation. And then you corrupt that input or transform it, run it through essentially what amounts to the same encoder with some minor differences and then train a predictor, sometimes the predictor is very simple, sometimes doesn't exist, but train a predictor to predict a representation of the first uncorrupted input from the corrupted input. But you only train the second branch. You only train the part of the network that is fed with the corrupted input. The other network you don't train. But since they share the same weight when you modify the first one, it also modifies the second one. And with various tricks, you can prevent the system from collapsing with the collapse of the type I was explaining before, with the system basically ignores the input. So that works very well. The technique, the two techniques we develop at Facebook.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  43. That's the hope. In fact, the techniques we're using are non contrastive. So not only is the architecture non-generative, the learning procedures we're using are non-contrastive. So we have two sets of techniques. One set is based on distillation and there's a number of methods that use this principle, one by deep mind could be way loud by FAR when called Craig. And another one called Ijepa. And Vraga, I should say, is not a distillation method, actually, but iJEPA and BYL certainly are. And there's another one also called Dino or Dino, also produced at FAIR. And the idea of those things is that you take the full input, let's say, an image, you run it through an encoder.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  44. How do we get machines to learn that before we combine that with language? Obviously, if we combine this with language, this is going to be a winner. But before that, we have to focus on how do we get systems to learn how the world works.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Well, eventually, yes, but I think if we do this too early, we run the risk of being tempted to cheat. And in fact, that's what people are doing at the moment with vision language model. We're basically cheating. We're using language as a crutch to help the deficiencies of our vision systems to kind of learn good representations from images and video. And the problem with this is that we might improve our vision language system a bit. I mean, our language models by feeding them images. But we're not going to get to the level of even the intelligence or level of understanding of the world of a cat or dog, which doesn't have language. They don't have language. And they understand the world much better than any LLM. They can plan really complex actions and sort of imagine the result of a bunch of actions.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Right. And the thing is those self-supervised algorithms that learn by prediction. Even in representation space, they learn more concept. If the input data you feed them is more redundant. The more redundancy there is in the data, the more they're able to capture some internal structure of it. And so there is way more redundancy in a structure in perceptual inputs, sensory input like vision than there is in text, which is not nearly as redundant. This is back to the question you were asking a few minutes ago. Language might represent more information really because it's already compressed. You're right about that. But that means it's also less redundant. And so self-supervision will not work as well.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  47. That is not predictable. And so we can get away without doing the transiting, without lifting the abstraction level and by directly predicting words.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  48. We don't always describe every natural phenomenon in terms of quantum field theory. That would be impossible. So we have multiple levels of abstraction to describe what happens in the world, starting from quantum field theory to atomic theory and molecules and chemistry, materials, and all the way up to kind of concrete objects in the real world and things like that. We can't just only model everything at the lowest level. And that's what the idea of JEPA is really on. Learn abstract representation in a self-supervised manner. And you can do it hierarchically as well. So that, I think, is an essential component of an intelligent system. And in language, we can get away without doing this because language is already to some level abstract and already has eliminated a lot of information.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  49. And that not only is a lot simpler, but also it allows the system to essentially learn and abstract representation of the world where what can be modeled and predicted is preserved and the rest is viewed as noise and eliminated by the encoder. So it kind of lifts the level of abstraction of the representation. If you think about this, this is something we do absolutely all the time. Whenever we describe a phenomenon,

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source

  50. So there's a lot of things in the world that we cannot predict. For example, if you have a self driving car driving down the street or road, there may be trees around the road. And it could be a windy day. So the leaves on the tree are kind of moving in kind of semi-chaotic random ways that you can't predict and you don't care. You don't want to predict. So what you want is your encoder to basically eliminate all those details. It will tell you there's moving leaves, but it's not going to keep the details of exactly what's going on. And so when you do the prediction in representation space, you're not going to have to predict every single pixel of every leaf.

    2024-03-07 · Lex Fridman Podcast · #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI · IDENTIFIED FROM THE TRANSCRIPT · source