YouSaid · the spoken record
Oriol Vinyals
- lines on the record
- 127
- first
- 2022-07-26
- most recent
- 2022-07-26
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“The know that you can get there. I mean, this is the beauty. Like, if you're not risking to trying to do something that feels impossible, you're not going to get there. But you need the way to measure progress. So the benchmarks that you build are critical. I've seen this beautifully pay out in many projects. I mean, maybe the one I've seen it more consistently, which means we establish the metric actually the community did, and then we leverage that massively as Alpha Fold. This is a project where the data, the metrics, where all there, and all it took was, and it's easier said than done, an amazing team working not to try to find some incremental improvement and publish, which is one way to do research that is valid, but aim very high and work literally for years to iterate over that process and working for years.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Everything. But then when you zoom into this age of these projects, then you realize the debil is indeed in the details. And then the teams. Have to work kind of together towards these goals. So engineering of data and obviously clusters and large scale is very important. And then one that is often not, maybe nowadays it is more clear, is benchmark progress, right? So we're talking here about multiple months of tens of researchers and people that are trying to organize the research and so on, working together.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“How important are all the pieces of these projects? How do they come together? So I'll give you maybe some of the ingredients of success that are common across these, but not the obvious ones in machine learning. I can always also give you those. Basically, there is engineering is critical. So very good engineering because ultimately we're collecting data sets, right? So the engineering of data and then of deploying the models at scale into some compute cluster that cannot go understated that is a huge factor of success. And it's hard to believe that details matter so much. We would like to believe that it's true that there is more and more of a standard formula, as I was saying, like this recipe that works for”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Look, I'll give you an answer that's very important because maybe people don't quite realize this, but the themes behind these efforts, the actual humans, that's maybe the surprising, obviously positive way. Anytime you see these breakthroughs, I mean, it's easy to map it to a few people. There's people that are great at explaining things and so on. That's very nice. But maybe the learnings or the method learnings that I get as a human about this is, sure, we can move forward. But the surprising bit is how”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, that's precisely, I mean, the beauty of research and the research business we're in, I guess, is to figure this out and ask the right questions and then iterate with the whole community publishing findings and so on. But yeah, this is a question. It's not the only question, but it's certainly, as you ask, is on my mind constantly, right? And so we'll need to wait for maybe the let's say five years. Let's hope it's not 10 to see what are the answers. Some people will largely believe in unsupervised or self-supervised learning of single modalities and then crossing them. Some people might think end-to-end learning is the answer modularity is maybe the answer. So we don't know, but we're just definitely excited to find out.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, those are open questions, I would say. I mean, the way you put, let me maybe further your example, right? If all you see is land images, but you're reading all about land and water worlds, but in books, right? Imagine, would that be enough? I mean, good question. We don't know, but I guess maybe you can join us if you want in our quest to find this. That's precisely.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“A bit of training, adding a few parameters, thinking of this as nearest neighbor, or just simply thinking of there's a sequence of words, it's a prefix. And that's the new classifier. We'll see, right? There's the beauty of research. But what's important is that is a good goal in itself that I see as very worthwhile pursuing for the next stages of not only meta-learning, I think this is basically what's exciting about machine learning period to me.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Nita's neighbor, which is very common in computer vision community, right? There's a very active area of research about how do you compute the distance between two images, but if you have a good distance metric, you also have a good classifier, right? All I'm saying is now these distances and the points are not just images. They're like words or sequences of words and images and actions that teach you something new, but it might be that technique-wise does come back. I will say that it's not necessarily true that you might not ever train the weights a bit further. Some aspect of meta-learning, some techniques in meta-learning do actually do a bit of fine-tuning, as it's called. They train the weights a little bit when they get a new task. So as I call the how or how we're going to achieve this, as a deep learner, I'm very skeptic, we're going to try a few things, whether”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“The metric is not the distance between the images or something simple, it's something that you compute that's much more advanced. But in a way, it's very similar, right? You simply are uploading some knowledge to this pre-trained system in nearest neighbor, maybe the metric is learned or not, but you don't need to further train it. And then now you immediately get a classifier out of this, right? Now it's just an evolution of that concept, very classical concept in machine learning, which is, yeah, just learning through what's the closest point, closest by some distance, and that's it. It's an evolution of that. And I will say how I saw meta-learning when we worked on a few ideas in 2016 was precisely through the lens of”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. By the way, this is not. If you look at way back at different ways to tackle even classification tasks, so this comes from longstanding literature in machine learning. What I'm suggesting could sound to some a bit like nearest neighbor. So Nita's neighbor is almost the simplest algorithm that does not require learning. So it has this interesting, like you don't need to compute gradients. And what Nearest neighbor does is you quote unquote have a data set or upload a data set. And then all you need to do is a way to measure distance between points and then to classify a new point, you're just simply computing what's the closest point in this massive amount of data. And that's my answer. So you can think of prompting in a way as you're uploading, not just simple points.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Percent performance in Go, in chess, in StarCraft, the next iteration might be 20% performance across quote unquote all tasks, right? And even if it's not as good, it's fine. We have ways to also measure progress because we have those special agents, specialized agents, and so on. So this is to me very exciting. next iteration models are definitely hinting at that direction of progress, which hopefully we can have there are obviously some things. If you ask me five to ten years, you might see these models that start to look more weights that are already trained, and then it's more about teaching or make their meta learn what you're trying to induce in terms of tasks and so on, well beyond the simple now task we're starting to see emerge like smaller methodic tasks and so on.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“But the way it were specialized, and what we're hoping is that we can teach a network to play games, to play any game, just using games as an example, through interacting with it, teaching it, uploading the Wikipedia page of StarCraft. This is in the horizon and obviously there are details need to be filled and research needs to be done. But that's how I see meta-learning above, which is going to be beyond prompting. It's going to be a bit more interactive. It's going to, you know, the system might tell us to give it feedback after it maybe makes mistakes or it loses a game. But it's nonetheless very exciting because if you think about this this way, the benchmarks are already there. We just repurpose the benchmarks, right? So in a way, I like to map the space of what maybe AGI means to say, okay, like we went 100.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“We have a system, a set of weights that we can teach it to play StarCraft. Maybe not at the level of Alpha Star, but play StarCraft a complex game. We teach it through interactions, to prompting. You can certainly prompt a system. That's what Gatto shows to play some simple Atari games. So imagine if you start talking to a system, teaching it a new game, showing it examples of in this particular game, this user did something good. Maybe the system can even play and ask you questions, say, hey, I played this game. I just played this game. Did I do well? Can you teach me more? So five, maybe to 10 years, these capabilities or what meta learning means will be much more interactive, much more rich, and through domains that we were specializing, right? So you see the difference, right? We build Alpha Star specialized to play StarCraft. The algorithms were generated.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, it expanded what it meant. So that's what you say. What does it mean? So it's an evolving term. But here is maybe now looking forward, looking at what's happening, obviously in the community with more modalities, what we can expect. And I would certainly hope to see the following. And this is a pretty drastic Hope, but in five years, maybe we chat again.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Are now doing a new task, right? So they meta-learn these new capabilities. Now that's where we are now. Flamingo expanded this to visual and language, but it basically has the same abilities. You can teach it, for instance, an emergent property was that you can take pictures of numbers and then do arithmetic with the numbers just by teaching it. Oh, when I show you three plus six, I want you to output nine and you show it a few examples and now it does that. So it went way beyond this image net sort of category of images that we were a bit stuck maybe before this revelation moment that happened in 2000, I believe it was 19, but it was after we chatted.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“That you just defined kind of on the fly. So fast forward, it was revealed that language models are future learners. That's the title of the paper. So very good title. Sometimes titles are really good. So this one is really, really good because that's the point of GPT-3 that showed that, look, sure, we can focus on object classification and what meta-learning means within the space of learning object categories. This goes beyond or before rather to also Omniglot, before ImageNet and so on. So there's a few benchmarks. To now all of a sudden, we a bit unlock from benchmarks. And through language, we can define tasks, right? So we literally telling the model some logical task or little thing that we wanted to do. We prompted much like we did before, but now we prompt it through natural language. And then not perfectly. I mean, these models have failure modes and that's fine, but these models then.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“A capability to learn about object identities. So it was very much overfitted to vision and object classification. And the part that was meta about that was that, oh, we're not just learning a thousand categories that ImageNet tells us to learn. We're going to learn object categories that can be defined when we interact with the model. It's interesting to see the evolution, right? The way this started was We have a special language that was a data set, a small data set that we prompted the model with saying, hey, here is a new classification task. I'll give you one image and the name, which was an integer at the time of the image and a different image and so on. So you have a small prompt in the form of a data set, a machine learning data set. And then you got then a system that could then predict or classify these objects.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, great question. Maybe it's good to give another data point looking backwards rather than forward. So when we talk in 2019, meta-learning meant something that has changed mostly through the revolution of GPT-3 and beyond. So what meta learning meant at the time was driven by what benchmarks people care about in meta-learning. And the benchmarks were about”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Right. So there's still big questions, but that idea is actually really akin to software engineering, which we're not re-implementing libraries from scratch. We're reusing and then building ever more amazing things, including neural networks with software that we're reusing. So I think this idea of modularity, I like it. I think it's here to stay. And that's also why I mentioned it's just the beginning, not the end.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so that vision is beautiful. I think there's still the question about within single modalities, like Chinchilla was reused. But now if we train an X iteration of language models, are we going to use chinchilla or not?”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Making the weight higher and so on and so forth. So it's a fascinating capability and it comes from this key idea of modularity where we took a frozen brain and we just added a new capability. So the question is, so in a way, you can see even from DeepMind, we have flamingo that this moderate approach and thus could leverage scale a bit more reasonably because we didn't need to retrain a system from scratch. And on the other hand, we had Gato, which used the same data sets, but then it trained it from scratch, right? And so I guess big question for the community is, should we train from scratch or should we embrace modularity? And this lies, like this goes back to modularity as a way to grow but reuse seems like natural and it was very effective, certainly.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“About Twitter was this image from Obama that is placing a weight and someone is kind of waiting themselves and it's kind of a joke style image. And it's notable because I think Andrew Carpathi a few years ago said no computer vision system can understand the subtlety of this joke in this image, all the things that go on. And so what we try to do, and it's by anecdotally, I mean, this is not a proof that we solved this issue, but it just shows that you can upload now this image and start conversing with the model, trying to make out if it gets that there's a joke because the person waiting themselves don't see that someone behind is”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Right, so it's total 80 billion, the largest one we released, and then you train it on a few data sets that contain vision and language. And once you interact with the model, you start seeing that you can upload an image and start sort of having a dialogue about the image, which is actually not something it's very similar and akin to what we saw in language only, this prompting abilities that it has. You can teach it a new vision task, right? It does things beyond the capabilities that in theory the data sets provided in themselves, but because it leverages a lot of the language knowledge acquired from Chinchilla, it actually has these few shot learning ability and this emerging abilities that we didn't even measure once we were developing the model, but once developed, then as you play with the interface, you can start seeing, wow, okay, yeah, it's cool. We can upload, I think, one of the tweets talking.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“For chinchilla. Yeah, chinchilla is 70 billion. And then the ones we add on top, which kind of almost is almost like a way to overwrite its little activations so that when it sees vision, it does kind of a correct computation of what it's seeing, mapping it back towards. So to speak, that adds an extra 10 billion parameters.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“A chatbot where you can also upload images and start conversing about images, but it's also kind of a dialogue style chatbot.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“So we created a small sub network initialized not from random but actually from self-supervised learning that a model that understands vision in general. And then we took data sets that connect the two modalities, vision and language. And then we froze the main part, the largest portion of the network, which was chinchilla, that is 70 billion parameters. And then we added a few more parameters on top train from scratch, and then some others that were pre-trained from with the capacity to see. It was not tokenization in the way I described for Gato, but it's a similar idea. And then we trained the whole system. Parts of it were frozen, parts of it were new. And all of a sudden, we developed flamingo, which is an amazing model that is essentially, I mean, describing it.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“All the way down. So, the way that we did modularity very beautifully in flamingo is we took chinchilla, which is a language only model, not an agent if we think of actions being necessary for agency. So we took chinchilla, we took the weights of chinchilla, and then we froze them. We said, these don't change. We train them to be very good at predicting the next word. It's a very good language model, state-of-the-art at the time you release it, et cetera, et cetera. We're going to add a capability to see we are going to add the ability to see to this language model. So we're going to attach small pieces of neural networks at the right places in the model. It's almost like injecting the network with some weights and some substructures in the good way, right? So you need the research to say what is effective, how do you add this capability without destroying others, et cetera.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, exactly. And it has to be, right? We need to have creativity. And sometimes you need to protect pockets of people, researchers, and so on.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“In an interesting way, kind of akin to what Agato did, but slightly different technique for tokenizing. But we don't need to go into that detail. But what Flamingo also did, which Gato didn't do, and that just happens because these projects, you know, they're different, it's a bit of like the exploratory nature of research, which is great”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“So the Meow network is not hard to grow if you retrain it What's hard is well, we have now one billion parameters. We train them for a while. We spend some amount of work towards building these weights that are an amazing initial brain for doing this kind of task we care about. Could we reuse the weights and expand to a larger brain? And that is extraordinarily hard, but also exciting from a research perspective and a practical perspective point of view, right? So there's this notion of modularity in software engineering. And we starting to see some examples and work that leverages modularity. In fact, if we go back one step from Gato to a work that I would say train much larger, much more capable network called Flamingo. Flamingo did not deal with actions, but it definitely dealt with images in a”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Space. I'm defining actions by words, right? So you could imagine a world in which I'm not learning that the action app in Atari is its own number. The action app in Atari maybe is literally the word or the sentence app in Atari, right? And that would mean we now leverage much more from the language. This is not what we did here, but certainly it might make these connections much easier to learn and also to teach the model to correct its own actions and so on. So all these to say that Gatto is indeed the beginning, that it is a radical idea to do this this way, but there's probably a lot more to be done. And the results, to be more impressive, not only through scale, but also through some new research that will come hopefully in the years to come.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“The same. We have word vectors or embeddings. We have image vector or embeddings and action vector of embeddings. And the beauty here is that as you train this model, if you visualize these little vectors, it might be that they start aligning, even though they're independent parameters. There could be anything. But then it might be that you take the word gato or cat, which maybe is common enough that actually has its own token. And then you take pixels that have a cat. And you might start seeing that these vectors look like they align, right? So by learning from this vast amount of data, the model is realizing the potential connections between these modalities. Now, I will say there would be another way, at least in part, to not have these different vectors for each different modality. For instance, when I tell you about actions in certain”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Is it focusing on one modality or not? The intermediate weights that are converting from this input of integers to the target integer you're predicting next, those weights certainly are common. And then the way the tokenization happens, there is a special place in the neural network, which is we map this integer, like number 1001, to a vector of real numbers, like real numbers we can optimize them with gradient descent, right? The functions we learn are actually surprisingly differential. That's why we compute gradients. So this step is the only one that this orthogonality you mentioned applies. So mapping a certain token for text or image or actions, each of these tokens gets its own little vector of real numbers that represents this. If you look at the field back, many years ago, people were talking about word vectors or word embeddings.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Say, okay, given the numbers 10,001, 10,004, 10,005, the next number that comes is 20,006, which is in the action space. And you're just optimizing these weights via very simple gradient, like mathematical is almost the most boring algorithm you could imagine. We settle the weights so that given this particular instance, these weights are set to maximize the probability of having seen this particular sequence of integers for this particular game. And then the algorithm does this for many, many, many iterations looking at different modalities, different games, right? That's the mixture of the data set we discuss. So in a way, it's a very simple algorithm. And the weights, right, they're all shared, right? So in terms of”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, it is one mind, and it's actually the simplest algorithm, which that's kind of in a way how it feels like the field hasn't changed since back propagation and gradient descent was purpose for learning neural networks. So there is obviously details on the architecture. This has evolved. The current iteration is still the transformer, which is a powerful sequence modeling architecture. The goal of this setting these weights to predict the data is essentially the same as basically I could describe. I mean, we described a few years ago Alpha Star, language modeling, and so on, right? We take, let's say, an Atari game, we map it to a string of numbers that will all be probably image space and action space interlived. And all we're going to do is...”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“The highest order of integers to actions, right? Which we discretize. And actions are very diverse in Atari. There's, I don't know if 17 discrete actions in robotics actions might be torques and forces that we apply. So we just use kind of similar ideas to compress these actions into tokens. And then we just, that's how we map now all the space to this sequence of integers. But they occupy different space and what connects them is then the learning algorithm. That's where the magic happens.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“That makes it orthogonal. So, what connects these concepts is the data, right? Once you have a data set, for instance, that captions images, that tells you, oh, this is someone playing a Frisbee on a green field, now the model will need to predict the tokens from the text green field to then the pixels, and that will start making the connections between the tokens. So these connections happen as the algorithm learns. And then the last, if we think of these integers, the first few”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“The statistics of the images and then quantize them based on the statistics much like you do in words, right? So common substrings are allocated a token and images is very similar. But there's no connection, the token space, if you think of the tokens are an integer and at the end of the day. So now we work on maybe we have about let's say, I don't know the exact numbers, but let's say 10,000 tokens for text, right? Certainly more than characters because we have groups of characters and so on. So from 1 to 10,000, those are representing all the language and the words we'll see. And then images occupy the next set of integers. So they're completely independent, right? So from 10,000 to 20,000, those are the tokens that represent these other modality images. And that is an interesting aspect.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, as I said, the algorithms look actually very similar to, they use the cosine transform in JPG. The approach we usually do in machine learning when we deal with images and we do this quantization step is a bit more data driven. So rather than have some sort of Fourier basis for how frequencies appear in the natural world, we actually just”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“No, you're only using this notion of compression. So you're trying to find common, it's like JPG or all these algorithms. It's actually very similar at the tokenization level. All we're doing is finding common patterns and then making sure. In a lossy way, we compress these images given the statistics of the images that are contained in all the data we deal with.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Like every pixel with every intensity, that would mean we have a very long sequence, right? Like if we're talking about 100 by 100 pixel images, that would make the sequences far too long. So what was done there is you just use a technique that essentially compresses an image into maybe 16 by 16 patches of pixels and then that is mapped again tokenized you just essentially quantize this space into a special word that actually maps to this little sequence of pixels and then you put the pixels together in some raster order and then that's how you get out or in or in the image that you're processing.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“The way we're doing these things, these things is they're actually mapped to seek small sequences of characters. So you can actually play with these models and input emojis. It will output emojis back, which is actually quite a fun exercise. You probably can find other tweets about this out there. But yeah, so anyways, text, there's like it's very clear how this is done. And then in Gato, what we did for images is we map images to essentially, we compressed images, so to speak, into something that looks more like...”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, yes, and you could think of very, very common words like the. I mean, that would be a single token, but very quickly you're talking two, three, four tokens or something.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, for a word in English, right? I mean, every language is very different. The current level or granularity of tokenization generally means it's maybe two to five. I mean, I don't know the statistics exactly, but to give you an idea, we don't tokenize at the level of letters, then it would probably be like, I don't know what the average length of a word is in English, but that would be the minimum set of tokens you could use.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, that's a great question. Tokenization is the entry point to actually make all the data look like a sequence because tokens then are just kind of these little puzzle pieces we break down anything into these puzzle pieces and then we just model what's the what's this puzzle look like, right, when you make it laid down in a line, so to speak in a sequence. So in Gato, the text, there's a lot of work, you tokenize text usually by looking at commonly used substrings, right? So there's ing in English is a very common substring, so that becomes a token. There's quite well studied problem on tokenizing text. And Gato just used the standard techniques that have been developed from many years, even starting from Ngram models in the 1950s and so on.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Right, my emerge with scale, and also, I believe there's some new research or ways in which you prepare the data that you might need to sort of make it more clear to the model that you're not only playing Atari and it's just you start from a screen. Hey, in this sequence that I'm showing, you're going to be playing an Atari game. So text might actually be a good driver to enhance the data, right? So then these connections might be made more easily, right? That's an idea that we start seeing in language. But obviously beyond this is going to be effective, right? It's not like I don't show you a screen and you from scratch, you're supposed to learn a game. There is a lot of context with my set. So there might be some work needed as well to set that context. But anyways, there's a lot of work.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, okay. That's a different conversation. But for neural networks, certainly size does matter. So it's the beginning because it's relatively small. So obviously scaling this idea up might make the connections that exist between text on the internet and playing Atari and so on more synergistic with one another. And you might gain. That moment we didn't quite see, but obviously that's why it's the beginning. That's synergy mode.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Single brain is not that big of a brain compared to most of the neural networks we see these days. It has one billion parameters. Some models we're seeing getting the trillions these days and certainly 100 billion feels like a size that is very common from when you train these jobs. So the actual agent is relatively small, but it's been trained on a very challenging diverse data set, not only containing all of internet, but containing all these Asian experience playing very different distinct environments. So this brings us to the part of the tweet of this is not the end, it's the beginning. It feels very cool to see Gato in principle is able to control any sort of environment that especially the ones that it's been trained to do these 3D games, Atari games, all sorts of robotics tasks and so on.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Agents that play in different environments. So we kind of created a data set of these trajectories, as we call them, or Asian experiences. So in a way, there are other agents we trained for a single mind purpose to, let's say, control 3D game environment and navigate a maze. So we had all the experience that was created through the one agent interacting with that environment. And we added this to the data set, right? And as I said, we just see all the data, all these sequences of words or sequences of this agent interacting with that environment or agents playing Atari and so on. We see this as the same kind of data. And so we mix these data sets together and we train Gato. That's the G part, right? It's general because it really has mixed. It doesn't have different brains for each modality or each narrow task.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, yeah, we can get back to that later. But, anyways, going back to the meow and the gato, right? So one of the leaps forward and what took the team a lot of effort and time was As you were asking, how has Gato been trained? So I told you, Gato is this transformer neural network, Molo's actions, sequences of actions, words, et cetera. And then the way we train it is by essentially pulling data sets of observations, right? So it's a massive imitation learning algorithm that it imitates obviously to what is the next word that comes next from the usual data sets we use before, right? So these are these web scale style data sets of people writing on webs or chatting or whatnot, right? So that's an obvious source that we use on all language work. But then we also took a lot of agents that we have at DeepMind. I mean, as you know, DeepMine, we're quite interested in learning reinforcement learning and learning.”
2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source