YouSaid · the spoken record

Oriol Vinyals

lines on the record
127
first
2022-07-26
most recent
2022-07-26
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I think we're all excited to build AGI to. Empower or make us more powerful as human species, not to say there might be some hybridization. I mean, this is obviously speculation, but there are companies also trying to, the same way medicine is making us better. Maybe there are other things that are yet to happen on that. But if the ratio is not at most one-to-one, I would not be happy. So I would hope that we are part of the equation, but maybe there's maybe a one-to-one ratio feels like possible, constructive and so on, but it would not be good to have a misbalance, at least from my core beliefs and the why I'm doing what I'm doing when I go to work and I research what I research.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Just work together to find what would be reasonable in terms of growth or how we coexist if that is to happen. I am very excited about obviously the aspects of automation that make people that obviously don't have access to certain resources or knowledge for them to have that access. I think those are the applications in a way that I'm most exciting to see and to personally work towards.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I'm afraid if there's a lot more. So I think maybe we'll need to think about if we truly get there just thinking of limited resources like humanity clearly hits some limits and then there's some balance hopefully that biologically the planet is imposing and we should actually try to get better at this as we know there's quite a few issues with having too many people coexisting in a resource limited way. So for digital entities it's an interesting question. I think such a limit maybe should exist but maybe it's going to be imposed by energy availability because this also consumes energy. In fact most systems are more inefficient than we are in terms of energy required. But definitely I think as a society we'll need to

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Define reward functions that from a seat of imitating human level intelligence that is general and then going beyond that bit is not so clear in my lifetime. But certainly human level, yes. And I mean, that in itself is already quite powerful, I think. So going beyond, I think it's obviously not, we're not going to not try that if then we get to superhuman scientists and discovery and advancing the world. But at least human level is also in general is also very, very powerful.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  5. I definitely think it's possible that it will go far beyond, but I'm definitely convinced that it will be human level intelligence. And I'm hypothesizing about the beyond because the beyond bit is a bit tricky to define, especially when we look at the current formula of starting from this imitation learning standpoint, right? So we can certainly imitate humans at language and beyond. So getting at human level through imitation feels very possible. Going beyond will require reinforcement learning and other things. And I think in some areas that certainly already has paid out. I mean, Go being an example that's my favorite so far in terms of going beyond human capabilities. But in general, I'm not sure we can.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  6. We have in practice, but then maybe we should also shape how the hardware looks like based on which methods might be needed to scale. And that's an interesting contrast of this GPU comment that is we got it for free almost because games were using this, but maybe now. We don't have the hardware, although, in theory, many people are building different kinds of hardware these days, but there's a bit of this notion of hardware lottery for scale that might actually have an impact at least on the year, again, scale of years on how fast we'll make progress to maybe a version of neural nets or whatever comes next that might enable truly intelligent agents.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  7. It is very appealing to search in domains like Go where you have a clear reward function that you can then discard some search traces. But then in some other tasks, it's not very clear how you would do that. Although recently one of our recent works, which actually was mostly mimicking or a continuation and even the team and the people involved were pretty much very like intersecting with Alpha Star was Alpha Code in which we actually saw the bitter lesson how scale of the models and then a massive amount of search yielded this kind of very interesting result of being able to have human level code competition. So I've seen examples of it being literally mapped to search and scale. I'm not so convinced about the search bit, but certainly I'm convinced scale will be needed, so we need general methods. We need to test them and we need to make sure that we can scale them given the hardware that

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I certainly think this essentially mimics a bit of the deep learning research almost like philosophy that on the one hand we want to be data agnostic, we don't want to pre-process data sets, we want to see the bytes, right? Like the true data as it is and then learn everything on top. very much agree with that. And I think scaling up feels at the very least, again, necessary for building incredible complex systems. It's possibly not sufficient bearing that we need a couple of breakthroughs. I think Rich Saturn mentioned search being part of the equation of skill and search. I think search, I've seen it, that's been more mixed in my experience. So from that lesson in particular, search is a bit more tricky because

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  9. The deep belief into some certain style of research pays out, right? It's good to be practical sometimes. And I actually think Ilya and myself are practical, but it's also good there's some sort of long-term belief and trajectory. Obviously, there's a bit of lack involved, but it might be that that's the right path. Then you clearly are ahead and hugely influential to the field, as he has been.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  10. In this case, very deeply into this style of research and clearly has had a tremendous track record of successes and so on. The funny bit about that talk is that we rehearse the talk in a hotel room before. And the original version of that talk would have been even more controversial. So maybe I'm the only person that has seen the unfiltered version of the talk. And maybe when the time comes, maybe we should revisit some of the skip slides from the talk from Ilya. But I really think.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Yeah, I mean, just on the point of, I mean, Ilya has been an inspiration, I mean, quite a few colleagues, I can think, shaped the person you are, like Ilya certainly gets probably the top spot, if not close to the top. And if we go back to the question about people in the field, like how the role would have changed the field or not, I think Ilya's case is interesting because he really has a deep belief in the scaling up of neural networks. There was a talk that is still famous to this day from the sequence to sequence paper where he was just claiming, just give me supervised data and large neural network, and then you'll solve basically all the problems, right? That vision was already there many years ago. So it's good to see someone who is.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Created debates on some of his tweets, right? That maybe it's good we have them early anyways, right? But yeah, then the reactions are usually polarizing. I think we're just seeing kind of the reality of social media a bit there as well, reflected on that particular topic or set of topics he's tweeting about.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I think it's Ilya's personality just knowing him for a while. Everyone in Twitter, I guess, gets a different persona. And I think Ilya's one does not surprise me, right? So I think knowing Ilya from before social media and before AI was so prevalent, I recognize a lot of his character. So that's something for me that I feel good about. A friend that hasn't changed or is still true to himself, right? Obviously, there is, though, a fact that your field becomes more popular and he is obviously one of the main figures in the field having done a lot of advancement. So I think that the tricky bit here is how to balance your true self with the responsibility that your worst carry. So in this sense, I think, yeah, like I appreciate the style and I understand it.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Yes, of course. Hopefully. I mean, we're, you know, ultimately we all have sharepaths and there's friendships that go beyond obviously institutional institutions and so on. So I hope he tells me the truth.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  15. As a learning system myself, I want to keep exploring. And I think it's great that to see parts of the debate and even I've seen a level of maturity in the conferences that deal with AI. If you look five years ago to now, just the amount of workshops and so on has changed so much is impressive to see how much topics of safety, ethics and so on come to the surface, which is great. And if we were too early, clearly it's fine. I mean, it's a big field and there's lots of people with lots of interest that will do progress or make progress. And obviously I don't believe we're too late in that sense, like I think it's great that we're doing this already.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  16. The people we talk to not include only our own researchers and so on. And in fact, places like DeepMind but elsewhere, there's more interdisciplinary. Groups forming up to start asking and really working with us on these questions because obviously this is not initially what your passion is when you do your PhD, but certainly it is coming. So it's fascinating kind of, it's the thing that brings me to one of my passions that is learning. So in this sense, this is kind of a new area that

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Yeah, absolutely. It's definitely a topic that is important we think about. And I think in a way, I always see not every movie is equally on point with certain things, but certainly science fiction in this sense at least has prepared society to start thinking about certain topics that even if it's too early to talk about, as long as we are like reasonable, it's certainly going to prepare us for both the research to come and how to, I mean, there's many important challenges and topics that come with building an intelligent system, many of which you just mentioned, right? So I think we're never going to be fully ready unless we talk about this and we start also, as I said, just kind of expanding the

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  18. So maybe, like, yeah, metering the question. And since you talk to a few people, then you do think that we'll need to figure something out in order to achieve intelligence in a grander sense of the world.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Exploration, exploitation, like all the how these things happen. I think this clearly can inform algorithmic level research. And I've seen some examples of this being quite useful to then guide the research. Even it might be for the wrong reasons, right? So I think biology and what we know about ourselves can help a whole lot to build essentially what we call AGI, these general, the real ghetto, right? The last step of the chain, hopefully. But consciousness in particular, I don't myself at least think too hard about how to add that to the system. But maybe my understanding is also very personal about what it means, right? I think this, even that in itself is a long debate that I know people have often, and maybe I should learn more about this.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  20. If consciousness or any other biological or evolutionary lesson can be repurposed to then influence our next set of algorithms, that is a great way to actually make progress, right? And the same way I try to explain transformers a bit how it feels we operate when we look at text specifically. These insights are very important, right? So there's a distinction between details of how the brain might be doing computation. I think my understanding is, sure, there's neurons and there's some resemblance to neural networks, but we don't quite understand enough of the brain in detail, right? to be able to replicate it but then more if you if you zoom out a bit how we then our thought process how memory works maybe even how evolution got us here what's

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Honestly, I think probably not to the degree of intelligence that There's this brain that can learn, can be extremely useful, can challenge you, can teach you. Conversely, you can teach it to do things. I'm not sure it's necessary, personally speaking.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Through interactions with the larger community, we can also have a certain level of education that in practice also will matter because, I mean, one question is how you feel about this. But then the other very important is you're starting to interact with this in products and so on. It's good to understand a bit what's going on, what's not going on, what's safe, what's not safe, and so on. Otherwise, the technology will not be used properly for good, which is obviously the goal of all of us, I hope.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  23. To see a bit more of the math and the fact that literally to create these models, if we had the right software, it would be 10 lines of code and then just a dump of the internet versus like then the complexity of the creation of humans from their inception, right? And also the complexity of evolution of the whole universe to where we are that is fills orders of magnitude more complex and fascinating to me. So I think, yeah, maybe part of the only thing I'm thinking about trying to tell you is, yeah, I think explaining a bit of the magic, there is a bit of magic. It's good to be in love, obviously, with what you do at work. And I'm certainly fascinated and surprised quite often as well. But I think hopefully as experts in biology, hopefully will tell me this is not as magic and I'm happy to learn that.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Not that fast progress, right? I mean, what we were thinking at the time, like almost 100 years ago, is not that dissimilar to what we're doing now. But at the same time, yeah, obviously others might experience the personal experience, I think no one should tell others how they should feel. I mean, the feelings are very personal, right? So how others might feel about the models and so on, that's one part of the story that's important to understand for me personally as a researcher. And then when I maybe disagree or I don't understand or see that, yeah, maybe this is not something I think right now is reasonable knowing all that I know. One of the other things and perhaps partly why it's great to be talking to you and reaching out to the world about machine learning is, hey, let's demystify a bit the magic and try.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Right. So, I mean, I guess the other side of this is that's how I feel personally. I mean, you asked me about the person, right? Now it's very interesting to see how other humans feel about things, right? Again, like I'm not as amazed about things that I feel this is not as magical as this other thing because of maybe how I got to learn about it and how I see the curve a bit more smooth because I've just seen the progress of language models since Shannon in the 50s and actually looking at that time scale.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Just makes me like, wow, there's such a level of complexity difference still, right? Like orders of magnitude complexity that, sure, these weights, I mean, we train them and they do nice things, but they're not at the level of biological entities, brains, cells. It just feels like it's just not possible to achieve the same level of complexity behavior and my belief when I talk to other beings is certainly shaped by this amazement of biology that maybe because I know too much, I don't have about machine learning, but I certainly feel it's very far-fetched and far in the future to be calling or to be thinking, well, this mathematical function that is differentiable is in fact sentient and so on.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  27. When I was a bit more involved in AlphaFold, learning a bit about proteins and about biology and about life, the complexity, it feels like it really is like, I mean, if you start looking at the things that are going on at the atomic level and also, I mean, there's obviously the We are maybe inclined to try to think of neural networks as like the brain, but the complexities and the amount of magic that it feels when, I mean, I'm not an expert, so it naturally feels more magic. But looking at biological systems as opposed to these computer computational brains,

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Sadly, no, I think, yeah, sadly, I have not, yeah, I think the current, any of the current models, although very useful and very good, yeah, I think we're quite far from that. And there's kind of a converse side story. So one of the my passions is about science in general. And I think I feel I'm a bit of a failed scientist. That's why I came to machine learning, because you always feel and you start seeing this, that machine learning is maybe the science that can help other sciences, as we've seen, right? Like it's such a powerful tool. So thanks to that angle, right? That, okay, I love science. I love, I mean, I love astronomy. I love biology, but I'm not an expert and I decided, well, the thing I can do better at these computers. But having especially with

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Yeah, so luckily we have quite a few benchmarks, some of which are simpler or maybe they're more like, I think people might call these systems one versus systems two style. Think what we're not seeing luckily is that extrapolations from maybe slightly more smooth or simpler benchmarks are translating to the harder ones. But that is not to say that this extrapolation will hit its limits. And when it does, then how much we scale or how we scale will suddenly be a bit suboptimal until we find better loss. And these laws, again, are very empirical loss. They're not like physical loss of models, although I wish there would be better theory about these things as well. But so far, I would say empirical theory, as I call it, is way ahead than actual theory of machine learning.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Maybe hundreds of millions up to one billion parameters. And then the cool thing is that you create some loss, right? Some loss that some trends, you extract trends from data that you see, okay, like it looks like the amount of data required to train now a 10x larger model would be this. And these loss so far, these extrapolations have helped us save compute and just get to a better place in terms of the science of how should we run these models at scale, how much data, how much depth, and all sorts of questions, we start asking, extrapolating from small scale. But then this emergence is sadly that not everything can be extrapolated from scale depending on the benchmark and maybe the harder benchmarks are not so good for extracting these laws. But we have a variety of benchmarks at this.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Barcy can only be done at certain scale. And one thing that conversely, I've seen great progress on in the last couple of years is this notion of science of deep learning and science of scale in particular, right? So on the negative is that there's some benchmarks for which progress might need to be measured at minimum at certain scale until you see then what details of the model matter to make that performance better, right? So that's a bit of a con. But what we've also seen is that you can sort of empirically analyze behavior of models at scales that are smaller, right? So let's say to put an example, we had this chinchilla paper that revised the so-called scaling laws of models. And that whole study is done at a reasonably small scale.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Only then you might start seeing performance going from random to non random. And this is more empirical. There's no formalism or theory behind this yet, although it might be quite important. But we're seeing these phase transitions of random performance and until some, let's say, scale of a model. And then it goes beyond that. And it might be that you need to fit a few low order bits of thought before you can make progress on the whole task. And if you could measure actually those breakdown of the task, maybe you would see more smooth, like, yeah, these, you know, once you get this and this and this and this and this, then you start making progress in the task. But it's somehow a bit annoying because then it means that certain questions we might ask about architectures

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Hop reasoning task, right? It's like here is an input and you think for a few milliseconds or 100 milliseconds, 300 as a human, and then you tell me, yeah, there's an alpaca in this image. In language, we are seeing benchmarks that require more pondering and more thought in a way, right? You need to look for some subtleties that it involves inputs that you might think of even if the input is a sentence describing a mathematical problem. There is a bit more processing required as a human and more introspection. How these benchmarks work means that there is actually a threshold just going back to how transformers work in this way of querying for the right questions to get the right answers that might mean that performance becomes random until the right question is asked by the querying system of a transformer or of a language model like a transformer. And then only only

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yeah, I mean, this is a property that we start seeing in systems that actually tend to be, so in machine learning traditionally, again, going to benchmarks, I mean, if you have some input output that is just a single input and a single output, you generally, when you train these systems, you see reasonably smooth curves when you analyze how much the data set size affects the performance or how the model size affects the performance or how much you train the system for affects the performance, right? If we think of image net, like the training curves look fairly smooth and predictable in a way. And I would say that's probably because of the, it's kind of...

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  35. It's still very valuable that someone put the thought and the time and the vision to build certain benchmarks. We've seen progress thanks to, but we're going to repurpose the benchmarks. The beauty of Atari is like we solved it in a way. We use it in Gato. It was critical, and I'm sure there's still a lot more to do thanks to that amazing benchmark that someone took the time to put, even though at the time maybe, oh, you have to think what's the next iteration of architectures. That's what maybe the field recognizes. But we need to, that's another thing. We need to balance in terms of humans behind. We need to recognize all these aspects because they're all critical. And we tend to think of the genius, the scientists, and so on. I'm glad you're, I know you have a strong engineering background.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Revolution, we are surfing thanks to that. Data centers as well. I mean, data centers are where, like, I mean, at Google, for instance, obviously they're serving Google, but there's also now, thanks to that, and to have built such amazing data centers, we can train these models. Software is an important one. I think if I look at the state of how I had to implement things to implement my ideas, how I discarded ideas, because they were too hard to implement. Yeah, clearly the times have changed and thankfully we are in a much better software position as well. And then, I mean, obviously there's research that happens at scale and more people enter the field. That's great to see, but it's almost enabled by these other things. And last but not least is also data, right? Curating data sets, labeling data sets, these benchmarks we think about maybe we'll want to have all the benchmarks in one system.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah, I mean, engineering, there's also kind of a historical, it might be a bit random, because if you think of the history of how especially deep learning and neural networks took off, feels like a bit random because GPUs happen to be there at the right time for a different purpose, which was to play video games. So even the engineering that goes into the hardware and it might have a time, like the time frame might be very different. I mean, the GPUs were evolved throughout many years where we didn't even we're looking at that, right? So even at that level, right, that revolution, so to speak, the ripples are like, we'll see when they stop, right? But in terms of thinking of why is this happening, right? There's, I think that when I try to categorize it in sort of things that might not be so obvious, I mean, clearly there's a hardware.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Need to then help them optimize that kind of research that might actually produce amazing change, perhaps is not as short term as some of these advancements or perhaps it's a different timescale. But the people and the diversity of the field is quite critical that we maintain it. And at times, especially mixed a bit with hype or other things, it's a bit tricky to be observing maybe too much of the same thinking across the board. But the humans definitely are critical. And I can think of quite a few personal examples where also someone told me something that had a huge effect on to some idea. And then that's why I'm saying at least in terms of years, probably some things do happen.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  39. But it is we have the particle accelerator here, so to speak, in physics. So we need to use it, we need to answer the questions that we should be answering right now for the scientific progress. But then at the same time, I look at many advances, including attention, which was discovered in Montreal initially because of lack of compute, right? So we were working on sequence to sequence with my friends over at Google Brain at the time. And we were using, I think, eight GPUs, which was somehow a lot at the time. And then I think Montreal was a bit more limited in the scale. But then they discovered this content-based attention concept that then has obviously triggered things like transformer. Not everything obviously starts transformer. There's always a history that is important to recognize because then you can make sure that then those who might feel now, well, we don't have so much compute.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  40. You become a bit more of a mentor to a large group of people, be it a project or the deep learning team or something, or even in the community when you interact with people in conferences and so on, you're identifying quickly, right? Some things that are explorative or exploitative. And it's tempting to try to guide people, obviously. I mean, that's what makes like our experience, we bring it and we try to shape things sometimes wrongly. And there's many times that I've been wrong in the past. That's great. But it would be wrong to dismiss any sort of the research styles that I'm observing. And I often get asked, well, you're in industry, right? So we do have access to large compute scale and so on. So there's certain kinds of research I almost feel like we need to do responsibly and so on.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  41. What idea works? I care about cracking protein folding. And at least these two kind of seem opposite sides. We need both. And we've clearly had both historically, and that made certain things happen earlier or later. So definitely humans involved in all of this endeavor have had, I would say, years of change or of ordering how things have happened, which breakthroughs came before, which other breakthroughs, and so on. So certainly that does happen. And so one other maybe one other axis of distinction is what I called, and this is most commonly used in reinforcement learning, is the exploration, exploitation trade-off as well. It's not exactly what I meant, although quite related. So when you start trying to help others.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Yeah, I mean, I do believe humans matter a lot. At the very least, at the Timescale of years on when things are happening and what's the sequencing of it, right? So you get to interact with people that, I mean, you mentioned this. Some people really want some idea to work and they'll persist. And then some other people might be more practical. Like, I don't care.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Yes. And I mean, hierarchies are very, I mean, it's a very nice word that sounds appealing. There's lots of work adding hierarchy to the memories. In practice, it does seem like we keep coming back to the main formula or main architecture. That sometimes tells us something. There's such a sentence that a friend of mine told me whether it wants to work or not. So Transformer was clearly an idea that wanted to work. And then I think there's some principles we believe will be needed, but finding the exact details, details matter so much, right? That's going to be tricky.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  44. I think it goes well beyond the current capabilities. So the question is how do we benchmark this and then how do we change the structure of the architectures? I think there's ideas on both sides, but we'll have to see empirically, right, obviously what ends up working in the...

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Networks that were more recently biased, which obviously works in some tasks, but it has major flaws. Transformer itself has flaws. And I think the main one, the main challenge is these prompts that we just were talking about, they can be a thousand words long. But if I'm teaching you StarGraft, I mean, I'll have to show you videos. I have to point you to whole Wikipedia articles about the game. We'll have to interact probably as you play, you'll ask me questions. The context requires for us to achieve me being a good teacher to you on the game as you would want to do it with a model.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  46. But it has the query as well. Look, if I'm thinking deeply about text, I need to go back to look at all of the text, attend over it. But it's not just attention. What is guiding the attention? And that was the key insight from an earlier paper. It's not how far away is it. I mean, how far away is it is important? What did I just write about? That's critical. But what you wrote about 10 pages ago might also be critical. So you're looking not positionally, but content-wise, right? Transformers have this beautiful way to query for certain content and pull it out in a compressed way so that you can make a more informed decision. I mean, that's one way to explain transformers. But I think it's a very powerful inductive bias. There might be some details that might change over time, but I think that is what makes transformers so much more powerful than the recurrent.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Intuition from how we do it as a human that is very nicely mimicked and replicated structurally speaking in the transformer, which is this idea of You're looking for something, right? So you're sort of when you just read a piece of text, now you're thinking what comes next, you might want to re-look at the text or look it from scratch. I mean, literally is because there's no recurrence. You're just thinking, what comes next? And it's almost hypothesis driven, right? So if I'm thinking the next word that I'll write is cat or dog, okay. The way the transformer works almost philosophically is it has these two hypotheses. Is it going to be cat or is it going to be dog? And then it says, okay, if it's cat, I'm going to look for certain words, not necessarily cat, although cat is an obvious word you would look in the past to see whether it makes more sense to output cat or dog. And then it does some very deep computation over the words and beyond, right? So it combines the words and

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Yeah, so a distinction between transformers and LSDMs, which were what came before. And, you know, there was a transitional period where you could use both. In fact, when we talked about AlphaStar, we used transformers and LSTMs. So it was still the beginning of Transformers. They were very powerful. But LSTMs were still very also very powerful sequence models. So the power of the transformer is that it has built in what we call an inductive bias of attention that makes the model, when you think of a sequence of integers, right? Like we discussed this before, right? This is a sequence of words. When you have to do very hard tasks over these words, this could be, we're going to translate a whole paragraph, or we're going to predict the next paragraph given 10 paragraphs before. There's some

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Is something there that will stick. And I think these advance in architectures, in kind of how neural networks are architecture to do what they do. It's been hard to find one that has been so stable and relatively has changed very little since it was invented five or so years ago. So that is a surprising that keeps recurring to other projects.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source

  50. With the team, I mean, it is tricky that also happened to happen partly during a pandemic and so on. So I think my meta learning from all this is the teams are critical to the success. And then if now going to the machine learning, the part that's surprising is, so we like architectures like neural networks. And I would say this was a very rapidly evolving field until the transformer came. So attention might indeed be all unit, which is the title. Also good title, although in hindsight it's good. I don't think at the time I thought this is a great title for a paper. But that architecture is proving that the dream of modeling sequences of any bytes.

    2022-07-26 · Lex Fridman Podcast · #306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence · IDENTIFIED FROM THE TRANSCRIPT · source