YouSaid · the spoken record

Ilya Sutskever

lines on the record
106
first
2020-05-08
most recent
2020-05-08
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Yeah, but in some sense, do you think in every sense, do you think there's a. It's all just a matter of coming up with a better mechanism of forgetting the useless stuff and remembering the useful stuff. Because right now, I mean, there's not been mechanisms that do remember really long-term information.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Yeah, I'm there with you. I've stopped betting against neural networks at this point because they continue to surprise us. About long term memory? Can you all networks have long term memory or something like Knowledge basis. So being able to aggregate important information over long periods of time that would then serve as useful sort of representations of state that you can make decisions by to have a long-term context based on what you make in the decision.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Do you think you kind of mentioned that the neural networks are good at finding small circuits or large circuits? Do you think then the matter of finding small programs is just the data? So not the size or character, the type of data, sort of ask giving it programs.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Being trainable means starting from scratch, knowing nothing, you can actually pretty quickly converge towards knowing a lot.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  5. So that's that the large circuit might be one that's helpful for the generalization. Do you see it important to be able to try to learn something like programs?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Right, that takes us to the. To one of the brilliant sort of ways you've described neural networks, which is you've referred to neural networks as the search for small circuits and maybe general intelligence as the search for small programs. Which I found as a metaphor very compelling. Can you elaborate on that difference?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Do you think the kind of stuff we've seen neural networks do is a So it's not a fundamentally different process. Again, this is stuff we don't, nobody knows the answer to.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  8. All right. So Do you think the architecture Will allow neural networks to reason, will look similar to the neural network architectures we have today

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Stepwise consideration of possibilities and sort of building on top of those possibilities in a sequential manner until you arrive at some insight. Sort of, yeah, I guess plain goes kind of like that. And when you have a single neural network doing that without search, that's kind of like that. So there's an existent proof in a particular constrained environment that a process akin to what many people call reasoning exists, but more general kind of reasoning so off the board.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Push back and disagree a little bit, we all agree that go is reasoning. I think I agree. I don't think it's a trivial. So, obviously, reasoning like intelligence is a loose gray area term a little bit. Maybe you disagree with that. Yes, I think it has some of the same elements of reasoning. Reasoning is almost like akin to search, right? There's a sequential element of

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  11. If we can't find back propagation in the brain. Well So, I guess your answer to that is back propagation is pretty damn useful. So, why are we complaining?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  12. He was saying that once we discover the mechanism of learning in the brain or any aspects of that mechanism, we should also try to implement that in your network.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Jeff Hinton suggested we need to throw back propagation. We already kind of talked about this a little bit, but he suggested that we need to throw away back propagation and start over. I mean, of course, some of that is a little bit... Wit and humor, but what do you think? What could be an alternative method of training neural networks?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So, okay, so at the very basic level to me, it is the most surprising thing that. Networks don't overfit every time very quickly. Before ever being able to learn anything. The huge number of parameters.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  15. You may be sort of on the topic of the science to deep learning. Talk about one of the recent papers that you've released, the Deep Double Descent, where bigger models and more data hurt. I think it's a really interesting paper. Can you describe the main idea?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  16. What about Vladimir Vapnik really insist on is taking Amnest and trying to learn from very few examples? So being able to learn more efficiently, do you think that there will be breakthroughs in that space that would may not need the huge compute?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Asking you all these questions that nobody knows the answer to, but you're one of the smartest people I know, so we're going to keep asking. So let's imagine all the breakthroughs that happen in the next 30 years in deep learning. Do you think most of those breakthroughs can be done by one person with one computer? Sort of in the space of breakthroughs, do you think compute will compute and large efforts will be necessary?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  18. In between, but it'd be nice if machine learning somehow helped us discover the unification of the two as opposed to serve the in between. You're right. You're kind of trying to juggle both. So do you think there's still beautiful and mysterious properties in neural networks that are yet to be discovered?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Geometric mean of biology and physics. I think I'm going to need a few hours to wrap my head around that. Because just to find the geometric represents.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Do you have insights of what so you just said empirical evidence is most of your Sort of empirical evidence kind of convinces you. It's like evolution is empirical. It shows you that, look, this evolutionary process seems to be a good way to design. Organisms that survive in their environment, but it doesn't really get you to the insights of how the whole thing works.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Have you built up an intuition of why are there little bits and pieces of intuitions of insights of why this whole thing works

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  22. So forgive the romanticized question, but looking back to you, what is the most beautiful or surprising idea in deep learning, or AI in general you've come across?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Yeah. No, I think it all boils down to the reason people click like on stuff on the internet, which is like it makes them laugh. So it's like humor or wit or insight.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  24. The ones, okay, so I'm a fan of monogamy, so I like the idea of marrying somebody being with them for several decades. So I believe in the fact that, yes, it's possible to have somebody continuously giving you pleasurable, interesting, witty new ideas. Friends, yeah, I think so. They continue to surprise you. Surprise, it's you know that injection of randomness seems to be a nice source of Yeah, continued Inspiration, like the wit, the humor. I think, yeah. that that would be a very subjective test, but I think if you have enough humans in the room

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Yeah, to me, so my definition is if a system looked at an image and then a system looked at a piece of text and then told me something about that. And I was really impressed.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Yeah, so Chomsky would say it starts at language. So vision is just a little example of the kind of structure and fundamental hierarchy of ideas that's already represented in our brain somehow that's represented through language. Where does vision stop and language begin That's a really interesting Question So, one possibility is that it's impossible to achieve really deep understanding in either images or language without basically using the same kind of system. So you're going to get the other for free.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Okay, the non interesting dumb answer to that is there's a benchmark and there's a human level performance on that benchmark and how is the effort required to reach the human level.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Problem do you think is harder so people like Noah Chomsky believe that language is fundamental to everything? So it underlies everything. Do you think language understanding is harder than visual scene understanding or vice versa?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Think it's a fundamentally different problem, or is it just a more difficult? It's a generalization of the problem of understanding.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Think action is fundamentally different. So, yet, what is interesting about what is unique about policy of learning to act?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Reinforcement learning has some aspects of language and vision combined almost. There's elements of a long-term memory that you should be utilizing and there's elements of a really rich sensory space. So it seems like the union of the two or something like that.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  32. So incredibly you've contributed some of the biggest recent ideas in AI in computer vision, language, natural language processing, reinforcement learning. Of everything in between, maybe not gans. There may not be a topic you haven't touched, and of course, the fundamental science of deep learning. Is the difference to you between vision, language, and as in reinforcement learning action as learning problems? And what are the commonalities? Do you see them as all interconnected? Are they fundamentally different domains that require different approaches?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And they represented this kind of where the big pillars of computer vision community kind of the wizards got together. And then all of a sudden there was a shift. It's not enough for the ideas to all be there and the compute to be there. It's for it to convince the cynicism that existed. It's interesting that people just didn't believe for a couple of decades.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Really interesting. So I guess the presence of compute and the presence of supervised data. Allowed the empirical evidence to do the convincing of the majority of the computer science community. So I guess there's a key moment with Jitendra Malik and Alex Alyosha Ephros, who were very skeptical, right? And then there's a Jeffrey Hinton that was the opposite of skeptical. And there was a convincing moment. And I think Amission had served as that moment. That's right.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Could you say that? Sort of, do you think there's a future of building large scale knowledge bases within the neural network? So we're going to pause on that confidence because I want to explore that. But let me zoom back out and ask. Back to the history of ImageNet. Neural networks have been around for many decades, as you mentioned. What do you think were the key ideas that led to their success that image net moment and beyond the success in the past ten years?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Typically thought of for processing sequences, I think it's also possible. What is to you a recurrent neural network? And generally speaking, I guess what is a recurrent neural network? You have a neural network which maintains a high-dimensional hidden state. And then when an observation arrives, it updates its high-dimensional hidden state through its connections in some way. So do you think... That's what expert systems did, right? Symbolic AI, the knowledge-based, growing a knowledge base is maintaining a hidden state, which is its knowledge base and is growing it by sequential processing. Do you think of it more generally in that way? Or is it simply the more constrained form of a hidden state with certain kind of gating units that we think of as today with LSDMs and that?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  37. That seems to be important for the brain in the firing of neurons in the brain neural networks are amazing and they can do anything we'd want them to do right now recurrent neural networks have been superseded by transformers but maybe one day they'll make a comeback maybe they'll be back we'll see Let me, on a small tangent, say, Do you think they'll be back? So, so much of the breakthroughs recently that we'll talk about on natural language processing and language modeling has been with transformers that don't emphasize recurrence. Think recurrence will make a comeback Well, some kind of recurrence, I think very likely. Current neural networks for as their

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  38. You sounded brilliant while saying it, but the timing that's one thing that's missing the temporal dynamics is not captured. I think that's like a fundamental property of the brain is the timing of the signals. What we have recurrent neural networks, but you think of that as this, I mean, that's a very crude, simplified, what's that called? There's a clock, I guess, to recurrent neural networks. This seems like the brain is the continuous version of that. The generalization where all possible timings are possible. And then within those timings this contains some information. You think recurrent neural networks, the recurrence in recurrent neural networks can capture the same kind of phenomena as the timing.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  39. I mean, one thing which may potentially be useful, I think people, neuroscientists have figured out something about the learning rule of the brain. Or I'm talking about spike time independent plasticity, and it would be nice if some people were to study that in simulation. Wait, sorry, spike time independent plasticity? Yeah, that's right. What's that? It's a particular learning rule that uses spike timing to figure out how to determine how to update the synapses. So it's kind of like if a synapse fires into the neuron before the neuron fires, then it strengthens the synapse. And if the synapse fires into the neuron shortly after the neuron fired, then it weakens the synapse. Something along this line, I'm 90% sure it's right. So if I said something wrong here, don't get too angry.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Self place starts to touch on that a little bit in reinforcement learning systems. That's right. Self play and also ideas around exploration where you're trying to take action that surprise a predictor. I'm a big fan of cost functions. I think cost functions are great and they serve us really well and I think that whenever we can do things with cost functions we should Know maybe there is a chance that we will come up with some yet another profound way of looking at things that will involve cost functions in a less central way. But I don't know. I think cost functions are, I Would not better guess against ghost functions. There are other things about the brain that pop into your mind that might be different and interesting for us to consider in designing artificial neural networks. So we talked about spiking a little bit.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Regions to which it will go towards, but I don't think I don't think the cost function analogy is the most useful. So if evolution doesn't, that's really interesting. So if evolution doesn't really have a cost function, like a cost function based on its Something akin to our mathematical conception of a cost function, then do you think cost functions in deep learning are holding us back? So you just kind of mentioned that cost function is a nice first profound idea. Do you think that's a good idea? Do you think it's an idea will go past?

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  42. And again, you have a game. So instead of thinking of a cost function where you want to know that you have an algorithm gradient descent, which will optimize the cost function. And then you can reason about the behavior of your system in terms of what it optimizes. With again, you say, I have a game and I'll reason about the behavior of the system in terms of the equilibrium of the game. It's all about coming up with these mathematical objects that help us reason about the behavior of our system. That's really interesting. Yeah, so Gen is the only one. It's kind of the cost function is emergent from the comparison. I don't know if it has a cost function. I don't know if it's meaningful to talk about the cost function of again. It's kind of like the cost function of biological evolution or the cost function of the economy. Can talk about

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Sorry, let me take a pause. Is supervised learning a difficult concept to come to? I don't know. All concepts are very easy in retrospect. Yeah, that's what it seems trivial now. Because the reason I ask that, and we'll talk about it, is there other things? Is there things that don't necessarily have a cost function, maybe have many cost functions, or maybe have dynamic cost functions, or maybe a totally different kind of architectures? Because we have to think like that in order to arrive at something new, right? So the good examples of things which don't have clear cost functions again.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  44. That's the big idea. The cost function is a way of measuring the performance of the system according to some measure. By the way, that is a big, actually let me think. Is that one, a difficult idea to arrive at and how big of an idea is that? That there's a single cost function.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  45. And that's how they're going to make them work. If you don't simulate the non spike in neural networks in spikes, it's not going to work because the question is why should it work? And that connects to questions around backpropagation and questions around Deep learning. You got this giant neural network. Why should it work at all? Why should the learning rule work at all? Not a self evident question, especially if you let's say if you were just starting in the field and you read the very early papers, you can say, hey, people are saying, let's build neural networks. That's a great idea because the brain is a neural network, so it would be useful to build neural networks. Now, let's figure out how to train them. It should be possible to train them properly, but how? And so the big idea is the cost function.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Number of very important advantages of the brain. Looking at the advantages versus disadvantages is a good way to figure out what is the important difference. So the brain uses spikes, which may or may not be important. That's a really interesting question. Do you think it's important or not? That's one big architectural difference between artificial neural networks. It's hard to tell, but my prior is not very high. And I can say why. You know, there are people who are interested in spike in neural networks. And basically what they figured out is that they need to simulate the non-spike neural networks in spikes.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  47. We're now at a time where deep learning is very successful. So let us squint less and say, let's open our eyes and say, what to you is an interesting difference between the human brain? Now, I know you're probably not an expert, neither in your scientists and your biologist, but loosely speaking, what's the difference between the human brain and artificial neural networks? That's interesting to you for the next decade or two. That's a good question to ask. What is an interesting difference between the neural between the brain and our artificial neural networks? So I feel like today, artificial neural networks, so we all agree that there are certain dimensions in which the human brain vastly outperforms our models. But I also think that there are some ways in which artificial neural networks have

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Deep learning researchers since all the way from Rosenblatt in the 60s. Like if you look at the whole idea of a neural network is directly inspired by the brain, you had people like McCollum and Pitts who were saying, hey, you got these neurons in the brain. And hey, we recently learned about the computer and automata. Can we use some ideas from the computer an automata to design some kind of computational object that's going to be simple, computational and kind of like the brain and they invented the neuron. So they were inspired by it back then. Then you had the convolutional neural network from Fukushima and then later Jan Lakhan who said, hey, if you limit the receptive fields of a neural network, it's going to be especially suitable for images. As it turned out to be true. So there were very small number of examples where analogies to the brain were successful. And I thought, well, probably an artificial neuron is not that different from the brain if it's quint hard enough. So let's just assume it is and roll with it.

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Would be enough compute to get a very convincing result. Then at some point, Alex Krzyszewski wrote these insanely fast UDA kernels for training convolutional neural nets. That was BAM, let's do this, let's get ImageNet, and it's going to be the greatest thing. Was your most of your intuition from empirical results by you and by others? So like just actually demonstrating that a piece of program can train a 10-layer neural network? Or was there some pen and paper or marker and whiteboard thinking, intuition? Because you just connected a 10 layer large neural network to the brain. So you just mentioned the brain. So in your intuition about neural networks, does the human brain come into play as an intuition builder? Definitely. I mean, you got to be precise with these analogies between artificial neural networks in the brain. But there is no question that the brain is a huge source of intuition and inspiration for

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Lots of supervised data, and then it must succeed because we can find the best neural network. And then there's also theory that if you have more data than parameters, you won't overfit. Today we know that actually this theory is very incomplete and you won't overfit even if you have less data than parameters. But definitely, if you have more data than parameters, you want to overfit. So the fact that neural networks were heavily overparametrized wasn't discouraging to you. So you were thinking about the theory that the number of parameters, the fact there's a huge number of parameters is okay. It's going to be okay. I mean, there was some evidence that it was okay-ish, but the theory was most, the theory was that if you had a big data set and a big neural narrative was going to work, the overparameterization just didn't really figure much as a problem. I thought, well, with images, you're just going to add some data augmentation and it's going to be okay. So where was any doubt coming from? The main doubt was, can we train a bigger, will we have enough computer train a big enough neural net with back propagation? Back propagation, I thought, was work. The thing which wasn't clear was whether

    2020-05-08 · Lex Fridman Podcast · #94 – Ilya Sutskever: Deep Learning · IDENTIFIED FROM THE TRANSCRIPT · source