YouSaid · the spoken record

Ishan Misra

lines on the record
175
first
2021-07-31
most recent
2021-07-31
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Exactly. So the main sort of enemy of self-supervised learning, any kind of similarity maximization technique is collapse. So collapse means that you learn the same feature representation for all daemers in the world, which is completely useless.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Neural network, which is the student neural network, and that also produces a feature. And now all you're doing is basically saying that the features produced by the teacher network and the student network should be very similar. That's it. There is no notion of a negative anymore. And that's it. So it's all about similarity maximization between these two features. And so all I need to now do is figure out how to have these two sorts of parallel networks, a student network and a teacher network. And basically researchers have figured out very cheap methods to do this. So you can actually have for free really two types of neural networks. They're kind of related, but they're different enough that you can actually basically have a learning problem set up.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Typically, these methods that are contrastive really require access to lots of negatives, which becomes harder and harder to sort of scale when designing an alerting algorithm. So that's been one of the reasons why non-contrastive methods have become popular and why people think that they're going to be more useful. So a non-contrastive method, for example, like clustering is one non-contrastive method. The idea basically being that you have two of these samples, so the cat and dog or two crops of this image, they belong to the same cluster. And so essentially you're basically doing clustering online when you're learning this network, and which is very different from having access to a lot of negatives explicitly. The other way which has become really popular is something called self-distillation. So the idea basically is that you have a teacher network and a student network. And the teacher network produces a feature. So it takes in the image and it basically the neural network figures out the patterns, gets the feature out.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  4. So, like I said about contrast with blow, I think you have this notion of a positive and a negative. Now, the thing is, this entire learning paradigm really requires access to a lot of negatives to learn a good sort of feature space. The idea is if I tell you, okay, so a cat and a dog are similar and they're very different from a banana. The thing is, this is a fairly simple analogy, right? Because, well, bananas look visually very different from what cats and dogs do. So very quickly, if this is the only source of supervision that I'm giving you, your learning is not going to be like after a point, the neural network is really not going to learn a lot. Because the negative that you're getting is going to be so random. So it can be, oh, a cat and a dog are similar, but they're very different from a Volkswagen beetle. Now, this car looks very different from these animals again. So the thing is in contrast to learning, the quality of the negative sample really matters a lot. And so what has happened is basically that

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  5. People are using that, like, there are actually folks right here in UT Austin, like Philip Grehanbull is a professor at UT Austin. He's been working on video games as a source of supervision. I mean, it's really fun. As a PhD student, you can basically play video games all day. Yeah.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Yes. I mean, there are certain tags which are going to be applicable pretty much to anything. So they're pretty useless for learning. But I mean, certain tags are actually like the Eiffelter, for example, or the Taj Mahal, for example. These tags are very indicative of what's going on. And they are. I mean, they are human supervision.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So, I mean, yes, tagging will help you a lot, it'll actually go a very long way in figuring out what images are related or not. And then the purists would argue that when you're using human tags because these tags are like supervision, is it really, really self-supervised learning now? Because you're using human tags to figure out which images are similar. Has

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Right. And so something because you've given me an arbitrarily large data set, I still need to use data augmentation to take that image, construct these two perturbations of it, and then learn from it. So the thing is our learning paradigm is very primitive right now. Even if you were to give me lots of images, it's still not really useful. A good data augmentation algorithm is actually going to be more useful. So reduce down the amount of data that you give me by 10 times. But if you were to give me a good data augmentation algorithm, that will probably do better than giving me 10 times the size of that data, but me having to rely on a very primitive data augmentation algorithm.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Thing is, like, because our learning algorithms for vision right now really rely on data augmentation, even if you were to give me an infinite source of image data, I still need a good data augmentation.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Now, this actually does take us to maybe a semi-supervised kind of a setting because you do want to understand what is it that you're trying to solve. So currently self-supervised learning kind of operates in the wild, right? So you do the self-supervised learning and the purists and all of us basically say that, okay, this should learn useful representations and they should be useful for any kind of end task, no matter it's like banana recognition or like autonomous driving. It's a tall order. Maybe the first baby step for us should be that okay if you're trying to loop in this data augmentation into the learning process, then we at least need to have some sense of what we're trying to do. Are we trying to distinguish between different types of bananas or are we trying to distinguish between banana and apple? Or are we trying to do all of these things at once? And so some notion of what happens at the end might actually help us do much better at this side.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Exactly. And it will be like for particular domains, you might actually see, like, if, for example, now we're doing medical imaging, there are going to be certain kinds of geometric augmentations which are not really going to be very valid for the human body. So, if you were to actually loop in data augmentation into the learning process, it will actually be much more useful.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  12. This data augmentation process is actually independent of the, like it has no notion of what is present in the image, so it can change this color arbitrarily. It can make it a red banana as well. And now what we're doing is we're telling the neural network that this red banana and so a crop of this image which has the red banana and a crop of this image where I change the color to a purple banana should be the features should be the same. Now bananas aren't red or purple mostly. So really the data augmentation process should take into account what is present in the image and what are the kinds of physical realities that are possible. It shouldn't be completely independent of the image.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I think so. I think so. And in fact, it will be really beneficial for us because a lot of these data augmentations that we use in Vision are very extreme. For example, when you have certain concepts, again, a banana, you take the banana and then basically change the color of the banana, right? So you make it a purple banana.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Course, this is not one, this doesn't feel like super satisfactory because a lot of our human knowledge or our human supervision is actually going into the data augmentation. So although we are calling it self-supervised learning, a lot of the human knowledge is actually being encoded in the data augmentation process. So it's really like we've kind of sneaked away the supervision at the input and we're really designing these nice list of data augmentations that are working very well

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Has been the key for visual self supervisor thing. And they play a fairly fundamental role to it. Now the irony of all of this is that for deep learning purists will say the entire point of deep learning is that you feed in the pixels to the network neural network and it should figure out the patterns on its own. So if it really wants to look at edges, it should look at edges. You shouldn't really go and handcraft these features, right? You shouldn't go tell it that look at edges. So data augmentation should basically be in the same category, right? Why should we tell the network or tell this entire learning paradigm what kinds of data augmentation that we are looking for? We are encoding a very sort of human specific bias there that we know things are like if you change the contrast of the image it should still be an apple or it should still see Apple not banana thank you basically if we change like colors it should still be the same kind of concept

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  16. So, data augmentation is key to self supervised learning. That has the kind of augmentations that we're using. And basically the fact that we're trying to learn these neural networks that are predicting these features from images that are robust under data augmentation

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Mean they don't, the augmentations that we work on aren't that involved, they're not going to be physically realistic versions of lighting. It's not that you're assuming that there's a light source up and then you're moving it to the right, and then what does the thing look like? It's really more about brightness of the image, overall brightness of the image, or overall contrast of the image, and so on.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  18. For vision, it's a lot of these image filtering operations, so, like, blurting the image, you know, all the kind of Instagram filters that you can think of. So arbitrarily make the red super red, make the green super greens, like saturate the image.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Think it's a kind of a chicken and egg problem, right? Because to have amazing data augmentation, you need to understand what the scene is And what we're trying to do data augmentation to learn what is seen is anyway. So it's basically just keeps going on.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  20. So, I mean, people do kind of like occlusion based augmentation as well. So you place in a random box, gray box to sort of mask out a certain part of the image. And the thing is basically you're kind of occluding it. For example, you place it, say, on half of a person's face. So basically saying that something below their nose is occluded, it's grayed out. No,

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So I think for some of them, we need to do it. So, like, babies, for example, pick up objects, move them, put them close to Dirai and whatnot. But for certain other things, actually, we are good at imagining it as well, right? So if you I have never seen, for example, an elephant from the top. I've never basically looked at it from top down But if you showed me a picture of it, I could very well tell you that that's an elephant. So I think some of it we're just like, we naturally build it or transfer it from other objects that we've seen to imagine what it's going to look like.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  22. That's what that means. So the neural network basically takes in the image and then outputs a set of basically a vector of numbers. And that's the feature. And you want this feature for both of these different crops that your computer to be similar. So you want this vector to be identical in its entries, for example.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Should the features from both of these images should belong in the same cluster because they're related, whereas image, like another image should belong to a different cluster. So there's a variety of different ways to basically enforce this particular constraint.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  24. And so it has played a fundamental role for computer vision, for self supervisor learning especially. The way most of the current methods work contrastive or otherwise is by taking an image in the case of images, is by taking an image and then computing basically two perturbations of it. So these can be two different crops of the image with like different types of lighting or different contrast or different colors. So you jitter the colors a little bit and so on. And now the idea is basically because it's the same object or because it's like related concepts in both of these perturbations, you want the features from both of these perturbations to be similar. So now you can use a variety of different ways to enforce this constraint, like these features being similar. You can do this by contrastive learning. So basically both of these things are positives. A third sort of image is negative. You can do this basically by like clustering. For example, you can say that both of these images.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So data augmentation is just a way, like you said, it's basically a way to augment the data. So you have, say, n samples, and what you do is you basically define some kind of transforms for the sample. So you take your, say, image and then you define a transform where you can just increase, say, the colors, like the colors or the brightness of the image or increase or decrease the contrast of the image, for example, or take different crops of it. So data augmentation is just a process to basically perturb the data or augment the data, right?

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Right. So, this is where all the sort of ingenuity or tricks comes in, right? So, for example, you can take the fill in the blank problem, or you can take in the context problem, and what you can say is two words that are in the same context that related, two words that are in different contexts are not related. For images, basically two crops from the same image are related, and whereas a third image is not related at all. For a video, it can be two frames from that video related because they're likely to contain the same sort of concepts in them. Whereas a third frame from a different video is not related. So it basically is, it's a very general term. Contrasted learning has nothing really to do with self-supervised learning. It actually is very popular, for example, like any kind of metric learning or any kind of embedding learning. So it's also used in supervised learning. And the thing is because we are not really using labels to get these positive or negative pairs, it can basically also be used for self-supervisor learning.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Right, exactly. So basically, the idea is that if you were to imagine the embedding as a manifold, a 2D manifold, you would get a hill or a high sort of peak in the energy manifold wherever two things are not related. And basically you would have like a dip where two things are related. So you'd get a dip in the manner.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  28. The idea basically is that you can explain a lot of the contrastive models, GANs, for example, which are like generative adversarial networks. A lot of these modern learning methods, or VIEs, which are variational autoencoders, you can really explain them very nicely in terms of an energy function that they're trying to minimize or maximize. And so by putting this common sort of language for all of these models, what looks very different in machine learning, that oh, VAEs are very different from what GANs are very different from what contrastive models are, you actually get a sense of like, oh, these are actually very, very related. It's just that the way or the mechanism in which they're sort of maximizing or minimizing this energy function is slightly different.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Now, energy based models are one way that Jan sort of explains a lot of these methods. So Jan basically, I think a couple of years, more than that, like when I joined Facebook, Jan used to keep mentioning this word energy-based models. And of course, I had no idea what he was talking about. So then one day I caught him in one of the conference rooms and I'm like, can you please tell me what this is? So then very patiently he sat down with a marker and a whiteboard. His idea basically is that rather than talking about probability distributions, you can talk about energies of model. So a model that's trying to minimize certain energies in certain space, or they're trying to maximize a certain kind of energy.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  30. And so you take both of these images and you take the image from the cat, the image from the dog, you get a feature from both of them. And now, what you're training the network to do is basically pull both of these features together while pushing them away from the feature of a banana. So, this is the contrastive part. So, you're contrasting against the banana. So there's always this notion of a negative and a positive

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Contrastive learning is sort of a paradigm of learning where the idea is that you are learning this embedding space or so you're learning this sort of vector space of all your concepts and the way you learn that is basically by contrasting so the idea is that you have a sample you have another sample that's related to it so that's called the positive and you have another sample that's not related to it so that's negative so for example let's just take an nlp or in a simple example in computer vision so you have an image of a cat you have an image of a dog and for whatever application that you're doing say you're trying to figure out what a pets are you're saying that these two images are related so an image of a cat and dog are related but now you have another third image of a banana because you don't like that word so now you basically have this banana thank you

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  32. I mean, I think what I'm saying is NLP is not easy. Of course, don't get me wrong, abstract thought expressed in knowledge or knowledge basically expressed in language is really hard to understand, right? I mean, we've been communicating with language for so long, and it is, of course, a very complicated concept. The thing is, at least getting somewhat reasonable, like being able to solve some kind of reasonable tasks with language, I would say slightly easier than it is with computer vision.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Like the way we capture images, lighting can be different. There might be different noise in the sensor. So the thing is you're capturing a physical phenomenon and then you're basically going through a very complicated pipeline of image processing and then you're translating that into some kind of like digital signal. Whereas with language, you write it down and you transfer it to a digital signal almost like it's a lossless transfer. Each of these tokens are very, very well defined

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Okay, the other reason is the distributional hypothesis that we talked about for NLP, right? So a word given its context, basically the context actually supplies a lot of meaning to the word. Now, because there are just finite number of words and there is a finite way in which we compose them, of course, the same thing holds for pixels. But in language, there's a lot of structure, right? So I always say whatever, the dash jumped over the fence, for example. There are lots of these sentences that you'll get. And from this, you can actually look at this particular sentence might occur in a lot of different contexts as well. This exact same sentence might occur in a different context. So the sheep jumped over the fence, the cat jumped over the fence, the dog jumped over the fence. So you immediately get a lot of these words, which are because this particular token itself has so much meaning, you get a lot of these tokens or these words which are actually going to have sort of this related meaning across given this context. Whereas for vision, it's much harder. Because just by pure.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  35. And so it's really, really large. And very quickly, the kind of prediction problems that we're setting up are going to be extremely intractable for us. And so the thing is for NLP, it has been really successful because we are very good at predicting doing this distribution over a finite set. And the problem is when this set becomes really large, we are going to become really, really bad at making these predictions and at solving basically this particular set of problems. If you were to do it exactly in the same way as NLP for vision, there is very limited success. The way stuff is working right now is actually not by predicting these masks. It's basically by saying that you take these two crops from the image, you get a feature representation from it. And just saying that these two features, so they're vectors, just saying that the distance between these vectors should be small. And so it's a very different way of learning from the visual.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  36. First thing is language is very structured. So you are going to produce a distribution over a finite vocabulary. English has a finite number of words. It's actually not that large. And you need to produce basically when you're doing this masking thing, all you need to do is basically tell me which one of these like 50,000 words it is. That's it. Now, for vision, let's imagine doing the same thing. Okay, we are basically going to blank out a particular part of the image and we ask the network or this neural network to predict what is present in this missing patch. It's combinatorily large, right? You have 256 pixel values if you're even producing basically a 7 cross 7 or a 14 cross 14 like window of pixels at each of these 169 or each of these 49 locations you have 256 values to predict

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  37. So, like the self supervised learning success was actually for Vision has not much to do with the transformers part, I would say it's actually been independent a little bit. I think it's just that the signal was a little bit different for vision than there was for NLP and probably NLP folks discovered it before. So for vision, the main success has basically been this crops so far, like taking different crops of images. Whereas for NLP, it was this masking

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  38. So I'm going to say computer vision is harder. My reason for this is basically that language, of course, has a big structure to it because we developed it. Whereas vision is something that is common in a lot of animals. Everyone is able to get by a lot of these animals on Earth are actually able to get by without language. And a lot of these animals we also deem to be intelligent. So clearly intelligence does have like a visual component to it. And yes, of course, in the case of humans, it of course also has a linguistic component. But it means that there is something far more fundamental about vision than there is about language. And I'm sorry to anyone who disagrees, but yes, this is what I feel.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Because of resolution, because of other things, it's just not easy always to just figure out by looking at just the neighborhood of pixels what these pixels are. And the same thing happens for language as well.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  40. And so this transformer ideas worked as well. So, basically, looking at all the elements to understand a particular element has been really powerful in vision. The reason is a lot of things when you're looking at them in isolation. So if you look at just a blob of pixels, so Antonio Teralbay at MIT used to have this really famous image, which I looked at when I was a PhD student, where he would basically have a blob of pixels and he would ask you, hey, what is this? And it looked basically like a shoe, or it could look like a TV remote. It could look like anything. And it turns out it was a beer bottle. But I'm not sure. It was one of these three things, but basically he showed you the full picture and then it was very obvious what it was. But the point is just by looking at that particular local window, you couldn't figure it out.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Right, so basically, if you have, say, again, a banana in the image, you're looking at the full image first. So whether it's all the pixels that are of a kitchen, of a dining table and so on, and then you're basically looking at the banana also.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Cross three or a seven cross seven neighborhood. And that's it. Whereas with the transformer, that self attention mainly, the sort of idea is that each element needs to pay attention to each other element.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  43. I mean, the core part of a transformer is something called the self attention model. So it came out of Google. And the idea basically is that if you have n elements, what you're creating is a way for all of these n elements to talk to each other. So the idea basically is that you are paying attention. Each element is paying attention to each of the other element. And basically by doing this, it's really trying to figure out you're basically getting a much better view of the data. So for example, if you have a sentence of like four words. The point is if you get a representation or a feature for this entire sentence, it's constructed in a way such that each word has paid attention to everything else. Now the reason it's like different from say what you would do in a conf net is basically that in the conf net you would only pay attention to a local window. So each word would only pay attention to its next neighbor or like one neighbor after that. And the same thing goes for images. In images you would basically pay attention to pixels in a

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  44. I mean, the idea of masking has been very powerful. It has been used in vision as well for predicting, like you say, the next, if you have in sort of frames and you predict what's going to happen in the next frame. So that's been very powerful. In terms of modeling, like in just terms in terms of architecture, I think you were asked about transformers a while back. That has really become, like it has become super exciting for computer vision now, like in the past, I would say year and a half, it's become really powerful.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  45. And so just the fact that by just scaling up the amount of data that we're training on and using better and more powerful neural network architectures has taken us from that to this is just showing you how maybe poor predictors we are. As humans, how poor we are at predicting how successful particular technique is going to be. So I think I can say something now, but like 10 years from now, I'll look completely stupid basically predicting this.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  46. So it has stood the test of time, right? I mean, so word to vec, the initial sort of NLP technique that was using this to now, for example, like all the BERT and all these big models that we get. Bert and Roberta, for example, all of them are still sort of based on the same principle of masking. It's taken us really far. I mean, you can actually do things like, oh, these two sentences are similar or not, whether this particular sentence follows this other sentence in terms of logic, so entailment. You can do a lot of these things with this, just this masking trick. So, I'm not sure if I can predict how far it can take us because when it first came out, when Word2C was out, I don't think a lot of us would have imagined that this would actually help us do some kind of entailment problems

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Essentially, the algorithm or the representation basically puts together these two concepts together. So it says, okay, dogs are going to be kind of slated to sheep because both of them occur in the same context. Of course, now you can decide depending on your particular application downstream. You can say that dogs are absolutely not related to sheep because, well, I really care about dog food, for example. I'm a dog food person and I really want to give this dog food to this particular animal. So depending on what your downstream application is, of course, this notion of similarity or this notion or this common sense that you've learned may not be applicable. But the point is basically that this just predicting what the blanks are is going to take you really, really far.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  48. I'm of course not the expert in NLP. I kind of follow it a little bit from the sides. The main sort of reason why all of this masking stuff works is I think it's called the distributional hypothesis in NLP. The idea basically being that words that occur in the same context should have similar meaning. So if you have the blank jumped over the blank, it basically whatever is in the first blank is basically an object that can actually jump is going to be something that can jump. So a cat or a dog or I don't know sheep, something, all of these things can basically be in that particular context. And now essentially the idea is that if you have words that are in the same context and you predict them, you're going to learn lots of useful things about how words are related because you're predicting by looking at their context what the word is going to be. So in this particular case, the blank jumped over the fence. So now if it's a sheep, the sheep jumped over the fence, the dog jumped over the fence.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  49. The general problem is hard. And I mean, self supervised earning is not the answer to everything. Of course, it's not. I think if you have machines that are going to communicate with humans at the end of it, you want to understand what the algorithm is doing, right? You want it to be able to produce an output that you can decipher, that you can understand, or it's actually useful for something else, which again is a human. So at some point in this sort of entire loop, a human steps in. And now this human needs to understand what's going on. And at that point, this entire notion of language or semantics really comes in. If the machine just spits out something and if we can't understand it, then it's not really that useful for us. So self-supervisor thing is probably going to be useful for a lot of the things before that part, before the machine really needs to communicate a particular kind of output with a human. Because, I mean, otherwise, how is it going to do that without language?

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  50. I mean, it won't do a great job, but it'll do something. It may actually be like it may find certain things which are not humorous, humorous as well, which is going to be bad for us. But I mean, it'll do a, it won't be random

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source