YouSaid · the spoken record

Ishan Misra

lines on the record
175
first
2021-07-31
most recent
2021-07-31
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. See that happen on a live Tesla, that basically just proves everyone wrong, I would say, in a way. And that's just working really well. I think there were also a lot of advancements in camera technology. Now there were like, I know at CMU when I was there, there was a particular kind of camera that had been developed that was really good at basically low visibility settings, so like lots of snow and lots of rain. It could actually still have a very reasonable visibility. And I think there are lots of these kinds of innovations that will happen on the sensor side itself, which is actually going to make this very easy in the future. And so maybe that's actually why I'm more optimistic about vision-based autonomous driving. I'm going to call it self-supervised driving.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Of course, looking so on that screen, it basically shows all the detections and everything that the car is doing as you're driving by. And that's super distracting for me as a person. Because all I keep looking at is like the bounding boxes and the cars it's tracking and it's really impressive. Like especially when it's raining and it's able to do that, that was the most impressive part for me. It's actually able to get through rain and do that. And one of the reasons why like a lot of us believed and I would put myself in that category is LIDAR based sort of technology for autonomous driving was the key driver, right? So Weymor was using it for the longest time. And Tesla then decided to go this completely other route that oh, we are not going to even use LIDAR. So their initial system I think was camera and radar based and now they're actually moving to a completely like vision based system. And so that was just like it sounded completely crazy. LIDAR is very useful in cases where you have low visibility. Of course, it comes with its own set of complications. But now to

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Okay, so first disclaimer, I'm not an expert in autonomous driving. So let me put it out there. I would say at least five to ten years. This would be my guess from now. I'm actually very impressed when I sat in a friend's Tesla recently and

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Okay, so what does solving autonomous driving mean? Does it mean solving it in the US? Does it mean solving it in India? Because I can tell you the very different types of driving happening.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  5. I think that really resonates with me. It's super smart to do it that way because, I mean, the thing is with any kind of practical system like autonomous driving, there are those edge cases are the things that are actually the problem, right? I mean, highway driving or like freeway driving has basically been, like, there has been a lot of success in that particular part of autonomous driving for a long time, I would say, like since the 80s or something. Now, the point is all these failure cases are the sort of reason why autonomous driving hasn't become like super, super mainstream available in every possible car right now. And so basically by really scaling this problem out, by really trying to get all of these edge cases out as quickly as possible, and then just using those to improve your model, that's super smart.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  6. So basically by forming these good predictive models, you are, I mean, these are kind of self supervised models, right? Prediction models are basically being trained just by looking at what's going to happen next and asking them to predict what's going to happen next. So I would say this is really one use of self-supervised learning. It's a predictive model, and you're learning a predictive model basically just by looking at what data you have.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So, I think so. I think for self supervised learning to be used in autonomous driving, there are lots of opportunities. Just like pure consistency in predictions is one way, right? So because you have this nice sequence of data that is coming in, a video stream of it, associated, of course, with the actions that say the car took. You can form a very nice predictive model of what's happening. So, for example, all the way one way possibly in which how they're figuring out what data to get labeled is basically through prediction uncertainty. So you predict that the car was going to turn right. So this was the action that was going to happen, say in the shadow mode. And now the driver turned left. And this is a really big surprise.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Here is like one simple way that you can use active learning. For example, you have your self supervised model, which is very good at predicting similarities and dissimilarities between things. And so if you label a picture as basically, say a banana, now you know that all the images that are very similar to this image are also likely to contain bananas. So, probably when you want to understand what else is a banana, you're not going to use these other images. You're actually going to use an image that is not completely dissimilar, but somewhere in between, which is not super similar to this image, but not super dissimilar either. And that's going to tell you a lot more about what this concept of a banana is

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  9. And I think that's the sort of key challenge. Now, self supervised learning by itself, like selecting data for it and so on, that's actually really useful. But I think that's a very narrow view of looking at active learning. If you look at it more broadly, it is basically about if the model has a knowledge about n concepts and it is weak basically about certain things. So it needs to ask questions either to discover new concepts or to basically increase its knowledge about these n concepts. So at that level it's a very powerful technique. I actually do think it's going to be really useful. Even in simple things such as data labeling, it's super useful.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So, there is some amount of understanding or knowledge that basically keeps getting built when you're doing active learning. So I think active learning by itself is really good. And the main thing we need to figure out is basically how do we come up with a technique to first model what the model knows And also model what the model does not know. I think that's the sort of beauty of it because when you know that there are certain things that you don't know anything about, asking a question about those concepts is actually going to bring you the most value.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  11. I actually really like active learning. So back in the day, we did this largely ignored CVPR paper called Learning by Asking Questions. So the idea was basically you would train an agent that would ask a question about the image, it would get an answer. And basically then it would update itself. It would see the next image, it would decide what's the next hardest question that I can ask to learn the most. And the idea was basically because it was being smart about the kinds of questions it was asking, it would learn in fewer samples. be more efficient at using data. And we did find to some extent that it was actually better than randomly asking questions. Kind of weird thing about active learning is it's also a chicken and egg problem because when you look at an image, to ask a good question about the image, you need to understand something about the image. You can't ask a completely arbitrarily random question. It may not even apply to that particular image

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Sound kind of does that too, right? So if you have the same sound, so that's the same context across different videos, you're very likely to be observing the same kind of concept. So that's the kind of reason why it figures out the guitar thing. It observes the same sound across multiple different videos and it figures out maybe this is the common factor that's actually doing it.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  13. So, in this case, for example, one thing is when you observe someone cutting something and you don't have any sort of sound there, whether it's an apple or whether it's an onion, it's very hard to figure that out. But if you hear someone cutting it, it's very easy to figure it out. Because apples and onions make a very different kind of characteristic sound when they're cutting. So, you really figure this out based on audio. It's much easier. So, your life will become much easier when you have access to different kinds of modalities. And the other thing is, so I like to relate it in this way, it may be like completely wrong, but the distributional hypothesis in NLP, right? Where context basically gives kind of meaning to that word.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Would say a little tending more towards the second part, so most of it can be sort of figured out with one modality, but having an extra modality always helps you

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Exactly. So, you don't need a lot of. Videos of humans doing actions annotated, you can just use a few of them to basically get

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  16. So, this is all what it had discovered. We never pointed out that this is a guitar and this is the kind of sound it produces. It can actually naturally figure that out because it's seen so many correlations of this sound coming with this kind of an object that it basically learns to associate this sound with this kind of an object.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  17. And the same thing, for example, for certain people's voices, like famous celebrities' voices, it can actually figure out where their mouth is. So it can actually distinguish different people's voices, for example, a little bit as well.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And basically, of a person just drumming a guitar, but of course, there is no audio in this. And now you give it the sound of a guitar. And you basically try to visualize where the network thinks the sound is coming from. It can kind of basically draw like when you visualize it, you can see that it's basically focusing on the guitar.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  19. So there is this data set called kinetics, for example, which has like 400 different types of human actions. So people jumping, people doing different kinds of sports or different types of swimming, so like different strokes in swimming, golf and so on. So there are just different types of actions right there. And the point is this kind of video network that you learn in a self-supervised way can be used very easily to kind of recognize these different types of actions. It can also be used for recognizing different types of objects. And what we did is we tried to visualize whether the network can figure out where the sound is coming from. So basically give it a video.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  20. And so basically that's it. It's just a simple application of contrastive learning. The main sort of finding from this work for us was basically that you can actually learn very, very powerful feature representations, very, very powerful video representations. So you can learn the sort of video network that we ended up learning can actually be used for downstream, for example, recognizing human actions or recognizing different types of sounds, for example. This was sort of the key finding.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So, the typical use case is basically, for example, someone playing a musical instrument. So guitars have a particular kind of sound and so on. So, because a lot of these things are correlated, the idea in multimodal learning is to take these two kinds of modalities, video and audio, and learn a common embedding space, common feature space where both of these related modalities can basically be close together. And again, you use contrasted learning for this. So in contrastive learning, basically the video and the corresponding audio are positives. And you can take any other video or any other audio and that becomes a negative.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Like this podcast, exactly. So, what we did in this work was basically trained two different neural networks, one on the video signal, one on the audio signal. And what we wanted is basically the features that we get from both of these neural networks should be similar. So it should basically be able to produce the same kinds of features from the video and the same kinds of features from the audio. Now, why is this useful? Well, for a lot of these objects that we have, there is a characteristic sound, right? So trains, when they go by, they make a particular kind of sound, both make a particular kind of sound, people when they're jumping around will shout, whatever. Bananas don't make a sound, so where you can't learn anything about bananas

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  23. And you want to use both of these signals, the audio signal and the video signal to learn a good representation for video. Good representation for audio

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  24. This paper, it actually came as a little bit of a shock to me at how well it worked, so I can describe what the problem setup was. So it's been used in the past by lots of folks, like, for example, Andrew Owens from MIT, Alyosha from Berkeley, Andrew Zisterman from Oxford. So a lot of these people have been sort of showing results in this. Of course, I was aware of this result, but I wasn't really sure how well it would work in practice for other sort of downstream tasks. The results kept getting better and I wasn't sure if a lot of our insights from self-supervised learning would translate into this multi- When you have multiple modalities, the particular modalities that we worked on in this work were audio and video. So, the idea was basically if you have a video, you have its corresponding audio track.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Really tricky to come up with a good small scale setup where a lot of your empirical observations will really translate to the other setup. So it's been really challenging. I've been trying to do that for a little bit as well because it does take time to train stuff on image net. It does take time to train on more images. But pretty much every time I've tried to do that, it's been unsuccessful because all the observations I draw from my set of experiments on a smaller data set don't translate into ImageNet or don't translate it to another sort of data set. So it's been hard for us to figure this one out, but it's an important problem.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  26. So we're trying to push something towards that. I think there are a few setups out there, but nothing super standard on the smaller scale. ImageNet in itself is actually pretty big, also, so that is not something which is feasible for a lot of people. But we are trying to push up with smaller sort of use cases. The thing is at a smaller scale, a lot of the observations or a lot of the algorithms that work don't necessarily translate into the medium or the larger scale.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  27. That as much as possible. And there was a paper like we did in 2019 just about benchmarking. And so whistle basically builds upon a lot of this kind of work that we did about benchmarking. And then every time we tried to, we come up with a self-supervised learning method, a lot of us try to push that into Vessel as well, just so that it basically is the central piece where all a lot of these methods can reside.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Whistle basically was born out of a lot of us at Facebook doing the self supervised learning research. So it's a common framework in which we have a lot of self-supervised learning methods implemented for Vision. It has in itself a benchmark of tasks that you can evaluate the self-supervised representations on. So the use case for it is basically for anyone who's either trying to evaluate their self-supervised model or train their self-supervised model or a researcher who's trying to build a new self-supervised technique. So it's basically supposed to be all of these things. So as a researcher before whistle, for example, or like when we started doing this work fairly seriously at Facebook, it was very hard for us to go and implement every self-supervised learning model, test it out in a sort of consistent manner. The experimental setup was very different across different groups, even when someone said that they were reporting image net accuracy. It could mean lots of different things. So with whistle, we tried to really sort of standardize.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  29. But in general, I think that's sort of the I guess it's outside my scope as well. The main thing is minimize the amount of synchronization steps that you have.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Think more and more there is this lot of systems are designed basically taking into account the machine learning needs, right? So because whenever you're doing this kind of distributed training, there is a lot of intercommunication between nodes. So like gradients or the model parameters are being passed. So you really want to minimize communication costs when you really want to scale these models up. You want basically to be able to do as much as limited amount of communication as possible. So currently like a dominant paradigm is synchronized sort of training. So essentially after every sort of gradient step all you basically have like a synchronization step between all the sort of compute chips that you're going on with. I think asynchronous training was popular, but it doesn't seem to perform as well.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  31. I mean, so the model was like a billion parameters And it was stained in the billion images. So basically the same number of parameters as the number of images. And it took a while. I don't remember the exact number. It's in the paper, but it took a while.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  32. I think data and data augmentation, the algorithm being used for the self supervised training matters a lot more than the particular kind of architecture. With different types of architecture, you will get different properties in the resulting sort of representation. But really, I mean, the secret sauce is in the data augmentation and the algorithm being used to train them. The architectures, I mean, at this point, a lot of them perform very similarly depending on the particular task that you care about. They have certain advantages and disadvantages.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Also, just in terms of pure hardware, they fit very well on GPU memory So, they can be really powerful neural network architectures with lots of parameters, lots of flops, but also because they're efficient in terms of the amount of memory that they're using, you can actually fit a lot of these on, you can fit a very large model on a single GPU, for example.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  34. And so designing one of the key findings from this paper was basically that you need to design network families or neural network architectures that are actually very efficient in the memory space as well, not just in terms of pure flops. So RegNet is basically a network architecture family that came out of this paper that is particularly good at both flops and the sort of memory required for it. And of course, it builds upon an earlier work, like ResNet being the sort of more popular inspiration for it, where you have residual connections. One of the things in this work is basically they also use squeeze excitation blocks. So it's a lot of nice sort of technical innovation in all of this. From prior work and a lot of the ingenuity of these particular authors in how to combine these multiple building blocks. But the key constraint was optimize for both flops and memory when you're

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  35. So, one of the sort of key takeaways from this paper, which the authors, like whenever you hear them present this work, they keep saying is Lot of neural networks are characterized in terms of flops, right? Flops basically being the floating point operations. And people really love to use flops to say this model is really computationally heavy or our model is computationally cheap and so on. Now it turns out that flops are really not a good indicator of how well a particular network is, like how efficient it is really. And what a better indicator is the activation or the memory that is being used by this particular model.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  36. For SER and for SWAB, we were using convolutional networks. But recently, in a work called Dyno, we've basically started using transformers for vision. Both seem to work really well.Nets and transformers, and depending on what you want to do, you might choose to use a particular formulation. So for Sier, it was a const net. It was particularly a regNet model, which was also work from Facebook. RegNets are really good when it comes to compute versus accuracy. So because it was a very efficient model, compute and memory-wise efficient, and basically it worked really well in terms of scaling. So we used a very large REGNET model and trained it on a billion images.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  37. So, these are not going to be pictures of a zoomed in table or a zoomed in wall. So it's not really completely uncurated because people do have their photographers' bias where they do want to keep things towards the center a little bit or really have nice looking things and so on in the picture. So that's the kind of bias that it typically exists in this data set. And also the user base, right? You're not going to get lots of pictures from different parts of the world because there are certain parts of the world where people may not actually be uploading a lot of pictures to the internet or may not even have access to a lot of internet.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Change respect. So it's not like uncurated can basically one way of imagining uncurated is basically you have cameras that can take pictures at random viewpoints. When people upload pictures to the internet, they are typically going to care about the framing of it. They're not going to upload, say, the picture of a zoomed in wall, for example.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  39. How well would this model do? So, is self supervised learning really overfit to image net, or can it actually work in the wild? And it was also out of curiosity what kind of things will this model learn? Will it actually be able to still figure out different types of objects and so on? Would there be particular kinds of tasks it would actually do better than an image net trained model? And so for CER, one of our main findings was that we can actually train very large models in a completely self-supervised way on lots of internet images without really necessarily filtering them out, which was in itself a good thing because it's a fairly simple process, right? So you get images which are uploaded and you basically can immediately use them to train a model in an unsupervised way. You don't really need to sit and filter them out. These images can be cartoons, these can be memes, these can be actual pictures uploaded by people. And you don't really care about what these images are. You don't even care about what concepts they contain. So this was a very sort of simple setup.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  40. These were basically about a billion or so images. And for context image net, the imagenet version that we used was one million images earlier. So this is basically going like three orders of magnitude more. The idea was basically to see if we can train a very large convolutional model in a self-supervised way on this uncurated but really large set of images.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  41. So for SER, our main goal was to fold. One, basically to move away from ImageNet for training. So the images that we used were uncurated images. Now, there's a lot of debate whether they're actually curated or not, but I'll talk about that later. But the idea was basically these are going to be random internet images that we are not going to filter out based on particular categories. So, we did not say that images that belong to dogs and cats should be the only images that come in this data set. Banana. And basically, other images should be thrown out. So we didn't do any of that. So these are random internet images. And of course, it also goes back to the problem of scale that you talked about.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Yes, I mean, it's basically because our data augmentations that we designed, all data augmentations that we designed for self-supervised learning and vision are kind of overfit to image net.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  43. So, these images are very curated, and what that means is these images, of course, belong to a certain set of noun concepts. And also ImageNet has this bias that all images contain an object which is very big and it's typically in the center. So when you're talking about a dog, it's a well-framed dog, it's towards the center of the image. So, a lot of the data augmentation, a lot of the hidden assumptions in self-supervised learning actually really exploit this bias of ImageNet. And so, I mean, a lot of my work, a lot of work from other people, always uses ImageNet sort of as the benchmark to show the success of self-supervised lear

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  44. So I'll first go to Swav because Swav is actually one of the key components for SEER. So Swab was when we used Swab, it was demonstrated on ImageNet. So typically self supervised methods, the way we sort of operate is in the research community, we kind of cheat. So we take image net, which of course I talked about as having lots of labels. And then we throw away the labels, like throw away all the hard work that went behind basically the labeling process, and we pretend that it is unsupervised. But the problem here is that we have when we collected these images, the image net data set has a particular distribution of concepts, right?

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Fixed game makes things simpler. Our clustering is not really hard clustering, it's soft clustering. So basically you can be point two to cluster number one and point eight. So essentially, even though we have 3,000 clusters, we can actually represent a lot of clusters.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  46. So, Swav basically ensures that at each point all these 3000 clusters are being used in the clustering process. And that's it. Basically, just figure out how to do this online. And again, basically, just make sure that two crops from the same image belong to the same cluster and others don't.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  47. So, the idea basically is that when you have n samples, we assume that we have access to, like there are always k clusters in a dataset. K is a fixed number. So for example, k is 3000. And so if you have any, when you look at any sort of small number of examples, all of them must belong to one of these k clusters. And we impose this equipartition constraint. What this means is that basically your entire set of n samples should be equally partitioned into k clusters. So all your k clusters are basically equal, they have equal contribution to these end samples. And this ensures that we never collapse. So collapse can be viewed as a way in which all samples belong to one cluster. So all this, if all features become the same, then you have basically just one mega cluster. You don't even have like 10 clusters or 3,000 clusters.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Suave is basically just a simple way of doing this online. So as you're going through the data, you're actually computing these clusters online. And so, of course, there is like a lot of tricks involved in how to do this in a robust manner without collapsing. But this is the sort of key idea to it.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Suave basically is clustering based technique, which is for, again, the same thing for self supervised learning in vision, where we have two crops and the idea basically is that you want the features from these two crops of an image to lie in the same cluster and basically crops that are coming from different images to be in different clusters. Now, typically in a sort of, if you were to do this clustering, you would perform clustering offline. What that means is you would, if you have a data set of n examples, you would run over all of these n examples, get features for them, perform clustering. So basically get some clusters, and then repeat the process again. So this is offline basically because I need to do one pass through the data to compute its clusters.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  50. A banana. Everything is a banana. Everything is a cat. Everything is a car And so, all we need to do is basically come up with ways to prevent collapse, contrast similar thing is one way of doing it. And then, for example, like clustering or self-distillation or other ways of doing it, we also had a recent paper where we used decorrelation between two sets of features to prevent collapse. So that's inspired a little bit by Horace Barlow's neuroscience principles.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source