YouSaid · the spoken record
Ishan Misra
- lines on the record
- 175
- first
- 2021-07-31
- most recent
- 2021-07-31
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Some extent, yes, I'm sure it'll work. I mean, it won't be as bad as randomly guessing. I'm sure it can still predict whether it's humorous or not in some way.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Think it is, yes, and getting to that kind of understanding, it's really out there. So if you ask me to solve just that particular problem, I can do it the supervised learning route. I can always construct a data set and basically predict, oh, is there humor in this or not? And of course I can do it.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Natan, like, aha, this actually makes sense because a pickup truck is not really like, what was I annotating? Was I annotating anything that is mobile? Or was I annotating particular sedans or was I annotating SUVs? What was I doing?”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Make it the cake, then we won't be able to set and annotate everything. That's as simple as it is. That's my very practical view on it. It's just, I mean, in my PhD, I sat down and annotated a bunch of cars for one of my projects. And very quickly, I was just like, it was in a video and I was basically drawing boxes around all these cars. And I think I spent about a week doing all of that and I barely got anything done. And basically this was, I think, my first year of my PhD or like second year of my master's. And then by the end of it, I'm like, okay, this is just hopeless. I can keep doing it. And when I've done that, someone came up to me and they basically told me, oh, this is a pickup truck. This is not a car.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“And I think the thing is, things are similar in very different contexts. So an elephant is similar to, I don't know, another sort of wild animal. Let's just pick, I don't know, lion in a different way because they're both four-legged creatures. They're also land animals. But of course, they're very different in a lot of different ways. So elephants are like herbivores, lions are not. So similarity and particularly dissimilarity also actually helps us understand a lot about things. And so that's actually why I think discrete categorization is very hard. Just like forming this particular category of elephant and a particular category of lion, maybe it's good for just like taxonomy, biological taxonomies. But when it comes to other things which are not as maybe, for example, like grill cheese, right? I have a grill cheese I dip it in tomato and I keep it outside. Now is that still a grill cheese or is that something else?”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah okay, wisdom space is good. I think I do think, right? So similarity does get you very, very far. Is it the answer to everything? I mean, I don't even know what everything is, but it's going to take us really far”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“And I think this kind of translation between experiences only happens because of similarity, because I'm able to relate it to a doorknob. If I related it to a hair dryer, I would probably be stuck still outside, not able to get in.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, lots of, for example, animals, they don't have necessarily a well formed syntactic language, but they're able to go about their day perfectly. The same thing happens for us. We probably look at things and we figure out, oh, this is similar to something else that I've seen before. And then I can probably learn how to use it. I haven't seen all the possible doorknobs in the world But if you show me I was able to get into this particular place fairly easily, I've never seen that particular doorknob. So I, of course, related to all the doorknobs that I've seen, and I know exactly how it's going to open. I have a pretty good”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Think so, yeah. So you don't necessarily need to name everything or assign a name to everything to be able to use it. There are lots of Shakespeare.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, if you were to enumerate all the foods up until, I don't know, whenever the crow nut was about 10 years ago or 15 years ago, then this entire thing called crownut would not exist.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Think it's hopeless in some way. So, the thing is for any particular categorization that you create, if you have a discrete sort of categorization, I can always take the nearest two concepts, or I can take a third concept and I can blend it in and I can create a new category. So, if you were to enumerate n categories, I will always find an n plus 1 category for you. That's not going to be in the n categories. And I can actually create not just n plus 1, I can very easily create far more than n categories. The thing is a lot of things we talk about are actually compositional. So it's really hard for us to come and sit and enumerate all of these out. And they compose in various weird ways, right? You have a croissant and a donut come together to form a cronut.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, the other thing is also that humans are not particularly good labelers, they're not very consistent. For example, what's the difference between a dining table and a table? Is it just the fact that one, like if you just look at a particular table, what makes us say one is dining table and the other is not? Humans are not particularly consistent. They're not like very good sources of supervision for a lot of these kind of edge cases. So it may be also the fact that if we want an algorithm or want a machine to solve a particular task for us, we can maybe just specify the end goal. And like the stuff in between, we really probably should not be specifying because we're not maybe going to confuse it a lot, actually.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“We don't know how exactly to do it But we are, I mean, a lot of us are actually convinced that it's going to be sort of a major thing in machine learning.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“You this is heavy, this is not something that can pour, this is something that cannot pour. This is somewhere that you can sit. This is not somewhere that you can sit.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Right. So I think self supervised learning, the way it's done right now is, I would say, like the first step towards what it probably should end up learning or what it should enable us to do. So the idea for that particular piece was self-supervised learning is going to be a very powerful way to learn common sense about the world or stuff that is really hard to label. For example, is this piece over here heavier than the cup? Now for all these kinds of things, you'll have to sit and label these things. So supervised learning is clearly not going to scale. So what is the thing that's actually going to scale? It's probably going to be an agent that can either actually interact with it or lift it up or observe me doing it. So if I'm basically lifting these things up, it can probably reason about, hey, this is taking him more time to lift up or the velocity is different. Whereas the velocity for this is different, probably this one is heavier. So essentially by observations of the data, you should be able to infer a lot of things about the world without someone explicitly telling.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Blog post was mainly about sort of just telling, I mean, this is really an accepted fact, I would say, for a lot of people now that self-supervised learning is something that is going to play an important role for machine learning algorithms that come in the future and even now.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, for example, if you have a chair and a table, basically these things are going to be closed by versus if you take, again, if you have a zoomed in picture of a chair, if you take in different crops, it's going to be different parts of the chair. So the idea basically is that different crops of the image are related. And so the features or the representations that you get from these different crops should also be related. So this is possibly the most widely used trick these days for self-supervised living in computer vision.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, sequence is possibly the most widely used one in NLP. For vision, the one that is actually used for images, which is very popular these days, is basically taking an image and now taking different crops of that image. So you can basically decide to crop, say, the top left corner and you crop, say, the bottom right corner. And asking a network to basically present it with a choice, saying that, okay, now you have this image, you have this image. Are these the same or not? And so the idea basically is that because different, like in an image,”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Basically, in the 10 second, can you predict what's going to happen? And the idea basically is because the model is predicting something about the data itself. Of course, you didn't need any human to tell you what was happening because the 10 second video was naturally captured. Because the model is predicting what's happening there. It's going to automatically learn something about the structure of the world, how objects move, object permanence, and these kinds of things. So if I have something at the edge of the table, it'll fall down. Things like these which you really don't have to sit and annotate. In a supervised learning setting, I would have to sit and annotate. This is a cup. Now I move this cup. This is still a cup. And now I move this cup. It's still a cup and then it falls down and this is a fallen down cup. So I won't have to annotate all of these things in a self-supervised setting.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Perhaps the most successful applications have been in NLP, not language processing. So the idea basically being that you can train models that can you have a sentence and you mask out certain words. And now these models learn to predict the masked out words. So if you have like the cat jumped over the dog, so you can basically mask out cat and now you are essentially asking the model to predict what was missing, what did I mask out. So the model is going to predict basically a distribution over all the possible words that it knows. probably it has like if it's a well-trained model it has a sort of higher probability density for this word cat for vision i would say the sort of more uh i mean the easier example which is not as widely used these days is basically say for example video prediction so video is again a sequence of things so you can ask the model so if you have a video of say 10 seconds you can feed in the first nine seconds to a model and then ask it hey what happens”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So self supervise the reason it has the term supervised in itself is because you're using the data itself as supervision. So because the data serves as its own source of supervision, it's self-supervised in that way. Now the reason a lot of people, I mean, we did it in that blog post with Jan, but a lot of other people have also argued for using this term self-supervised. So starting from like 94 from Virginia Desa's group at I think UCSD, now she's at UCSD. JTendra Malik has said this a bunch of times as well. So you have supervised and then unsupervised basically means everything which is not supervised. But that includes stuff like semi-supervised, that includes other transductive learning, lots of other sort of settings. So that's the reason now people are preferring this term self-supervised because it explicitly says what's happening. The data itself is the source of supervision and any sort of learning algorithm which tries to extract just sort of data supervision signals from the data.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“These concepts And now we come to the other extreme, which is like self supervised learning. The idea basically is that the machine or the algorithm should really discover concepts or discover things about the world or learn representations about the world which are useful without access to explicit human supervision.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“To collect to annotate. And it's not even that many concepts, right? It's not even that many images. 14 million is nothing, really. Like you have about, I think, 400 million images or so or even more than that uploaded to most of the popular sort of social media websites today. So now supervised learning just doesn't scale. If I want to now annotate more concepts, if I want to have various types of fine-grained concepts, then it won't really scale. So now you come up to these sort of different learning paradigms. For example, semi-supervised learning, where the idea is, of course, you have this annotated corpus of supervised data, and you have lots of these unlabeled images. And the idea is that the algorithm should basically try to measure some kind of consistency or really try to measure some kind of signal on this sort of unlabeled data to make itself more confident about what it's really trying to predict. So by access to this lots of unlabeled data, the idea is that the algorithm actually learns to be more confident and actually gets better at predicting.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So now on Unseen or like new kinds of data, this model can automatically learn to predict these concepts. So, this is a standard sort of supervised setting. For semi-supervised setting, the idea typically is that you have, of course, all of the supervised data, but you have lots of other data which is unsupervised or which is not labeled. Now, the problem basically with supervised learning and why you actually have all of these alternate sort of learning paradigms is supervised learning just does not scale. So if you look at for computer vision, the sort of largest, one of the most popular data sets is ImageNet. So the entire ImageNet data set has about 22,000 concepts and about 14 million images. So these concepts are basically just nouns and they're annotated on images. And this entire data set was a mammoth data collection effort. It actually gave rise to a lot of powerful learning algorithms. It's credited with sort of the rise of deep learning as well. But this data set took about 22 human years.”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Let's start with supervised learning. So, typically for machine learning systems, the way they're trained is you get a bunch of humans. The humans point out particular concepts. So if it's in the case of images, you want the humans to come and tell you what is present in the image, draw boxes around them, draw masks of things, pixels, which are of particular categories or not. For NLP, again, there are lots of these particular tasks, say about sentiment analysis, about entailment and so on. So typically for supervised learning, we get a big corpus of such annotated or labeled data. And then we feed that to a system. And the system is really trying to mimic. So it's taking this input of the data and then trying to mimic the output. So it looks at an image and the human has tagged that this image contains a banana and now the system is basically trying to mimic that. So that's its learning signal. And so for supervised learning, we try to gather lots of such data and we train these machine learning models to imitate the input or output. And the hope is basically by doing”
2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source