YouSaid · the spoken record

Ishan Misra

lines on the record
175
first
2021-07-31
most recent
2021-07-31
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Thanks, Lex. I mean, I've listened to you. I told you it was unreal for me to actually meet you in person, and I'm so happy to be here. Thank you.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  2. I would say it's a feature because then everyone actually has very different kinds of objective functions that they're optimizing, and those objective functions evolve and change dramatically through their course of their life. That's actually what makes us interesting, right? Otherwise, like if everyone was doing the exact same thing, that would be pretty boring. We do want people with different kinds of perspectives, also people evolve continuously. That's, I would say, the biggest feature of being humans.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I think it's an endless sort of quest for us. I don't think AI will help us there. This is like a very hard, hard, hard question, which so many humans have tried to answer.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Because all of these were failed experiments. And then he says, Oh, these 990 things don't work, and I know that. Did you know that?

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  5. I know this is like super cheesy that failure is something that you should be prepared for and so on, but I do think, I mean, especially in research, for example, failure is something that happens almost every day is like experiments failing and not working. And so you really need to be so used to it. You need to have a thick skin and only basically through when you get through it is when you find the one thing that's actually working. Lexomus Edison was like one person like that, right? So I really like when I was a kid I used to really read about how he found the filament, the light bulb filament. And then I think his thing was like he tried 990 things that didn't work or something of the sort. And then they asked him, so what did you learn?

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Think there's not going to be any single thing that you're going to need. There are going to be different types of things that you need. But whenever you need something, you just go push for it. And of course, once you may not get it, or you may find that this was not even the thing that you were looking for, it might be a different thing. But the point is you're pushing through things and that actually gives a lot of skills and brings a lot of like. Build a certain kind of attitude which will probably help you get the other thing once you figure out what's really the thing that you want.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I would say just be hungry. Like, always be hungry for what you want. And I think I've been inspired by a lot of people who are just driven and who really go for what they want no matter what you shouldn't want it. You should need it. So if you need something, you basically go towards the ends to make it work.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  8. So, like in Pittsburgh, where I did my PhD, the thing was it used to snow a lot And so when it was snowed, you really couldn't do much. So, the thing that a lot of people said was snow builds character. Because when it's snowing, you can't do anything else.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Googling is, of course, like a good way to solve it when you need a quick answer. But I think initially, especially when you're starting out, it's much nicer to figure things out by yourself. And I just say that from experience, because when I started out, there were not a lot of resources. So we would in the lab, a lot of us, we would look up to senior students. And then the senior students were, of course, busy and they would be like, hey, why don't you go figure it out? Because I just don't have the time. I'm working on my dissertation or whatever. Find her a PhD students. And so then we would sit down and just try to figure it out. And that, I think, really helped me. That has really helped me figure a lot of things out.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So, for example, if an algorithm, if you try to train a network and it's not converging, whatever, rather than trying to Google the answer or trying to do something, really spend those five, eight, ten, 15, 20, whatever number of hours really trying to figure it out yourself. Because in that process, you'll actually learn a lot more.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Don't be afraid to get your hands dirty. I think that's the main thing. So if something doesn't work, really drill into why things are not working.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So, I think in terms of having these two frameworks or multiple, I think, of course, there are different use cases, so there are going to be benefits to using one or the other framework. And like you said, I think competition is just healthy because both of these frameworks keep, or like all of these frameworks really sort of keep learning from each other and keep incorporating different things to just make them better and better.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I think the downside is that for a lot of research code that's released in one framework, and if you're using the other one, it's really hard to really build on top of it. But thankfully, the open source community in machine learning is amazing. Whenever something pops up in TensorFlow, you wait a few days and someone who's super sharp will actually come and translate that particular code base into PyTorch and basically have figured that allooks and crannies out. So the open source community is amazing and they really figure out this gap.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Makes debugging easier for them. So, like, I learned programming in C, and so for me, imperative way of programming is more natural

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Yeah, I mean, so Luitorch was really good because it actually allowed you to do a lot of different kinds of things. So which cafe was very rigid in terms of its structure. You would create a neural network once and that's it. Whereas if you wanted very dynamic graphs and so on, it was very hard to do that. And Lower Torch was much more friendly for all of these things. Okay, so in terms of PyTorch and TensorFlow, my personal bias is PyTorch just because I've been using it longer and I'm more familiar with it. And also that PyTorch is much easier to debug is what I find because it's imperative in nature compared to like TensorFlow, which is not imperative. But that's telling you a lot that basically the imperative design is sort of a way in which a lot of people are taught programming. And that's what actually

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  16. So then I moved to Lua Torch, which was the torch in Lua, and then in 2017, I think basically pretty much to PyTorch completely.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  17. So, a disclaimer to this is that the last time I used TensorFlow was probably like four years ago. And so it was right when it had come out. I started on deep learning in 2014 or so and the dominant sort of framework for us then for Vision was CAFE, which was out of Berkeley. And we used cafe a lot, it was really nice. And then TensorFlow came in, which was basically like Python first. So Cafe was mainly C++, and it had very loose kind of Python binding. So Python wasn't really the first language you would use. You would really use either MATLAB or C++ to get stuff done in Cafe. And then Python, of course, became popular a little bit later. So TensorFlow was basically around that time. So 2015, 2016 is when I last used it. It's been a while.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  18. I would say Python just because it's the easiest one to learn and also a lot of programming in machine learning happens in Python so if you don't know any other programming language Python is actually going to get you a long way.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  19. And so a lot of famous researchers I know actually would start off first they would, even before the experiments were in, a lot of them would actually start with writing the introduction of the paper. With zero experiments in Because that at least helps them figure out what they're trying to solve and how it fits in the context of things right now. And that would really guide their entire research. So, a lot of them would actually first write in intros with zero experiments in, and that's how they would start projects.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  20. So as a grad student, I used to always wade toward maybe in the last week or whatever to write the paper because I used to always believe that doing the experiments was actually the bigger part of research than writing. And my advisor always told me that you should start writing very early on and I thought, oh, it doesn't matter. I don't know what he's talking about. But I think more and more I realized that's the case. Like, whenever I write something that I'm doing, I actually think much better about it. And so if you start writing early on, you actually, I think, get better ideas, or at least you figure out holes in your theory or particular experiments that you should run to block those holes.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Really, need to believe in that to be able to do that for that long. In terms of papers, I think the one thing that I've learned is I've like in the past whenever I used to write things and even now whenever I do that, I try to cram in a lot of things into the paper. Whereas what really matters is just pushing one simple idea. That's it. That's all because the paper is going to be like whatever, eight or nine pages. If you keep cramming in lots of ideas, it's really hard for the single thing that you believe in to stand out. So if you really try to just focus especially in terms of writing, really try to focus on one particular idea and articulate it out in multiple different ways. It's far more valuable to the reader as well. And basically to the reader, of course, because they get to, they know that this particular idea is associated with this paper. And also for you, because you have, like when you write about a particular idea in different ways, you think about it more deeply.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  22. So I think it's just the fact if you're starting a PhD, for example, what is one problem that you want to focus on that you do think is interesting enough and you will be able to make a reasonable amount of headway into it that you think you'll be doing a PhD for. So in that kind of a timeframe. So that's one. Of course, there's the second part, which is what excites you genuinely. So you shouldn't just pick problems that you are not excited about because as a grad student or as a researcher, you really need to be passionate about it to continue doing that because there are so many other things that you could be doing in life.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  23. So, I think both of these I've picked up from lots of people I've worked with in the past. So one of them is picking the right problem to work on in research is as important as finding the solution to it. So, I mean, there are multiple reasons for this. So one is that there are certain problems that can actually be solved in a particular timeframe. So now say you want to work on finding the meaning of life. This is a great problem. I think most people will agree with that. But do you believe that your talents and like the energy that you'll spend on it will make a meaning, like make some kind of meaningful progress in your lifetime? If you are optimistic about it, then go ahead.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Very, very hard question to answer For certain people, it might be something that they really derive a lot of value out of, derive a lot of enjoyment and happiness out of. And maybe the real world wasn't giving them that. That's why they did that. So maybe it is good for certain people.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So now, if you were to build a simulator based on just right now to build them today, you don't have that many autonomous cars on the road. So you will try to make all of the other agents in that simulator behave as humans. But that's not really going to hold true 10, 15, 20, 30 years from now.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Think another problem is for autonomous driving, right? It's a constantly changing world. So say autonomous driving in 10 years from now, there are lots of autonomous cars, but they're still going to be humans. So now there are 50% of the agents which are humans, 50% of the agents that are autonomous like car driving agents. So now the mixture is changed. So now the kinds of behaviors that you actually expect from the other agents or other cars on the road are actually going to be very different. And as the proportion of the number of autonomous cars to humans keeps changing, this behavior will actually change a lot.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Lighting all those kinds of things. I mean, all these companies that you have, right? So Pixar and like. Whatever. All these companies are all this computer graphic stuff is really about accurately. A lot of them is about accurately trying to figure out how the lighting is and how things reflect off of one another and so on and how sparkly things look and so on. So it's a very hard problem. So do we really need to solve that first to be able to do computer vision? Probably not.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Right. Yeah. For example, one example would be to train an autonomous car driving system. You basically first build a simulator, which builds the environment of the world, and then you basically have a lot of, you train your machine learning system in that. So I believe it is possible, but I think it's a really expensive way of doing things. And at the end of it, you do need the real world. I'm not sure. So maybe for certain settings, like maybe the payout is so large, like for autonomous driving, the payout is so large that you can actually invest that much money to build it. But I think as a general sort of principle, it does not apply to a lot of concepts. You can't really build simulations of everything, not only because one, it's expensive, because second, it's also not possible for a lot of things. So in general, there is a lot of work on using synthetic data and synthetic simulators. I generally am not very, I don't believe.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  29. This is going to be a little controversial, but okay, sure. I don't believe in simulation, like actually using simulation to do things very much.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Counting, I do believe, I mean, should be possible. Don't know yet, but I do think it's not that far in the realm of possibility.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Something like object permanence can definitely emerge, right? So that's like a fundamental concept which we have. Maybe not through images through video, but that's another concept that should be emergent from it because it's not something that we don't teach humans that this is about this concept of object permanence. It actually emerges. And the same thing for animals like dogs, I think, actually permanence automatically is something that they are born with. So I think it should emerge from the data. It should emerge basically very quickly.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Something we are taking for granted, maybe like a lot in terms of how we're setting up these algorithms, but it's just a very beautiful and powerful idea. So it's really fundamentally telling us something about. There is so much signal in the pixels that we can be super dumb about it, about how we are setting up the self super exerting problem. And despite being super dumb about it, we'll actually get very good surprising things.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  33. The fact that object level, what objects are in some notion of objectness emerges from these models by just like self-supervised sorting. So, for example, one of the things like the dyno paper that I was a part of at Facebook is the object sort of boundaries emerge from these representations. So if you have like a dog running in the field, the boundaries around the dog, the network is basically able to figure out what the boundaries of this dog are automatically. It was never trained to do that. It was never trained to, no one taught it that this is a dog and these pixels belong to a dog. It's able to group these things together automatically. So that's one. I think in general that entire notion that this dumb idea that you take like these two crops of an image and then you say that the features should be similar, that has resulted in something like this, like the model is able to figure out what the dog pixels are and so on. That just seems like so surprising. And I mean, I don't think a lot of us even understand how that is happening really.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So I think of self awareness a little bit more than the broader goal of it because I think self awareness is pretty critical for any kind of AGI or whatever you want to call it that we build because it needs to contextualize what it is and what role it's playing with respect to all the other thing that exist around it. I think that requires self-awareness. It needs to understand that it's an autonomous car and what does that mean? What are its limitations? What are the things that it is supposed to do and so on? What is its role in some way? Or I mean these are the kind of things that we kind of expect from it, I would say. And so that's the level of self-awareness. That's, I would say, basically required at least, if not more than that.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  35. I don't think so. I think at some point we will need to bite the bullet and actually interact with the physical as much as I like working on passive computer vision where I just sit in my armchair and look at videos and learn. I do think that we will need to have some kind of embodiment or some kind of interaction to figure out things about the world.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  36. But they do display emotion, right? I mean, dogs do display emotion. So they don't have to be anthropomorphic for them to display the kind of emotions that we want

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  37. I think emotion I think emotion is something which is It's not really attributed typically in standard machine learning. It's not something we think about. There is NLP, there is vision, there is no emotion. Emotion is never a part of all of this. And that just seems a little bit weird to me. I think the reason basically being that there is surprise and basically emotion is one of the reasons emotions arise is like what happens and what you expect to happen, right? There is a mismatch between these things. And so that gives rise, I can either be surprised or I can be saddened or I can be happy and all of this. And so this basically indicates that I already have a predictive model in my head and something that I predicted or something that I thought was likely to happen. And then there was something that I observed that happened and there was a disconnect between these two things.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  38. I think whenever we try to understand that we are putting our own subjective human bias into it. And I think that's the sort of problem with self supervised thing. The goal is that it should learn naturally from the data. So now, if you try to understand it, you are using your own preconceived notions of what this model has learned. That's the problem.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Right. So I think we're like continual learning is sort of the paradigm for this in machine learning, and I don't think it's a very well explored paradigm. We have things in deep learning, for example, right? Catastrophic forgetting is one of the standard things. The thing basically being that if you teach a network to recognize dogs and now you teach that same network to recognize cats, it basically forgets how to recognize dogs. So it forgets very quickly. I mean, and whereas a human if you were to teach someone to recognize dogs and then to recognize cats, they don't forget immediately how to recognize these dogs. I think that's basically sort of what you're trying to get. Yeah, I just.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  40. So, I think it is, of course, learnable because we are prime examples of machines that have, or individuals that have learned this. Humans have learned this. So it is, of course, a technique that is very easy to learn. I think where we are kind of hitting a wall basically with current machine learning is the fact that when the network learns all of this information, we basically are not able to figure out how well it's going to generalize to an unseen thing. And we have no a priori, no way of characterizing that. And I think that's basically telling us a lot about the fact that we really don't know what this model has learned and how well it's basically, because we don't know how well it's going to transfer.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  41. But if you were to give them a very complicated thing that they've not seen before, they have very limited ability right now to compose different things. Like, oh, I've seen this particular part before, I've seen this particular part before, and now probably like this is how they're going to work in tandem. It's very hard for them to come up with these kinds of things.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  42. So I think of it slightly differently. So for me, learning is whenever I can make a snap judgment. So if you show me a picture of a dog, I can immediately say it's a dog. But if you give me a puzzle, whatever a Goldsberg machine Of things going to happen, then I have to reason because I've never, it's a very complicated setup, I've never seen that particular setup, and I really need to draw and like imagine in my head what's going to happen to figure it out. So I think, yes, neural networks are really good at recognition, but they're not very good at reasoning because they have seen something before or seen something similar before they're very good at making those sort of snap judgments.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  43. I think most people would agree with that statement. But we are still okay with it. We still don't call this as a bug. As in traditional computer science or traditional science, like if you have this kind of failure case existing, then you think of it as something is wrong. Think there is this sort of notion of nebulous correctness for machine learning, and that's something we just need to be very comfortable with. And for deep learning, or like for a lot of these machine learning algorithms, it's not clear how do we characterize this notion of correctness. I think limitation in our understanding, or at least a limitation in our phrasing of this. And if we were to come up with better ways to understand this limitation, then it would actually help us a lot.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Just that I have been very used to seeing handwritten digits in my life. The other sort of problem with any deep learning system or any kind of machine learning system is it's guarantees, right? There are no guarantees for it. Now you can argue that humans also don't have any guarantees. There is no guarantee that I can recognize a cat in every scenario. I'm sure there are going to be lots of cats that I don't recognize, lots of scenarios in which I don't recognize cats in general. But I think from just a sort of application perspective, you do need guarantees. We call these things algorithms. Now, algorithms, like traditional CS algorithms have guarantees. Sorting is a guarantee. If you were to call sort on a particular array of numbers, you are guaranteed that it's going to be sorted. Otherwise, it's a bug. Now for machine learning, it's very hard to characterize this. We know for a fact that like cat recognition model is not going to recognize cats, every cat in the world in every circumstance.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  45. I think we are good at transferring knowledge a little bit. We are just better at a lot of these problems where we are generalizing from a single sample, recognizing from a single sample, we are using a lot of our own domain knowledge and a lot of our inductive bias into that one sample to generalize it. So I've never seen you write the number nine, for example. And if you were to write it, I would still get it. And if you were to write a different kind of alphabet and write it in two different ways, I would still probably be able to figure out that these are the same two characters

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  46. And you're actually going to be able to recognize this kind of instance in a very different scenario. So, like, when it was snowing, so you got that thing labeled, when it was snowing, but now when it's raining, you're actually not able to get it. Or you basically have the same scenario in a different part of the world. So the lighting was different or so on. So it's just really hard for these models, like deep learning, especially to do that.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  47. In general, I think for deep learning, the biggest challenge is just data efficiency, even with self-supervised learning, even with anything else. If you just see a single concept once, like one image of a, like, I don't know, whatever you want to call it, any concept, it's really hard for these methods to generalize by looking at just one or two samples of things. And that has been a real, real challenge. And I think that's actually why these edge cases, for example, for Tesla are actually that important. Because if you see just one instance of the car failing, and if you just annotate that and you get that into your data set, you have like very limited guarantee that it's not going to happen again.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  48. I think for deep learning in particular, because self supporting the thing is, I would say a little bit more vague right now, so I wouldn't like for something that's so vague, it's hard to predict what its limits are going to be. But like I said, I think anywhere you want to interact with human self-supervisor, I think kind of hits a boundary very quickly because you need to have an interface to be able to communicate with the human. So really, if you have just vacuous concepts or just nebulous concepts discovered by a network, it's very hard to communicate those with the human without inserting some kind of human knowledge or some kind of human bias there.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So, I think for me, it's five if I'm being optimistic and it's going to be five for a lot of cases and ten plus, yeah, I agree with you. 10 plus basically if we want to recover most of the contiguous United States or something.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Vision based autonomous driving, that's the reason I'm quite optimistic about it because I think there are going to be lots of these advances on the sensor side itself. So acquiring this data, we are actually going to get much better about it. And then, of course, once we're able to scale out and get all of these edge cases in, as Andre described, I think that's going to make us go very far away.

    2021-07-31 · Lex Fridman Podcast · #206 – Ishan Misra: Self-Supervised Deep Learning in Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source