YouSaid · the spoken record

Jitendra Malik

lines on the record
97
first
2020-07-21
most recent
2020-07-21
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. The other thing which I have, which I think I bring to the table is a certain intellectual breadth. I've spent a fair amount of time studying psychology, neuroscience, relevant areas of applied math and so forth. So I can probably help them see some connections to disparate things. they might not have otherwise. So the smart students coming into Berkeley can be very deep in the sense they can think very deeply, meaning very hard down one particular path, but where I could help them is the shallow breadth, whereas they would have the The narrow depth. But that's of some value.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So, I think that is important. So if I have that and if I can convey that to students, it's not just that they do great research while they're working with me, but that they continue to do great research. So in a sense, I'm proud of my students and their achievements and their great research even 20 years after they've ceased being my student.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I think I have a sense of what is a good problem, which is there is this British scientist, in fact he won a Nobel Prize, Peter Medeware, who has a book on this. And basically he calls it the research is the art of the soluble. So we need to sort of find problems which are Which are not yet solved, but which are approachable. And he sort of refers to this. Sense that there is this problem which isn't quite solved yet, but it has a soft underbelly. There is some place where you can spear the beast And having that intuition that this problem is ripe is a good thing because otherwise you can just beat your head and not make progress.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  4. So, I think that in science, in the short run, success is always based on technical competence, you're quick with math or you're whatever. I mean, there's certain technical capabilities which make for short range progress. Long-range progress is really determined by asking the right questions and focusing on the right problems. I've been able to bring to the table in terms of advising these students is some sense of taste of what are good problems, what are problems that are worth attacking now as opposed to waiting 10 years.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Yeah, I think what I feel I've been lucky to have had very, very smart and hardworking and creative students. I think some part of their credit just belongs to being at Berkeley. I think those of us who are at top universities are blessed because we have very, very smart and capable students coming and knocking on our door. So I have to be humble enough to acknowledge that. But what have I added? I think I have added something. What I have added is I think what I've always tried to teach them is a sense of picking the right problems.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Solving a lot of practical problems has offered a lot of tools for scientific research. Because computer vision is impactful for images in biology or astronomy and so on and so forth. And we have, so we have made great scientific progress, which has had real practical impact in the world. And I feel lucky that I got in at a time when the field was Very young, and at a time when it is now Mature but not fully mature. It's mature but not done. I mean, it's really still in a productive phase.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I don't think of single moments, but I look over the long haul I feel that I've been very lucky because I feel that I think that in scientific research, A lot of it is about being at the right place at the right time. And you can work on problems at a time when they're just too premature. You butt your head against them and nothing happens because it's Prerequisites for success are not there. And then there are times when you are in a field which is all pretty mature and you can only solve curly cues upon curlics. I've been lucky to have been in this field which... 34 years. Well, actually, 34 years as a professor at Berkeley, so longer than that, which when I started in it was just like some little crazy, absolutely useless field. Couldn't really do anything to a time when it's really, really

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  8. So I would argue that when these systems make mistakes, there are consequences. And we are in a certain sense responsible for those consequences. So I would argue that this is a continuous effort. And this is something that in a way is not so surprising. It's about all engineering and scientific progress, which great power comes great responsibility. So as these systems are deployed, we have to worry about them. And it's a continuous problem. I don't think of it as something which will suddenly happen on some day in 2079 for which I need to design some clever trick. I'm saying that these problems exist today. And we need to be continuously on the lookout for. Worrying about safety biases, risks, right? I mean, a self-driving car kills a pedestrian and they have, right? I mean, this Uber incident in Arizona.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Think we need to be worried about AI today. I think that it is not just a worry we need to have when we get that AGI. I think that AI is being used in many systems today. And there might be settings, for example, when it causes biases or decisions which could be harmful. I mean, decisions which could be unfair to some people, or it could be a self-driving cars which kills a pedestrian. AI systems are being deployed today. And they are being deployed in many different settings, maybe in medical diagnosis, maybe in a self-driving car, maybe in selecting applicants for an interview.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  10. The community as a whole. Certainly not. And I think to me that was the surprise that they actually worked robustly for a wide range of problems from a wide range of initializations and so on. And so that was certainly more rapid progress than we expected. But then there are certainly lots of times, in fact, most of the history and AI is when we have made less progress at a slower rate than we expected. So we just keep going. I think what I regard as really unwarranted are these fears of AGI in 10 years and 20 years and that kind of stuff because that's based on completely unreal Models of how rapidly we will make progress in this field.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Over parametrized People who used to play with them a lot, the ones who are totally immersed in the lore and the black magic, they knew that they worked well, even though

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  12. That is possible because certainly I have been very positively surprised by how effective these deep learning systems have been because I certainly would not have believed that in 2010. I think what we knew from the mathematical theory. Was that convex optimization works when there's a single global optima, then there's gradient descent techniques would work. Now these are nonlinear systems with non-convex systems.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Feel with respect to natural language, understanding, and high-level cognition, it's not just known unknowns, but also unknown unknowns. So it is very difficult to put any kind of a timeframe to that.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So I think there are two answers here. One answer is, in principle, can we do this at some time? And my answer is yes. The second answer is a pragmatic one. Do you think we will be able to do it in the next 20 years or whatever? And to that man says no. So, and of course, that's a wild guess. I think that Donald Rumsfeld is not a favorite person of mine, but one of his lines was very good, which is about knowns, known unknowns, and unknown unknowns. So in the business we are in, there are no unknowns and we have unknown unknowns. So I think with respect to a lot of what the case in vision and robotics, I feel like we have known unknowns. So I have a sense of where we need to go and what the problems that need to be solved are.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  15. But this has an impact on what treatments should be picked, right? So there are settings where I want to know more than just this is the answer. But what I acknowledge is that in that sense explainability and interpretability may matter. It's about giving error bounds and a better sense of the quality of the decision. Where I'm willing to sacrifice interpretability is that I believe that there can be systems which can be highly performant but which are internally black boxes.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  16. So now is the computer program's diagnosis based on data which was data collected for American males who are in their 30s and 40s and maybe not so relevant to me. Maybe it is relevant, you know, et cetera, et cetera. In medical diagnosis, we have major issues to do with the reference class. So we may have acquired statistics from one group of people and applying it to a different group of people who may not share all the same characteristics. The data might have, there might be error bars in the prediction. So that prediction should really be taken with huge grain of salt.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  17. One human to another human is not fully explained. I think there are settings where explainability matters and these might be, for example, questions on medical diagnosis. So I am in a setting where maybe the doctor, maybe a computer program has made a certain diagnosis. And then depending on the diagnosis, perhaps I should have treatment A or treatment B.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  18. I think that's fine To me, that is, let me put some caveats on that. So it depends on the setting. So first of all, I think humans are not explainable.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  19. It's hidden in the representation, the ability to predict new views. And what I would see if I went to such and such position.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  20. And from those, as part of that, I might see a chair from different viewpoints or a table from different viewpoints and so on. Now, as part that enables me to build some internal representation. And then next time I just see a single photograph and it may not even be of that chair, it's of some other chair. And I have a guess of what its 3D shape is like.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Okay, then we have some techniques that we have developed in the computer vision community which try to guess 3D from single views. And these techniques are based on a supervised learning and they are based on having a training time, 3D models of objects available. This is completely unnatural supervision CAD models are not injected into your brain Okay, so what would I like? What I would like would be a kind of Learning as you move around the world notion of 3D. So we have our Succession of visual experiences.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  22. And if their behavior might be, these are agents. So they are not just like passive objects, but the agent. So therefore, they would exhibit goal-directed behavi Okay, so this is one area. Then I will talk about understanding the world in 3D. Now this may seem paradoxical because in a way we are able to do 3D understanding even like 30 years ago, right? But I don't think we currently have the richness of 3D understanding in our computer vision system that we would like. So let me elaborate on that a bit. So currently we have two kinds of techniques which are not fully unified. So they are the kinds of techniques from multi-view geometry that you have multiple pictures of a scene and you do a reconstruction using stereoscopic vision or structure for motion. But these techniques do not, they totally fail if you just have a single view because they are relying on this multiple view geometry.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Me pick one which I regard as clearly unsolved, which is what I would call long-form video understanding. So we have a video clip and we want to understand the behavior in there in terms of agents, their goals, intentionality. and make predictions about what might happen. So that kind of understanding which goes away from atomic visual action. So in the short range, the question is, are you sitting, are you standing? Are you catching a ball? Are we, even if we can't do it fully accurately, if we can do it at 50%, maybe next year we'll do it at 65 and so forth. I think the long range video understanding, I don't think we can do today.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  24. No, there aren't any. I think essentially to me a really good aid to the blind. So suppose there was a blind person and I needed to assist the blind person.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  25. I think that to me, and this is not an exhaustive list by any means. So I would, I think that that's where we need to be going to. And each of these, on each of these axes, there's a fair amount of work to be done.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  26. And I think it made a lot of sense then where we are today, 70 years later, I think we should not worry about that. I think the Turing test is no longer the right way to channel research in AI because it takes us down this path of this chat bot which can fool us for five minutes or whatever. I think I would rather have a list of 10 different tasks. I mean, I think there are tasks in the manipulation domain, tasks in navigation, tasks in visual scene understanding, tasks in reading a story and answering questions based on that. I mean, so my favorite language understanding task would be reading a novel and being able to answer arbitrary questions from it. Okay.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  27. The type of problem. I think I'm not a fan of the I think Turing testing as he proposed the test in 1950 was trying to solve a certain problem.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  28. I think I don't think we should have create a single test of intelligence. So, just like I don't believe in IQ as a single number, I think generally there can be many capabilities which are correlated, perhaps. So I think that There will be accomplishments which are visual accomplishments, accomplishments which are accomplishments in manipulation or robotics and then accomplishments in language. I do believe that language will be the hardest nut to crack

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  29. So the world, so these ancestors of ours, three, four million years ago, they had spatial intelligence. So they knew that the world consists of objects. They knew that the objects were in certain relationships to each other. They had observed causal interactions among objects. They could move in space, so they had space and time and all of that. So language builds on that substrate. So language has a lot of, I mean, all human languages have constructs which depend on a notion of space and time. That notion of space and time come from. It had to come from perception and action in the world we live in.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  30. But after that And other primates don't have that. But so it developed somewhere in this era, but it developed I mean, argue that it probably developed after we had this stage of. Human species already able to manipulate hands-free, much bigger brain size.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  31. No, no, sorry. The first multicellular animals, which you can say had some intelligence. 500 million years Okay, and now let's fast forward to say the last seven million years, which is the development of the hominid line, right? From the other primates, we have the branch which leads on to modern humans. Now many of these hominids, but the ones which people talk about Lucy because that's like a skeleton from three million years ago and we know that Lucy walked. At this stage you have that the hand is free for manipulating objects. And then the ability to manipulate objects, build tools. And the brain size grew in this era. So now you have manipulation. Now we don't know exactly when language arose.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  32. And you can think of it either in phylogeny or in ontogeny. So phylogeny means if you look in evolutionary time. So we have vision that developed 500 million years ago. Okay, then something like when we get to maybe like 5 million years ago, you have the first bipedal primate. So when we started to walk, then the hands became free. And so then manipulation, the ability to manipulate objects and build tools and so on and so forth.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Yeah, I mean, we have a little bit of work at this, but so much more needs to be done. So this is a good example be physical, that's to do with the one thing we talked about. Earlier, that there's an embodied world.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yes, because they are occurring at the same time in time. You have time which links the two. So at a certain moment T1, you've got a certain signal in the auditory domain and a certain signal in the visual domain. But they must be causally related

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Okay, an audio visual. So there is some event that happens in the world, and that event has a visual signature, and it has auditory signature. So there is this glass bowl on the table and it falls and breaks, and I hear the smashing sound and I see the pieces of glass. Okay, I've built that connection between the two, right? People, I mean, this has become a hot topic in computer vision in the last couple of years. There are problems like separating out multiple speakers, right? Which was a classic problem in audition. They call this the problem of source separation or the cocktail party effect and so on. But just try to do it visually when you also have, it becomes so much easier and so much more useful.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  36. And then we have the pixels corresponding to the visual view, but we know that they correspond to the same object. Right? So that's a very, very strong cross calibration signal. And it is self supervisory, which is beautiful, right? There's nobody assigning a label. The mother doesn't have to come and assign a label. The child doesn't even have to know that this object is called a ball. But the child is learning something about the three dimensional world from this signal. I think tactile and visual, there is some work on. There is a lot of work currently on audio and visual.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Bridge the two worlds. So let's take an example of a multimodal. I like that. So multimodal, a canonical example is a child interacting with an object. So then the child, so the child holds a ball and plays with it. So at that point, it's getting a touch signal. So the touch signal is getting the notion of 3D shape, but it is sparse. And then the child is also seeing a visual signal. And these two, so imagine these are two in totally different spaces, right? So one is the space of receptors on the skin of the fingers and the thumb and the palm. And then these map onto these neuronal fibers are getting activated somewhere. These lead to some activation in somatosensory cortex. I mean, a similar thing will happen if we have a robot hand.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  38. So, I mean, I should say to give due credit, this is from a paper by Smith& Gasser. And it reflects essentially, I would say common wisdom among child development people. It's just that these are, this is not common wisdom among people in... Computer vision and AI and machine learning. So, I view my role as trying to

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  39. So I think that currently what that end to end learning means nowadays is end-to-end supervised learning. And that, I would argue, is too narrow a view of the problem. I like this child development view, this lifelong learning view, one where there are certain capabilities that are built up, and then there are certain capabilities which are built up On top of that, so that's what I believe in. So I think End-to-end learning in the supervised setting for a very precise task to me is Kind of sort of a limited view of the learning process.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  40. So therefore, I try to promote this particular framework as a way of considering the problems that people in computer vision were actually working on and trying to be more explicit about the fact that they actually are connected to each other And I was at that time just doing this on the basis of information flow. Now it turns out in the last five years or so in the post the deep learning revolution that this architecture has turned out to be very conducive to that because basically in these neural networks we are trying to build multiple representations There can be multiple output heads sharing common representations. So, in a certain sense, today, given the reality of what solutions people have to these, I do not need to preach this anymore. It is just there. It's part of the solution space.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  41. I mean, I started pushing this kind of a view in around 2010 or something like that. Because at that time, in computer vision, the distinction that people... We're just working on many different problems, but they treated each of them as a separate isolated problem. With each with its own data set, and then you try to solve that and get good numbers on it. So I wasn't, I didn't like that approach because I wanted to see the connection between these. And if people divided up vision into various modules, the way they would do it is as low level, mid-level, and high-level vision. Corresponding roughly to the psychologist's notion of sensation, perception, and cognition. And that didn't map to tasks that people cared about

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Okay, reorganization is to do with essentially finding these entities. So it's organization. The word organization implies structure. perception, in psychology, we use the term perceptual organization. The world is not just an image is not just seen as not internally represented as just a collection of pixels, but we make these entities. We create these entities, objects, whatever you want to call them.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  43. That's recognition, reconstruction is essentially You can think of it as inverse graphics. I mean, that's one way to think about it. So graphics is you have some internal computer representation and you have a computer representation of some objects arranged in a scene. And what you do is you produce a picture. You produce the pixels corresponding to a rendering of that scene. So let's do the inverse of this. We are given an image and we try to We say, oh, this image arises from some objects in a scene looked at with a camera from this viewpoint. And we might have more information about the objects like their shape, maybe their textures, maybe color, etc. So that's the reconstruction problem. In a way, you are in your head creating a model of the external world.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  44. How they interact? Yeah. So recognition is the easiest one because that's what I think people generally think of as computer vision achieving these days, which is labels. So is this a cat? Is this a dog? Is this a Chihuahua? I mean, it could be very fine-grained. you know, specific breed of a dog or a specific species or bird or it could be very abstract like animal

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  45. And it is not semantics free. So I think it blends into, it involves perception and cognition. It is not, I think the mistake that we used to make in the early days of computer vision was to treat it as a purely bottom-up perceptual task. It is not just that. We do revise our notion of segmentation with more experience, right? Because, for example, there are objects which are non-rigid, like animals or humans. And I think understanding that all the pixels of a human are one entity is actually quite a challenge because the parts of the human, they can move independently. The human wears clothes, so there might be differently colored. So it's all sort of a challenge.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Weak supervision works in the context that you have the ability to create objects. So I think that to me that's a very fundamental capability. There are applications where this is very important. For example, medical diagnosis. So in medical diagnosis, you have some brain scan, I mean, this is some work that we did in my group where you have CT scans of people who have had traumatic brain injury and what the radiologist needs to do is to precisely delineate various places where there might be bleeds, for example. And there are clear needs like that. So they're certainly very practical applications of computer vision where segmentation is necessary. But philosophically, segmentation Enables the task of recognition to proceed with much weaker supervision than we require today.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  47. So, I think that we have that capability, and that enables us to, as we are growing up, to acquire names of objects with very little supervision. So suppose the child, let's posit that the child has this ability to separate out objects in the world. Then when the mother says, pick up your bottle or the cat's behaving funny today. The word cat suggests some object, and then the child sort of does a mapping. The mother doesn't have to teach specific object labels by pointing to them.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Yeah, so I think so. So, segmentation enables us to say. That some set of pixels are an object without necessarily even being able to name that object or knowing properties of that object

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  49. And the question that you can ask is so for me, I'm inspired a lot by human vision and I care about that. You could be just a hard-boiled engineer, not give a damn. So to you, I would then argue that you would need far less training data if you could make my research agenda.

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source

  50. So there was a bottom up process by which we were able to segment out these objects. And we have totally focused on this top-down training signal. So in my view, we have currently solved it in machine vision, this top-down, bottom-up interaction. But I don't find the solution fully satisfactory, and I would rather have a bit of both at both stages

    2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source