YouSaid · the spoken record
Jitendra Malik
- lines on the record
- 97
- first
- 2020-07-21
- most recent
- 2020-07-21
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Okay, the other, so that's, so I want to have a little bit more top down in the at test time. Okay, then at training time, we make use of a lot of top-down knowledge right now. So basically to learn to segment an object, we have to have all these examples of this is the boundary of a cat and this is the boundary of a chair and this is the boundary of a horse and so on. And this is too much top-down knowledge. How do humans do this? We managed with far less supervision and we do it in a sort of bottom-up way because, for example, we're looking at a video stream and the horse moves. And that enables me to say that all these pixels are together. The gestural psychologist used to call this the principle of common fate.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“And this might enable us to deal with the more ambiguous stimuli, for example. So the biological solution seems to involve feedback. The solution in artificial vision seems to be just feed forward, but with a much deeper network. And the two are functionally equivalent because if you have a feedback network which just has like three rounds of feedback, you can just unroll it and make it three times the depth and create it in a totally feedforward way. So this is something which, I mean, we have written some papers on this theme, but I really feel that this theme should be pursued further”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“And they involve much shallower neural networks. So, the kinds of neural networks we are using in computer vision, say a ResNet 50 has 50 layers. Well, in the visual cortex, going from the retina to IT, maybe we have like seven. So they're far shallower, but we have the possibility of feedback. So there are backward connections”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Starting from the raw pixels trying to, yeah, they start from the raw pixels and they end up with some something like cat or not a cat, right? So our systems are running totally feed forward. They're trained in a very top-down way. They are trained by saying, okay, this is a cat, this is a cat, there's a dog, there's a zebra, etc. And I'm not happy with either of these choices fully. We have gone into, because we have completely separated these processes, right? So I would like the process. So what do we know compared to biology? So in biology, what we know is that the processes at test time, at runtime, those processes are not purely feed forward, but they involve feedback.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Computer vision, right? So today what we have are very interesting systems because they work completely bottom up. How are they bottom-up?”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, so I mean, the way to think about it is just how successful is image compression, right? And that's been done with older technologies, but it can be done with there are several companies which are trying to use sort of these more advanced neural network type techniques for compression, both for static images as well as for video one of my former students as a company which is trying to do stuff like this. And I think that they are showing quite Interesting results, and I think that that's all the success statistics and video statistics.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so what we seem to have learned is that there's a lot of redundancy in these images. And as a result, we are able to do a lot of compression. And this compression is very important in biological settings, right? So you might have 10 to the 8 photoreceptors and only 10 to the 6 fibers in the optic nerve. So you have to do this compression by a factor of 100 is to 1. And so there are analogs of that which are happening in our artificial neural network.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't see any fundamental problems there. I mean, and look, the computer graphics community has come a long way. In the early days, going back to the 80s and 90s, they were focusing on visual realism. And then they could do the easy stuff, but they couldn't do stuff like hair or fur and so on. Well, they managed to do that. Then they couldn't do physical actions, right? Like there's a bowl of glass and it falls down and it shatters. But then they could start to do pretty realistic models of that. And so on and so forth. So, the graphics people have shown that they can do this forward direction, not just for optical interactions, but also for physical interactions. So I think, of course, some of that is very compute intensive, but I think by and by we will find ways of making our models ever more realistic.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“In a small scale. But that is a way that the child gets to build and refine its causal models of the world. And my colleague, Alison Gopnik, together with a couple of authors, co-authors, has this book called The Scientist in the Crib, referring to children. So, I like the part that I like about that is the scientist wants to build causal models. And the scientist does control experiments. And I think the child is doing that. So to enable that, we will need to have these active experiments. And I think this could be done some in the real world and some in simulation.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Causality is important, but causality is not like a single silver bullet. It's not like one single principle. There are many different aspects here. And one of the ways in which our most reliable ways of establishing causal links, and this is the way, for example, the medical community does this, is randomized control trials. So you have you pick some situation and now in some situation you perform an action for certain others you don't right so you have a control experiment well the child is in fact performing controlled experiments all the time right”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So you can imagine that subsequent generations of these simulators will be accurate not just visually, but with respect to forces and masses and haptic interactions and so on. And then we have that environment to play with. I think let me state one reason why I think this active being able to act in the world is important. I think that this is one way to break the correlation versus causation barrier. This is something which is of a great deal of interest these days. I mean, people like Judea Pearl have talked a lot about that we are neglecting causality and he describes the entire set of successes of deep learning as just curve fitting, right? Because it's, but I don't quite agree.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think that I believe in that. And I think that we could achieve it in two ways. And I think we should use both. So one is actually real robotics, right? So real physical embodiments of agents who are interacting with the world and they have a physical body with a dynamics and mass and moment of inertia and friction and all the rest and you learn your body, the robot learns its body by doing a series of actions. The second is that simulation environments. So I think simulation environments are getting much, much better. In my life in Facebook AI research, our group has worked on something called Habitat, which is a simulation environment, which is a visually photorealistic environment of places like houses or interiors of various urban spaces and so forth. And as you move, you get a picture which is a pretty accurate picture.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“We don't build flapping wings. So, yes, that's one of the points of debate. In my mind, I would bet on this learning like a child approach.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think we should have those and we don't have those today. And I think the part of the challenge is that we should really be collecting data of the type that a child experiences. So that gets into issues of privacy and so on and so forth. But there are attempts in this direction to sort of try to collect the kind of data that a child encounters growing up. So what's the child's linguistic environment? What's the child's visual environment? So if we could collect that kind of data and then develop learning schemes based on that data, that would be one way to do it. I think that's a very promising direction myself. There might be people who would argue that we could just short circuit this in some way. And sometimes we have imitated, we have not had success by not imitating nature in detail. So the usual example is airplanes, right?”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“It's possible, but I think it's not so typical that somebody that a mother coaches a child through all the stages of what happens in a restaurant, they just go as a family, they go to the restaurant, they eat, come back, and the child goes through 10 such experiences and the child has got a schema of what happens when you go to a restaurant. So we somehow need to provide that capability to our systems.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Now, how will we bring that in? So we could either revert back to the 70s and say, okay, I'm going to handcode in a script, or we might try to learn it. So I tend to believe that we have to find learning ways of doing this. Because I think learning ways land are being more robust. And there must be a learning version of the story because children acquire a lot of this knowledge by sort of just observation. So at no moment in a child's life does a”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I think this kind of knowledge is absolutely essential. So I think that when we are going to do long form video understanding, we are going to need to do this. I think the kinds of technology that we have right now with 3D convolutions over a couple of second of clip or video, it's very much tailored towards short-term video understanding, not that long-term understanding. Long-term understanding requires a notion of this notion of schemas that I talked about, perhaps some notions of goals, intentionality, functionality, and so on and so forth.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Similar ideas okay, and in the 70s, the way the AI of the time dealt with it was by hand coding this. So they hand coded in this notion of a script and the various stages and the actors and so on and so forth and use that to interpret, for example, language. I mean, if there's a description of a story involving some people eating at a restaurant, there are all these inferences you can make because you know what happens typically at a restaurant.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So I completely agree that knowledge is called for and that knowledge can be quite sophisticated. So the way I would say it is that perception blends into cognition and cognition brings in issues of memory and this notion of a schema from psychology, which is, let me use the classic example, which is you go to a restaurant, right? Now the things happen in a certain order, you walk in, somebody takes you to a table, waiter comes, gives you a menu, takes the order, food arrives, eventually bill arrives, etc. This is a classic example of AI from the 1970s. It was called, there was the term frames and scripts and schemas. These are all quite”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So I think that this is part of the reason why we have so emphasized static images. I think that this is changing. And over the next few years, I see a lot more progress happening. Video, so Have this generic statement that to me video recognition feels like 10 years behind object recognition You can quantify that because you can take some of the challenging video data sets and their performance on action classification is like, say, 30%, which is kind of what we used to have around 2009 in object detection. So it's like about 10 years behind. And whether it'll take 10 years to catch up is a different question. Hopefully it will take less than that.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“No, that's when you want to do things at scale. So if you want to operate at the scale of all the content of YouTube, it's very challenging. And there are similar issues in Facebook. But as a researcher, you have more opportunities.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“And yes, and this will save us a computation. So many of the choices were dictated by that. I think today we are no longer detecting edges. We process images with com nets because we don't need to. We don't have those compute restrictions anymore. Now, video is still understudied because video compute is still quite challenging if you are a university researcher. I think video computing is not so challenging if you are at Google or Facebook or Amazon.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Exactly. Not enough compute, not enough storage. So think of these choices. So one of the choices is focusing on single images rather than video. Questions, storage and compute. We had to focus on, we used to detect edges and throw away the image. So you have an image which is, say, 256 by 256 pixels. And instead of keeping around the grayscale value, what we did was we detected edges, find the places where the brightness changes a lot. So now that throw away the rest. So this was a major compression device. And the hope was that this makes it that you can still work with it and the logic was humans can interpret a line drawing.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“That's a fair question. I think sometimes we can simplify our problem so much that we essentially lose part of the juice that could enable us to solve the problem. And one could reasonably argue that to some extent this happens when we go from video to single images. Now, historically, you have to consider the limits imposed by the computation capabilities we had. So many of the choices made in the computer vision community through the 70s, 80s, 90s can be understood as Choices which were forced upon us by the The fact that we just didn't have access to compute enough compute.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Certainly with respect to interactions with firms. So in this homo economicus role, when you're interacting with firms, it does become that.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, I think that's the most fundamental purpose. We have by now hyper evolved. So we have this visual system which can be used for other things. For example, judging the aesthetic value of a painting. And this is not guiding action. Maybe it's guiding action in terms of how much money you will put on your auction bid, but that's a bit stretched. But the basics are, in fact, in terms of action. But we have Evolved really this hyper evolved our visual system.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Do you have that capability? That model is not perfect, and psychologists have great fun in pointing out the ways in which the model in your head is not a perfect model of the external world. They create various illusions to show the ways in which it is imperfect. But it's amazing how far it has come from a very simple perception action loop that you exist in an animal 500 million years ago. Once we have these very sophisticated visual systems, we can then impose a structure on them. It's we as scientists who are imposing that structure where we have chosen to characterize this part of the system as this quote module of object detection or this module of 3D reconstruction. What's going on is really all of these processes are running. Simultaneously.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Or seeing, I mean, vision is perhaps the single most perception sense, but all the others are equally also important. So perception and action kind of go together. So earlier it was in these very simple feedback loops, which were about finding food or avoiding becoming food if there's a predator running trying to eat you up and so forth. So we must at the fundamental level connect perception to action. Then as we evolved, perception became more and more sophisticated because it served many more purposes. And so today we have what seems like a fairly general purpose capability, which can look at the external world and build a model of the external world inside the head”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, but it's not in isolation. So we have to, so for all intelligence tasks, I always go back to sort of biology or humans. And if we think about vision or perception in that setting, we realize that perception is always to guide action. Perception in a for a biological system does not give any benefits unless it is coupled with action. So we can go back and think about the first multicellular animals which arose in the Cambrian era, you know, 500 million years ago. And these animals could move and they could see in some way. And their two activities help each other because how does movement help movement helps you? Get food in different places. But you need to know where to go, and that's really about”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“But it's not of the same style Of a very different style. So, I mean, for example, the style of computing that we have in our GPUs is far, far more power hungry than the style of computing that is there in the human brain or other biological entities.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Yes, so if we go back. So, this is a point I've been making for 20 years now. And I think once upon a time, the way I used to argue this was that we just didn't have the computing power of the human brain. Our computers were not quite there. I mean, there is a well-known trade-off which we know that neurons are slow compared to transistors, but we have a lot of them and they have a very high connectivity. Whereas in silicon, you have much faster devices, transistors switch on the order of nanoseconds The connectivities usually smaller At this point in time, I mean, we are now talking about 2020, we do have, if you consider the latest GPUs and so on, amazing computing power. And if we look back at Hans Morowicz type of calculations, which he did in the 1990s, we may be there today in terms of computing power comparability.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So there are many aspects of sort of human learning, and these have been studied in child development by psychologists. And what they tell us is that supervised learning is a very small part of it. There are many different aspects of learning. And what we would need to do is to develop models of all of these and then train our systems in that with that kind of protocol.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I think I don't see any in-principle problem with neural networks doing it, but I think the learning techniques would need to evolve significantly. So the current learning techniques that we have Are supervised learning. You're given lots of examples XIY pairs, and you learn the functional mapping between them. I think that human learning is far richer than that. It includes many different components. There is a child explores the world and sees, for example, a child takes an object and manipulates it in his or her hand and therefore gets to see the object from different points of view. And the child has commanded the movement. So that's a kind of learning data, but the learning data has been arranged by the child. And this is a very rich kind of data. The child can do various experiments with the world.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Have to deal with it. But I can deal with this as a driver, even though I did not encounter this in my driver ed class. And the reason I can deal with it is because I have all this general visual knowledge and expertise.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“In early childhood, and of course, reinforced through their growing up to age 16. So then at age 16, when they go into driver ed, what are they learning? They're not learning afresh the visual world. They have a mastery of the visual world. What they are learning is control. They are learning how to be smooth about control, about steering and brakes and so forth. They're learning a sense of typical traffic situations. Now that education process can be quite short because they are coming in as visual geniuses. And of course, in their future, they're going to encounter situations which are very novel, right? So during my driver ed class that I may not have had to deal with a skateboarder, I may not have had to deal with a truck driving in front of me where the back opens up and some junk gets dropped from the truck.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“And now at that point they learn, but at the age of 16, they are already visual geniuses. Because from zero to 16, they have built a certain repertoire of vision. In fact, most of it has probably been achieved by age two. In this period of age up to age two, they know that the world is three dimensional. They know how objects look like from different perspectives. They know about occlusion. They know about common dynamics of humans and other bodies. They have some notion of intuitive physics. So they have built that up from their observations and interactions.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“That's a scientific hypothesis as to which way is it going to go? I will tell you what I would bet on. And this is my general philosophical position on how these learning systems have been. What we have found currently very effective in computer vision in the deep learning paradigm is sort of tabular Asa learning and tabular RASA learning in a supervised way with lots and lots of examples. In the sense that blank slate. We just have the system which is Given a series of experiences in this setting, and then it learns there. Now, let's think about human driving. It is not tabular assay learning. So at the age of 16 in high school, Teenager goes into driver ed class.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“What this skateboarder was going to do. Okay. And because it really required that higher level cognitive understanding of what skateboarders typically do as opposed to a normal pedestrian. So what might have been the correct behavior for a pedestrian? A typical behavior for a pedestrian was not the typical behavior for a skateboarder. And so therefore, to do a good job there, you need to have enough data where you have pedestrians, you also have skateboarders, you've seen enough skateboarders to see what kinds of patterns or behavior they have. So it is. In principle, with enough data, that problem could be solved. But I think our current systems, computer vision systems, they need far, far more data than humans do for learning those same capabilities”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I think we have to build predictive models of behaviors, of people. And those can get quite complicated. So, I mean, I've seen examples of this in actually, I mean, I own a Tesla and it has various safety features built in. And what I see are these examples where let's say there is some skateboarder. I mean, and I don't want to be too critical because obviously these systems are always being improved and any specific criticism I have maybe the system six months from now will not have that particular failure mode. So it had the wrong response and it's because it couldn't predict”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I include that in vision because to me, perception blends into cognition and building predictive models of other agents in the world, which could be other agents, could be people, other agents could be other cars. That is part of the task of perception. Because perception always has to not tell us what is now, but what will happen. Because what's now is boring, it's done, it's over with. If we care about the future because we act in the future.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, the way I think about it is that there are certainly subsets of the visual based driving problem which are quite solvable. So for example, driving in freeway conditions. Is quite a solvable problem. I think there were demonstrations of that going back to the 1980s by someone called Ernst Tikmans in Munich in the 90s there were approaches from Carnegie Mellon. There were approaches from our team at Berkeley in the 2000s. There were approaches from Stanford and so on. Autonomous driving in certain settings is very doable. The challenge is to have an autopilot work under all kinds of driving conditions. At that point, it's not just a question of vision or perception, but really also of control and dealing with all the edge cases.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Very tolerant of errors there, right? I mean, when Google Image Search gives you some images back and a few of them are wrong, it's okay. It doesn't hurt anybody. There's not a matter of life and death, but making mistakes when you are driving at 60 miles per hour and you could potentially kill somebody is much more important.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I think the higher levels of the problem, the cognitive levels of the problem, are there in many real applications we have to confront them. Now, how much that is necessary will depend on the application. For some problems, it doesn't matter. For some problems, it matters a lot. I am, for example, a pessimist on fully autonomous driving in the near future. And the reason is because I think there will be that 0.01% of the cases where quite sophisticated cognitive reasoning is called for. However, there are tasks where you can, first of all, they are much more robust. So in the sense that error rates, error is not so much of a problem. For example, let's say you're doing image search. You're trying to get images based on some visual description.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“So, Vision operates at all levels, and there are parts which can be solved with what we could call maybe peripheral processing. So in the human vision literature, there used to be these terms sensation, perception, and cognition, which roughly speaking referred to like the front end of processing, middle stages of processing and higher level of processing. And I think they made a big deal out of this and they wanted to study only perception and then dismiss certain problems as being cognitive. But really, I think these are artificial divides. The problem is continuous at all levels and there are challenges at all levels. The techniques that we have today, they work better at the lower and mid levels of the problem.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I think in the early days it could have been excused because in the early days all aspects of AI were regarded as too easy. But I think today it is much less excusable. And I think why people fall for this is because of what I call the policy of the successful first step. Are many problems in vision where Getting 50% of the solution, you can get in one minute. Getting to 90% can take you a day, getting to 99% may take you five years. And 99.99% may be not in your lifetime.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“I think vision appears to be easy because most of what visual processing is subconscious or unconscious. So we underestimate the difficulty. Whereas when you are Like proving a mathematical theorem or playing chess, the difficulty is much more evident because it is your conscious brain which is processing various aspects of the problem solving behavior. Whereas envision, all this is happening, but it's not in your awareness, it's in operating below that”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source
“Human vision. So that gives us that effortlessness, gives us the sense that, oh, this must be very easy to implement on a computer. This is why the early researchers in AI got it so wrong. However, if you go into neuroscience or psychology of human vision, then the complexity becomes very clear. The fact is that a very large part of the cerebral cortex is devoted to visual processing. And this is true in other primates as well. So once we looked at it from a neuroscience or psychology perspective, it becomes quite clear that the problem is very challenging and it will take some time.”
2020-07-21 · Lex Fridman Podcast · #110 – Jitendra Malik: Computer Vision · IDENTIFIED FROM THE TRANSCRIPT · source