YouSaid · the spoken record

Jeremy Howard

lines on the record
135
first
2019-08-27
most recent
2019-08-27
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. So, yeah, model looking at data, particularly from the lens of which parts of the data the model says is important, is super important.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So, to me, one of the really cool things about machine learning models in general is that you can, when you interpret them, they tell you about things like what are the most important features, which groups are misclassifying, and they help you become a domain expert more quickly because you can focus your time on the bits that the model is telling you is important So it lets you deal with things like data leakage. For example, if it says, oh, the main feature I'm looking at is customer ID. When you're like, oh, customer ID shouldn't be predictive. And then you can talk to the people that manage customer IDs and they'll tell you like, oh, yes, as soon as a customer is application is accepted, we add a one on the end of their customer ID or something.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Yeah, I mean, it's a key part of our course is like before we train a model in the course, we see how to look at the data. And then the first thing we do after we train our first model, which we fine-tune an image net model for five minutes. And then the thing we immediately do after that is we learn how to analyze the results of the model by looking at examples of misclassified images and looking at a classification matrix and then doing research on Google to learn about the kinds of things that it's misclassifying.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Cheaper and also with less hyperparameters to set, it means you don't need deep learning experts to train your deep learning model for you, which means that domain experts can do more of the work, which means that now you can focus the human time on the kind of interpretation, the data gathering, identifying model errors and stuff like that

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Hopefully, the input of the human expert will be almost entirely unneeded from the deep learning point of view. So again, Google's approach to this is to try and use thousands of times more compute to run lots and lots of models at the same time and hope that one of them's good. A lot of mail kind of stuff, which I think is insane. When you better understand the mechanics of how models learn, you don't have to try a thousand different models to find which one happens to work the best. You can just jump straight to the best one, which means that it's more accessible in terms of compute.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  6. So we do something we call discriminative learning rates, which is really important, particularly for transfer learning. So really, I think in the last 12 months, a lot of people have realized that all this stuff is important. There's been a lot of great work coming out. And we're starting to see algorithms appear which have very, very few dials, if any, that you have to touch. So I think what's going to happen is the idea of a learning rate, it almost already has disappeared in the latest research. And instead, it's just like, you know, we know enough about how to Interpret the gradients and the change of gradients we see to know how to set every parameter. There you can all wait.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Well, there's been a lot of great work in the last 12 months in this area, and people are increasingly realizing that we just have no idea really how Optimizers work. And the combination of weight decay, which is how we regularize optimizers and the learning rate, and then other things like the epsilon we use in the atom optimizer, they all work together in weird ways. And different parts of the model, this is another thing we've done a lot of work on, is research into how different parts of the model should be trained at different rates in different ways

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  8. So I've been kind of studying that ever since. And eventually Leslie kind of figured out a lot of how to get this done. And we added minor tweaks. And a big part of the trick is starting at a very low learning rate, very gradually increasing it. So as you're training your model, you take very small steps at the start and you gradually make them bigger and bigger until eventually you're taking much bigger steps than anybody thought was possible. A few other little tricks to make it work, but basically we can reliably get super convergence, and so for the dawn bench thing, we were using just much higher learning rates than people expected to work.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Every other scientific field I know of works that way. I don't know why ours is uniquely disinterested in publishing unexplained experimental results, but there it is. So it wasn't published. Having said that, Read a lot more unpublished papers and published papers because that's where you find the interesting insights. So I absolutely read this paper. And I was just like, this is. Astonishingly mind blowing and weird and awesome, and like, why isn't everybody only talking about this? Because if you can train these things 10 times faster, they also generalize better because you're doing less epochs, which means you look at the data less, so you get better accuracy.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Because it's not an area of kind of active research in the academic world. No academics recognize that this is important. And also deep learning in academia is not considered an experimental science. So unlike in physics, where you could say like, I just saw a subatomic particle do something which the theory doesn't explain Could publish that without an explanation. And then in the next 60 years, people can try to work. It's literally impossible for Leslie to publish a paper that says, I've just seen something amazing happen. This thing trained 10 times faster than it should have. I don't know why. And so the reviewers were like, well, you can't publish that because you don't know why. So anyway.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Interesting. Yeah, so this is all work that came from a guy called Leslie Smith. Leslie's a researcher who, like us, cares a lot about. The practicalities of Training neural networks quickly and accurately, which you would think is what everybody should care about, but almost nobody does. And he discovered something very interesting, which he calls super convergence, which is there are certain networks that with certain settings of hyperparameters could suddenly be trained 10 times faster. Using a 10 times higher learning rate. No one published that paper.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  12. But I mean, you can still do it even with one. Again, it's not much work's been done in this area, so we're actually going to be releasing an audio library soon, which hopefully will encourage development of this because it's so underused. The basic approach we used for our super resolution, in which Jason uses for de-oldify of generating high quality images, the exact same approach would work for audio. No one's done it yet, but it would be a couple of months work

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Yeah, basically everybody now is doing most of the fanciest stuff on their phones with computational photography and also increasingly people are putting more than one lens on the back of the camera. So the same will happen for audio for sure.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Photography is basically standard now. So the Google Pixel Night Light, I don't know if you've ever tried it, but it's astonishing you take a picture in almost pitch black and you get back a Very high quality image, and it's not because of the lens. Same stuff with adding the Bocke to the background blurring done computationally.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  15. I mean, it's eminently doable and it should have been done by now. I felt the same way about computational photography four years ago. Why are we investing in big lenses when Three cheap lenses plus actually a little bit of intentional movement, so like take a few frames gives you enough information to get excellent sub pixel resolution, which particularly with deep learning, you would know exactly what you meant to be looking at. We can totally do the same thing with audio. I think madness that it hasn't been done yet.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Yeah, so we've found loss functions that work super well without the adversarial part. And then one of our students, a guy called Jason Antich, has created a system called Deldify, which uses this technique to colorize old black and white movies. You can do it on a single GPU, colorize a whole movie in a couple of hours. And one of the things that Jason and I did together was we figured out how to Had a little bit of GAN at the very end, which it turns out for colorization makes it just a bit brighter and nicer. And then JSON did masses of experiments to figure out exactly how much to do, but it's still all done on his whole machine on a single GPU in his lounge room. And like if you think about like... Colorizing Hollywood movies. That sounds like something a huge studio would have to do, but he has the world's best results on this.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  17. And we've actually recently shown that you don't even need GANS. So we've developed GAN level outcomes without needing GANS, and we can now do it with, again, by using transfer learning, we can do it in a couple of hours on a single GPU.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Well, let's put it this way none of the big breakthroughs of the last 20 years have required multiple GPUs. So, like batch norm, relue, dropout.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  19. And even worse than that, Lex, I keep hearing from people who say, I decided not to get into deep learning because I don't believe it's accessible to people outside of Google to do useful work. So like I see a lot of people make an explicit decision to not learn this incredibly valuable tool because they've drunk the Google Kool-Aid, which is that only Google's big enough and smart enough to do it. And I just find that's so disappointing and it's so wrong.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  20. A hard one, and I've discovered that if you just look at these two subsets, you can train things on a single GPU in 10 minutes. And the results you get are directly transferable to ImageNet nearly all the time. And so now I'm starting to see some researchers start to use these.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Yeah, particularly multi machines because it's just clunky. Multi GPUs is less clunky than it used to be, but to me, anything that slows down your iteration speed is a waste of time. So you could maybe do your very last Perfecting of the model on multi-GPUs if you need to. So for example, I think doing stuff on ImageNet is generally a waste of time. Why test things on 1.3 million images? Most of us don't use 1.3 million images. And we've also done research that shows that doing things on a smaller subset of images gives you the same relative answers anyway. So from a research point of view, why waste that time? So actually I released a couple of new data sets recently. One is called ImageNet. The French ImageNet, which is a small subset of ImageNet which is designed to be easy to classify.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Doing a distributed version over multiple machines a couple of months later and ended up at the top of the leaderboard. We had 18 minutes.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Yeah, 93%, they picked a good threshold. It was a little bit higher than what the most commonly used ResNet 50 model could achieve at that time. So, yeah, so it's quite a difficult problem to solve. But yeah, we realized if we actually just use 64 by 64 images. It trained a pretty good model And then we could take that same model and just give it a couple of epochs to learn 224 by 224 images. And it was basically already trained, which makes a lot of sense. Like if you teach somebody, like, here's what a dog looks like and you show them low res versions and then you say here's a really clear picture of a dog. They already know what a dog looks like. So that just jumped to the front and we ended up winning. Parts of that competition, we actually ended up

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  24. So screwed, but we kind of thought we'll keep trying. You know, if Google can do it in, I mean, Google did on five hours on like a TPU pod or something, like a lot of hardware. But we kind of had a bunch of ideas to try a really simple thing was why are we using these big images? They're like 224, 256 by 256 pixels. Why don't we try smaller ones?

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Know a very small number of minutes. I can't remember exactly how many minutes it was, but it might have been like 10 minutes or something. And so, yeah, we found ourselves at the top of the leaderboard easily for both time and money, which really shocked me because the other people competing this were like Google and Intel and stuff were like know a lot more about this stuff than I think we do. So then we were emboldened, we thought Try the image net one too. I mean, it seemed way out of our league, but our goal was to get under 12 hours. And we did, which was really exciting, and but we didn't put anything up on the leaderboard, but we were down to like 10 hours. But then Google put in like five hours or something and we just like.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Always, but it's always for me, it's always how do I make it much faster on a single GPU that a normal person could afford in their day-to-day life? It's not how could I do it faster by. Having a huge data center because to me it's all about as many people should be able to use something as possible without Around with infrastructure. So, anyway, so in this case, it's like, well, we can use 8 GPUs just by renting an AWS machine. So we thought we'd try that. And yeah, basically using the stuff we were already doing, we were able to get the speed within a few days we had the speed down to

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  27. And yeah, I really never done that before. Like, I'd never really, like things like using more than one GPU at a time was something I tried to avoid because to me it's very against the whole idea of accessibility is to be able to do things with one GPU.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  28. As possible, too. That's also another one for as cheap as possible. And there's a couple of categories image net and so far 10. So image net's this big 1.3 million image thing that took a couple of days to train member of a friend of mine, Pete Warden, who's now at Google. I remember he told me how he trained ImageNet a few years ago and he basically had this All granny flat out the back that he turned into Wiz ImageNet Training Center and he figured, you know, after like a year of work, he figured out how to train it in like 10 days or something. That was a big job. Whereas Cyfar 10 at that time you could train in a few hours, you know, it's much smaller and easier. So we thought we'd try sci-fi 10.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Case study for the course, like it's all the stuff we're already doing. Why don't we just put together our current best practices and ideas? So me and I guess about four students just decided to give it a go and we focused on this small one called SciFar 10, which is little 32 by 32 pixel images.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  30. So it was during one of these times that somebody in the group said, Oh, there's a thing called Dawn Bench that looks interesting. And I was like, what the hell's that? And they said, oh, it's some competition to see how quickly you can train a model seems kind of not exactly relevant to what we're doing, but it sounds like the kind of thing which you might be interested in. I checked it out and I was like, oh, crap, there's only 10 days till it's over. It's pretty much too late. And we're kind of busy trying to teach this course We're like, it would make an interesting.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Sure. So something which I really enjoy is that I basically teach two courses a year, practical deep learning for coders, which is kind of the introductory course and then cutting edge deep learning for coders, which is the kind of research level course. And while I teach those courses, I basically have a big office at the University of San Francisco. It'd be big enough for like 30 people, and I invite anybody, any student who wants to come and hang out with me while I build the course. And so generally it's full. And so we have 20 or 30 people in a big office with nothing to do but study deep learning.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Actually, helped start a startup called Platform AI, which is really all about active learning. And yeah, it's been interesting trying to kind of What research is out there and make the most of it, and there's basically none, so we've had to do all our own research

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Although, I mean, the nice thing is nowadays everybody is now working on NLP transfer learning because since that time we've had GPT and GPT2 and BERT and it's like, so yeah. Show that something's possible everybody jumps in, I guess.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Junior researchers or like I don't care whether I get citations or papers or whatever. There's nothing in my life that makes that important, which is why I've never actually bothered to write a paper myself. But for people who do, I guess they have to pick the kind of Option, which is like, yeah, make a slight improvement on something that everybody's already working on.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  35. It actually wrote it for the course for the fast AI course. I wanted to teach people NLP and I thought I only want to teach people practical stuff. And I think the only practical stuff is transfer learning. And I couldn't find any examples of transfer learning in NLP. So I just did it. And I was shocked to find that soon as I did it, you know, the basic prototype took a couple of days. smashed the state of the art on one of the most important data sets in a field that I knew nothing about. And I just thought, well, this is ridiculous. And so I spoke to Sebastian about it and he kindly offered to write it up, the results. And so it ended up being published in ACL, which is the top link computational linguistics conference. So people do actually care once you do it, but I guess it's difficult for maybe like...

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Yeah, it's just like transfer learning, it's understudied and the academic world just has no reason to care about practical results. The funny thing is, like I've only really ever written one paper, I hate writing papers. And I didn't even write it. It was my colleague Sebastian Ruder who actually wrote it. The research for it, but it was basically introducing transfer learning, successful transfer learning to NLP for the first time. The algorithm is called ULMF.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Everybody kind of reinvents active learning when they actually have to work in practice because they start labeling things and they think, gosh, this is taking a long time and it's very expensive and then they start thinking. Why am I labeling everything? I'm only, the machine's only making mistakes on those two classes. They're the hard ones. Maybe I'll just start labeling those two classes. And then you start thinking, well, why did I do that manually? Why can't I just get the system to tell me which things are going to be hardest? It's an obvious thing to do, but...

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  38. But almost nobody works on that. Or another example, active learning, which is the study of how do we get more out of the human beings. In the loop.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  39. What I was getting at, yeah. It's a problem in science in general. Scientists need to be published, which means they need to work on things that their peers are extremely familiar with and can recognize and advance in that area. So that means that they all need to work on the same thing. And so it really, and the thing they work on, there's nothing to encourage them to work on things that are practically useful. So you get just a whole lot of research, which is minor advances in stuff that's been very highly studied and has no Significant practical impact. Whereas the things that really make a difference, like I mentioned transfer learning, like if we can do better at transfer learning, then it's this world-changing thing where suddenly lots more people can do world-class work with less resources and less data.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Yeah. And like partly it's because Unlike most people in this field, my background is very applied and industrial, like my first job was at McKinsey Company. I spent 10 years in management consulting. I spent a lot of time with domain experts. So, I kind of respect them and appreciate them and know I know that's where the value generation in society is. And so I also know how most of them can't code and most of them don't have the time to invest three years in a graduate degree or whatever. So it's like, how do I? Upskill those domain experts. I think that would be a super powerful thing, you know, bigger societal impact I could have. Yeah, that was the thinking

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Solve, what kind of data do they have? Who has that data So I kind of felt like I need to approach this differently if I want to maximize the positive impact of deep learning rather than me picking an area and trying to become good at it and building something, I should let people who are already domain experts in those areas and who already have the data. It themselves That was the reason for fastai to basically try and figure out how to get deep learning into the hands of people who could benefit from it and help them to do so Quick and easy and effective way as possible.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  42. So before I started FastAI, I spent a year researching where are the biggest opportunities for. Learning because I knew from my time at Kaggle in particular that deep learning had kind of hit this threshold point where it was rapidly becoming the state of the art approach in every area that looked at it. And I'd been working with neural nets for over 20 years. I knew that from a theoretical point of view, once it hit that point, it would do that in kind of just about every domain. And so I kind of spent a year researching what are the domains. It's going to have the biggest low-hanging fruit and the shortest time period. I picked medicine, but there were so many I could have picked. And so there was a kind of level of frustration for me of like, okay, I'm really glad we've opened up the medical deep learning world and today it's huge, as you know, but we can't do, you know, I can't do everything. I don't even know, like in medicine it took me a really long time to even get a sense of what kind of problems do medical practitioners.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  43. But for something like medicine, yeah, I mean, the hospital has my medical imaging, my pathology studies, my medical records. And also I own my medical data. So you can, so I help a startup called DocAI. One of the things DocAI does is that it has an app you can connect to Sutterhealth and LabCore and Walgreens and download your medical data to your phone and then upload it again at your discretion to share it as you wish. So with that kind of approach we can share our medical information with the people we want to.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Saving me five minutes on answering some questions versus the negative externalities of the privacy issue doesn't add up. So I think a lot of the time the places where people are Invading our privacy in order to provide convenience is really about just trying to make them more money and they move these negative externalities to places that they don't have to pay for them. So when you actually see Regulations appear that actually cause the companies that create these negative externalities to have to pay for it themselves. They say, well, we can't do it anymore. So the cost is actually too high.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  45. On the whole, I'd say yes. I mean, there are. Things like Recommender systems have this cold start problem where Jeremy is a new customer. We haven't seen him before, so we can't recommend him things based on what else he's bought and liked with us. And there's various workarounds to that. A lot of music programs will start out by saying which of these artists you like, which of these albums do you like, which of these songs do you like? Netflix used to do that nowadays, they tend not to. People kind of don't like that because they think, oh, we don't want to bother the user. So you could work around that by having some kind of data sharing where you get my marketing record from Axiom or whatever. To guess from that. To me, The benefit to me and to society of

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Lot of people are looking for ways to share data and aggregate data, but I think often that's unnecessary. They assume that they need more data than they do because they're not familiar with the basics of transfer learning, which is this critical technique for needing orders of magnitude less data.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  47. And Google's very upfront about this, like Jeff Dean has gone out there and given talks and said our goal is to require a thousand times more computation, but less people. Our goal is to use the people that you have better and the data you have better and the computation you have better. So one of the things that we've discovered is, or at least highlighted, is that you very, very, very often don't need much data at all. And so the data you already have in your organization will be enough to get data the art results. My starting point would be to kind of say around privacy is

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  48. One of my areas of focus is on doing more with less data. So most vendors, unfortunately, are strongly incented to find ways to require more data and more computation. So Google and IBM being the most obvious So, Wat So Google and IBM both strongly push the idea that you have to be more data and more computation and more intelligent people than anybody else. And so you have to trust them to do things because nobody else can do it.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Right. Well, so it also saves lives in a very abstract way, which is like, oh, we've been able to release these 100,000 anonymized records. I can't point at the specific person whose life that saved. I can say like, oh, we ended up with this paper which found this result, which all diagnosed a thousand more people than we would have otherwise, but it's like which ones were helped. It's very abstract.

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Yeah, I mean, funnily enough, the problem is less the regulation and more the interpretation of that regulation by lawyers in hospitals. So HIPAA is actually... Designed to P and HIPAA does not stand for privacy, it stands for portability. It's actually meant to be a way that data can be used. And it was created with lots of gray areas because the idea is that would be more practical and it would help people to use this. Legislation to actually share data in a more thoughtful way. Unfortunately, it's done the opposite because when a lawyer sees a gray area, they see, oh, if we don't know we won't get sued, then we can't do it. HIPAA Not exactly the problem. The problem is more that. Hospital lawyers are not incented to Make bold decisions about data portability

    2019-08-27 · Lex Fridman Podcast · Jeremy Howard: fast.ai Deep Learning Courses and Research · IDENTIFIED FROM THE TRANSCRIPT · source