YouSaid · the spoken record

Michael Kearns

lines on the record
116
first
2019-11-19
most recent
2019-11-19
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. And once, you know, I kind of made that decision on that particular day at that particular moment in Boston Common, I'm glad I made that decision.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Trying to think about some problems with very little success. But I knew that I hadn't really tried to do the thing that I knew I'd come to do. And so I thought, you know, I'm going to stick through it for the summer and that was very formative because I went from kind of contemplating quitting to a year later it being very clear to me I was going to finish because I still had a ways to go but I kind of started doing research. It was going well. It was really interesting and it was sort of a complete transformation. You know, it's just that transition that I think every doctoral student makes at some point, which is to sort of go from being like a student of what's been done before to doing your own thing and figuring out what makes you interested and what your strengths and weaknesses are as a researcher.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  3. I could go back to the Bay Area and to California. And this was from one of the early periods where there was like you could definitely get a relatively good job paying job at one of the tech companies back the big tech companies back then. And so I distinctly remember kind of a late spring day when I was kind of sitting in Boston Common and kind of really just kind of chewing over what I wanted to do in my life. And then I realized like, okay, and I think this is where my academic background helped me a great deal. I sort of realized, you know, yeah, you're not having a great time right now. This feels really narrowing, but you know that you're here for research eventually and to do something original and to try to carve out a career where you kind of choose what you want to think about and have a great deal of independence. And so, you know, at that point, I really didn't have any real research experience yet.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah, they're kind of all over the map. And I was a grad student here just up the river at Harvard and came to study with Les Valley. And I was unhappy because at Berkeley, as an undergraduate, yeah, I studied a lot of math and computer science, but it was a huge school, first of all. And I took a lot of other courses, as we've discussed, I started as an English major and took history courses and art history classes and had friends that did all kinds of different things. And Harvard's a much smaller institution than Berkeley, and it's computer science department, especially at that time, was a much smaller place than it is now. And I suddenly just felt very I'd gone from this very big world to this highly specialized world. And now all of the classes I was taking were computer science classes. And I was only in classes with math and computer science people. And so I was, you know, I thought often in that first year of grad school about whether I really wanted to stick with it or not. And, you know, I thought like, oh, I could, you know, stop with a master's.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  5. I'll answer a slightly different question, which is like what's a day in my life or my career that was kind of a watershed moment. I went straight from undergrad to doctoral studies, and that's not at all atypical. And I'm also from an academic family. Like my dad was a professor, my uncle on his side is a professor. Both my grandfathers were professors. Okay.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  6. We jointly sponsored a workshop at Penn with the Federal Reserve Bank of Philadelphia a little more than a year ago on, you know, I think the title was something like Machine Learning for Macroeconomic Prediction, macroeconomic referring specifically to these longer timescales. And it was an interesting conference, but it left me with greater confidence that you have a long way to go to, you know, and so I think that people that in the grand scheme of things, if somebody asked me like, well, whose job on Wall Street is safe from the bots? I think people that are at that longer time scale and have that appetite for all the risks involved in long-term investing and that really need kind of not just algorithms that can optimize from data, but they need views on stuff. They need views on the political landscape, economics.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I really, my main source of data is just the data from the exchanges themselves about the activity in the exchanges, right? And maybe I need to pay, you know, I need to keep an eye on the news because that can cause sudden CEO gets caught in a scandal or gets run over by a bus or something that can cause very sudden changes. But I don't need to understand economic cycles. I don't need to understand recessions. I don't need to worry about the political situation or war breaking out in this part of the world because all I need to know is as long as that's not going to happen in the next 500 milliseconds, then my model's good. When you get to these longer timescales, you really have to worry about that kind of stuff. And people in the machine learning community are starting to think about this.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Human nature at a level. Yeah, and you need to just be able to ingest many, many more sources of data that are on wildly different time scales. So if I'm an HFT, my high frequency trader, like I don't, I don't.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Apple really should be. And he doesn't even look at what Apple's doing today. He just decides, I think that this was what its long-term value is and it's far from that right now. And so I'm going to buy some Apple or, you know, short some Apple and I'm going to sit on that for 10 or 20 years. So, when you're at that kind of time scale, or even more than just a few days, All kinds of other sources of risk and information. So now you're talking about holding things through recessions and economic cycles. Wars can break out.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Yeah. So I think we are in an era where clearly there have been some very successful quant hedge funds that are in what we would traditionally call still in the Stat R regime. Stadar referring to statistical arbitrage, but for the purposes of this conversation, what it really means is making directional predictions in asset price movement or returns, your prediction about that directional movement is good for, you know, you have a view that it's valid for some period of time between a few seconds and a few days. And that's the amount of time that you're going to kind of get into the position, hold it, and then hopefully be right about the directional movement and buy low and sell high as the cliche goes. So that is kind of a sweet spot, I think, for quant trading and investing right now and has been for some time. When you really get to kind of more Warren Buffett style timescales, right? Like my cartoon of Warren Buffett is that, you know, Warren Buffett sits and thinks what the long-term value of app.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  11. That hurts your execution. So this is the kind of, you know, this is an optimization problem. This is a control problem. And so Machines are better. We know how to design algorithms that are better at that kind of thing than a person is going to be able to do because we can take volumes of historical and real-time data to kind of optimize the schedule with which we trade. And, you know, similarly high frequency trading, which is closely related, but not the same as optimized execution, where you're just trying to spot very, very temporary mispricings between exchanges or within an asset itself or just predict directional movement of a stock because of the kind of very, very low level granular buying and selling data in the exchange. Machines are good at this kind of stuff.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Yeah. And so, you know, I think the places where algorithmic trading have had the greatest inroads and had the first inroads were in kind of execution problems, kind of optimized execution problems. So what I mean by that is at a large brokerage firm, for example, one of the lines of business might be on behalf of large institutional clients taking what we might consider difficult trade. So it's not like a mom and pop investor saying, I want to buy 100 shares of Microsoft. It's a large hedge fund saying, you know, I want to buy a very, very large stake in Apple and I want to do it over the span of a day. And it's such a large volume that if you're not clever about how you break that trade up, not just over time, but over perhaps multiple different electronic exchanges that all let you trade Apple on their platform, you will move, you'll push prices around in a way.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  13. It's a good question. I mean, in the time I've spent on Wall Street and in finance, I've seen a clear progression, and I think it's a progression that kind of models the use of algorithms and automation more generally in society, which is the things that kind of get taken over by the algos first are sort of the things that So, first of all, there needed to be this era of automation where just financial exchanges became largely electronic, which then enabled the possibility of trading becoming more algorithmic because once exchanges are electronic and an algorithm can submit an order through an API just as well as a human can do at a monitor.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Yeah. Even though I kind of laugh at those efforts, they were more sensible than they would be now, right? Because there were sort of only two nuclear powers at the time and you didn't have to worry about deterring new entrants and who was developing the capacity. And so we have many, you know, we have, it's definitely a game with more players now and more potential entrants. I'm not in general somebody who advocates using kind of simple mathematical models when the stakes are as high as things like that and the complexities are very political and social, but we are still here.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Just to prime your viewers a little bit. I mean, I think you're referring to the fact that game theory was taken quite seriously back in the 60s as a tool for reasoning about kind of Soviet U.S. nuclear armament disarmamentive detente, things like that. I'll be honest, as huge of a fan as I am of game theory and its kind of rich history, it still surprises me that you had people at the RAND Corporation back in those days kind of drawing up two by two tables and one, the row player is the US and the column player is Russia and that they were taking seriously I'm sure if I was there maybe it wouldn't have seemed as naive as it does at the time

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  16. But similarly, on social media platforms or on Amazon, all these algorithms that are essentially trying to optimize our behalf, they're driving us in a colloquial sense towards some kind of competitive equilibrium. And one of the most important lessons of game theory is that just because we're at equilibrium doesn't mean that there's not a solution in which some or maybe even all of us might be better off. And then the connection to machine learning, of course, is that all these platforms I've mentioned, the optimization that they're doing on our behalf is driven by machine learning, like predicting where the traffic will be, predicting what products I'm going to like, predicting what would make me happy in my newsfeed.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Now, you might ask, like, well, that sounds great. Why is that a bad thing? Well, you know, it's known both in theory and With some limited studies from actual traffic data. Of us being in this competitive equilibrium might cause our collective driving time to be higher, maybe significantly higher, than it would be under other solutions. And then you have to talk about what those other solutions might be and what the algorithms to implement them are, which we do discuss in the kind of game theory chapter of the book.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  18. The car, which literally just told you what roads were available. And then you might have like half-hourly traffic reports just about the major freeways, but not about side roads. So you were pretty much on your own. And now we've got these apps. You pull it out and you say, I want to go from point A to point B. And in response kind of to what everybody else is doing, if you like what all the other players in this game are doing right now, here's the route that minimizes your driving time. So it is really kind of computing a selfish best response for each of us in response to what all of the rest of us are doing at any given moment. And so I think it's quite fair to think of these apps as driving or nudging us all towards the competitive or Nash equilibrium of that game.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  19. I mean, I think we've kind of talked about these ideas already in kind of a non-technical way, which is maybe the more interesting way of understanding them first, which is, you know, we have many systems, platforms, and apps these days that work really hard to use our data and the data of everybody else on the platform to selfishly optimize on behalf of each user. So, you know, let me give, I think, the cleanest example, which is just driving apps, navigation apps like Google Maps and Ways, where miraculously compared to when I was growing up at least, the objective would be the same when you wanted to drive from point A to point B, spend the least time driving, not necessarily minimize the distance, but minimize the time, right? And when I was growing up, like the only resources you had to do that were maps and

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Yeah. Maybe answering a slightly less personally than you asked the question. I think within the field of algorithmic game theory, perhaps the single most important kind of technical contribution that's been made is the realization between close connections between machine learning and game theory and in particular between game theory and the branch of machine learning that's known as no-regret learning. And this sort of provides a very general framework in which a bunch of players interacting in a game or a system, each one kind of doing something that's in their self-interest will actually kind of reach an equilibrium and actually reach an equilibrium in a rather short amount of steps.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Not that it will necessarily emerge, just that it's possible, right? Like the existence of equilibrium doesn't mean that sort of natural iterative behavior will necessarily lead to it.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  22. There's a lot of them. I'm a big fan of the field. Mean, you know, technical answers to that, of course, would include Nash's work just establishing that there's a competitive equilibrium under very, very general circumstances, which in many ways kind of put the field on a firm conceptual footing because if you don't have equilibriates kind of hard to ever reason about what might happen

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  23. So you need at least two people to get started in game theory. And many people are probably familiar with prisoner's dilemma as kind of a classic example of game theory and a classic example where everybody looking out for their own individual interests leads to a collective outcome that's kind of worse for everybody than what might be possible if they cooperated, for example. But cooperation is not an equilibrium in prisoners' dilemma. And so my work and the field of algorithmic game theory more generally in these areas kind of looks at settings in which the number of actors is potentially extraordinarily large and their incentives might be quite complicated and kind of hard to model directly, but you still want kind of algorithmic ways of predicting what will happen. Or influencing what will happen in the design of platforms.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Yeah. Game theory, of course. Let us give credit where it's due. They don't come from the economist first and foremost. But as I've mentioned before, computer scientists never hesitate to wander into other people's turf. And so there is now this 20-year-old field called algorithmic game theory. But game theory first and foremost is a mathematical framework for reasoning about collective outcomes in systems of interacting individuals.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  25. And this is not a weird idea, right? Because there are markets for data already. It's just that consumers are not participants in them. There's like, you know, there's sort of publishers and content providers on one side that have inventory and then they're advertised on the others. And Google and Facebook are running their entire revenue stream is by running two-sided markets between those parties, right? And so it's not a crazy idea that there would be like a three-sided market or that on one side of the market or the other we would have proxies representing our interests. It's not a crazy idea, but it would crazy technical idea, but it would have Pretty extreme economic consequences.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Yeah, I mean, the other thing that has to be said, right, is that it's a cliche, but the users of many systems, platforms, and apps, we are the product. We are not the customer. The customer are advertisers, and our data is the product. So it's one thing to kind of suggest more individual control of data and privacy and uses, but this, you know, If this happens in sufficient degree, it will upend the entire economic model that has supported the Internet to date. And so some other economic model will have to be replace it.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Again, privacy guarantees are few far between and weak, and users have very, very little control. And I'm optimistic that we'll land in something that provides better privacy overall and more individual control of data and privacy. But I think to get there, it's, again, just like fairness, it's not going to be enough to propose algorithmic solutions. There's going to have to be a whole kind of regulatory legal process that prods companies and other parties to kind of adopt solutions.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Bit if I'm trying to build a predictive model for some rare disease and I'm trying to use machine learning to do it, it's easy to get negative examples because the disease is rare. But I really want to have lots of people with the disease in my data set. And so somehow those people's data with respect to this application is much more valuable to me than just the background population. And so maybe they should be compensated more for it. And so, you know, I think these are kind of very, very fledgling conceptual questions that maybe will have kind of technical thought on them sometime in the coming years. But I do think we'll, you know, to kind of give more directly answer your question, I think I'm optimistic at this point from what I've seen that we will land at some better compromise than we're at right now, where

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  29. So I'm optimistic about the possibility of balancing the desire for individual privacy and individual control of privacy with kind of societally and commercially beneficial uses of data, not unrelated to differential privacy or suggestions that say like, well, individuals should have control of their data, they should be able to limit the uses of that data, they should even, you know, there's fledgling discussions going on in research circles about allowing people selective use of their data and being compensated for it. And then you get to sort of very interesting economic questions like pricing, right? And one interesting idea is that maybe differential privacy would also be a conceptual framework in which you could talk about the relative value of different people's data. To demystify this a little bit.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Privacy. It's also, I was actually talking to somebody at one of the large tech companies recently about the fact that just this kind of thing that there are sometimes when the response to my data needs to be very specific to my data, right? Like I type mountain biking into Google. I want results on mountain biking and I really want Google to know that I typed in mountain biking. I don't want noise added to that. And so I think there's sort of maybe even interesting technical questions around notions of privacy that are appropriate where it's not that my data is part of some aggregate like medical records and that we're trying to discover important correlations and facts about the world at large, but rather there's a service that I really want to pay attention to my specific data, yet I still want some kind of privacy guarantee. And I think these kind of obfuscation ideas.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Which they basically built a browser plugin that Tried to essentially obfuscate your Google searches. So to the extent that you're worried that Google is using your searches to build predictive models about you to decide what ads to show you, which they might very reasonably want to do. But if you object to that, they built this widget that you could plug in. And basically whenever you put in a query into Google, it would send that query to Google. But in the background, all of the time from your browser, it would just be sending this torrent of irrelevant queries to the search engine. So, you know, it's like a weed and chaff thing. So, you know, out of every thousand queries, let's say that Google was receiving from your browser, one of them was one that you put in, but the other 999 were not. It's the same kind of idea kind of, you know, privacy by obfuscation. So I think that's an interesting idea. It doesn't give you different...

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  32. No, it's not a stretch. I've not seen That is not a technique that, to my knowledge, will provide differential privacy. But to give an example, one very specific example about what you're discussing is there was a very interesting project at NYU, I think led by Helen Nissenbaum there.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Propagation in neural networks, you know, CART for decision trees, support vector machines boosting, you name it, as well as classic hypothesis testing and the lichen statistics. None of those algorithms are differentially private in their original form. All of them have modifications that add noise to the computation in different places and different ways that achieve differential privacy. So this really means that to the extent that we've become a scientific community very dependent on the use of machine learning and statistical modeling and data analysis, we really do have a path to kind of provide privacy guarantees to those methods. And so we can still enjoy the benefits of kind of the data science era while providing rather robust privacy guarantees to individuals.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yeah, but now it's pretty mature. But I must admit the first time I saw the definition of deferential privacy, my reaction was like, wow, that is a clever definition and it's really making very strong promises. And I first saw the definition in much earlier days. And my first reaction was like, well, my worry about this definition would be that it's a great definition of privacy, but that it'll be so restrictive that we won't really be able to use it. We won't be able to compute many things in a differentially private way. So that's one of the great successes of the field, I think, is in showing that the opposite is true and that most things that we know how to compute absent any privacy considerations can be computed in a differentially private way. So for example, pretty much all of statistics and machine learning can be done differentially privately.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Yeah, so I'm a relatively recent member of the Differential Privacy Community. My co-author, Aaron Roth, is really one of the founders of the field and has done a great deal of work. And I've learned a tremendous amount working with him on it. It's a pretty grown-grown.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  36. That average to numerical precisions, and then you'd add some noise to it, right? You'd add some kind of zero mean Gaussian or exponential noise to it so that the actual value you output is not the exact mean, but it'll be close to the mean. But the noise that you add will sort of prove that nobody can kind of reverse engineer Any particular value that went into the average?

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah, so it's basically by adding noise to computations. So the basic idea is that every differentially private algorithm, first of all, or every good differentially private algorithm, every useful one, is a probabilistic algorithm. So it doesn't on a given input, if you gave the algorithm the same input multiple times, it would give different outputs each time from some distribution. And the way you achieve differential privacy algorithmically is by kind of carefully and tastefully adding noise to a computation in the right places. And to give a very concrete example, if I want to compute the average of a set of numbers, right, the non-private way of doing that is to take those numbers and average them and release like a numerically precise value for the average. In differential privacy, you wouldn't do that. You would first compute

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  38. But the point is that if the same analysis has been done with all the other n-1 medical records and just years missing, the outcome would have been the same. Your data was an idiosyncratically crucial to establishing the link between smoking and lung cancer because the link between smoking and lung cancer is like a fact about the world that can be discovered with any sufficiently large

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Done without your medical record included. So, in other words, this doesn't say that bad things cannot happen to you as a result of data analysis. It just says that these bad things were going to happen to you already, even if your data wasn't included. And to give a very concrete example, right, we discussed at some length the study that in the 50s that was done that created the link between smoking and lung cancer. And we make the point that, well, if your data was used in that analysis and the world kind of knew that you were a smoker because there was no stigma associated with smoking before those findings, real harm might have come to you as a result of that study that your data was included in, in particular your insurer now might have a higher posterior belief that you might have lung cancer and raise your premium. So you've suffered economic

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  40. One is when I do this, I build this model on the database of medical records, including your medical record. And the other one is where I do the same exercise with the same database with just your medical record removed. So basically, you know, it's two databases, one with N records in it and one with n-1 records in it. The N-1 records are the same and the only one that's missing in the second case is your medical record. So differential privacy basically says that any harms that might come to you from the analysis in which your data was included are essentially nearly identical to the harms that would have come to you if the same analysis had done been

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Yeah, so differential privacy basically is an alternate, much stronger notion of privacy than these anonymization ideas. And it's a technical definition, but the spirit of it is we compare two alternate worlds. Let's suppose I'm a researcher and I want to do, you know, there's a database of medical records and one of them is yours. And I want to use that database of medical records to build a predictive model for some disease. So based on people's symptoms and test results and the like, I want to build a model predicting the probability that people have disease. So, you know, this is the type of scientific research that we would like to be allowed to continue. And in differential privacy, you ask a very particular counterfactual question. We basically compare two alternatives.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  42. The things on which you clicked on the thumbs up button on the platform, not using any information, demographic information, nothing about who your friends are, just knowing the content that you had liked. Was enough to, in the aggregate, accurately predict things like sexual orientation, drug and alcohol use, whether you were the child of divorced parents. So we live in this era where even the apparently irrelevant data that we offer about ourselves on public platforms and forums often unbeknownst to us more or less acts as a signature or fingerprint. And that if you can kind of do a join between that kind of data and allegedly anonymized data, you have real trouble.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Where lots of Netflix users publicly rate their movie preferences. And so the anonymized data in Netflix, it's just this phenomenon, I think, that we've all come to realize in the last decade or so is that just knowing a few apparently irrelevant innocuous things about you can often act as a fingerprint. Like if I know what rating you gave to these 10 movies and the date on which you entered these movies, this is almost like a fingerprint for you in the sea of all Netflix users. There were just another paper on this in science or nature about a month ago that kind of 18 attributes. I mean, my favorite example of this was actually a paper from several years ago now where it was shown that just from your likes on Facebook, just from the

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  44. So, this is exactly the problem with these notions that notions of anonymization, removing personally identifiable information, the kind of fundamental conceptual flaw is that these definitions kind of pretend as if the data set in question is the only data set that exists in the world or that ever will exist in the future. And of course, things like the Netflix prize and many, many other examples since the Netflix Prize, I think that was one of the earliest ones, though You can re-identify people that were anonymized in the data set by taking that anonymized data set and combining it with other allegedly anonymized data sets and maybe publicly available information about you.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  45. So, by anonymization, I'm kind of referring to techniques like I have a database, the rows of that database are, let's say, individual people's medical records. I want to let people use that. Sensitive information about specific people's medical records. So anonymization broadly refers to the set of techniques where I say like, okay, I'm first going to delete the column with people's names. I'm going to not put, you know, so that would be like a redaction, right? I'm just redacting that information. I am going to take ages and I'm not going to say your exact age. I'm going to say whether you're zero to 10, 10 to 20, 20 to 30. I might put the first three digits of your zip code, but not the last two, et cetera, et cetera. And so the idea is that through some series of operations like this on the data, I anonymize it another term of art that's used is removing personally identifiable information. And this is basically the most common way of providing data privacy, but that's in a way that still lets people access the some.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Differential privacy seems to be a much, much better notion of privacy that kind of avoids a lot of the weaknesses of anonymization notions while still letting us do useful stuff with data.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Algorithmic privacy more broadly is just the study or the notion of privacy definitions or norms being encoded inside of algorithms. And so I think we count among this body of work just the literature and practice of things like data anonymization, which we kind of at the beginning of our discussion of privacy say like, okay, this is sort of a notion of algorithmic privacy. It kind of tells you something to go do with data. But our view is that it's, and I think this is now quite widespread, that it's despite the fact that those notions of anonymization kind of redacting and coarsening are the most widely adopted technical solutions for data privacy. They are deeply fundamentally flawed. And so, you know, to your first question, what is differential privacy?

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Yeah, I mean, I don't think I disagree enough, but I think that that trend of more and more people and more and more disciplines Adopting ideas from computer science learning how to code. I think that that trend seems firmly underway. I mean, an interesting digressive question along these lines is maybe in 50 years there won't be computer science departments anymore because the field will just sort of be ambient in all of the different disciplines. And people will look back and having a computer science department will look like having an electricity department or something. It's like, you know, everybody uses this. It's just out there. I mean, I do think there will always be that kind of canoe style core to it. But it's not an implausible path that we kind of get to the point where the academic discipline of computer science becomes somewhat marginalized because of its very success in kind of infiltrating all of science and society and the humanities, et cetera.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  49. I think one of the reasons we decided to write this book is we thought 10 years ago I wouldn't have tried this just because I just didn't think that sort of people's awareness of algorithms and machine learning, you know, the general population would have been high. I mean, you would have had to first write one of the many books kind of just explicating that topic to a lay audience first. Now I think we're at the point where lots of people without any technical training at all know enough about algorithms and machine learning that you can start getting to these nuances of things like ethical algorithms. I think we agree that there needs to be much more mixing. But I think a lot of the onus of that mixing needs to be on the computer science community.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Was about to agree with everything you said except that last point. I think that the other way of looking at it is that I think computer scientists and many of us are. But we need to wait out into the world more, right? I mean, just the influence that computer science and therefore computer scientists have had on society at large just has exponentially magnified in the last 10 or 20 years or so. And before when we were just tinkering around amongst ourselves and it didn't matter that much, there was no need for sort of computer scientists to be citizens of the world more broadly. And I think those days need to be over very, very fast. And I'm not saying everybody needs to do it, but to me, like the right way of doing it is not to sort of think that everybody else is going to become a computer scientist. But, you know, I think people are becoming more sophisticated about computer science even lay people.

    2019-11-19 · Lex Fridman Podcast · Michael Kearns: Algorithmic Fairness, Bias, Privacy, and Ethics in Machine Learning · IDENTIFIED FROM THE TRANSCRIPT · source