YouSaid · the spoken record

Dawn Song

lines on the record
146
first
2020-05-12
most recent
2020-05-12
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. So, in this case, for example, you don't want some malicious nose to be able to change the transaction logs. And in certain cases, it's called double spending. You can also cause different views in different parts of the network and so on.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Right, yes. In a community of nose that come together. And even though Ichiwan may not be trusted, and as long as a certain threshold of the set of nodes behaves properly, then the system can essentially achieve certain properties. For example, in the distribute ledger setting, you can maintain an immutable log and you can ensure that for example the transactions actually are agreed upon and then it's immutable and so on.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Really has brought the issue of data privacy and even this consideration of data ownership to the forefront to really much wider community. And I think more of this voice is needed, but I think it's just that we want to have a more constructive dialogue to bring the both sides together to figure out a constructive solution.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  4. I think also I can understand. I think actually in the past, especially in the past couple years, this rising awareness has been helpful. Users are also more and more recognizing that privacy is important to them. maybe right, they should be owners of their data. I think this definiteness is very helpful. And I think also this type of voice also and together with the regulatory framework and so on also help the companies to essentially put these type of issues at a higher priority and knowing that also it is their responsibility too to ensure that users are well protected. So I think definitely the rising voice is super helpful. And I think that actually

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Right, right. Yes, at high level. So essentially also knowing that there are technical challenges in addressing the issue to like basically you can't have just like the example that I gave earlier, it's really difficult to balance the two between utility and privacy. And that's also a lot of things that I work on in my group worked on as well is to actually develop these technologies that are needed to essentially help this balance better and essentially to help data to be utilized in a privacy preserving and responsible way. And so we essentially need people to understand the challenges and also at the same time to provide the technical abilities and also regulatory frameworks to help the two sides to be more More in a woman situation instead of a fight.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  6. But also, of course, at the same time, we want to ensure that however that data is being handled, it's done in a privacy-preserving way so that, for example, that recommendation system doesn't just go around and sell your data and then cause all the bad consequences and so on.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Right. And on the other hand, of course, they also feel, oh, they are providing a lot of services to users. And users are getting it all for free. So I think I actually talk a lot to different companies and also like basically ample size. So one thing I hope also, like this my hope for this year also is that we want to establish a more constructive dialogue And to help people to understand that the problem is much more nuanced than just this two sides fighting. Because naturally there is a tension between the two sides, between utility and privacy. So if you want to get more utility, essentially like the recommendation system example I gave earlier, if you want someone to give you good recommendation, essentially whatever that system is, the system is going to need to know your data to give you good recommendation.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Yeah, so that's a very good question. I think, so one thing I also want to mention is that it seems that especially in press, the conversation has been very much like two sides fighting against each other. One hand, users can say that, right? The Don't Trust Facebook delete Facebook.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Right. So then users essentially can have choices. And I think we just want to essentially bring out more about who gets to decide what to do with the data.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So that's a very good question. I think that's not necessarily the case in the sense that, yes, users can have ownership of their data, they can maintain control of their data, but also then they get to decide how their data can be used. So that's why I mentioned earlier, so in this case, if they feel that they enjoy the benefits of social networks and so on, and they're fine with having Facebook having their data, but utilizing the data in certain way that they agree, then they can still enjoy the free services. But for others, maybe they would prefer some kind of private vision. And in that case, maybe they can even opt in to say that I want to pay to have, so for example, it's already fairly standard, like you pay for certain subscriptions So that you don't get to be shown as.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  11. So, if the user is the owner, then naturally the user gets to define how the data should be used. But if you even say that users are actually not the owner of this data, whoever is collecting the data is the owner of the data, then of course they get to use the data however way they want. So to really address these complex issues, we need to go at the root cause. So it seems fairly clear that first we really need to say who is the owner of the data and then the owners can specify how they want their data to be utilized.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Users in joy and can really benefit from good recommendation systems, are recommending you better music, movies, news, even research papers to read. But of course, then in these targeted ads, especially in certain cases where people can be manipulated by these targeted ads, they can have really bad severe consequences. So essentially users wonder data to be used. to better serve them and also maybe even get paid for or whatever like in different settings but the thing is that the first of all we need to really establish like who needs to decide who can decide how the data should be used and typically the establishment and clarification of the ownership will help this and it's important first step

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  13. There's no clear notion of ownership of such data. And also, we talk about privacy and so on, but I think actually clearly identifying the ownership is a first step. Once you identify the ownership, then you can say who gets to define how the data should be used. So maybe some users are fine with internet companies serving them as using their data as long as the data is used in a certain way that actually the user consents with are allowed. For example, you can see the recommendation system in some sense. We don't call it an as, but a recommendation system similarly is trying to recommend you something.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  14. And also, more and more, I would say even assets of a person is more and more into the real world, the physical world as well. It's the data that the person's generated. Essentially, it's like in the past what defines a person. You can say, right, like oftentimes besides the innate capabilities, actually it's the physical properties that defines a person. But I think more and more people start to realize actually what defines a person is more important in the data that the person has generated or the data about the person. All the way from your political views, your music taste and financial information, a lot of these and your health. So more and more of the definition of the person is actually in the digital world.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  15. I think it's a first of all, it's a good lesson for us to. To recognize that these rights and the recognition and enforcement of these type of rights is very, very important for economic growth. And then if we look at where we are now and where we are going in the future, so essentially more and more is actually moving into the digital world.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Institutional rights, like governmental enforcement of this actually has been a key driver for economic growth. And there have been even research proposals saying that for a lot of the developing countries, they're essentially the challenge in growth is not actually due to the lack of capital. It's more due to the lack of this notion of property rights and enforcement of property rights.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  17. I think this is a fascinating topic and also a really complex topic. Right. I think there are these natural questions. Who should be owning the data? And so I can draw one analogy. So for example, for physical properties like your house and so on. So really this notion of property rights is not just from day one we knew that there should be like this clear notion of ownership of properties and having enforcement for this. So actually People have shown that this establishment and enforcement of property rights has been a main driver for the economy earlier. And that actually really propelled the economic growth, even in the earlier stage.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Right. So then the finally trained learning, the learned model is differentially private. And so it can enhance the privacy protection.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  19. So, right. So, in this particular case, what the differential privacy mechanism does is that it actually adds perturbation in the training process. As we know during the training process, we are learning the model, we are doing gradient updates, the width updates, and so on. Essentially, differential private machining algorithm in this case will be adding noise and adding various perturbation during this training process.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  20. For the salanguage model, language model, okay. So instead of just training a vanilla language model, instead if we train a differentially private language model, then we can still achieve similar utility. But at the same time, we can actually significantly enhance the privacy protection. And of the learned model and our proposed attacks actually are no longer effective.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  21. That's an example showing that's why even as we train machine learning models, we have to be really careful with protecting users' data privacy.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  22. It's called an NRAL email dataset. And the NRAM email datasets naturally contain uses social security numbers and credit card numbers. So we train the language model over the datasets, and then we show that an attacker by devising some new attacks by just querying the language model, and without knowing the details of the model, the attacker actually can extract the original social security numbers and credit card numbers that were in the original training data.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  23. So, in our work, in collaboration with researchers from Google, we actually studied the following question. So at high level of the question is, as we mentioned, the neural networks can have very high capacity and they could be remembering a lot from the training process. Then the question is, can an attacker actually exploit this and try to actually extract sensitive information in the original training dataset through just querying the learned model without even knowing the parameters of the model, like the details of the model or the architectures of the model and so on? So that's a question we set out to explore. And in one of the case studies, we showed the following. So we trained a language model over an email datasets.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  24. So, I can give you some examples. And another type of attack which is even easier to carry out is not a web box model. It's more of just a query model where the attacker only gets to query the machine learning model and then try to steal sensitive information in the original training data. So, right, so I can give you an example. In this case, training a language model.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  25. And then from that, a smart attacker potentially can try to figure out information about the training data set. They can try to figure out what type of data has been the training data set. And sometimes they can tell whether a person has been a particular person's data point has been used in a training datasets as well.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  26. And also, we can talk about essentially how the attacker may try to learn information from the, right? And also there are different types of attacks. So in certain cases, again, like in white box attacks, we can see that the attacker actually gets to see the parameters of the model.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  27. And hence, just from the learned model in the ants, actually, attackers can potentially infer information about their original training data set.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Right, so especially in the machine learning setting, so in this case, as we know how the process goes is that we have the training data and then the machine learning system trains from the screening data and then builds a model and then later on inputs are given to the model to inference time to try to get prediction and so on. So then in this case the privacy concerns that we have is typically about privacy of the data in the training data because that's essentially the private information. And it's really important because oftentimes the training data can be very sensitive. It can be financial data, it's your health data or like in IoT case it's the sensors deployed in Real world environment and so on, and all this can be collecting very sensitive information. And all the sensitive information gets fed into the learning system.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Right. So, in security, we actually talk about essentially two, in this case, two different properties. One is integrity and one is confidentiality. So we have been talking earlier is essentially the integrity property after the learning system, how to make sure that the learning system is giving the right prediction, for example. And privacy essentially is on the other side is about confidentiality of the system, is how attackers can, when the attackers compromise the confidentiality of the system, that's when the attacker steals sensitive information about individuals and so on.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  30. And also, that's why I was also saying earlier when defense is this multi model defense and more of these consistency checks and so on. So in the future, I think also it's important that for this autonomous vehicles, they have lots of different sensors and they should be combining all these sensory readings to arrive at the decision and the interpretation of the world and so on. And the more of these sensory inputs they use and the better they combine these sensory inputs, the harder it is going to be attacked. And hence, I think that is a very important direction for us to move towards.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  31. That's one way to demonstrate. And this is also like so far we've many talks about work in this adversary setting showing that today's learning system, they are so vulnerable to the adversarial setting, but at the same time, actually, we also know that even in natural settings, these learning systems, they don't generalize well, and hence they can really misbehave under certain situations.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Breaking into I mean, so one way to look at this in terms of how real these attacks can be, one way to look at it is that actually you don't even need any sophisticated attacks. Already we've seen many real world examples of incidents where showing that the vehicle was making the wrong decision.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  33. I think there are two separate questions. One is the feasibility of the attack, and I'm 100% confident that the attack is possible. And there's several questions whether someone will actually go. You know, deploy that attack. I hope people do not do that. That's a two separate quest

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So actually, there has already been research shown that, for example, actually even with Tesla, if you put a few stickers on the road, it can actually, once arranging certain ways, it can fool the...

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  35. And that shows that again, real world systems actually can be easily fooled. And in our previous work, we also showed this type of black box attacks can be effective cloud vision APIs as well.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  36. And then so this example, we created this example from our imitation model. And then this work actually transfers to the Google Translates.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  37. As a target model. And then once we have the invitation model, we can then try to create adversary examples on these invitation models. So for example, giving in a work one example is translating from English to German, we can give it a sentencing, for example, I'm feeling freezing, it's like six Fahrenheit. And then translating to German. And then we can actually generate adversary examples that creates a target translation by very small perturbation. So in this case, let's say we want to change the translation itself six Fahrenheit to 21 Celsius. And in this particular example, actually, we just changed six to seven. In the original sentence, that's the only change we made, it caused the translation to change from the six Fahrenheit into 21 Cels

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  38. So, in this work, which shows so far, I talked about adversary examples mostly in the vision. Category. And of course, adversary examples also work in other domains as well. For example, in natural language. So in this work, my students and collaborators have shown that one, we can actually very easily steal the model from, for example, Google Translates. By just doing queries from the APIs, and then we can train an imitation model ourselves. Using the curries. And then once we and also the invitation model can be very, very effective, essentially achieving similar performance.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  39. I wouldn't call that a hope. I think it's more of a wishful thinking. I'll try to be lucky. So actually, in our recent work, my students and collaborators have shown some very effective attacks on real-world systems. For example, Google Translate.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Right now, of course, it's attack site. It's much easier to develop attacks, and there are so many different ways to develop attacks. Even just us, we develop so many different methods for doing attacks. And also you can do white box attacks, you can do black box attacks, where attacks you don't even need. The attacker doesn't even need to know the architecture of the target system and now knowing the parameters of the attacker system and all that. So there are so many different types of attacks.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Right, right, right. Speech data? Right. And then we can actually combine spatial consistency and temporal consistency to help us to develop more resilient methods in video, so to defend against attacks for video also.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  42. So it turns out actually to be really hard for the attacker to do, we try the best we can, the CTF attacks actually show that this defense method is actually very, very effective. And this goes to, I think also what I was saying earlier is essentially we want the learning system to have to have richer transition also to learn from more you can add the same multi-model essentially to have more ways to check whether it's actually having the rep prediction. So for example in this case doing the spatial consistency check and also actually so that's one paper that we did and then this is spatial consistency this notion of consistency check it's not just limited to spatial properties it also applies to audio so we actually had follow-up work in audio

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Right, right, right, exactly, exactly. But then this actually poses a challenge for adversarial examples. Because for the attacker to add perturbation to the image, then it's easy for us to fool the segmentation system into, for example, for a particular patch or for the whole image to cause the segmentation system to get to some wrong results, but it's actually very difficult.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Exactly, exactly. So then in our work and experiments, we show the following. So when we take like normal images, this actually holds pretty well for the segmentation systems that we experimented with.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  45. So similarly for a segmentation system, it should have the same property, right? So in the image, if you pick two patches, the header intersection, you feed each patch to the segmentation system, you get a result. And when you look at the results in the intersection, the results, the segmentation results should be very similar.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Right, yes. And then if you pick two patches of the scene that has an intersection. And for humans, if you segment A and patch B, and then you look at the segmentation results, and especially if you look at the segmentation results at the intersection. The two patches, they should be consistent in the sense that what the label, what the pixels in this intersection, what their labels should be, essentially from these two different patches, they should be similar in the intersection. So that's what we call spatial consistency.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Right. So that's on the attack side showing that the segmentation system, even though they have been effective in practice, but at the same time, they're really, really easily fooled. So then the question is how can we defend against this? How can we build more resilient segmentation system? So that's what we try to do. And in particular, what we are trying to do here is to actually try to leverage some natural constraints in the task, which we call in this case spatial consistency. So the idea of the spatial consistency is a following. So again, we don't really know how human vision works, but in general, at least what we can say is, so for example, as a person looks at a scene, We can segment the scene easily.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Defense there? Yeah, sure, sure. So in that paper, what we look at is the semantic segmentation task. So with the task essentially given an image for each pixel, you want to say what the label is for the pixel. So just like what we talked about for adversary example, it can easily full image classification systems. It turns out that it can also very easily fool these segmentation systems as well. So given image, I essentially can add adversary perturbation to the image to cause the segmentation system to basically segment it in any pattern I wanted. So you know people also showed that you can segment it even though there's no KT in the image. We can segment it into a KT pattern, a hello KT pattern. We segment it into like ICCV.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Right, so you want to learn the right things. You don't want to, for example, learn this spurious correlations and so on. But at the same time, an example of a richer information representation is, again, we don't really know how human vision works. But the one way to look at the visual world, we actually can identify counters, we can identify much more information than just what's, for example, image classification system is trying to do. And then these two, I think the question you asked earlier about defenses. So that's also in terms of more promising directions for defenses. That's where some of my work is trying to do and trying to show as well.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source

  50. From the world. And we use all this information together in the end to help us to do motion planning and to do other things, but also to classify what the object is and so on. So we are linear much richer representation. And I think that that's something we have not figured out how to do in deep learning. And I think the ritual representation will also help us to build a more generalizable and more resilient learning system.

    2020-05-12 · Lex Fridman Podcast · #95 – Dawn Song: Adversarial Machine Learning and Computer Security · IDENTIFIED FROM THE TRANSCRIPT · source