YouSaid · the spoken record
Devi Parikh
- lines on the record
- 44
- first
- 2023-07-20
- most recent
- 2023-07-20
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Just apply for it. And yeah, like don't assume, don't question, oh, am I good enough? Am I not? It's on the world to say no to you if you are not a good fit, the world will tell you that. And so, yeah, there's nothing to lose by just kind of giving it a shot. So don't self-select.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, so time management is something that I am quite excited about. And so, yeah, if I have a blog post on that, it's sort of main philosophy is that you should be writing everything that you want to do down on your calendar. So it's kind of the point is that it should be on your calendar. It shouldn't be your to-do list. And the reason it should be on your calendar is it forces you to think through how much time everything is going to take. If it's just a list, you have no idea how long it's going to take and that's not a good way to plan your time out. So that's kind of the main thesis of it that I hadn't anticipated, but it resonated a whole lot with many people, which was kind of surprising when I put it out there. So yeah, if anyone's interested, you should check that out. In terms of advice outside of what I've written, one advice that has stuck with me over the years is don't self-select, that if you want something, go for it. If you want a job, apply for it. If you want a fellowship for any students who might be listening, just apply for it. You can't intern.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, there's going to be more of it. That's one prediction that I can very confidently make that we are going to see all of these tools sort of show up where millions and billions of people can be using these in various forms. And I'm excited about that. Like I said, I think it kind of enhances creative expression and just sort of communication. And we're going to have these entirely new ways of interacting with each other. And even the social entities on these networks might change, right? Like when we talk about AI agents and you think about them being sort of part of the social graph, what that does and how that changes, how we connect with each other, all of that is fascinating. And I can't wait to see how that evolves.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yes, I was at CVPR. I was on a few different panels and I was giving some talks. So one of it was on vision language and creativity at the main conference. And so yeah, that was kind of what I was representing there. In terms of something exciting that I saw there not necessarily a paper, but there was a workshop there called Scholars and Big Models where the topic of discussion was as these models are getting larger and larger, making a lot of progress in what way can sort of academics or labs that don't have sort of as many compute resources, what should their strategy be, how should they be approaching these things. And that I thought was a really nice discussion in general. I tend to enjoy venues that talk about kind of the meta, like the human aspects of the work that we do. We have a lot of technical conversations, but we don't tend to talk about these other components. And so that workshop is something that I enjoyed.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, so one is the control piece that we already talked about quite a bit. And I think the other is multimodality, like bringing all of these modalities together. Right now we have models that can generate text, models that can generate images, models that can generate video. But there's no reason these all need to be independent. You can envision systems that are sort of ingesting all of these modalities, understanding all of it, and generating all of these modalities. And I'm starting to see some work in that direction, but I haven't seen a whole lot of it that goes across many different modalities.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah. And there are some people who have a very specific vision and they just want the tool to kind of help them get there. And then there are others whose process involves sort of bringing the model along where the unpredictability and sort of not necessarily knowing what this model is going to generate is a part of their process and is a part of the final piece that they create. So some view them view these models very much as tools. And then others tend to view them as more of a collaborator in this process of creating. And it's always interesting to see what end of the spectrum different people lie on.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Gives you more control, but it also means that there is more space to be creative. I can now pick interesting images or interesting videos or interesting pieces of audio and pair that up with like this really interesting text prompt and just kind of see what happens. Like if I put all of this in, you don't know what the model is necessarily going to do. And so it's also just more norms to play with as you're trying to interact with these that yeah there's just more space to be creative if there's more ways of more norms to control these models”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Looks great. Thank you. Thank you. And yeah, so to be honest, I don't know if it's hard to kind of look back and get a sense for did that play a certain role in it or not. I know for sure that it plays a role in just how excited I am about this technology, that anytime there's some new model out there, whether it's from the teams that I'm working with or if it's something external, I'm definitely very enthusiastic to want to try it out and see what it can do, what it can't do and sort of tell people about it. And so just kind of my baseline level of excitement around this technology is in part because of all these other interests that I have. I'm pretty sure that my emphasis on control is probably also coming from that, where I feel like I want to be using these tools to kind of have them do the thing that I want to do and sort of text prompts are restrictive in that way. And I mean, we talked about it in the context of control that if you can bring in multiple modalities as input that”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“I kind of always hesitate a little bit to call myself an artist. I feel like somebody else should be deciding whether I'm an artist or not, but then there's this whole community.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, yeah, and I think people across the entire spectrum, right? Like on one hand, you can talk. Hard and then the other end of the spectrum are people who don't necessarily have the skills, may not have had the training, but are still interested in being able to express their voices a little bit more creatively than they would have otherwise. And so I do think that there is one question of whether or not artists want to be engaging with this technology. And that is the other question of does it kind of just lift the tide for all of the rest of us to be able to be more expressive in what we can create and what we can communicate? And so I think both of those ends are relevant here. And with artists, there are artists who are like whose sort of brand is AI artists where they are explicitly using AI as the tool of choice for expressing themselves and their entire practices around that. Someone like Sophia Crespo or Scott Eaton and others. So there's also, and this was before Midjourney or anything like that, right? Like they've been doing this.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, yeah. And to your point that this technology is brand new, right? So it's not like there are existing product lines or ways of thinking about product that we can kind of directly plug into and kind of see, oh, did the metric go up? Did the metric go down? I think there's a lot of just kind of thinking about where do we anticipate people will be excited to use this. And as you said, I think there's a very good chance that there will be things that we don't necessarily foresee which is kind of come up as very exciting spaces.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think I'm not too much of a product person, so I feel like I don't know if I have the strongest intuitions there. But I think for kind of like I was touching on earlier, that a lot of these situations where we find ourselves searching for things to express ourselves, I think thinking of whether that can be generated so that it's a closer reflection of what you're trying to communicate is likely things that we'll see. And I know we're not talking about sort of LLMs and conversational agents and all of that too much, but I think AI agents is going to be a thing that we'll see a whole bunch of across many different surfaces. And then thinking about what media creation looks like in the context of AI agents is another dimension to this.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“And even with audio similar to what we were talking about with video, there is the same kinds of challenges and dimensions exist that you want the piece to be longer. You may want compositionality, right? I might want to be able to say that, well, first, it's the car driving down the street and then there is a sound of, I don't know, a baby crying and then something else. And maybe I'm saying that two of these sounds are happening simultaneously, which is not like that's something that can happen in audio where you can have the socor imposition. But in video is not something that would where it's quite as natural. And so all of that isn't stuff that these models can do very well right now. If I described a complex sequence of sounds or if I tried to talk about these different sounds simultaneously, these models can't do that very well.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Both for audio, similarly for music. I think it just makes the content much more expressive, much more delightful. But I feel like we don't do enough of that.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“I think I haven't tracked the text to speech quite as much. What I have tracked a little bit more closely is things like text to audio, where you might say that the sound of a car driving down the street and what you expect is sort of a sound of a car driving down a street to be generated. And so there, the state-of-the-art right now is sort of roughly a few second to tens of seconds long audio. And I would say that roughly it probably works reasonably well one in five times or so. It's like because there aren't concrete metrics, it's kind of hard to articulate where state of the art is, but hopefully this is helpful. And I do think that audio added to visual content makes it much more expressive and much more delightful. And I do think that it tends to be under investment.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Exactly, exactly, exactly. Like at least get random good first, then maybe let me give it text, then let me give it these other prompts. So I do think we'll force probably see more progress in just the core capabilities of sort of text to video generation before we look at prompting. Although we are, and this is in the context of sort of me generating something from scratch, right, which is where I might want this iterative control and things like that. A parallel scenario is where I already have a video and I'm trying to edit it in an interesting way. I might want to stylize it. And all of that, I think we're already seeing that even in products with runway, for example, right? So I think that we'll probably see much more of. We're already seeing that and I think we'll see more of where you already have a bit here that you're starting with and then you're trying to edit it, which has similarities too, but is a little bit different in my mind compared to sort of generating something from scratch and wanting control over that.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“I think control sort of tends to lag behind the core capability, like even with images, I feel like we first had to get to a point where these models can actually generate nice looking images before we start worrying about whether it's really doing what I wanted it to do. And I feel like we're not quite there with video yet.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“If the model then goes off and kind of does its own thing It would be ideal if there's some way of having iterative editing mechanisms where whatever I get back, I have a way of communicating to the model what it is that I want changed in what way so that over iterations I can get to the content that I intended in sort of a fairly reasonable way without having to sort of spend hours learning a new tool or something like that, right? So that can be done in a very intuitive interface. I think that would be pretty awesome.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“So I think of more control, at least in two different ways. One is to allow for prompts that are not just text, but are multimodal themselves. So for image generation, for example, instead of just text, it would be nice if I can kind of sketch out what I want the composition of the scene to look like. And the model would be expected to kind of respect that. For video, instead of just text as input, maybe I can also provide an image as input so that I can tell the system that this is the kind of scene that I want. Maybe I can provide sort of a little audio clip as input to convey that this is the kind of audio or sound that I want associated with it. Maybe I also bring in a short video clip and expect the model to sort of bring in all of these different modalities in a reasonable way to generate a video. So that's one piece where I can bring in more inputs as a way of more control. And the second piece is sort of the predictability part that even if I bring in all of these modalities as input,”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think that's very important exactly to your point that if we want these generative models not just for video but for any modality to be tools for creative expression, then it needs to be generating content that corresponds to what someone wants to express, right? Like it has to bring somebody's voice to life and that is not possible if there aren't good enough ways of controlling these models. And so text is one way that's better than random samples. That's one way in which I can say what I want. But right now, for the most part, you type in a text prompt, you get an image back, a video back, and either you take it or leave it, right? Like if you like it, that's great. If not, you just kind of try again and maybe you tweak the prompt a little bit. You sort of try a whole bunch of these prompt engineering tricks and hope that you get lucky, but it's not really a very direct form of control.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“It's that the next visual signal that the robot will see will be a consequence of the action that the robot had taken, right? So if it chose to move a certain way, that's going to change what the video looks like in the next few seconds. And so there's that interesting feedback loop there where it knows what action it had taken. It sees how that changed the visual signal that it is now getting as input. And so that connection makes it adds a layer of interestingness to how it can process the video sort of in contrast to with sort of regular computer vision disembodied tasks where we think of it just streaming a video as just kind of happening and you're not controlling what you're seeing.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“I think there, the video understanding piece is probably more relevant than the video generation piece. And video is like if you think of embodied agents, right? They're sort of moving around and consuming visual content, which inherently is video, right? They're not looking at static images. And so I think that video understanding piece is very relevant there. What's also interesting in the context of embodied agents or sort of robotics, physical robots that are moving around is that it's not passive consumption of videos, right? It's not like how you and I might be watching videos on YouTube or anything else. It's that the”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, yeah, and same thing for computer vision, right? Like here we're talking about generation, but even just understanding with images, with image understanding, there was so much progress that was being made. And videos was always kind of not only trailing behind, but just sort of continued to be hard at even sort of the rate of progress was slower, not just the absolute progress. And I think we're seeing some of that for generative models as well.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think what's lacking may not be so much the data source itself, although that is certainly a challenge as it is with other modalities. But I think it might also be the data recipes that do we want to start with training with sort of these very short videos where not much is happening, the scene isn't really changing, but then that also tends to limit the motion. There's just not much happening. And so you can end up with these kind of animated image looking things. And on the other hand, you have sort of very complex video, might be multiple minutes long with all sorts of scene transitions. And that's ideally what you want to shoot for. That's where you want to get. But if you're just going to directly throw all of that into the network, it's unclear if the models will be able to learn all of that complexity well. So I think thinking through some sort of a curriculum may be valuable here. And I don't think we've quite nailed that recipe down.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“And the third is this hierarchical architecture that if you want longer videos, there are just so many pixels that you're trying to generate, right? It's a very, very high dimensional signal compared to anything else that we're doing. And so just thinking through how do we even approach that, what sort of hierarchical representation makes sense, especially if you want these scene transitions, if you want to have this consistency, which may be a form of memory, figuring those architectural pieces out, I think may be another piece of this puzzle. And then finally data, right? Data is kind of gold in anything that we're trying to do. And I don't know if as a community, we've quite built the muscle of thinking through data sort of massaging the data appropriately and all of that in the context of video. We have that muscle quite a bit with language, quite a bit with images, but with video, we're perhaps not quite there yet.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think there's a few different things. One is videos are just sort of from an infrastructure perspective harder to work with, right? They're just sort of larger, more storage and sort of more expensive to process and more expensive to generate and all of that. So there's just that iteration cycle that is much slower with video than it would be with other modalities. So that is one. The second is I don't think we've still figured out the right representations for video, right? There is a lot of redundancy in video, one frame to the next frame. There's not a whole lot that changes. We still kind of approach them fairly independently as sort of independent images even if you're generating it as kind of one after the other or even if you're generating in parallel and then making it finer grained. So I think maybe that could be something that helps with a breakthrough that if we really figure out how to represent videos efficiently.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“In terms of how we approach video generation. So it's not quite answering what you asked me, but I do think that it might be a little bit slower than what we might have guessed just based on progress and other modalities.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, that is hard to say. And to be honest, I've actually been surprised that we haven't seen more of this already. Like make a video was, I think, what nine months or so ago, maybe approaching one year. And it's not like even from other institutions, it's not like we are seeing amazingly longer videos or significantly higher resolution or much more complexity. We're still kind of in this videos equals animated images. And yeah, maybe the resolution is a little bit bigger, quality is a little bit higher. But it's not like we've made significant breakthroughs, unlike, for example, what we've seen with images. So I do that has given me a sense that maybe this is harder than what we might think and sort of our usual curves of like with language models or image models. We're like, oh, just six months and there's going to be something else that's an entirely different step change over this. I think that might be harder in video and I wondered if there is something that we are kind of fundamentally missing.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“I think there's a ton to do in the context of video generation. Like, if you look at make a video, it was very exciting. It was sort of forced off its kind capabilities at the time. But it's a four second video. It's essentially an animated image, right? It's kind of the same scene, the same set of objects that are moving around in reasonable ways, but you're not seeing objects appear, objects disappear, you're not seeing objects reappear, you're not seeing scene transitions. None of this is in there. And so if you look at if you think about the complexity of videos that you just regularly come across on various surfaces, this is far from that. And so there is a ton to be done in terms of making these videos longer, more complex, having memory so that if an object reappears, it's actually consistent. It doesn't now look entirely different. Things of that sort sort of being able to tell more complex stories through videos, all of this is interactive.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, so one way of thinking of it could be that you may not have seen a flying corgi, but you've probably seen flying airplanes or flying birds and other things that fly in images and in video. And so from images, you will have text associated with it. So you will have a sense for what things tend to look like when someone is saying, oh, this is X flying or Y-flying. And then in videos, you will have seen the motion of what stuff looks like when it flies. And in images, you will have seen what Gorgis look like. It's hard to kind of know for sure what it is that these models learn sort of interpretability is not a strength of many of these deep large architectures, but that could be one intuitive explanation for how the model is managing to figure out what applying Gorgi might look like.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Not temporally coherent. So there is going to be independent images like the corgi playing with the ball is just going to be independent images of corgi's playing with blue balls, but they're not going to be temporally coherent. And then what the network is trying to do as it goes through the learning process is to make these images temporally coherent so that at the end of training it is generating a video rather than just unrelated images. And that's where the videos come in as training data.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“That image models already know all sorts of fantastic depictions of like dragons and unicorns and things like that where you may not have as much video data easily available. All of that is inherited. So all the diversity of the visual concepts can come in through images, even if your video data set doesn't have all of that. And the third benefit, maybe the biggest one, is that because of the separation, you don't need video and associated text as paired data. You have images and text as paired data and then you just have unlabeled video to learn motion from. And so these were kind of three things that we thought were quite interesting in how we approached Make a Video. And so concretely, the way it works is that when you initialize the model, you're starting off with image generation sort of parameters that have already been learned. So before you do any training for make a video, we set it up so that it can generate a few frames that are”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Leveraging all the progress that's happened with images in a very direct way to sort of make video generation possible. And so that led to this intuition that what if we try and use images and associated text as a way of learning what the world looks like and how we, how people talk about the visual content and then separate that out from trying to learn how things in the world move. So separate out appearance and language and that correspondence from motion of how things move. And so that is what led to make a video. And so there are sort of multiple advantages of thinking of it that way. One is there's less for the model to learn because you're directly bringing in everything that you already know about images to start with. The second is all of the diversity that we have in our image data sets.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, yeah. So make a video, it started because, I mean, this was a couple of years ago where this was before Dali 2, by the way. So this was like after Dali 1 had happened. And I feel like a lot of people don't even remember Dali 1 anymore. Like people don't even talk about that. It's fun to go check out what those images look like. And that had blown our minds at the time. But now when you go back, you're like, wait, that's not even interesting. But anyway, so we had seen a lot of progress in image generation. And so it seemed like the next kind of entirely open question where we hadn't seen much work at all was to see what can we do with video generation. And so that was kind of the inspiration behind that. And for make a video, the approach specifically, the thinking was we have these image generation models. By this time, we had seen a lot of progress with diffusion-based models from a variety of different institutions. And so the idea was, is there a way of”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Think of images, video, you can ask this question, like for any situation where you're searching for something, trying to find something, it's relevant to ask, well, could I just create what it is that I have in my head? So when you think of it that way, you can see how it can touch a lot of different things across a variety of products and services.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, yeah. I mean, yeah, that's a very exciting space right now. There's a lot happening both within Meta and outside, as I'm sure many of the people listening to this are aware of. But yeah, so that the new organization was created a few months ago. So not a long time ago. And it's looking at things like large language models, image generation, video generation, generating 3D content, audio, music, all sort of all sorts of modalities that you might think of. And why is it interesting? I mean, like right now, if you think about all the content, there's so much content that we consume in all modalities and all sorts of surfaces. And it makes a lot of sense to ask that instead of maybe not instead of, but in addition to all of this consumption, can more of us be creating more of this content, right? And so almost everything that you”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“This is going for longer. And so for many years after that, for five years or so, I was splitting my time per every fall. I would go back to Georgia Tech, to Atlanta, spend the fall semester there, teach, and then the rest of the year I would be in Mendel Park at Facebook, now Meta.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“So, this was. I've been at Meta for about seven years now, and so this had started when I was transitioning from Virginia Tech to Georgia Tech. I was an assistant professor at Virginia Tech, then I was getting started at Georgia Tech. And in that transition, I decided to spend a year at FAIR. At the time, it was called Facebook Air Research, now fundamentally AI research at Meta. And I knew colleagues there from some of them had been at Microsoft research before that. I had interned at MSR. I had spent summers at MSR, even as a faculty member. And so I had a lot of colleagues who I knew. And so I thought it would be fun to kind of spend a year collaborate with them, get to know what FAR is like. So that's what that was. It was supposed to be a one-year stint. And then I was going to go back to Georgia Tech and kind of continue with my academic position. But in that one year, I enjoyed it enough, I think Fair enjoyed having me around enough where we tried to figure out, is there a way to keep?”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Full fled research agenda, and that's how I started working more seriously on generative modeling, including transformer-based approaches, diffusion models for images, for video, for 3D video, things of that sort.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Than explaining why they're making the decisions that they're making. And that slowly led to the sort of more into natural language processing, where instead of these kind of just adjectives and attributes, looking at more natural language as a way of interacting. So a lot of my work in visual question answering, where you're answering questions about images, image captioning was coming from there. And then over time, I sort of started thinking of are there ways to go even deeper in this interaction? There are there ways where AI tools can enhance sort of creative expression for people, give them more tools for expressing themselves. And that's how I got interested in AI for creativity. And I was dabbling in kind of a few fairly random projects a few years ago, kind of my bread and butter research was still multimodal vision and language. But then sort of I enjoyed what I was doing with AI for creativity in a couple of years ago that took sort of a little bit more of a serious turn where I made it more of my sort of”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“I think, I mean, you can always kind of look back and try and find patterns. Like when you're actually doing it, you don't necessarily have a grand strategy of anything in mind. But when I look back, I think one common theme that led to me transitioning across topics a little bit was that I was interested in seeing how we can get humans to interact with machines in more meaningful ways. And so kind of even my transition from kind of non-visual to visual modalities in hindsight, I feel like was essentially that. I felt like you can't interact with these systems too much if it's sort of these abstract modalities that you're looking at. And then when I was working in computer vision, I wanted to find ways for humans to be able to interact with these systems more. So I started looking at kind of these attributes and adjectives of like, oh, something is furry or something is shiny. And using that as a mode of communication between humans and machines, both for humans to teach machines new concepts and for machines to be more interpretable.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“So at first, I wasn't working on projects that didn't have too much of a visual element to them. But when I got to CMU, my advisors lab was working in image processing and computer vision and I always thought that it was pretty cool that everybody gets to kind of look at the outputs of their algorithms and see what they're doing. Whereas if it's kind of non-visual, then yeah, you see these metrics, but you don't really have a sense for what's happening, if it's working, if it's not. And so that's how I got interested in computer vision. And that then defined the topic of my thesis over the course of my PhD.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Or you go to a PhD, and so they kind of slotted me onto the PhD track, which I wasn't so sure of, but my advisor there was reasonably confident that I'm going to enjoy it and I'm going to want to keep going. So yeah, that's how I got started in this space at first. I was doing projects that didn't have a visual element to it.”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT
“Kind of, kind of, yeah. So, my background is that I grew up in India and then I moved to the US after high school. And I went to a small school called Rovan University in southern New Jersey for my undergrad. And that is where I first got exposed to what at the time was being called pattern recognition. We weren't even calling it machine learning and got exposed to some research projects. There was a professor there who kind of showed some interest in me, thought I might have potential to contribute meaningfully to research projects. And that's how I got exposed. And I really, really enjoyed what I was doing there. Decided to go to grad school to Carnegie Mellon. I knew I was enjoying it, but I wasn't sure if I wanted to do a PhD. So at first, I wanted to just kind of get a master's degree with a thesis, but I can do some research. But the year that I applied that the ECE department at CMU decided that there wasn't going to be a master's track for thesis. Like either you can just take”
2023-07-20 · No Priors · The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality · IDENTIFIED FROM THE TRANSCRIPT