YouSaid · the spoken record

Shomi Patwary

lines on the record
53
first
2025-08-16
most recent
2025-08-16
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Oh, yeah, that's every angry yesterday. My thinking about that is actually, yeah, I thought about it a bit. I think if we live in a simulation, my take is that it doesn't run on our current hardware because it's analog and not like, you know, it's continuous. All of the observations are continuous and there is nothing like, but maybe the quantum level is some limitation of, you wanted to go philosophical. It's some kind of a hardware limitation of the simulation we run on. Take it or leave

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  2. World, for example, it's actually quite different to that. When you do take a minute to look away from who to screen, it's quite a bit richer out there. And that's just for the real world. We also want this ability to generate completely new things, right? I think we've got a huge gap to close, right, with the new capabilities that we want to add. But I think it's maybe a bit different to language models, or actually maybe it is similar to language models, but with language models, there's been like lots of new steps that have actually come on top, right? That maybe we didn't think were possible. We thought things were plateauing. And then a new idea came that made a significant change. And that has happened a couple of times in the past few years. So I think that there's a few more of those left for sure.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  3. That's a really good question. I actually have a super hand wavy somewhat swerving answer, right? And I think it's actually both. So I think the current capabilities are actually already quite compelling. And so you could make the case that like if what you wanted was a minute of photorealistic any world generation with memory, that could actually be the end goal, right? And two or three years ago, I probably would have said that was a five-year goal. And so at that point, if you just wanted to improve that, I think you probably end up with this maybe like, I think the jump from Genie 2 to Genie 3 was absolutely massive and went from being like kind of a cool bit of research that was like showing signs of life something that could already be very compelling. But I think there's a lot more that you can do with this. And Shomi kind of referenced this to himself, right? Like it's not the case that you're dropping yourself in the world, right? And like, it's like the real being in the real.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  4. One of the things I've been thinking about a lot is we see sort of with every modality, maybe first LLMs and then image and video and audio, there was like early kind of glimmers of something really exciting in a project or a research preview. And then there's like a ton of data and compute and researchers kind of poured out the problem and you hopefully see this sort of like exponential progress till you eventually get to the point where like you're out of data or the improvements don't come as easily. I'm wondering for your thoughts like where we are on sort of that curve for world models.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  5. So, as you can see, we are very excited about having more people accessing it. So, we're definitely want to make it happen. There is no kind of concrete timeline at the moment. But I'm sure once we have more to share, we'll do. Awesome.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  6. We'll convince her. Yeah, just to touch on maybe a final point on the robot's kind of robotics part. I think it's definitely robotics means more than visual, right? Like we need to be able to, I think this is an important point. We can drive the decisions of a robot by looking around, but still it has to kind of do actuations, decide where to move, how to respond to the environment. So I think there are definitely some gaps, but still at the core of the problem being able to reason about the environment, we think this is something that world models or general purpose world models such as Gini-Free can really help with and maybe with future research we can actually bridge those gaps of physical understanding and actually getting responses, physical responses from the world, which is a very interesting direction to explore.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Someone who's scared of dogs know to go around then, see someone with a ball, change directions, like all these challenging situations in the real world, right? And of course, you still have gripping, you still have these other tasks, but you need to really discover your own behaviors from your own experience, right? And doing that in physical embodied worlds is super challenging because there's so many reasons why, firstly, that could be expensive to collect data in those settings. You'd have to keep moving the robot back to where it started every time it doesn't do something right. And also could be unsafe, right? So there's many reasons why we can't really do learning from experience in the physical world, right? So we do it in simulation. But really what we think with Genie 3 is it's the best of both, right? Because you're taking a real world data-driven approach, right? But then you've got the ability to learn the simulation. So it kind of combines the good parts of each of those. And so that's why I think it could be.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Collect data in a quite laborious way, but it looks like the downstream task. So it looks real. And there's not so much of a mismatch between the two domains. Or you can learn in simulation, right? But the robotic simulations are even the best ones. And we have some of the best ones, Deep My Majoko, right, which we work with. They're still quite far away from the real world, right? And so you have the Sim to Real Gap. But even the Sim to Real Gap itself I think is kind of like poorly named because what people consider to be real and robotics is typically still a lab or some very constrained environment where you've got a bunch of spotlights on a robot and then tons of researchers crowding around watching where it's really real for me is making any references it's the ability to walk my dog when I'm too busy to hold the lead cross the street you know

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  9. So we designed it to be an environment rather than an agent, right? So Genie 3 is very much like an environment model. Like we don't see it as like an agent itself that can think and act in the world. It's more just a general purpose sort of simulator in a sense, right? That can actually simulate experiences for agents. And we know that learning from experience is a really important paradigm for agents, right? That's how we got alpha go because the agent alpha go learned by playing Go by itself trying new things, right? And then learning from feedback with reinforcement learning, learning to improve itself and actually discover new things like it discovered new moves that move 37 that humans didn't think was a worthwhile move, right? But actually AlphaGol learned that it was because it could experience and try things for itself. And in robotics, we have this paradigm right now where there's some data-driven approaches, right, where you can

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Breaking my mind, which is that you had one simulation agent asking the world asking the Genie agent to essentially create a real time environment for it to interact in, right? Which was when I realized, oh, the way you guys have built it, it's composable with other agents. Can you talk a little bit about why that's so important for robotics? Like Marco was saying and what are the major limitations today that you think we'd have to overcome as a space to make the robotics sort of progress, the rate of progress in robotics much faster than it is now.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  11. The robotics application, there was a conversation that I was listening to from Demos yesterday where he was talking about your guys' work on Genie 3. And he mentioned that there's an agent. I think you guys call it SIMA, which can then interact with the Genie agent. And as I was hearing him describe it, which was kind of...

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  12. We were just talking about before we started that, we might see applications like in robotics. I mean, Jack, you were talking about embedded AI and now limitation in robotics is the data, right? Like how much data you can collect. And now probably you can just generate a lot of different scenes that you were not able to do before, purely from just recording videos or so. So I think that's another thing that is pretty exciting. I mean, congrats on the model. It's phenomenal.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  13. I am actually personally petrified of skiing, and the model's already quite good at that. So I might, when things quieten down, spend some time. Because I promised my wife that our children would grow up knowing how to ski. And we're getting close to the age where I have to live up to my promise. And I'm not sure if I want to do it yet

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Yeah, I'm really excited. I think we are only as impressive, maybe the model is. I think they're very far from actually simulating the world accurately and being able to do kind of put a person in there and then do whatever they want. And when I say far, it doesn't mean it's far in terms of a calendar time because we live in an accelerated timeline, but it feels like there is more work to do to get there. And I think I just imagine once we can actually whatever the form factor would be, but stepping to this world and just kind of like maybe tell it how we want to, what you want to experience. There's so many applications. Imagine, for example, someone is afraid of talking to people on a stage or in a podcast, right? They can simulate that, right? Or you can have someone who is like afraid of spiders. They can maybe.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  15. More excited about applications I never thought of that come up from other people seeing the model, right? So I think it's kind of this trade-off of, you know, obviously you want to focus on some applications, but then you want to be open-minded about others. And I think that's the real joy of building models like this, right? Is you get to see all of these people, you can be way more creative than me with it. So I think that there's always really cool things that we can do. And I honestly don't really can't really tell you in one year what the biggest application will be, but we'll definitely be trying to build better models.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Top of mind, I think, for the next few days might be a vacation. After that. Maybe walking my dog in the real world. And then I think you mentioned a bunch of really interesting things, to be honest. And I think we're still collecting a lot of feedback on this current model, right? And I think that in general, are most interested in building the most capable models, right? And so we would hope to have even broader impact in future and really enable other teams to do cool things with it, right? both internally and externally. And for me, it's like I just started this with like a very, very focused vision about AGI. And I still think, honestly, for my, what I'm excited about for AGI, and which is more embodied agents, I really believe this is the fastest path to getting these agents in the real world. And I think we made a big step towards that. But and still, sometimes even.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  17. I guess somewhat related to that, like how do you think going forward like Gini 4 5 or any other models, like what is like top of mind right now? Like if you wonder, for example, to focus on, I don't know, like seems like gaming could be one of the applications, having multiplayer type of games. We have two special memories or two different completely views, but at some point they merge. How are you thinking on going forward? What's next? Is it like scaling these models just on more data, more compute? Is it creating this sort of like multi-universe type of things where you have multiple players, multiple people looking at the same model, putting different views? What's top of mind for you guys?

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yeah, I'll say that basically we have some applications in mind, but that's not what's driving the research. It's more about can we, how far can we push in this particular direction? Can we make all of that work, like really great quality, really fast generation real time, very controllable? I think that's kind of what drived us. I think there are two to develop Gini3. follow and I don't think to be honest, I don't know what would be the applications for like i think we're very surprised um you know i like to mention like vehicle we all people find new new ways in how it can be useful and to prompt it we have like visual stuff you know people just discover it right we didn't even think about it initially so i i expect kind of the same thing and i think that's why i am excited for more people to be to be able to accept in the future and in general our approach

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  19. So, I guess more like worlds where tasks can be achieved but doesn't require the high quality cinema style videos you could generate with BO model, right? It's quite different. And then on the filmmaking element, I mean, I'm not so sure that Genie 3 is really there at this point. And that would be necessarily the goal.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  20. I mean, VI3 has played a higher quality threshold than Genie 3, right? And it has very different priorities. So then the natural things you could say, oh, well, what if we just took these together and combined them? But that may not be the best next step for either of those two models, right? So it may not be the case that the thing that the other one has is actually the most compelling thing for a completely different experience. And I think that given the breadth of interest in both models, right, there's actually quite a small set of people that are like really actively using both. And they tend to be more folks like yourself who are just more broadly interested in AI, right? Rather than really downstream use cases. So like you mentioned agent training for one, which is very sort of like high action frequency requires more ecocentric.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  21. I think this is a really interesting point, Mike, and ultimately, it has to be driven by technical decisions. And also the goals, right? So if you look at the models right now, we obviously made a choice that we want VO3. And Genie 3 to be separate projects this year, right? And if you look at them both as they are right now, they have very different capabilities that the other model does not have. And technically to combine all of that already into one model would be, I think, very challenging to.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Products, different models can go in different direction. I think the space is pretty big and there are a lot of trade-offs to be made. So yeah, I don't know. I think it really depends. Some people really, there is one model that's with everything. Or I think there is still open-ended, what's the best way? We're in a place where engineering is a big.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  23. From my perspective, there are different, so I would say modalities is one thing, right? We have text, we have audio, even without different type of submodalities. Speech is not the same as music. We have different products for music generation, and we have other models for speech generation, speech understanding. So even within one modality, you can have different flavors. And then, of course, you have video and other things. I think basically, I would say the modality is one dimension and another is how fast or how quickly we can create new samples and completely orthogonal maybe the direction is or dimension is how much control we have right so i think we kind of picked specific direction or a specific vector in the space for gini free um i think different

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Something we think about a lot is what are the edges of a modality? We're talking about this all the time, which is the lines start blurring pretty quickly between real-time image and video and then real-time video and interactive world generation, world model. I don't think we have a good word for what Genie 3 is yet, but you guys called that world model, which is, I think, a great term. But in your mind, where does the video generation modalities stop and real time worlds start? And do you think in the future, are these converging into basically one modality? Or if you had to predict over the next few years, do you guys think actually, yeah, these will diverge into completely different disciplines? It seems like they share kind of one parent today, which is video generation. But where is the world going, do you think? Are these two completely different fields?

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So, I think definitely a bit different, right? Like Gini allows you to navigate environment and then maybe take actions, right? And that's not something that they at this point can do. But there are other aspects that are different that Genie doesn't have, right? Doesn't Jinnie doesn't have audio, for example, right? So we just think it's while definitely there are potential similarities, it's sufficiently different. Also another thing is that at this point, Jenny Free is not available as a product. And we do think about it as a product that is made mainstream and became very, very popular. And what the future holds, I don't know. But I mean, at this point, we just felt it's sufficiently different in terms of what capabilities and how kind of like we think about this. So Ginnie Fries pretty much a research preview, right? It's not something we are really.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Well, I mean, our team had never really worked on this. So Genie 1 and 2 both worked with image prompting. And so obviously this next phase, we leveraged a lot of the research done internally on other projects. And personnel-wise, I mean, Shomi has obviously worked in co-leading the VO project. And so we were able to kind of build on a lot of other working ideas internally. that basically allowed us to kind of like turbocharge progress, right? So if we done this sort of by incrementally building ourselves on an island, it would have taken, I think, a lot longer than being part of Google DeepMind where we have these teams that have a lot of knowledge in different areas and sort of lean and build on, which I think is super exciting about our being in the company right now is that we have so many experts in different areas that we can seek out.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  27. And yeah, I don't know if that's a big secret, but it looks exactly like her. And the model just kind of knows, right? I think that's pretty amazing. So I think that that's actually a really important capability that we didn't have with G2 as well, right? Because we relied on image prompting. And so there was some transfer issue, like where you rely on imagine to generate the image. And that often does look really good, but it's not necessarily a good image for starting the world. Whereas like going directly from text, you get the controllability of pretty much anything you want. Plus, it just kind of naturally works because it's in the like correct space for the model to do its thing. And that's something really powerful.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Text following is really amazing in this model. And that does feel really magical. I think this is something that the VO does really well as well, right? Like pretty much what you ask for, it's really well aligned with text. And we have that with Genie 3. So you could describe very specific worlds and really kind of like arbitrary silly things. And it pretty much works. We actually had this discussion because people were very disappointed to find out that the video I made of my dog actually was not my dog's photograph. I just described her in text.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Because they don't want to just look at the real that looks like they're on this room, but something a bit more exciting. And that's where I think this is the magic of the models that they can take you to places that maybe are not so likely to be in reality.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  30. On top of that, one kind of trade off that typically we have is that we want the models to do two things. We want the model to create the world in a way that looks consistent. So Jack said, like, if you walk in drain or in pulse, then probably wearing boots. But if we provide it with a different description or like the prompt is saying something else, we want it to still follow the prompt. And there is some tension here because some things are very unlikely, right? You might say, I want to wear flip-flops and jump in the rain or whatever. Then the model still has to try and create something that is very unlikely. And that's where typically video models maybe find it more challenging. And that's where our models might find it more challenging, but still it's still successful to it to a surprising degree to go into this kind of like low probability area. And I think that's really, in a way, that's what we want, right?

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  31. When you look down near a puddle, hopefully you're wearing Wellington boots, like this kind of stuff does just kind of make sense. And I think it feels pretty magical because it very much aligns with what you were thinking about the world and the model's just generated at all. So yeah, that's also one of the really exciting things for sure.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  32. What you're basically describing is like the real breadth of different kind of environment terrains and worlds and things like that, like water or walking on sand versus going downhill and snow and how the agent's sort of interactions should differ given the terrain that they're in. And I think that that really is a property of scale and breadth of training. So this is very much like an emergent thing. I don't think there's anything like really specific we do for this, right? Again, like you hope the model has learned this because it should have a general world knowledge doesn't always work perfectly, but in general it's pretty good. So for the skiing examples, you do go fast when you go downhill and then when you turn try and go back uphill, it's very slow, if not at all possible. When you go into water, obviously, you hope, as you said, that the agent will start swimming and splashing. And this does typically happen.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Yeah, one of the things that was really cool in all the examples was the water is sort of a great way to see does it understand like what the world is and how objects interact? And that example someone posted of the feet going in the puddle was amazing. But then there was also that example of like a cartoon character. It was more of an animated style who was like running across this kind of green patch of land and then ran into this blue kind of wavy thing that looked like water and he started swimming, which I thought was really interesting. Like, were there particular things you had to do around that for the model to be able to understand how characters should interact in different environments and different styles?

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  34. And from Genie 2 to 3, I think the real world capabilities really increased, right? So on the physics side, some of the water simulations, you can see some of the lighting as well, like a really breathtaking. I think we have this example of the storm on the blog. And that one, I think, is super cool. And it's at the point where like a human who is not an expert will watch it and think it looks real, right? And I think that's pretty incredible. Whereas Virginie 2, it was like, it kind of understands roughly what these things should do, but you know it's not real, right? You can look at it and you can clearly see that it's sort of not completely photorealistic. So I think that's quite a big leap on the quality in that side.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  35. The right term, but we do see that some definitely things like it can infer from if you approach like a door and it makes sense for the agents to maybe open it. So you might see that it's starting to do that, for example. Or there's some better word understanding that happens over time and it just things look Better and more realistic. So, I think these are the trends that we've still observed

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  36. One more question related on Genie 1 to, for example, NLMs, like you have deep sea car one, like they saw in this paper, the longer they keep it running, they suddenly will see these interesting behaviors, like the model will start reasoning or would give like a, oh, I'm wrong on this. I should self-correct. Do you see anything in kind of like this scaling from? To three, do you see any sort of interesting behavior that you were not expecting that suddenly just appeared by increasing the amount of data and the amount of compute? Yeah, I'll just say, I think there is a bit of overall definitely like many generative models we see that improvements happen with scale. I think that's not a secret. And I don't think it's not the same type of intelligence, I would say, like an LLM has. I'm not sure if reasoning.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah, it's also a real time trade off for the guests as well. We felt that because of the breadth and the other capabilities that like a minute was sufficient for this version, like it's quite a significant leap. But obviously eventually you'd want to.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Very cool. And how long is this special memory? I don't know if you can talk about it. You mentioned a minute plus, but is there some sort of like measure that you have? Can you keep it for half an hour? Or what is the limit on that? Is no fundamental limitation, but currently the current design will be limited to one minute of this type of memory.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Every time someone interacts with it for the first time and they test, they look away and then look back, I'm always like holding my breath. And then it looks back into saying, I'm like, whoa. Really, it's really cool

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  40. We know How the Static pretty much then. Build representation. Then what you're looking at Great. But we didn't want to go down this path because we felt it somewhat So we can definitely say that And it does generate frame by frame.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Other models around the same time or more recently didn't have this feature, right? So people kind of indexed to that because they didn't notice that early signs of it in the GE2 work And then for Genie 3, we basically went much more ambitious on the same sort of approach, right? And we made it like a headline goal for ourselves is can we make the memory be what it is, right? We said we want minute plus memory and real time and this higher resolution, all in the same model. And those are kind of conflicting objectives, right? So we sell ourselves this kind of technical challenge. And we said if we target this, then It's just about feasible and it'll be pretty incredible. And then you still don't know obviously it's going to pan out. So then when you get to the end of the research, one, seven months later, to see the samples still is quite mind-blowing, to be honest. So yeah, it's kind of planned for, but still pretty cool and exciting when you see it because it's like research projects aren't sure things are they?

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Yeah, so that's a great question. I'll say a few things. So, the TLDR is it was totally planned for, but still incredibly surprising when it worked that well, right? So that specific sample, when I saw it, it was hard to believe. I actually wasn't sure that the model generating per second. I was like, that took me to watch it a few times and like really check and freeze the frames and look back and check that it was the same, but go back a few steps. So obviously Genie 2 had some memory, right? So this got kind of lost because I mean, Genie 2 came at a time when there were lots of announcements, very exciting announcements. I mean, VO2, only a few days later, it was a busy time of the year. And the main headline act was that we could generate new worlds at all, right? So that was the thing that we wanted to emphasize. But it did have a few seconds of memory. And we had a couple of examples, like I created a robot near a pyramid, looked away, looked back in the pyramid's there, but it's like kind of blurry. It's not perfect.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  43. When I saw the Genie 3 post, it was like, oh, okay, they actually went and did it. But the special memory, the persistence was when I kind of sat up in my chair and I was like, how did that happen? Could you talk a little bit about when did you discover that as an emergent property or was that a specific design goal? What's the backstory on that? Because that feels like a big unlock, Jack. Why don't we start with you?

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  44. A different part of the wall paints and then moves back. And the original paint is still there. And I didn't believe it. There's no way. And then I read, and you read, it was described as a special memory. So the persistence part for me, I'm not taking away from all the other stuff. The interactivity is amazing. But I think, broadly speaking, folks expected that at some point, video generation, for example, would become real time.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Think long term, what's the way to really unlock unlimited environments? That being said, over the course of the project, and originally we started it, I guess in 2022, it was very focused on that one application, but it seems quite clear now that this could have a big impact on all those other areas you mentioned, right? So I think it's like language models in 2021 maybe, you probably wouldn't have guessed like an IMO gold medal a few years later would come that fast, but as a direct application of that technology, right? It was probably, oh, it can help me with my emails or whatever it was. And I think it's really cool to build these kind of new class of foundation models and then see what people can imagine doing with it. And that's one of the really exciting things about sharing the research preview, right? So you've got this kind of feedback. So hoping a lot of these things can happen.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Yeah, I would get basically the same answer in the end with a different journey to get there, right? Which is I'm pursing myself worked in reinforcement learning for a few years before starting the GE project in 2022. And the motivation originally was like that in RL at the time, we had this problem where we'd say, which environment should we try and solve, right? Because once you've already done go, which people thought was years or decades away, and then that was solved in 2016, what's solved, but we reached Super Human 11, 2016. And then StarCraft, three years later, which is not a particularly long time for something incrementally significant. So it was 2021 time, it was a big question of what should we try and do with RL. We know that the algorithms can learn superhuman capabilities if they have the right environment, but we don't know what the environment would be. And so we were working on designing our own ones, right, with colors. But then instead it seemed like the more promising path when you had the first text to image models coming out, whereas like, what if we just

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  47. I think a lot of the applications basically stem from the ability to generate a world just from a few words. And I think for me, kind of like this potential when I started looking at video models, I think it was pretty early when I think it was one of the models where like imagine video, which was modeled by Google Research. But there are a lot of models that they were very basic compared to what we have today. But the ability to simulate something like you look at it and like there's a world that's generated in front of your eyes. And it's amazing that it's happening. And I think at this point, I was very excited about how far can we push that, right? So I think there was one way to do it. And Ginnie is definitely another way to make it a bit more interactive. So I think all of the applications basically stem from this core capabilities. So it can be entertainment, of course, as you said. It can be training agents. It can be helping agents to reason about the world, education. So I don't think any particular application.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Yeah, I'll just add to this that I think the real time, so a component is really important. And not many people experience it firsthand, but we really tried in the release to at least have a few trusted testers interacted with it and also get the feel of it by adding these overlays that show what happens, how people can use the keyboard to control it. And I think there is something magical about the real-time aspect. I felt it for the first time when our model, like the game engine model, started working fast enough and we were just like, oh my God, it's actually, I can actually walk around. And it was the mid of a wow moment. And yeah, I think there is something when it responds immediately that is really magical. I think that's kind of sparked the imagination of many people when the Doom kind of simulation came out. And here we really wanted to push it to somewhere. We weren't sure it's going to work. So it was definitely at the edge of what's possible, I think. That's how we felt. So we just said, yeah, let's try and see.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  49. And so that we felt that across these different projects we had quite a lot of interesting things that would naturally kind of combine and we could basically take the most ambitious version of the combined project and see if it was possible. And fortunately it was. And quite, I think the timeline is probably the bit that surprised many of us because obviously we set ourselves these goals and like we tried very hard to achieve them. You can never be totally sure how it's going to actually feel when you've got to that point. I think it ended up being something that resonated with people a lot more than maybe we expected, but we were always believers.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Sure, yeah. So I think you kind of highlighted a few capabilities, sort of the length of the generation, the consistency of the world, maybe diversity as well of the kind of things you can generate. I think the main thing is that obviously we made progress in quite a few different fronts, right? In separate efforts. So we had this Gene 2 project that was much more sort of like 3D environments that it could generate. And it wasn't super high quality. It felt like it coming from Genie 1, but it wasn't the same quality as things like VO2, which the state of the art video model at the time came out in December exactly the same time. It came out a week later than Genie2. And obviously internally there was a lot of discussion between the two projects about the different directions we were pursuing. And then Joomi had also worked on game and gen, right, which is the Doom paper, as people know it, which I think you guys also wrote a nice piece on straight after that came out. So I think that also attracted a lot of attention.

    2025-08-16 · a16z Podcast · Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building · IDENTIFIED FROM THE TRANSCRIPT · source