YouSaid · the spoken record

Brian He

lines on the record
77
first
2025-09-15
most recent
2025-09-15
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I think lots of kind of biological evals that you can kind of add onto these models over time that are really tangible textbook examples as opposed to I think what the kind of early generation of models do today, which is very quantitative things like mean absolute error over the differential express genes and stuff like that. Those are ML benchmarks and we want to increase the sophistication into something that you could explain to an old professor who has never touched a terminal in their life.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  2. That would be sort of one really kind of classic example. And then you could go do the inverse if you have a stem cell, can it discover neurogen 2, ASCL1, Myod, can it find differentiation factors? We'll turn that into a neuron or into a muscle cell or so on. And, you know, these are kind of classic examples in developmental biology, but you could also use this to try to discover or kind of recapitulate the mechanism of action of FDA approved drugs, right? And so you could say, for example, you know, if you kind of inhibit her too in, you know, breast cancer cell states, right? It would be, you know, you would get this type of response. Or it could predict the certain clones that will be able to kind of be more metastatic or, you know, they'll be more resistant and they'll lead to minimal residual disease.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Well, the good thing with biology is we have a lot of ground truth. There are entire textbooks that describe cell signaling and cell biology and how these things work. And so even without a virtual cell model at all, right? If you went into ChatGPT or Claude and you basically, you asked us some question about receptor tyrosine kinase signaling, it would have an opinion on how that works, right? And so I think you would want the model to be able to predict perturbations that are kind of famous canonical examples of biological discovery. So I'll give you an example. If you load it into the model, an IPSC, kind of an induced Plipon stem cell stay or human embryonic stem cell state and fibroblast cell state, right? Could it predict that the four Yamanaka factors would reprogram the fibroblast into a stem-like state, right? And essentially rediscover from the model something that won the Nobel Prize in 2012.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  4. The GPT, I will say GPT 3 moment going to look like, and by that I mean sort of a public release that alters the public's conception of just what's possible here from a capabilities perspective and also inspires a whole new generation of talent to like rush into biology.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Need to curate public data. You need to generate massive amounts of internal private data, build the benchmarks, and train new models and build new sort of architectures. I'm kind of doing these things full stack. And we'll just kind of attack this hill climb over time.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I find it helpful to frame these in terms of GPT 1, 2, 3, 4, 5 capabilities, right? And I think most people would agree where somewhere between GPT one and two, right? A lot of the excitement was that we could achieve GPT-1 in the first place, that you could see a path with scaling laws of some kind to kind of make successive generations where capabilities would improve. But these are, you know, with like our Evo kind of DNA foundation models that we developed at ARC with Brian He, right? One of the things that we've seen is that, you know, these are really kind of these genome generations are like quote unquote blurry pictures of life, right? We don't think if you synthesize these novel genomes, they would be alive, but we don't think that's actually also impossibly far away. We'll just have to kind of follow these capabilities. We're generating, we're taking a very integrated approach to attack this problem, right?

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So, if the goal is I'm oversimplified for you, like if we wanted to get to the alpha pho moment where it kind of gives you a useful structure, folded structure 90% of the time to use your data point, you wanted to take that comparison in the virtual cell model. And we said, okay, 90% of the time, if I ask the model, I want to. And it's going to give me a list of perturbations. And let's say that at 90% of the time, those perturbations in fact result in the shifting experimentally, in the shifting from cell state A to cell state B. How far away are we from that alphafold moment for virtual cells?

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Make a new AI like vertically integrated AI enabled pharma company, right? Which, you know, I think is obviously a very exciting idea today, but I think in many ways the kind of pitch and the framing of these companies precedes the fundamental research capability breakthroughs. And that's what we're really invested in at ARC is kind of just making that happen along with many other amazing colleagues in the field to just make this possible for the community.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  9. State A to sell state B. There are these three changes I need to make first, then these two changes, and then these six changes over time. And we kind of want models to be able to suggest this. And the reason why we scoped Virtual Cell this way is because we felt it was just experimentally very practical. You want something that's going to be a co-pilot for a wet lab biologist to decide, what am I going to do in the lab? We're not trying to do something that's like a theory paper that's really interesting to read where the numbers go up on a ML benchmark, but you practically can decide what are the 12 things that you're going to do in the lab in 12 different conditions, right? And actually just test them, right? And then that's how we kind of enter the kind of the lab and the loop aspect of model predictions to experimental measurements to, you know, kind of improved or RL'd or whatever model kind of predictions again. And the goal is to be able to do in silico target ID where you can basically figure out new drug targets, figure out then the compositions, the drug compositions you and you need to actually make those changes. I think if we could do that, we could.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Ways to zero shot these binders. But ultimately, what you're trying to do with these binders is to inhibit something. And then by doing so, kind of click and drag it from kind of toxic gain of function, disease-causing state to a more quiescent homeostatic, healthy one, right? And the thing that is very clear in complex diseases, right, where you don't have a single cause of that disease is there are some complex set of changes. There's a combination of perturbations, if you will, that you would want to make to be able to move things around. Now, you know, people talk about this classically as things like polypharmacology, right? But, you know, I think we're moving from a, oh, this thing happens to have, you know, a whole bunch of different targets kind of by accident to we have the ability to manipulate these things commentorily in a purposeful way, right? That to go from

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Cell and so on, and you know that you can kind of move cells across this manifold, right? Sometimes they become inflamed, sometimes they become apoptotic, sometimes they become cell cycle rested, they become stressed, they're metabolically starved, they're hungry in some way. And so if you have this sort of this representation of universal sort of cell space, right, can you figure out what are the perturbations that you need to move cells around this manifold? And this is fundamentally what we do in making drugs, right? Whether we have small molecules which started out as natural products from boiling leaves or antibodies when we injected proteins into cows and rabbits and sheep and took their blood to get those antibodies. We were basically trying to get to more and more specific probes, right? And we had experimental ways to kind of cook these up. Now we have computational.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  12. I would say the most kind of famous success of ML in biology is Alphafold, right? And this solved the protein folding problem of when you take a sequence of any amino acid, what does the protein look like, right? And it's pretty good. It's not perfect. It certainly doesn't simulate the biophysics and the molecular dynamics, but it gives you a sense of what the end state is with 90% plus accuracy, right? And that's the alpha-fold moment that people talk about, right? Where anytime you want to, you know, work with a protein, if you don't have an experimentally solved structure, you're just going to fold it with this algorithm. And we kind of want to get to that point with virtual cells as well and the way that at arc we're operationalizing this is to do perturbation prediction, right? Where the idea is you have some manifold of cell types and cell states, right? That can be a hard cell, a blood cell, a lung.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  13. You have to bet on what you can skill today, right? We're able to scale single cell and transcriptional information today. We're able to add on protein level information over time. We'll need spatial information, spatial tokens, and we'll need temporal dynamics as well. And we'll, you know, I kind of bucket things into three tiers. There's invention, engineering, and scaling. And there are certain things today biotechnologically that are scale ready. And then there are things that we still need to invent, right? And that's part of why we felt like we needed a research institute to be able to tackle these types of problems, that we weren't just going to be an engineering shop, that's just trying to scale single cell perturbation screens, right? That would be interesting, but in three years would feel very dated, I think, right? And so there's a lot of novel technology investment that we're making that we think will bear fruit over time.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Tremendous amounts of RNA data that will read in kind of like what's happening at the protein level at some sort of mirror echo. And then that can be the case for metabolic metabolic information as well and so on.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Many ways are the scaling laws that you get in data and in modeling. I'll give you an example, right? There's a lot of discussion in molecular biology about how RNAs don't reflect protein and protein function. And so, well, we don't have proteomic measurement technologies that are nearly as scalable as transtroptomic measurement technologies today, like that's the single cell resolution, certainly, but we're getting there and you can layer on certain nodes of protein information that you can add on top of the RNA information. But in many ways, the RNA representation is a mirror, right? It might be a lower resolution mirror for what's happening at the protein layer, but eventually what is happening in protein signaling will get reflected in a transcriptional state. For an individual cell, this may not be very accurate, but when you imagine the massive data scale that we're generating in genomics and functional genomics, right, you start to gather

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  16. People talked a lot about this in NLP as well. There's this long academic tradition in natural language processing, right? And then it was just weird and non-intuitive and intensely controversial that you could just feed all this unstructured data into a transformer and it would just work. Now we're not saying this will just work in all the other domains, including in biology, but I think there is this controversy.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Much of this is the fact that you talk about we speak biology poorly or with a very thick accent, how much of this is like if you're training on an image, we can see the image. And so we can see how good the output is. What about all the things in biology that we can't see or don't even know exist yet? Like how can we create a virtual cell and maybe we should come back to what a virtual cell model is, by the way, for the lay audience. Like how can we create a virtual cell model? We're not even sure if we understand all of the components that are in a cell and how they function.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Goal is to figure out ways that you can actually interpret the weird fuzzy outputs that the model is giving you. And I think that's what slows down the iteration cycle is you have to do these lab and the loop things where you have to run actual experiments to actually test with experimental ground truth. And, you know, I think increasing the speed and dimensionality of that is going to be really important.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Natural language and video modeling is easier than modeling biology, right? And to some degree, if you Understand and learn machine learning and how to train these models, you have already learned how to speak. You already know how to look at pictures. And so your ability to evaluate the generations or the predictions of these models are very native. We don't speak the language of biology. At very best with an incredibly thick accent, right? So you're training these DNA foundation models. I don't speak DNA natively. So I only have a sense of the types of tokens that I'm feeding into the model and what's actually coming out, right? Similarly, these virtual cell models, you know, I think a lot of the

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Yeah, these are absolutely part of it. We have two flagship projects, one trying to find Alzheimer's disease drug targets, the other two to make these virtual cells. And I think it's not just the people and the infrastructure, but also the models will hopefully literally make science faster that you could do experiments at the speed of forward passes of a neural network if these models could become accurate and useful.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  21. That's cool. So, that's sort of the original hypothesis for the Arc Institute is if you can bring multiple disciplines together to increase the collision frequency, as you said. And if one could remove some of the cross incentives that may exist in sort of traditional structures, the combination of those two things will make science faster.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  22. That's part of it. And I think the other part is folks have their own incentive structures, right? They need to publish their own papers. They need to do their own thing and make their own discovery. And you're not really incentivized to work together, I think, in many ways in the current academic system. And a lot of what we've done is to try to have people work on bigger flagship projects that require much more than any individual person or group or idea.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  23. To try to see what happens when you bring together neuroscience and immunology and machine learning and chemical biology and genomics all under one physical roof, right? If you increase the collision frequency across these five distinct domains, there would hopefully be a huge space of problems that you could work on that you wouldn't be able to. Now, obviously in any university or any kind of geographical region, you have all of these individual fields represented at large across these different campuses. But people are distributed and you want everyone together.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Whose fault is that? Now, that is a long one. We should get into it. We should get into it. It's really multifactorial. It's this weird Gordian knot that ultimately comes down to incentives, right? Comes down to people talk a lot about science funding and how science funding can be better, but it's also about how the training system works, right? How we incentivize long-term career growth, how we try to separate basic science work from commercially viable work. And generally the space of problems that people are able to work on today. I think things are increasingly multidisciplinary. It's very hard for individual research groups or individual companies to be good at more than two things, right? You might be able to do computational biology and genomics or chemical biology and molecular glues. But how do you do five things at once? It's increasingly hard. And we really built archives and organizational experiment.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Want to make science faster, right? You know, we can frame this in high-level philosophical goals like accelerating scientific progress. Maybe that's not so tangible for people. I think the most important thing is science happens in the real world. If it's not AI research, which moves as quickly as you can iterate on GPUs, right, you have to actually move things around. Adams clear liquids from tube to tube to actually make life-changing medicines. And these are things that take place in real time. You have to actually grow cells, tissues, and animals. And I think the promise of what we're doing today with machine learning in biology is that we could actually accelerate and massively paralyze this. And so our moonshot is really to make virtual cells at Arc and simulate human biology with foundation models. And, you know, we'd like to figure out something that feels useful for experimentalists, people who are skeptical about technology.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  26. My goal is to really try to figure out ways that we can improve the human experience in our lifetime. There are a few things that if we get them right in our lifetime, we'll fundamentally change the world.

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source

  27. I want to make science faster. Our moonshot is really to make virtual cells at Arc and simulate human biology with foundation models. Why are we so worried about modeling entire bodies over time when we can't do it for an individual cell?

    2025-09-15 · a16z Podcast · Faster Science, Better Drugs · IDENTIFIED FROM THE TRANSCRIPT · source