YouSaid · the spoken record

Misha Laskin

lines on the record
66
first
2025-07-17
most recent
2025-07-17
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. And there was some big noticeable difference between the two. Now, that's great, but I think the downside of that is that does humanity's last exam actually matter in any meaningful way for an end user? And I would argue that some weak correlation, but the answer is most likely no. And so you have to build the tools and train for the things that users actually want. I think that there's sort of no way around that.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  2. That kind gives you a nice general base to start from. But then to drive a capability kind of depth-wise, like if you really want this reasoner that has search tools and ability to call these long reasoning context models and other tools that might want to interact with, like, oh, when do I read from JIRA? When do I read from another tool? This is kind of a reasoning problem. If you train with those specific tools in mind, that's typically what people refer to when they say tool use. Like they actually train for a specific set of tools and really drive the capabilities for those tools. So these are the kinds of research problems that you need to solve in order to build the overall system that's the best in the world. It's not any one thing.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  3. When I say long context reasoning, I don't mean, I actually mean kind of small models with very long contexts that are able to go into giant code bases, sort of suck up as much information as they can, and reason over an output-relevant stuff, basically. So it's almost like neural retrieval. There are capabilities like Tool Use and Multi-Hop reasoning. So this is more for a, you have your agent and it's designed with some tools. And there are two ways of training agentic models. One is in this very general way where you just train it on thousands of environments and make it like the most general agent possible. And that is kind of almost like the pre-training of agents. And that's sort of what that's what a Frontier Lab does. That's what there's a new release from Kimi2. That's kind of what that model does. And that's definitely part of it. But in order to...

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  4. Customer really closely and develop differentiated product almost independently of the models that are towering it. But then you also need to innovate on the research in terms of agent design and model training to actually drive the capabilities that you want to see out of the system. And this becomes an evaluation problem, which is basically at the heart of any frontier lab as well. This is, I think, the least spoken about part of what Frontier Labs do, but possibly the most important, which is figuring out how they evaluate. like what makes Claude magically feel better at code than another model out there. They did something right in their evaluations. So when you look at this problem specifically, there are different capabilities that you need to train. And what we do is really post-train models where we really focus on post-training today. Some of these things are long context reasoning.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  5. There are a few things. So, I think this is kind of where it is so important to co-design research and product. Because as a researcher, you'd go in and say the answer is entirely in the agent design or the model or something like this. And as a product person, you would say, well, it's in these product differentiators like being able to draw not just from your code base, but knowledge that lives in other sources of information or being able to learn from the engineering team to offload their tribal knowledge. So an engineer can go in and teach Asimov like, hey, we deploy our, you know, when we say environment jobs on our team, we mean this specific thing, which we mean kind of Google Bath jobs. So now another engineer asks a question about environment jobs in the future, the system just knows what they're talking about. A lot of knowledge is stored in engineers' heads. And I think you need both of these things. You need to understand your

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  6. Would prevent a superintelligence from actually working within an organization? It's really this kind of understanding and being able to ingest from a lot of sources of information and from the team. And once you have that, then the action part, I think, becomes, I don't want to say trivial, but a lot easier. Like to me, it seems like really 20% of the problem is teaching these agents how to act and it's more or less solved

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  7. Complexity and it will provide you an answer at the level of what that principal level engineer would have given you or in the future as the product expands to other categories what the person who's most embedded in the organization understands. And of course, once you have that solved, it begets much more reliable agents that act for you as well. But I think the world today is focused on, I would say, 80% kind of action, 20% understanding. So 80% code generation, 20% comprehension. The actual problem is exactly the opposite, that when you look at what an engineer does in an organization, 80% of the time they're spending trying to comprehend complex systems and collaborating with teammates. And what is collaboration? It's usually someone asking someone else a question about a system that they don't know. And so that, I think, is kind of the problem at the heart of what

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  8. Yeah, the meter report was very close to what I've been hearing when talking to engineering leaders within larger organizations. And it's not just enterprises. It's, I would say gross stage startups. It's any kind of engineering organization that has a sufficiently complex code base and sufficiently large team that no one engineer can have the entire code base kind of in their heads. And so reflection is one of those places as well. We use our product actively because the training large language models is complex and there's the large language model code base, there's the product code base. Knowledge is kind of scattered across engineers. It's not just in the code base that exists in your chats and project management tools and other places where Novel lives. And so what we're effectively building towards is this kind of omniscient Oracle for organizations that you can go in, ask any question at any level.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  9. So Asimov is the best code research agent in the world. Coding tools, and this is enterprise specific. So I think the world is different with startups. But within enterprises, when they're adopting coding tools and you see the impact that this is having on their actual productivity, and I think it's much lower than people expect. So in fact, it's sometimes negative, sometimes negligible.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  10. Right. So Yanis, my co-founder was the overall RL lead for Gemini at the time, for one and 1.5. I was working very closely with him on his team. Yeah, it was a really exciting kind because we went both of us from being reinforcement learning researchers to training large language models at scale. And we kind of saw at the end of that project of what's to come, which was Gemini 1, 1.5 lands. And it became pretty clear to us that the next paradigm and effectively the final paradigm that we need to have in place before what people used to call AGI or now I think the goalposts have shifted to ASI is reached is just figuring out how to scale reinforcement learning on top of large language models. And the first instances of that have been happening over the last year. I think we're still actually a lot earlier than people think.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  11. Is such a thing as kind of the root science of our time. I think a lot of physics has a field. It's very interesting, but it's crystallized a lot more than a new dynamic field that was being born out of nothing. And AI to me felt like it was going through the moment that physics went to maybe 100 years ago that when I do problem sets, but I did problem sets in physics. And the most exciting stuff that I was working on there was basically the things that people were discovering 100 years ago. I saw it kind of happening in front of my eyes and I just decided that that was the science to bet on. And in particular, because it was Alphabo that inspired me because it was just unbelievable to me you could train a neural network to have such immense kind of basically reasoning capabilities. This thing was able, it was super intelligent within the realm of Go. Yeah, I decided that I needed to kind of get myself into the best reinforcement learning.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  12. So, when my interest in Pixie started, that was probably around middle school, and it really, I think, became the thing I wanted to do in high school. And the reason physics was so interesting was because it kind of seemed like the science that was at the root of many of the things that became impactful. So I was reading about the history of the transistor and it was invented by a group of theoretical physicists. I was reading about how GPS works. So it turns out you need special relativity in order to accurately account for spatial coordinates using GPS. And so I felt that physics was kind of the root science to pursue. I went in and studied it, got my PhD in it. At the same time, I started seeing kind of deep learning takeoff and really saw AlphaGo happen. And my sense was that I want to pursue the kind of the root science.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  13. Yeah, as a kid, I became really interested in physics, theoretical physics. I mean, probably a byproduct of a Russian kind of Israeli-American and moved around. And then when I landed in the States, it was kind of in a desert in Washington State, learning a new language. And so I had a lot of time in my hands and bumped into my parents had defined them lectures in their library. And so I spent a lot of time just reading what was on the shelf and bumped into that and got really interested in physics.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  14. ASI complete. I think, and this is kind of our approach and reflection, is it makes a lot more sense to be focused and co-design those two things together, the product of the research.

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  15. What is superintelligence more concretely? How is it going to be deployed? What is it actually going to look like in people's hands and build backwards from there? So I would kind of. Say that that approach is more kind of co designing product and research together. Now, the kind of benefits of that approach is that you're kind of optimizing for real problems. The const to it is that you have to be a lot more focused, right? Because your product kind of defines the sort of capabilities that you want to draw out of the system. And you have to start out a lot more focus before expanding across other product categories and other capabilities. So I would say that on the spectrum of companies that are kind of super intelligence and just a research lab and then figure out what the product is once it's built as opposed to co-designing product and research together to build very powerful systems in what I would call kind of ASI complete categories. You can pick something that is maybe too small of a category to draw out a super intelligence. As long as you pick a category that I would say is kind of big enough to

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT

  16. At a high level, it's fairly synonymous, but maybe there are different ways of thinking about how to build super intelligence and what that might look like. I think on one spectrum there is an academic way to look at it, which is, in some sense, to some extent superintelligence in that sense has already been achieved. So AlphaGo was a superintelligent system, and there were other systems during that time that were built that were super intelligent in narrow domains. And I think you can go for the goal of building a very broad superintelligence by kind of locking yourself up in an academic. Or it's not really an academic, but kind of an industrial lab with that is sort of kind of decoupled from product or customers and kind of maps out all the benchmarks that are out there and build superintelligence that way. I think that is one approach. I think the other approach is to kind of think about

    2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT