YouSaid · the spoken record
Misha Laskin
- lines on the record
- 66
- first
- 2025-07-17
- most recent
- 2025-07-17
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“I think coding is a Sarah as well. This one, I think, will take longer than people thought as well, because again, enterprise, organizational problems are just much different than the benchmarks that we have today. But I think it will be one of the faster ones. So I don't think that that's kind of a decade out. That's within the next, you know. A dozen, dozens of months kind of thing. So I think the next sort of generational companies in coding are definitely being built today.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Of models. It's a model called Codex, which was post trained for that environment. The deep research models, like that's a specific environment. They're also post-trains for that environment. And I think we'll basically see more and more of that, that any category that has a sufficiently large business around it that requires an intelligence core to power it. There will be all sorts of interesting design decisions at the research and product level of how do you actually gain the most performance out of this particular category. I think we'll kind of see a lot more kind of depth first players emerge over the coming decade or so.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“You go and try to solve it with some combination of imitation learning and reinforcement learning. And when you look at all those projects, these were basically things that were called strikes within DeepMind. And each strike within and outside of DeepMind was a bit of a snowflake. Like the reinforcement learning methods and environment setup for Go was at a high level, conceptually similar, but in the detailed implementation level, very different from StarCraft, very different from Dota 5. And so I think that that sort of, we're going into every big category having a different environment, right? And different kinds of agents with different tools. And that means that you'll need to, you'll have general base models that you can start with, but you'll need to post train things in specific ways for those categories. And we're starting to see that already in the sense that the model that powers OpenAI's codecs is not the O series.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Breakthroughs that need to happen, but more or less there will be a blueprint for how do you build a super intelligence in a particular category. Actually going in and deploying it and building it for specific categories of work, there are going to be a lot of product and kind of research innovations specific to those categories that will probably make us a multi-decade thing. So I don't think that it's a couple of years from now and GDP starts growing 10% year over year globally. I think we're actually going to get there, but it's going to be a kind of multi-decade endeavor. I tend to kind of see a lot of patterns now in kind of real world deployment with reinforcement learning research as it worked again before large language models. And before large language models, it used to be kind of you pick an environment, like you pick Go, you pick StarCraft, you pick something else.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“I think we're a lot earlier than most people think that this is going to be one of those areas where the technological building blocks outpace their deployment. And so, yeah, within the next couple of years, the blueprint roughly for how to build ASINs will have been set more or less. Maybe there are still some efficiency.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Already places where customers are pulling us in different directions. It's just kind of a matter of whether you engage on that today or not. And I think that the risk that a startup has is that you see a lot of shiny areas where you can go and you start kind of going diffuse before you've really nailed a category. So I think it's really important to be focused and not diffuse in the short term. And that if you kind of build the right as we kind of think about as a contextual core for an organization, in this case an engineering organization, then you can naturally start expanding that into adjacent areas of work in that enterprise.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“The way I think about it is more first just build, not trying to get too ahead of yourself of kind of just first build the kind of most depth-wise comprehension system for software engineers. This will naturally induce more reliable coding agents. You can plug that in as an MCP to your favorite IDE or coding agent, or use one of our own, right? You can kind of plug that into whatever surface area makes sense for the customer. And then sort of naturally start seeing where you're getting pulled from there. And the reason I think this will work is because this is kind of what we're already seeing, right? How do you make the system useful for product managers or technical support people? And then I think moving on to things like sales or something like this, but there are”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Thing that makes coding as a category special is that it's not synonymous with software engineering. It's just kind of how we think about the market today. The reason coding is special is if you believe that the way a language model will interact with almost any piece of software is through function calls and therefore code, then if you build very capable reasoners coding reasoners that are sort of purpose built for organizations. So you've solved the kind of long context. How do I reason over a bunch of disparate source of information problem? And I can act on pieces of software through code, then you've kind of built a system like the technology that will generalize at least operationally across other categories of work.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, maybe. I mean, that's kind of actually when Janis and I were starting the company, we were thinking about, well, what like”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“You can do it in like this teleop way. Like it's sort of these are things like the things that we try to train robots to do are very intuitive for humans as well. I mean, actually more intuitive for humans, right? People are master manipulators. So you can have a lot of kind of teleop like data collection. Things that we want language models to do are sort of at the level of it's really hard to collect data of the chain of thought process that goes on in like a human's head when they're trying to solve some task. And that's kind of the data that you need. And so for that reason, I think language models favor this more like synthetic data RL-like approach where, well, it's easier for us to verify whether the thing was done or not than it is to actually generate all that data from a person specifically.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Quadruped is moving at this velocity without damaging its body, kind of thing. Maybe it's a bit of a roundabout answer to the question, but it's that I think these two fields are very different in the data distribution that they support and the kind of imitation learning data for language models is, of course, the internet. It's, of course, people who've gathered all this data on how we write and so forth. And so aside from that, when we're generating synthetic data, There is the only scale the path is really reinforcement learning. The other thing that I'll say here is that. You're collecting data for robotics.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“More hackable than anything you have in language models. So the same problems blow up and become much larger. And so that's actually why I changed to language models because I felt that this was a fundamental problem, but we now have these confounding factors of these noisy signals coming in. I think that in at least in a generalizable way, that's why it's really hard to get reinforcement learning to work with robotics. The one place where it really does work.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Design reward functions or models for robotics or for 3D video games like Minecraft or something like this that have, I think, similar challenges scientifically. The challenge is that if you think that language model rewards are hackable, vision language model rewards or other sensory signal rewards are infinitely more hackable. They're much more short-lived than rewards. You can think of language as just a compressed representation of the world that we have that we are kind of magically have to start with. Whereas if you're processing pixels or sensory motor signal, this is raw signal that has a lot more noise in it. And so if you train a neural network that is sort of trying to detect whether this thing was manipulated correctly or this thing was moved correctly, then that thing is just infinite.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“So I was actually a robotics researcher in reinforcement learning, the Peter at Beals Lab is a robotics lab. And it was, you know, it was a mixture. Peter's lab was always around the intelligence problem and robotics is being a domain where you study it. And one of the, you know, the reason I came to lead reward models for Gemini was because that's the question I was studying with robotics. We have these RL algorithms for getting robots to do some very narrow tasks like moving blocks and various kind of narrow tasks and simulation. And the question was, well, how do we get generalized manipulators and just how do we build this all into one system? And it seemed like the rewards were bottlenecks. So a lot of what I was studying before starting getting into language models was how do we”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Great user experience on top of someone else's model. And I think some of those dynamics are probably starting to play out as well. Like I think that there are some question marks around if you're on this critical path category and you don't have your own intelligence, how do you compete when your competitor can just basically subsidize their product a lot more than you can, right? Because you're effectively as a startup that's building on top of these things to grow quickly, you're subsidizing the margin that an anthropic or Gemini or whatever is making. And Google and Anthropic and OpenAI can subsidize their products a lot more than you can. So I think that companies that don't own their intelligence or are not kind of deeply integrated into a customer in some way that makes them hard to remove find themselves in this pretty existential place as it becomes clear.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Think that's guaranteed that a big lab can buy their way to the end user because the fundamental problems of your research team being far away from your product team will still be true and the company having 100 different focus areas will still be true. So I don't think that acquiring an asset will change that fundamentally. But it does underscore the importance of verticalization. And then from the startup side, I think it actually puts companies that are in these kind of critical path categories like search and coding in a pretty existential place if they can't build their own frontier models. Not all frontier labs will be able to verticalize correctly, but some will, maybe one will, and that's going to be enough, I think, to kind of take the thunder out from a company that's built a”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“We've seen this verticalization basically happen across categories that are material to frontier intelligence. And one could argue that the first verticalized category was actually search, like through ChatGPT. That's sort of a place where OpenAI verticalized first. And coding has obviously emerged as another kind of frontier level category that could like all these companies have aspirations of Yeah, ASI, and I think being basically trillion dollar companies or more, I don't think that it's really the economics that are the driving factor, but it's more that if you want to sustain frontier research, that's kind of what you have to become. And so coding has clearly become one of these categories where verticalization is extremely important. And I think that there are kind of two sides of the story. One on the frontier lab side.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“The reality is that. Deployment of it is half the problem, which goes back to kind of evaluating on customer problems and building product together with the models.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, right. So in the parts of work that are really meaningful that you want to see these things driving meaningful increase in GDP. And I think the only way you'll see that is if you go into a company and there's kind of a universal understanding that, yeah, my engineers are double digit percentage points as a whole, every single one of them more productive. That's the kind of thing that if that starts happening across every field, then you'll see double digit increases in GDP. I think that the kind of benchmark matching that's, and it's a bit different than benchmark maxing used to be before because you have benchmark maxing that is weakly correlated to customer outcomes, but it still looks very similar to taking a board game, training RNRL agent on it, getting kind of a landmark result in superintelligence, and then making a claim that superintelligence is solved, I think.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Where you have A handful of these superintelligences for categories that matter, maybe subsumed into one model at some point, but at first they'll probably be, again, I think there will be a few companies that have kind of product model coupling that is super intelligent in different categories. I think an example of, again, starting to see the first glimpses of superintelligence, but in a way that hasn't really transferred to anything meaningful yet is, well, we have these like super intelligent test acres now. Amy, the Amy benchmark is completely saturated. Code forces and other competitive coding environments, the models are almost best in the world and within the year will probably be just the best in the world. And yet competitive coding agents. Then you go into a company and you ask them, have these things been helpful? And they say,”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“We have these systems that we can make amazing at very narrow tasks. We have absolutely no answer for generalization, like zero. And we went from that to things that feel like they're generalizing. They're certainly generalizing much better than anything we had before. But it's likely because the training distributions are so broad. So at least the way I think about it is more kind of output as a user is the system super intelligent in some meaningful categories of work. And then from a research perspective, is it obvious how to make it general for anything that you might care about? And at that point, again, it's just a matter of economics. Maybe there are some categories where collecting the data is so expensive and the return on investment is low, where effectively just better to have craftspeople than super intelligent AIs. So I think we're moving into this kind of jag world of jagged super intelligence.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“I think I kind of. Have a similar viewpoint to the people you describe. I think the generalization capabilities of these things has been weaker. First of all, it's all mind blowing if this exists. So we went from fundamental existential crises in generalization. Like this was the feel of reinforced learning before language models was”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Near superintelligent systems. And if OpenAI and Deepind had sunk more compute into them, they would have definitely become super intelligent. It's just that at that point, it didn't really make sense. Economically, like, why would you do that?”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“The reason why I would say that the problem ASI would have been solved by then is because you've kind of at that point it's just a matter of operationalizing like what you know you know it's just so happened that these particular categories like you might have a superintelligent front end developer because there's so much data distribution for that on the internet and it's easier to make synthetic data for that but at that point you have the recipe and it's just a matter of making kind of economic decisions of is it worth syncing in X amount of dollars to get the data in this category to get kind of something close to super intelligence there an example of that is what happened with reinforcement learning before language models effectively the blueprint for building superintelligent systems was developed it happened with the Atari games alpha go you know then Dota5 and AlphaStar were”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“I still do believe that's true. I think that where I think we'll be in a couple of years from now is that there will be kind of definitive superintelligence in. Some meaningful categories of work. And so, for example, when I say coding, I don't mean all of coding, but there will be a super intelligence within some kind of slivers, some meaningful slivers of coding that are driving, I'd say immense progress in the companies that can benefit from that.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“So as long as your train distribution looks something like what you Users will experience it as generalization. I think there is some generalization that happens in these models, but we probably as users overestimate it because we don't actually see how they were made. But then, you know, yeah, if you saw, oh, this synthetic environment was actually very similar to the thing I was asking about. So it makes sense by the model would be good at that.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Personally, maybe the take is not very hot. I'm very bullish on it because how else are you going to, maybe the hot take is that there's no such thing as generalization. There's just bringing the test distribution into train.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“They don't discern at all along your, say, reasoning chain which part of the reasoning was correct and which part was incorrect. And so that's why you get these reasoning chains that are kind of garden path meandering. Like they'll explore all sorts of things that are completely unnecessary and don't look at all like the kind of structured thinking that a person would have. That's how the algorithm works. It doesn't actually look at there's no credit assignment step on any atomic level. And so that I would say falls into more algorithmic progress bottlenecks.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Bound field. And then there's also kind of algorithmic progress in terms of the RL methods we have today are quite bad, I would say, at exploration and credit assignment. They're sort of addressed like the fundamental algorithms are take the things that work and make them happen more frequently and the things that don't work and make them happen less frequently.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“In the methods that leverage the rewards I have today. And so examples of that are basically every synthetic generation pipeline is some example of this. So it's a messy problem, but I think it's fundamentally like we're in a reward-bound world. I don't think there's going to be any breakthrough that all of a sudden we go from we didn't have rewards for everything to we do because the reward problem in itself is at the time i called i thought it was agi complete now i'd say it's a complete but by the time you have a neural network that can accurately verify any outcome that is probably a super intelligence and so then it goes back to again evaluations what if you're training your rewards your reward models on something like what are you evaluating against what are the tasks that um you want it to be good at so that's kind of um how i think about it i think it's a fundamentally reward model uh or rewards”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“By their exploration abilities. That's the only thing, right? But if today We certainly are not in this world where we have clean rewards for every task we could imagine. And so we're kind of making as a field, have to make sort of various shortcuts and compromises to that. So you'll have things like LLM is judge with different rubrics. And that works to some extent, but it inevitably annoisy or like stochastic reward inevitably gets hacked. So you kind of need a lot of these. And there's only so much you can extract out of them. Then you have sources that do have ground truth rewards, but there are not many of them. And so you have to hope that by optimizing against those, you'll get some generalization effects. And so I think that the fundamental problem is like the reward problem. You can either go in and say, I'm just going to, all I'm going to focus on is kind of rewards, or you can say I'm going to take things as they are and just be more creative.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“I'd say that there are two categories, or one would think that things fall into. One is more around the limitation to the problem structure, and the other one is, well, maybe the structure is fine, but you need algorithmic advances to really drive the next frontier forward. I'd say some mixture of both, but the biggest way I could is on the problem structure. The thing that I lent for Gemini was reward models. I built out the reward models that were used to post-train Gemini 1.5. And I thought is that if you have a reward that accurately basically describes the outcome of any arbitrary task that you throw at it, then that's it. At that point, it's just algorithmic advances, but even the very simple RL methods we have today will be able to get a lot out of this. Like they'll only be bound.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“But that was very much the ethos that we can come in. We don't need to pre train. You can get by with two orders of magnitude less compute and really get something out there that's really good. I think that roughly speaking, you won't need the amount of compute that I think Frontier Lab needs today as you're focused, but you'll still need kind of an order of magnitude less. I think that the capitalization requirements are still high. There's no way of avoiding that. But I'd say they're an asymptotically they're probably the same. But asymptotically, the idea is that at that point you just have a generational business that can raise capital off of that.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Exactly, right? So you can get into it and kind of build out both kind of the product and a research arm. Our thought was that this was the time where you can actually start a generational frontier lab that does not need to be coupled to a big cloud provider because if you do it right, you'll actually be able to generate sufficient revenues to not have to be acquired or find some strange deal where the cloud provider kind of owns you. And that was kind of the model, I think, of a lot of what Frontier Labs looked like pre-LLMs. I think we're already starting to see that this kind of more of a field-wide thing independently of reflection, right? When you look at how fast anthropics revenue is growing, I think they're kind of in this spot where it's like a massive revenue generating business that's growing at an unprecedented.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Surprise me. They're actually the models are better than I thought they would be And we thought that you can just focus on, you know, we're in this brief period in history right now where the RL flops are still manageable. Like you can really have a best in class product if you're focused. And yes, you'll need to put, you know, you still need a decent amount of GPUs. But from a Flop's perspective, it's nowhere near where pre-training is.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“I think that's roughly correct for these sort of why you can get into the game sort of short term. I bet that we made when you were starting a company a year and a half ago was that very pretty decent open weight models out there that pre-training, you know, we kind of saw it pre-training as starting to more or less converge on kind of a known paradigm. There's sort of a there's a known big data set on the internet. Yes, there are going to be some algorithmic innovations, but you're basically extracting signal from an extremely noisy data set. And we felt like there's only so much signal that one would be able to extract without getting into just absurd dollars for scaling this in terms of what you're trying to get out of it. So what we thought would happen is that there'd be decent open weight models. I think”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, exactly. But I think this is also how it was very common at Google and I think other places as well for different parts of the code base to have owners. And so there are like these ownership files that we have as well. And basically if you're on the ownership file, then the review has to go through you or through it has to be approved by at least one of the members of the ownership file. And as people move around teams and so forth, the ownership files themselves get updated. So I think a pretty similar structure is probably going to hold here, but it's a lot more nuanced than building kind of an individual memory, which is just kind of personal to you and lives on your computer in your agents MB file or something.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“And I think actually it's going to look not too dissimilar from that, right? Where if you want to change the team-wide memory, then it probably is going to look something like a pull request where the person who really understands that system approves or edits it or something like this. I don't think it's going to look too dissimilar.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“We built that. I think it's a thing to iterate a lot until you kind of get the right design here because you're effectively building a new git from scratch.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“That are wrong. The way it's worked with customers, we've started working with is they typically have, they want to start off with kind of a group of trusted kind of senior staff level plus engineers who are kind of the gatekeepers, which is a very, I think, common notion. You have permissions, right? And ownerships structure and code bases. And they basically are the ones who kind of populate the memory first and then sort of expand the scope. But I think it works. It's actually a much more complex feature to build because it touches on org-wide permissions. There's some parts of the code where a certain engineer should be able to edit the memory, but other engineers shouldn't. And so it actually starts looking like the new way of versioning code effectively, right? It's kind of a GitHub++ because you're not versioning the code. You're kind of versioning the meta knowledge around it that helps language models understand it better. But definitely that is something.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Yes, that's so, this is actually one of the more fun things to, I think, work on in product today. And I think it's one of the more fun kind of features to work on at the company is how do you design a teamwide memory? Because there are all sorts of details around who can edit the memory, who can view different parts of the memory. How do you maintain a kind of repository of this memory for people to edit and view?”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Tend to be the place where a product like this shines in the same way that when you think of what kind of query would you ask ChatGPT to when it just needs to use kind of the browser tool? So it's like a quick factual thing. You wouldn't invoke the deep research experience. But when you wanted to compile kind of a lot of information around some more nebulous query, I think that's where people seem to find a lot of value at Deep Research. So I think a similar kind of mindset holds here.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“We've had the way we've used it Asking things like my jobs are running slowly five times more slowly than usually. Why is that? That's kind of a vague query that would be very hard to answer with existing systems, especially since the knowledge around that query might live not just from the code base. So in the example that I just brought up, when this was happening that our environment jobs were slowing down, it turned out that two different teams, kind of infrastructure and research team submitted pull requests that were, they passed tests. It wasn't that they were wrong, but they kind of conflicted together in a way that caused this kind of effectively a race condition and slowed everyone's jobs down. And these are the kinds of bugs that actually engineers spend, you know, that's where you have like two or three engineers who spend a few days trying to solve one of these. So I think these kinds of semantic queries.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“I think the kinds of queries that it tends to be better at are, I guess, what we would call semantic queries. So let's say like an example of a query where this is not the best system to use. It's like file level. If you're looking at a file and there's like a specific thing in that file and you're just trying to get quick answer to it, you don't really need the hammer of like a deep research experience. You don't need to wait tens of seconds or a minute or two to get that answer because that should just be delivered snappily. But if you don't exactly know where you're looking for and you don't know the function name or you don't, you know, something, and this is kind of the hard problems that engineers are usually in, like there's a flaky test. I mean, you know that this test is flaky, but that's where your knowledge stops, right? And that's when you usually go to Slack and ask some engineers like, this test is flaky. What's going on? Does anyone know?”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“The model was post trained for the prompts that are coming in from users for chat GPT. There's a reason why it was, you know, when I saw the first coding blog post that ChatGPT producer me, that was just insane. That was an insane magical moment. And they post trained specifically for that. And I think there's another example of that happening right now with pod code that's kind of tight model to product coupling. And so I really think that it's important to really be able to do both at a great degree of excellence.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“So, the thing that's extremely satisfying about working on Gemini is that you're driving research in the frontier, and there's something very gratifying about that. The downside was that you were so far away removed from product that it was kind of a broken telephone game of talking to there are four different people that information flowed through before the model got into a customer's hands. That coupling was very loose. And I think it's very true that just because a company might have the best model in some general set of academic benchmarks doesn't actually mean they have the best product. And I think what we're seeing is when things really fit together, it's usually that there's a type coupling between a product and a model, it's a whole system. It's not just a model alone. Obviously, the first big example of that was ChatGPT, ChatGPT is kind of an incredible product that was coupled with the model.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Well, then you need to figure out some way to reason over that giant code base. So you have kind of a long context reasoning capability. Or you kind of look at your agent and seeing what's preventing it from satisfying the square from a user. And so you kind of work backwards and reverse engineer from what a user is asking for to what capabilities you want to drive in your system. But the important part, I think, is to be able to tweak every part of the system from the product features to the agent design to the model training in order to build the best overall system. And if you are caft in which parts you can change, if you can only change the product and agent design, then you actually are pretty limited in what you can do because you're kind of at the mercy of what kind of these general third-party models can do.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Pain points that we've identified like onboarding being one of them in a big company, it takes months to onboard an engineer. So how do you develop evals that accelerate the onboarding of an engineer from months to hopefully just a couple of weeks now that all the questions that they had, they can just ask Asimov and be able to onboard much faster. So I think there's no silver bullet other than coupling to the information coming from customers but then being very scientific in the evals that you develop across them. So you have these let's say customer needs let's say onboarding and a bunch of others and then you have your system capabilities which is well what do you need in order to provide a good experience there? Well this customer is being onboarded onto a giant code base like it has you know it might be a code base that on its phone is like 100 million tokens or something well”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“Model card for, let's say, the 01 paper that came out, I think, last year. If you look at the distribution of what most people worked on on that paper, it was evals. So you're one of many people doing all sorts of evals and spreading yourself in that sense, you get something that's general, but it's spread fairly thin. As a startup and a startup that has a very focused product that didn't, you know, that's not kind of being too diffuse and it's pretty opinionated about what it is it's building. Your evals are basically what, you know, in the startup lore when, I don't know, Paul Graham would tell you to kind of go talk to customers, like half the time build product, half the time, talk to customers. I think in the AI age, it's develop your evals based on what customers are saying and what they're doing. So you have to work with your customers to look at what prompts it trying to solve, what general questions are they trying to unlock. So there's very specific.”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT
“This is sort of why it makes sense to do something like this as a startup. So the only advantage that you'll ever have as a startup over a big incumbent, especially when there are such talented teams out there, is kind of focus and velocity against the thing that you're focused on. Now, I think you need, if you want to be playing in what is arguably, I think, the biggest category in AI, which is coding, then you need to have the talent as well to do it. But, you know, what do you do if you don't have the billions of dollars to pre-trained models? The only way we can win, I think, is by being very focused. So the way I would describe what does it look like to work on a big model within a incumbent lab is that you are one of hundreds of evals. There are teams, you know, when you look at...”
2025-07-17 · No Priors · Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin · IDENTIFIED FROM THE TRANSCRIPT