YouSaid · the spoken record

Edwin Chen

lines on the record
85
first
2025-12-07
most recent
2025-12-07
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. So, I would say, definitely tell me blog topics that you'd like me to write about. And then I'm always fascinated by all of these. AI failures that happen in the real world. So whenever you come across a really interesting failure that I think illustrates some deep question about how we want models to behave, there's just so many different ways a model can respond. I just oftentimes think there's just not a single right answer. And so whenever there's one of these examples, I just love saying them.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So I used to love writing a blog, but I haven't had time in the past few years. But I am starting to write again. So definitely check out the Surge blog. Surgehq.ai slash blog. And yeah, hopefully I'll be running a lot more dare. And I would say we're definitely always hiring. So, for people who just love data and people who love this intersection of math and language and computer science, definitely reach out in time.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  3. And it's kind of funny. I hadn't thought about that because I never quite knew what a CO did. I always thought a CEO was kind of generic and like, okay, you're just doing whatever your VPs and your board and whatever tell you to do and you're just saying yes decisions. Been said this idea where when I think about certain big hard decisions we have to make, I don't think what would company do. I don't think what metrics are we trying to optimize. I just think what do I personally care about? Like what are my values and what do I want to see happen in the world? And so I think following that idea about, okay, so ask yourself, what are the values you care about? What are the things you're trying to shape and not what we'll look at on a dashboard? I think that results for pretty important.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah, so I think it would always be to really follow your interest and do what you love. And it's almost like a lot of decisions I make about surge. I think one of the things that I Think about a couple years ago, but then someone said it to me. It's that. Companies, in a sense, are an embodiment of their CEO.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  5. So, I think I mentioned this idea that founders should build a company that only they could build, almost like it's this destiny that their entire life and experiences and interests shaped them towards. And so I think that principle applies pretty broadly, not just the founders, but the people creating anything.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Funny, but I was in SF earlier this week and I finally took away Mo for the first time. Honestly, it was magical and it really felt like living in future.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  7. One of my new all time favorite TV shows is something I found recently. It's called Travelers. It's basically about a group of travelers from the future who are sent back in time to prevent alcohol clips. I just really like science fiction. Then I actually just rewatched contact, which is one of my all-time favorite movies. So, yeah, I think one of the things you'll notice about me is that, yeah, I love any kind of book or film that involves scientists deciphering alien communication. Just this dream I always had as a kid.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  8. And then second, mythosyphus by Camus. I actually can't really explain why I love this, but I always find a final chapter somehow overly inspiring. And then third, Latombo demorrow by Douglas Hofstadter. And so I think Gerdo Escherbach is his more famous book, but I've actually always loved this one better. It basically takes a single French poem and translates it 89 different ways and discusses all the motivations behind each translation. And so I've always loved the way embodies this idea that translation isn't this robotic thing that you do. Instead, there's a million different ways to think about what makes a high quality translation, which may look a lot of the ways I think about data and quality in LMs.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Yes, so three books I often recommend are first Story of Your Life by Ted Chang. It's my all-time favorite short story, and it's about a linguist learning and alien language. And I obviously reread it every couple years.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  10. The thing I would end with is I think a lot of people think of data labeling as really simplistic work. Like labeling cat photos and drawing bounty boxes around cars. And so I've actually always hated word data labeling because it just paints this very simplistic picture when I think what we're doing is completely different. I think a lot of it. About what we're doing as a lot more like raising a child. You don't just feed a child information, you're teaching them values and creativity and what's beautiful. And these infinite subtle things about what makes somebody a good person. And that's what we're doing for AI. So I just often think about what we're doing as Almost Future of humanity, or how are we raising humanities, children? So I'll leave it at that.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  11. And it's basically applied research where we're booting all these amazing data systems that really push the frontier of AI. So yeah, I wish I known that you don't need to spend all your time fundraising. You don't need to constantly generate hype. You don't need to become someone you're not. You can actually build successful company by simply building something so good that it cuts through all that noise. And I think if I known this was possible, I would have started even sooner. So I option on that.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So I thought if I start a company, I'd have to become a business person looking at financials all day and being in meetings all day and doing all this stuff that sounded incredibly boring. And I always hated. So I think it's crazy that didn't end up being true at all. I'm still in the weeds and the data every day. And I love it. I love that I get to do all these analyses and talk to researchers.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  13. So, I definitely wish I had known that you could build a company by being heads down and doing great research and simply building something amazing and not by constantly tweeting and hyping and fundraising. It's kind of funny, but I never thought I wanted to start a company. Like I love doing research. And I was actually always a huge fan of DeepMind because they were this amazing research company that got bought and still managed to keep on doing amazing science. But I always thought that they were just magical ILR in corn.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  14. And yeah, I think it's really relevant to what we do because it's very hard and difficult to measure and define whether something is genuinely advancing in humanity. It's very easy to measure all these proxies instead, like clicks and likes. But I think that's why our work is so interesting. We want to work on hard, important metrics that require the hardest types of data and not just easy ones. So I think one of the things I often say is you are your objective function. So we want a rich complex objective functions and not these simplistic proxies. And our job is to figure out how to get the data to match this. So yeah, we want data. We want metrics that measure whether AI is making our life richer. We want to train our systems this way. And we want tools that make us more curious and more creative, not just lazier. And it's hard because, yeah, humans are kind of inherently lazy. So AI stop radios are the easiest way to get engagement, make all your metrics wall. So I think this question about choosing the right objective.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Broader question is Are we building these systems that actually advance humanity? And if so, how do we build the data sets to train towards that and measure it? Are we optimizing for all these wrong things, just systems that suck up more and more of our time and make us lazier on laser?

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  16. If you take that notion, it's like, okay, how do you define happiness? How do you, how do you measure what you're happy? How do you measure whether they're financially successful? Like, it's a lot harder than something measuring whether or not you're getting a high score on SAT. And what we're doing is we want to help our customers reach, again, their dream nor stars and figure out how to measure them. So I talked about this example of What you want models to do when you're asking them to write 50 different email iterations, you just continue them for 50 more, or do you just say, no, just move on to the day because it's perfect enough?

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Get a bit philosophical here, but I think the question is always a bit philosophical, so bear with me. So the most straightforward way of thinking about what we do is we train and evaluate AI. There was a deeper mission that I often think about, which is helping our customers think about their dream objective functions. Like, yeah, what kind of model do they want their model to be? And once we helped them do that, we'll help them train their model to freeze their star. We'll help them measure that progress. That's really hard because objective functions are really rich and complex. It's kind of like the difference between having a kid and asking them, okay, what test do you want to pass? Do you want them to get a high score in SAT and write a really good college essay? Like that's a simplistic version. Versus what kind of person do you want them to grow up to be? Will you be happy if they're happy no matter what they do? Or are you hoping they'll go do a good school and be financially successful?

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  18. I think this is this really powerful ecosystem where honestly people just don't know where models are headed and how they want to shape them yet. How do you want? Humanity kind of kind of plays a role in the future of all this. And so I think there's a lot of opportunity to just continue shaping this discussion.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  19. And we don't care as much about quarterly metrics and what's going to look good in a board deck. And so my goal is to take all these unique things about us as a company and use that to make sure that we're shaping AI in a way that's really beneficial for species in a water.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  20. And I think what drives me is that I want Serge to play this critical role in the future of AI, which I think is also the future of humanity. Like we had these really unique perspectives on data and language and quality and how to measure all this and how to ensure it's all going on the right path. And I think we're uniquely unconstrained by all of these influences that can sometimes steer companies in a negative direction. Like what I was saying earlier, we built Surge a lot more like a research lab than a Tipple startup. So we care about curiosity and long-term incentives and intellectual rigor.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  21. And I think I could do this all day. Like, I have a very hard time being a meeting this all day. I'm terrible at sales. I'm terrible at doing the typical CO things that people expect you to do. But I love writing these analyses. I love jamming with a research team on what they're seeing. Sometimes I'll be like up until 3 a.m. Just talking on a phone with somebody on a research team and taking stream model. So I love that I still get to be really hands-on working on the data and the science all day.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Think I'm a scientist at heart. I always thought I was going to become this math or CS professor and work on trying to understand a universe and language and the nature of gain migration. Like it's kind of funny, but I always had this fanciful dream where if aliens ever came to visit Earth and we need to figure out how to commune with them, I wanted to be the one to government recall. And I'd use all this fancy math, think computer science, linguistics to decipher it. So even today, what I loved me most is every time a new model is released, we'll actually do a really deep dive into the model itself. I'll play around with it. I'll run evals. I'll compare where it's improved, where it's a rest. I'll create this really deep dive analysis that we send our customers. And that's kind of funny because a lot of times we will save it from a data science team, but often it's actually just for me.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Build something focused on all these advanced complex use cases instead that would really help us build an explanation models. So yeah, I think my background in kind of cross math and computer science and linguistics really, really informed what I always wanted to do. And so I started surgery a month later with our one mission to basically build the use cases that I thought were going to be needed to push the frontier AI.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  24. I was always fascinated by math and language when I was a kid. I went to MIT because it's obviously one of the best places for math and CS, but also because the homo Noam Chomsky. My dream in school was actually to find some underlying theory connecting all these different fields. And then I became a researcher at Google and Facebook and Twitter. And I just kept running into the same problem over and over again. It was impossible to get the data that we needed to train our models. So I was always a huge believer in the need for high quality data. And then GB3 came out in 2020. And I realized that, yeah, if we want to take things to the next level and build models that could code and use tools and tell jokes and write poetry and solve every hypothesis and cure cancer, and yeah, we were going to need a completely new solution. The thing that always drove me crazy when I was out of these companies was we had the full power of the human mind in front of us. And all the data students out there were focused on really simple things like image labeling.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Yeah, yeah. I think there's a very, very powerful notion where it helps people just achieve their ideas in a much more way.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  26. In terms of overhyped areas, I definitely think that vibe coding is overhyped. I think people don't realize how much it's going to make your systems unmaintainable in the long term and do something dumped this code into your code bases. This seems to work out right now. So I kind of worry about future coding. It just keeps on happening.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  27. I think one of the things that was underhyped is the Built in products that all of the chatbots are going to start having. Like, I've always been a huge fan of college artifacts. And I think it just works really, really well. And actually the other day, I don't know if it's a new feature or not, but it's asking me to help me create an email. And then it just creates, so it didn't quite work because it didn't allow me to send an email. But what it created instead was like a little, I don't know what we call it, like a little box where I could click on it and it would just text someone that did this message. I think that concept of Taking artifacts to the next level where you just have these like mini apps, mini UIs within the chats themselves. I feel like people aren't talking enough about that. So I think that's one underhyped area.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  28. And again, again, just because in the same way, I do it kind of like a fork in a road between how you could choose how your model behaves for disk question. It's like for every other question that models have. Kind of behavior that you want will fundamentally affect it. It's almost like in the same way that when Google builds a search engine, it's very, very different from how Facebook would build a search engine, which is very, very different from how Apple would build a search engine. Like they all have their own principles and values and things that they're trying to achieve in a world that shape all the products that they're going to build. And in the same way I think all the dellums will start behaving very, very differently too.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Then I realized I spent 30 minutes doing something that didn't matter at all. Like, sure, now I got the perfect email, but I spent 30 minutes doing something I wouldn't have worried at all before. And this email probably didn't even move the needle or anything anyways. So I think there's a deep question here, which is, if you could choose the perfect model behavior, which model would you want? Do you want a model that says, you're absolutely right, there are definitely 20 more ways to improve this email. And it continues for 50 more iterations, and it sucks up all your time and engagement. Or do you want to model that's optimizing for your time and productivity and just says no, you need to stop your email's great. Just send it and move on with your day.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Will shape the model. So let me give an example. So I was asking Claude to help me draft an email the other day. And it went through 30 different versions. And after 30 minutes, yeah, I think it really crafted me the perfect email and I sent it.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Think one of the things that's going to happen in the next few years is that the models are actually going to become increasingly differentiated because of the personalities and behaviors that the different labs have and the kind of objective functions that they are optimizing their models for. I think it's one thing I didn't appreciate a year or so ago. Like a year or so ago, I thought that all of the AI models would essentially become very, very commoditized. They would all behave like each other. And sure, one of them might be slightly more intelligent. And one way today, but sure, deller ones would catch up in the next few months. I think over the past year, I've realized that the values that the companies have

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  32. So we look for people who are just fundamentally interested in data all day. So types of people who could literally spend 10 hours digging through a data set and playing around with models and thinking, okay, yeah, this is where I think the model is failing. This is kind of a behavior you want the model to have instead. And just this aspect of being very, very hands-on and thinking about the qualitative aspects of models and not just the quantitative parts. So again, it's like this aspect of being hands-on with data. And not just caring about these kind of abstract algorithms.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Yeah, yeah, I think it's just because it's something I've fundamentally always cared about. Like, I often think about us more like a research lab than a startup. Because dad is my goal. It's kind of funny, but I've always said I would rather be Terence Taud than Warren Buffett. So that notion of creating research that pushes the frontier forward and not just getting some valuation, like that's always been what drives me.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  34. And then we also have our internal researchers. So our internal researchers are focused on slightly different things. So they are focused on building better benchmarks and better leaderboards. So I talked a lot about how I worry that the leaderboards and benchmarks out there today are steering models in the wrong direction. So, yeah, so the question is, how do we fix that? And so that's what our research aim is focused on really, really focused really heavily on right now. So they're working a lot on that. And they're also working on these other things like, okay, we need to train our own models to see what types of data performs the best, what types of people perform the best. And so they are also working on all these kind of training techniques and evaluation of our own data sets to prove, improve our data operations. And the internal data products that we have that determine what makes something good quality.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  35. And we're going to design these data sets, these evaluation methods, these training techniques to make your models better. So just like this very, very kind of collaborative notion of working with our customers, like being researchers themselves, just a little bit more focused on the data side and where you handle them to do whatever it takes to make them the best.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Yeah, so I think that stems from my own background. Like, my own background is as a researcher. And so I've always cared fundamentally about pushing the industry and pushing the research community and not just about revenue. And so I think what a research team does is a couple different things. So we almost have two types of researchers at our company. One is our four deployed researchers who are often working hand in hand with our customers to help them understand their models. So we will work very closely with the customers to help them understand, okay, this is where your model is today. This is where you're

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  37. It's almost like maybe the end goal is just throwing you into the environment and Seeing how you evolve. But within that evolution, there's all these different sub-learning mechanisms.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Yeah, yeah. So, I mean, I really do think that we are going to need to build a suite of products that reflect the million different ways that humans learn. For example, think about becoming a great writer. You don't become great by memorizing a bunch of grammar rules. You become great by reading great books and you practice writing and you get feedback from your teachers and from the people who buy your books in a bookstore and leave reviews. And you notice what works and what doesn't. And you develop taste by being exposed to all these masterpieces and also just terrible writing. So you learn through this endless cycle of practicing reflection and each type of learning that you have, again, like these are all very, very different methods of learning to become a great writer. So just in the same way that there's a thousand different ways that the great writer becomes great. I think there's going to be a thousand different ways that AILs need to learn.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Yeah, yeah. So I think Evals often covers two terms. One is you're using the evaluations for training because you're evaluating whether or not the model did a good job. And when it does do a good job, you're rewarding it. And then there's this other notion of evals where you're trying to measure the model's progress. Like, okay, yeah, I have five different candidate checkpoints and I want to pick the one that's best in order to raise it to the public. So I'm going to run all these evals on these five different checkpoints in order to decide which one, which one is best. Now we have R environment, so it's kind of like a hot

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  40. So SFT stands for supervised fine tuning. And it's a lot like, so again, I think often in terms of these human analogies, and so SFT is a lot like mimicking a master and copying what they do. And then our LHF became very dominant. And analog there would be like, sometimes you learn by writing five different essays and someone telling you which one they like the most. And then I think over the past year or so, rubrics and verifiers have become very important. rubrics and verifiers are like learning by being graded and getting detailed feedback on where you went wrong.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  41. And I think it's also really important because some of these trajectories can be very, very long. And so if all you're doing is checking whether or not the model reaches the final answer, it's like there's all this information about how the model behaved in the immediate step that's missing. Like sometimes you want models to get to the correct answer by reflecting on what it did. Sometimes you wanted to get the correct answer by just one-shotting it. And if you ignore all of that, it's like teaching it's just missing a lot of information that you could be teaching Moloto to do.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  42. One of the things that people don't realize is that sometimes even though the model reaches the correct answer, it does so in all these crazy ways. So it may have in the intermediate directory, it may have tried 50 different times and failed, but eventually it just kind of like randomly lands on a correct number. Maybe it sometimes it just does things very, very inefficiently, or it almost reward hacks a way to get at the correct answer. And so I think paying attention to the directory is actually a really, really important point.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Yeah, exactly. So that financial illness might create a spreadsheet. They may create certain tools that the model needs to call in order to help fill out a spreadsheet. Like it might be, okay, the model needs to access Bloomberg terminal. We needs to learn how to use it. And it needs to learn how to use this calculator. And it needs to learn how to perform this calculation. So it has all these tools that it has access to. And then the reward might be, okay, it's like maybe I will download that spreadsheet. And I want to see does sell B22 contain the correct profits and profit and loss number or does tab number two contain this specific information.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Yeah, yeah. So just in the same way that there were all these different methods for models learning in the past, originally we had SFT and RHF. And then we had rubrics and verifiers. This is the next stage. And it's not the case that the previous methods are obsolete. This is, again, just a different form of learning that complements all the previous types. So it's just like a different skilled model, not only to learn how to do.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Yeah, so dejective function might be, or the goal of the task might be, okay, go figure out why and fix it. And so the objective function might be passing a series of unit tests. It might be writing a document, like maybe the retro containing certain information that matches exactly what happened. There's all these different rewards that we might give it that determine whether or not it's succeeding. And so the models were basically teach the models to achieve that reward.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  46. And I think one of the interesting things is that these environments really showcase where models are end to end weak at end to end tasks in the real world. You have all these models that seem really smart on isolated benchmarks. Like they're good at single-step tool calling. They're good at single step instruction following. But suddenly you dump them into these messy worlds where you have confusing stock messages and tools they'd never seen before and they need to perform right actions and modify the databases and interact over longer time horizons where what they do in step one affects what they do in step 50. And that's very, very different from these kind of academics, single step environments that they've been in before. And so the model just fails catastrophically in all these crazy ways. So I think these R environments are going to be really interesting playgrounds for the models to learn from that will essentially be simulations and mimics in real world. And so they'll hopefully get better and better at real tasks compared to all these contrived environments.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Enforcement learning is essentially training your model to reach a certain reward. And let me explain what an RRN is. An R environment is essentially a simulation of real world. So think of it like building a video game with a fully fleshed out universe. Every character has a real story. Every business has tools and data you can call and you have all these different entities interacting with each other. So for example, we might build a world where you have a startup with Gmail messages and Slack threads and Jira tickets and GitHub PRs and a whole code base. And then suddenly AWS goes down and Slack goes down. And so, okay, model, what do you do? Like the modeling to figure it out. So we give the models tasks in these environments. We design interesting challenges for them. And then we run them to see how they perform. And then we teach them. We get them these rewards when they're doing a good job or a bad job.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  48. I'm in a camp where I do believe that something new will be needed. Like the way I think about it is I think about your NAI, I take a very, I don't know if I would say biological point of view, but I believe that in the same way that there's a million different ways that humans learn, we need to build models that can mimic all those ways as well. And maybe don't have a different distribution of the focuses that they have on it. I know they'll be different for humans, so maybe you'll have a different distribution. But we want to be able to mimic the learning abilities of humans. Make sure that we have the algorithms and the data for models to learn in the same way. And so to the extent that LOMs have different ways from humans, then yeah, I think something new needed.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Yeah, I absolutely think that you have to have huge ambitions and you have to have a huge belief in your idea that's going to change the world. And you have to be willing to double down and keep on doing whatever it takes to make it happen.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  50. And if you fail because the market isn't ready yet, I actually think that's way better. At least you took a swing at something deep and novel and hard instead of pivoting into another LLM wrapper company. So yeah, I think the only way you build something that matters since that's going to change the world is if you find a big idea you believe in and you say no to everything else. So you don't keep on pivoting when it gets hard. You don't hire a team of 10 product managers because that's where every other cookie cutter startup does. You just keep building that one company that wouldn't exist without you. And I think there are a lot of people. It's like I'm now sick of all the grift, who want to work on big things that matter with people who actually care. And I'm hoping that that will be the future of how we go with technology.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source