YouSaid · the spoken record

Edwin Chen

lines on the record
85
first
2025-12-07
most recent
2025-12-07
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I've always hated this because Silicon Valley loves to score on Wall Street for focusing on money, but honestly, most of the Silicon Valley is chasing the same thing. And so we stayed focused on our mission from day one, pushing that frontier of high quality complex data. And I've always loved that because I think startups, I have this very romantic notion of startups. Like startups are supposed to be about taking big risks to build something that you really believe in. But if you're constantly pivoting, you're not taking risks. You're just trying to make a quick walk.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Yes. So I've always really hated a lot of Silicon Valley intros. This standard playbook is to get product market fit by pivoting every two weeks and to chase growth and chase engagement with all of these dark patterns and to blitz scale by hiring as fast as possible. And I've always disagreed. So, yeah, I would say don't pivot. Don't blit scale. Don't hire des Stanford grad who simply wants to add a hot company to a resume. Just build the one thing only you could build, the thing that wouldn't exist without the insight and expertise that only you have. You see these buy-table companies everywhere now. Some founder who was doing crypto in 2020 and then pivoted NFTs in 2022. And now during your AI company, there's no consistency. There's no mission. They're just chasing evaluations.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Yeah, I think it's almost like Do you care about how you get there? And in the same way, so I made his tabloid analogy earlier, but. Would you sell tabloids in order to fund? I don't know, some other newspaper? Like sure. In some sense. You don't care about the path, then you'll just do whatever it takes, but it's possible that it has negative consequences in of itself that will harm the long-term direction of what you're trying to achieve. And maybe it'll distract you from all the more important things. So yeah, I think the path you take matters a lot as well.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  4. When what it entails? And so it's kind of interesting. It's like which companies would build Sora and which wouldn't. And I think that answer that question. I mean, I don't know if the answer is myself. I have any idea in my head, but I think the answer to that question maybe reveals certain things about what kinds of AI models those companies want to build and what direction in what future they want to achieve. So I think

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Think there is a question of what products they're building and whether those products themselves are something that kind of help or hurt humanity. Like I think a lot about Sora.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Would say have always been very, very impressed by anthropic. Like, I think Anthropic takes a very principled view about what they do and don't care about and how they want their models to behave in a way that feels a lot more, a lot more principled to me.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I think I worried that the same thing's happening with AI. If you think about all the sick fancy issues with ChatGPT, oh, you're absolutely right. What an amazing question. The easiest way to hook users is to tell them how amazing they are. And so these models, they constantly tell you you're a genius. They'll feed into a delusions and conspiracy theories. They'll pull you down these rabbit holes because Silicon Valley loves maximizing time spent and just increasing the number of conversations dropping with it. And so yeah, companies are spending all their time hacking these leaderboards and benchmarks. And the scores are going up. But I think it actually masks that the models with the best scores, they are often the worst or just have all these fundamental failures. So I think I'm really worried about that. All of these negative ascendants are pushing AGI into the wrong direction.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  8. The only way I'm going to get promoted at the end of the year is if I climb this leaderboard, even though I know that climbing is probably going to make my model worse and accuracy and structure falling. So I think there's all these negative incentives that are pushing work in the wrong direction. I'm also worried about this trend towards optimizing AI for engagement. I used to work on social media. And every time we optimize for engagement, terrible things happened. You'd get Clickbait and pictures of bikinis and Bigfoot and horrifying skin diseases just filling your feeds.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Optimizing your models for the types of people who buy tabloids at the grocery store. Like, we've seen a scenario with data ourselves. The easiest way to climb El Marina, it's adding crazy boating. It's doubling the number of emojis. It's tripling the length for your model responses, even if your model starts for loose naming and getting the answer completely wrong. And the problem is, again, because all of these frontier labs, they kind of have to pay attention to PR because they're a sales team when they're trying to sell all these enterprise customers. Those enterprise customers will say, well, well, but your model is only number five on Elmoreno. So why should I buy it? They have to, in some sense, pay attention to these leaderboards. And so what do researchers all tell us is like if they say,

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  10. I'm worried that instead of building AI that will actually advance us as a species, cure in cancer, solving poverty, understanding universe, all these big grand questions, we are optimizing for AI slop instead. We're basic teaching our models to chase dopamine instead of truth. And I think this relates to what we're talking about regarding these benchmarks. So let me give you a couple examples. So right now the industry is played by these terrible data boards like LM Arena. It's this popular online leaderboard where random people from around the world vote on which AI response is better. But the thing is, like I was saying earlier, they're not carefully reading or fact-checking. They're skimming these responses for two seconds and picking whatever looks flashiest. So a model can hallucinate everything. It'll completely hallucinate, but it will look impressive because it has crazy emojis and boating and markdown headers and all these superficial things that don't matter at all, but it catch your attention. And these alimating users love it.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  11. I'm certainly on a longer time horizon front. I think people don't realize that there's a big difference between moving from 80% performance to 90% performance to 99% performance to 99.9% performance and so on and so on. And so in my head, I probably bet that within the next one or two years, yeah, the models are going to automate 80% of the average L06 software engineer's job. It's going to take another few years to move to 98% and another few years to 99% and so on and so on. So I think we're closer to a deck.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Yeah, I think that will not happen until we've reached AGI. It's almost like by definition, if we haven't reached AGI yet, then there's more for the models to learn from. And so, yeah, I don't think that's going to happen anytime soon.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  13. They are, yeah, they're going to evaluate the code editor rights. They're going to double check the physics equations that it writes. They're going to evaluate the models in a very, very deep way, sort of pay attention to accuracy and instruction following and all these things that casual users don't when you suddenly get a pop-up on your chat GPT response asking you compare these two different responses like people like that, they're not evaluating models deeply, they're just vibing and taking whatever response looks flashiest orientators are looking closely at responses and evaluating them for all of these different dimensions. And so I think that's a much better approach than these benchmarks or kind of these random online ABS.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So, the way we really care about measuring model progress is by running all these human evaluations. So, for example, what we do is, yeah, we will take core human annotators and we'll ask them, okay, go have a conversational model. Maybe you're having a small conversation model across all of these different topics. So you are a Nobel Prize-winning physicist. So you go have a conversation about pushing different tier of your own research. You are a teacher and you're trying to create lesson plans for your students. So go talk to model about these things. Or you are a, yeah, you're a coder and you're working at one of these big tech companies and you have these problems every day. So go talk to the model and see how much it helps you. And because or sergers or annotators, they are experts at the top of their fields and they are not just giving your responses. They're actually working through the responses deeply.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Yes, there's again, maybe two parts to this. So, one is sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the Frontier Labs will tweak the way they evaluate their models on these benchmarks, like they'll tweak their system prompt, or they'll tweak the number of times they run their model and so on and so on in a way that games these benchmarks. The other part of it, though, is Like by optimizing for the benchmark, instead of optimizing for the real world, you will just naturally climb on the benchmark. And yeah, it's basically another form of gaming.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  16. And that's because, yeah, even though IMO own gold medals seem hard to the average person, yeah, like they are hard at the end of the day, but they have this notion of objectivity that, okay, yeah, parsing MP def sometimes doesn't have. And so it's easier for the frontier labs to hill climb on all these than to solve all these messy ambiguous problems in a real world. So I think there's a lack of direct correlation there.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Yeah, so I don't trust the benchmarks at all. And I think that's for two reasons. So, one is I think a lot of people don't realize even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wrong answers, they're full of all this kind of messiness. And people trust for the popular ones. People have maybe realized this to some extent, but the vast majority just have all these flaws that people don't realize. So that's one part of it. And the other part of it is these benchmarks at the end of the day, they are often, they often have well-defined objective answers that make them very easy for models to hill climb on in a way that's very, very different from the messiness and ambiguity to real world. I think one thing that I often say is that it's kind of crazy that these models can win IMO gold medals, but they still have trouble parsing PDFs.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yep, yep, exactly. Like, again, going back to the example I said earlier, certain companies, if you ask them what is good poem, they will simply robotically check off all of these instructions on our list. But again, I don't think that makes very good poetry. So certain frontier labs, the ones with more taste in sophistication, they will realize that it doesn't reduce to this fixed set of checkboxes and they'll consider all of these kind of implicit, very subtle qualities instead. And I think that's what makes them better at descendant day.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Taste and sophistication, like, okay, do I think that these, so going back to the example of how good the model is at visual design? Like, okay, maybe you have a different notion of visual design than what I do. Maybe you care more about minimalism and you care more about, I don't know, like 3D animations than I do. Maybe you saw a person prefers things that look a little bit more broke. There's all these notions of taste infusion that you have to decide between when you're signing your post training mix. And so that matters as well. Long story short, I think there's all these different factors. And certainly the data is a big part of it, but it's also like, what is the objective function that you're trying to optimize your model towards?

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  20. It's almost like there's a trade off between all of these different things, and one of the things I often think about is that there's a, it's almost like there's an art to post training. It's not purely a science. Like when you are deciding what kind of model you're trying to create and what it's good at, there's this notion of

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Like some companies see these benchmarks and you're like, okay, for PR purposes, even though we don't think that these academic benchmarks matter all that much, maybe we just need to optimize for them anyways because our marketing team needs to show certain progress on certain standard evaluations that every other company talks about. And if we don't show good performance here, it's going bad for us, even if like ignoring these academic benchmarks makes us better at every real tasks. Other companies are going to be principal and like, okay, yeah, no, I don't care about marketing. I just care about how my market will performs on these real-world tasks at the end of the day. And so I'm going to optimize.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  22. I think there are multiple parts to it. So, a big part of it certainly is the data. I think people don't realize that there's almost like this infinite amount of choices that all the frontier labs are deciding between when they're choosing what data goes into the models. It's like, okay, are you purely using human data? Are you gathering the human data in XYZ way when you are gathering the human data? What exactly are you asking that people were creating it to create for you? Like maybe you create maybe you care more, for example, in the coding realm. Maybe you care more about front encoding versus back encoding. Maybe when you're doing front encoding, you care a lot about the visual design of the front end applications that you're creating, or maybe you don't care about it so much and you care more about, I don't know, deficiency of it or the pure correctness over that visual design. Then other questions like, okay, are you carrying? Are you like how much synthetic data are you throwing into the mix? How much do you care about these 20 different benchmarks?

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  23. And you also want to discover the best of the best. Okay, like this is the best webpage, or just the best person for this job. They are not just somebody who writes the equivalent of high school level poetry. Again, they're not just robotically writing poetry that checks all these boxes, checks all these explicit instructions, but rather, yeah, they're writing poetry that makes you emotional. And so we have all these signals as well that, again, like completely differently from moving to worse, we are finding the best of the best. And so we have all these signals again, just like Google Search uses all these signals and feeds them into their ML algorithms and uses them predicts certain types of things. We do the same with all of our workers and all of our tasks and all of our projects. And so it's almost like a complicated machine learning problem at the end of the day.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Way it works is we essentially gather thousands of signals about everything that you're doing when you're working on a platform. So we are looking at your keyboard strokes. We are looking how fast you answer things. We are using reviews. We are using gold standards. We are using, like we're training models ourselves on the outputs that you create. And we're seeing whether they improve the model's performance. And so in a very similar way to how Google search like when Google search is trying to determine what is good webpage, there's almost two aspects of it. One is you want to remove all of the worst of the worst web pages. So you want to remove all the spam, all the just like low quality content, all the pages that don't load. And so there's like a, it's almost like a condom moderation problem. You just want to remove the worst forest.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Insights into language and imagery and human expression. I think think about quality in this way is really hard. It's hard to measure. It's really subjective and complex and rich. And that's a really high bar. And so we have to build all of this technology in order to measure it, like thousands of signals on all of our workers, thousands of signals on every project, every task. Like we know at the end of the day, if you are good at writing poetry versus good at writing essays versus good at writing technical documentation. And so we have to gather all these signals on what your background is, what your expertise is, and not just that, like how you're actually performing when you're writing all these things. And we use those signals to inform whether or not you are a good net forker for these projects and whether or not you are improving the models. And it's really hard. And so to build all this technology to measure it, but I think that's exactly what we want AI to do. And so we have these really, really deep notions about quality that we're always trying to try and achieve.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Think most people don't understand quality even means in this space. They think you can just throw bodies at a problem and get good data, and that's completely wrong. Let me give you an example. So imagine you wanted to train a model to write an A line poem about the moon. What makes it a good high quality poem? If you don't think deeply about quality, you'll be like, is this a poem? Does it contain eight lines? Does it contain a word moon? You check all of these boxes. And if so, sure, yeah, you say it's a great poem. But that's completely different from what we want. We are looking for Nobel Prize winning poetry. Like, is this poetry unique? Is it full of subtle imagery? Does it surprise you and tug out your heart? Does it teach you something about the nature of moonlight? Does it play through motions and does it make you think? That's what we are thinking about when we think about high quality bulb. So it might be like a haiku about moonlight on water. It might use internal rhyme and meter. There are a thousand ways to write a poem about the moon, and in each one gives you all these different.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Essentially teach AI models what's good and what's bad. So we train them using human data and there's a lot of different products that we have like SFT, RHF, Rubrics, Verifiers, R environments and so on and so on. And then we also measure how well they're progressing. So essentially we're a data company.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  28. I think it also meant that our customers were people who really understood data and really cared about it. I always thought it was really important for us to have customers, early customers who are really aligned with what we were building and who really cared about having really high quality data and really understood how that data would make their AI models so much better because they're the ones helping us. They were the ones giving us feedback on what we're producing. And so just having that kind of like very, very close mission alignment with our customers actually helped us early on. So these are people who basically just buying our product because they knew how different it was and because it was helping them rather than because they saw suffering that kind of chatline. So it made things harder for us, but I think in a really good way.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Basically, never wanted to play the Silicon Valley game. I always start ridiculous. What did you dream of doing when you were a kid? Was it building a company from scratch yourself and getting into weeds of your code and your product every day? Explaining all your decisions to VCs and getting on this giant PR and fundraising hamster wheel. And it definitely made things more difficult for us because, yeah, when you fundraise, you just naturally get part of this kind of Silicon Valley industrial complex where people will, your VCs will tweet. Get announced in all of the newspapers because you raise a dismantled valuation. And so it made things more difficult because the only way we were going to succeed was by building a 10 times better product and getting word of mouth from researchers.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Just going to lead to a really amazing time in company building. Like the thing I'm most excited about is that the types of companies are going to change to you. Won't just be that they're smaller. We're going to see fundamentally different companies emerging. If you think about it, fewer employees means less capital. Less capital means you don't need a raise. So instead of companies started by founders who are great at pitching and great at hyping, you'll get founders who are really great at technology product. And instead of products optimized for revenue and what VCs want to see, you'll get more interesting ones built by these tiny Obsessed teams. So people building things they actually care about, real, real technology of real innovation. So I'm actually, really hoping that the Silicon Maddie Sarup team will actually go back to being updates for hackers again.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  31. So we hit over a billion of revenue last year with under 100 people. And I think we're going to see companies with even crazier ratios, like 100 million per employee in the next few years. AI is just going to get better and better and make things more efficient. So debt ratio just becomes inevitable. Like I used to work at a bunch of the big tech companies. And I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions. And so when we started surge, we wanted to build it completely differently with a super small, super awake team. And yeah, what's crazy is that we actually succeeded. So, I think two things are colliding. One is that people are realizing that you don't have to build giant organizations in order to win. And two, yeah, all these efficiencies.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Instead of building AI that will actually advance us as a species, curing cancer, solving poverty, understand the universe, we are optimizing for AI slop instead. Could we be optimizing our models for the types of people who buy tabloids at a grocery store? We're basically teaching our models to chase dopamine instead of truth.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Over the past year, I've realized that the values that the companies have will shape the model. I was asking Claude to help me drop an email the other day. And after 30 minutes, yeah, I think it really crafted me the perfect email and I sent it. But then I realized I spent 30 minutes doing something that didn't matter at all. If you could choose the perfect model behavior, which model would you want? Do you want a model that says you're absolutely right? There are definitely 20 more ways to improve this email. And it continues for 50 more iterations. Or do you want a model that's optimizing for your time and productivity and just says no? You need to stop. Your email's great. Just send it and move on.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Essentially, teach AI models what's good and what's bad. People don't understand what quality even means in a space. They think you could just throw bodies at a problem and get good data. That's completely wrong.

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Never wanted to play the Silicon Valley game. I always was ridiculous. I used to work at a bunch of the big tech companies, and I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions. So when we started Surge, we wanted to build a completely differently with a super small, super elite team

    2025-12-07 · Lenny's Podcast · The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI) · IDENTIFIED FROM THE TRANSCRIPT · source