YouSaid · the spoken record
Chip Huyen
- lines on the record
- 84
- first
- 2025-10-23
- most recent
- 2025-10-23
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Mixture was it's like it can retrieve the information that's relevant to the query even though it might not immediately like obvious that it's related so it will come up with like only thing of like I think like contextual retrieval like giving HTML data when maybe in a summary metadata so that it knows or like some people use like as a hypothetical question it's very interesting like for a given term of like documents I might get it a bunch of questions that the chunks can help answer so that when they have a query it's like okay does it match any of the like hypothetical questions I can it can fetch it so it's a very interesting approach okay so maybe before I go to the next thing I just want to say this like data preparations for rack is really really important and I would say that's like in the a lot of the companies that I have seen that's like the biggest performance in the rack solutions coming from like better data preparations not”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Traditionally, when it started out, a rack is mostly like text. So we talk about a lot of ways how to prepare data so that the model can retrieve effectively. Let's say that not everything is a Wikipedia page, right? Wikipedia page is pretty contained and you know, okay, everything about it is about a topic. But I want to talk about documents extremely a lot, right? And they have a weird way of structure the documents. Let's say that you have documents about Lenny podcast, right? And in the future, in the beginning of documents, from now on, podcasts wouldn't refer to Lenny's podcast, right? So let's say somebody in the future is like, okay, tell me about Lenny, right? Lenny's work. And because the residence document does not have the term Lenny, you just don't know, you might not read through it. And the document is long enough that it chunked into a different part. So like the second part doesn't have the word mammic. So you cannot reach it. So you have to find a way to process data so that.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“So Ragus Danford with river augmented gyrations in a sort not a specific shoe JW AI. So the idea is just like follow questions we need context to answer. So I think it came pretty, I think it's from the paper 2017. So someone was like, so they realized it's like for a bunch of like benchmark when the question answering benchmarks they realize it's like, okay, if we give the model information about the questions, then the answer can be much, much better. So what they do was actually try to retrieve information from Wikipedia. So for questionable topics, it's like retrieve that and then put into the context and like answer it does much better. So I feel like it sounds like a no-brainer, right? I mean, like obviously. So I think that's what dragged it as a simplest sense. It's just like providing some model with a relevant context so that it can answer the questions. And that's where things get really more interesting because”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“More diverse, right? And then look at the results of the search grad and SIU enter the search query, like Lenny Parscat data labeling. And they come up with like 10 pages, 10 results. And then you come up with like, oh, Lenny podcast on, I don't know, I don't know, like Frontier Labs and have like 10 resorts. I mean, Locke was a different web page, like how much of some overlapping. Are we doing both the breadth, like getting a lot of page, but also like, do we have death? And also they have relevance because if we come up with a search queries, they're completely irrelevant to the original prompt. So I feel like every aspect of it would need a way of evaluating, right? So I don't think it's like how many eval should I get, but like how many Evar should do I need to get a good coverage, a high confidence in my application's performance, and also to help me understand like where it is not performing well.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Ways, it's like. How do you produce the result of the summary? At first, you need to be able to do gather information. And to gather information, you need to do a lot of search queries. You gather search results and then some of the search results aggregate and then maybe say, okay, I'm still missing on this. You have to go another route, on like another round, ACN has the summary. So every step of the way, you need evaluations, right? You don't need to end-to-end. So maybe for the search query, you may first think about like, okay, now I write five search queries. Am I looking to like, how good is this search queries? Like, do they?”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“User sensitive data. And another is like for lang, but has a number of like Okay, let's just speak an example, concrete example, like DIF research. So you have the application, you have to use a model to do deep research for you, right? Like, okay, like have a prompt. And I may say, okay, do me your comprehensive research on online podcasts and help me like propose, show me report on what kind of topics he's interested in, what kind of videos could get the most views or like what topics that he's missing on that he should be covering, right? Like, have that kind of like prompt, then how do you evaluate the result, right? I don't think there's like one metrics that would help. Maybe just like maybe you have like 100, I think somebody has a benchmark and is it get like 100 expert, like write a bunch of prompts and they go through like on the answers on AI and I do it. And it's like it's extremely costly and slow, right? But if you might have something else, for example, like one way was thinking about it, I was talking to a friend about it.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't think of like just a fixed number on the evolves. Like what was it going to evolve, right? The go to eval is to guide the product development. So like EC eval, because I think I'm a big fan of Eval is that it helps you uncover opportunities where the products are doing well. So I sometimes seen it very obvious. Okay, when you look at the Eva, I'm going to realize it's like, okay, perform really poorly on this specific segment of users. And then we look into it. It's like, okay, what's wrong with it? And it turns out it's like we just don't have a good messaging to it. So like people should just focus on the things of reading polynomial can improve significantly. Yeah, so I kind of as a number of eval is really depends. Like we have seen product with like hundreds of different metrics. This is because like that product is like Gerald, right? Have different except like on Evarm for like I don't know like verbosity have like one eval for like”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Very important if you have operate a scale and where like failures can have like catastrophic consequences, then you do need to be very tyrannical about like what you print in front of the users, understand different failure modes, like what could go wrong, and also maybe in the space when it's a feature of the product is as a competitive advantage, right? It wants to be the best at this. You want to have like a very strong understanding of where you are and like where you are with the competitors. But it's just something that's like more like a low-key, okay, something is like, okay, that's not so core, but like it helps with our users, then maybe you don't need to be so, so obsessed or lyrical about it. It's like, okay, that's good enough for now. And if it fails and it fails like, okay, I know it's like it's such terrifying, but like, yeah, yeah, I think it's only about like the question of like return investment. I'm a big fan of ABA. I love reading Eva. And that says like I understand.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Retake that two engineers and we want to launch a new feature, then you could give me like so much more, like improvement, right? So I think it's like one of them is like, Eva, sometimes people think of Eva. It's like, okay, this is good enough to touch it. Like if you do spell a lot of energy on Eva, it would only incremental improvement where it spends the energy on another use case. And maybe I scatter good enough that you guys vibe check it, right? So I do think it's like maybe just a debate is about. I do think that's like a lot of time people just like get things to the place where it's like, okay, good enough. People run. And then, but of course, it's like there's a lot of risk associated with it because we don't have a clear metric. You have good visibility to household applications and some models of performing. It might do something very dumb or it can cause you like, I know something like crazy can happen. So yeah, so I do think Eva is very”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“You don't have to be absolutely perfect at things to win. You just need to be good enough and be inconsistent about it. Okay, this is not the philosophy I follow, but like I have worked with enough companies to see that play out. So what I say, why companies don't evaluate? Let's say you are like an executive, right? And you want to have a new use case. So here's a use case you started out. We built and it's like it works well, right? The customers are somewhat happy. You don't have the exact method for it. But like, so the traffic keeps increasing, like people seem happy, people keep buying stuff, right? And now here's our engineer coming like, okay, we need Eva for it. And so, and it's not like, okay, how much effort do we need to go into Eva? And they were like, okay, maybe like two engineers, as much as much. And they could maybe would improve. And was like, okay, so how much expected gain can I get from it? And the NGO could be like, oh, maybe you can improve it from like 80% to like 80%, 85%, right? And I was like, okay, but it was.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I think if people approach Eva, I think there are like two very different problems. One is a app builder, right? And like, can I say have an app that do like maybe a chatbot? Very simple announcement. It was the first thing that came to my mind. And I want you to know each chatbot is good or bad.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“You have some bearish on it. I think I'm curious because I think things have had a way of workout in ways that I don't expect. So I think that maybe these companies, they have a lot of data, maybe they wouldn't be able to use that to have some insights that helps them stay ahead of the curve. So I don't know.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I'm like a bit like, look me, I'm easy, right? Have companies are growing crazy, but it's like heavily dependent on like two or three companies. And at the same time, if I was this company frontier, what would be the right economical things for me to do? Right now, I want a lot of startups. I want to have a lot of providers so I can pick and choose. And then these providers can also compete each other to lower the price. And it's so dependent on me. It would sell social media regardless. So I feel like, yeah, so this economics, a whole economy is very interesting to me. And I'm curious to see how it plays out.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Data to trade them. And it's a massive thing. It's like people because everyone wants a lot of data and only slaps at like unlimited budget. But whether agencies is also a little bit of low key interesting economics. So I'm not sure you've talked to like the guests about. I thought it's very interesting if I think about because it's very lopsided, right? Because like there are only like a very small numbers of frontier labs, right? And they want a lot of data. And there's like a massive amount of startups or companies providing data. So like you can see these companies like this startup like doing data labeling. They have like maybe some like massive AR. But you also like, okay, so how many customers you have? And they could be like a very small numbers. I'm not sure. I'm not sure I saw you smiling.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think that's a way of looking into it. And I think that's a space is so exciting nowadays because Has so many domain exports tasks that the model developers want model to do well on, right? Let's say you're like accountant, right? Like maybe once you use a model to have accounting tasks, I need a lot of like accounting data, like examples from my accountants. So you need to hire a lot of them. Should I do it? Or if you want to physics problems, I want to do, I don't know, like legal questions and stuff, like engineering questions or like somebody was telling me like one should do like using like coding for to solve scientific problems and not just like coding to build product, which is another different whole realm of things. And I also like using very specific toolings like, yeah, like I'm not sure what apps you use, but maybe like a 40-day app or like QuickBooks or like Google Excel. They are very specific to specific expertise that you want the model to learn. So they need a lot of humans expert in this area to like create.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so again, it's like it's general, it's a way of learning. It's like training is controversial learning and whether it learn from human feedback or like AI feedback or like verifiable rewards. I think they say you say it's just different way of like collapse signals. Awesome.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“A human feedback, and then you use this human feedback to trade a reward model. So you might tell, and then the reward model will help you like, okay, the model produces response. It's the reward model can score. Is this good or bad? New charge and bias for what producing better model, the better responses. Another way that you can instead of using a human, so you can use like AI, right? Like, look at the response, say yes or good, good or bad, right? Orange, I think, is like people are very big on nowadays, like verifiable rewards, which is like natural. So basically they give it a math problem and a math solutions. Like as a model, I put a solution, you know, okay, it's a expected response shit in it for you too. And it doesn't provide 482s and it's wrong, right? It's not a good response. So yeah, so like a lot of time people using this human labor like human labor should produce how to say expert questions. And I said,”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Reinforce or encourage the model to produce an output that is better, right? So select it and have like now, how do we know that the answer is good or bad? So usually we realize on signals. So one way to get a fresh one good or bad is like human feedback, right? They happen. We have two responses. You can, okay, this one's better than the other. And we do that is because as humans, we tend to, it's very hard to give concrete score, but it's easier to do comparisons, right? Like if you ask me, okay, give this song a score. I'm not a musician and don't know how hard it is. And it's like, yeah, I don't know what 10, I'm going to six, you know, and if you ask me again a month from now on, I completely forgot it's like, okay, maybe now seven, only four. I don't know. But then if you ask me, okay, here are two songs and which one could you prefer to play for the birthday party? I was like, okay, I can pray it before this song. So like comparison is a lot easier. So yeah.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“It is in a way, but I think it's more like a part of a big equations, so there are a lot more different components than that. So that's why I was talking about reinforcement learning. I'm not sure if you're a CEO that you interview bring up that term. So the idea is that once you both should like, so like let's say you have a model, give the model a prop, right? And it produce an output, right? You want to buy, like you want to.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Like, what is the new source of data? But we're like portraying, but like a cost of like, is this more of like everyone have very similar pre-trading data? It's like post-trading is where they make a big difference nowadays.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I think my HTML is a bit of fun process when my friends training what was like try to play with the appreciation model and they're horrendous. They're like saying Taylor's like, oh my gosh, it's like, yeah, it's crazy. So it is very interesting to look at how much of post-training can change the motor behavior. Yeah, and I think that's where like a lot of time is that a lot of people are spending energy on nowadays. They found Shab is on like post-training because pre-training, I think, so pre-training have been used to like increase the gerroc capacity of a model capabilities of a model. And it depends on, it means a lot of data and like model size to increase the model capabilities. And at some point, we are actually happy Max out on the internet data, right? And people like text it max out. I think a lot of people are doing with other data like audios and videos and everyone trying to think of.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think you break widths as functioned, right? So let's say just like maybe it has a function of like maybe Lenny's height is maybe like one x plus something like 2x like one and plus something is a weight, right? So you change it until you fit the correct data, which is like my height and your height, right? So you can think of the width is just like a weight like they function. So you like chain, adjust the width so they can fit the data, which is the training data.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“I think most likely is a most simple way of doing it because it's more like building a distribution of like, okay, Sony is a next token could be more like 90% of the GM. It could be like a color, like 10% of the time could be something else, right? So you base the distribution, so language could pick depending on your sampling strategy. Like, do you want it to always pick the most likely token or do you want it to pick something more creative? So I think my sampling strategy, I think, is something extremely important. So you can have your booster performance in a huge way and very, very underrated.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Right, whereas tokens you can like Be able to like get the sweet spot between the two. So let's say that we have the new wordcasting, right? Let's say it's a new word, but it can divide any podcast and ink. So people understood, okay, podcast when it was a meaning, you know that ing is like a verb, like gerant, whatever it is. So we know the word like podcasting. So that's why it's a token comes in. But yeah, that's like the pre-training is basically like encoding statistical information of language to have you predict what is most likely”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so this is a story of when Shirley Combs was using this statical information to have shown a case. So he was getting, so this is this story, somebody left a message with a lot of stick figures. So Shutohom was like, okay, he knows that in English. The most common letter is E. Then the most common stick figure must be E, right? And then he goes, he starts like that. It was only so the code. So I think that's language. in a way that's like simple language modeling, right? But instead of at the word level, he does this as like character level. And token is something in between, right? Token is not quite a word, but it's because I'm a character. So let's say we say token because it helps us reduce vocabulary because with character is the smallest amount of vocabulary right now. So M5 has eternal character, but words can have like millions and millions.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it's like Claus Shannon. It's a great paper. And I think it reveals a story I really like is from Didier Rich, Sherlock, by the way.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“So I think of language modeling as a way of encoding statistical information about language, right? So let's say that we both speak English. So we kind of get a sense of like what is more statistically likely. Like if I say my favorite color is then it would take okay that should be another color like the word blue would be much more likely to appear than the word like end of table, right? Because statistically blue is more like Asian color is. So it's a sign get this is a way of anchoring statical information. So like when language modeling when you train a large amount of data like you see a lot of languages, a lot of domains. So it can tell like, okay, you basically say this standard, then user do the prompts and it could come like with the next most likely token. So by the way, it's not a new idea. Actually, I realized the idea comes very, very old from the 1951 papers, like English and”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“So that's because I really appreciate open source community, by the way, but like going from the ability trainer models, I can emulate an existing good model. It's very different from being able to train good models like an output for existing good model. So it's a big step there. So yeah, so we have my supervised fi tuning. And another thing that's very big, I'm not sure you have guests talking about it already, but like reinforcement learning is like everywhere.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“Disclaimer, I don't have like phone visibility into what this big secretive frontier labs are doing. But right from what I heard, right? So I think it's like one is like supervised fine tuning when you have demonstration data and you have like a bunch of like experts like, okay, here's a propt, right? And here is what the answer should be like. And you just train it on simulate, like emulate.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“I think you could be a lot of work switching it out. And I was just like, hmm, let's say he's a new technology. It hasn't been tested by a lot of people. And if you adopt it, it would be like stuck with it forever. Like, do you actually want to adopt it? Right. And maybe you want to think twice about over commit to new technologies that hasn't been battle tested.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“On my shoulder get us a lot and a lot is that Do I keep up to date with the latest AI news? And I'm like, why? Why do you need to keep up to date with the latest AI news? And I always have very cultural interest, but I guess so much news out there. A lot of people also ask me questions like how do I choose between two different technologies? Like maybe like recently like MCB versus like Asian Asians, right? Like protocol. And it was like, which one is better or like this or that? And I think it's a serious question you should ask them. It's like, first like if how much of the improvement would you get like from like optimal solutions versus non-optimal solutions, right? And sometimes they were like, actually, it's not much, right? And I was like, okay, if it's not much improvement, then why do you want to spend so much time debating something that doesn't make that much difference to your performance? And another question they ask is like, if you adopt a new technology, like how hard it could be switch that out to another. And sometimes it will like.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“It's really hard to measure productivity. So, I do ask people to ask their managers would you rather give everyone on the team very expensive cooling agent subscriptions or you get an extra headcount? Almost everyone, the managers who say headcal. But if you ask VP level or someone who manages a lot of teams, they would say AI assistant. Because as managers, you are still growing. So for instance, having one extra haircut is big. Whereas for executive, maybe we have more business metrics that you care about. So you actually think about what actually drives productivity metrics for you.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“We are in an ideal crisis now. We have all these really cool tools you have to do everything from scratch. It can have your design, it can have your record, you can have your website. So in theory, we should see a lot more. But at the same time, it's all like somehow stop. They don't know what to build.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source
“How do we keep up to date with the latest AI news? Why do you guys keep up to date with the latest AI news? If your talk to the users understand what they want or they don't want looking to the feedbacks, then you can actually improve application, way, way, way, way, way, way more.”
2025-10-23 · Lenny's Podcast · Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) · IDENTIFIED FROM THE TRANSCRIPT · source