YouSaid · the spoken record

Lip-Bu Tan

lines on the record
110
first
2026-07-15
most recent
2026-07-15
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Is super interesting. I think this is the first time that we've seen that there's a launch of a new model where, to some degree, cost is the headline item, right? I mean, so far I was always like, who is the smartest model? Who has the highest MLU score? Now we have suddenly people who use models for coding for eight hours a day and surprised that if you take a large context window and the best model creates thousands of dollars of cost a month. So cost matters. And so to some degree, sort of way on the Pareto frontier between cost and performance is the new benchmark for model competitive no longer cost alone. Is that what we're seeing here?

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Yeah, of course they're going to take a cut, right? But this is a much better way of monetizing the free user. It's like, you know, it's like Etsy, 10% of their traffic now comes from chat. And OpenAI makes nothing off of that. But they really, really will soon, right? And partially that's because Amazon blocks chat. But there's a way to make money from shopping decisions, whether it's booking flights or looking for items. And those you now say, free user, I don't care. I'm going to send you to my best model. I'm going to send you to agents. I'm going to spend ungodly amounts of compute on you because I can make money off of this. But if it's a query that's like, help me with my homework, I'll send you like a decent model, right? I don't need to spend money on you. And so this is how I think OpenAI can finally make money off of the free user. And I think that's the biggest thing about the router, right?

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  3. For shopping, right? And now this immediately clicks like, oh, if the user asks a low value query, hey, why is this guy blue? Just route them to mini, right? The model can answer perfectly fine. And that is a chunk of queries, right? But if they ask, what's the best DUI lawyer near me, right? All of a sudden this is like, you know, you're in jail, you have one shot, you're like, screw it, let me ask ChatGPT what the best DUI lore is. And now all of a sudden, the model is not capable of it today, but soon enough it'll be able to contact all the lawyers in the area and figure out what their results are and maybe search their court filings and whatever, right? Book the best lawyer for you or an airplane negotiator.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  4. If they need to, right? And I think the router points to the future of OpenAI from a business, right? Like you can look at sort of the model companies, right? Anthropic is fully focused on B2B, right? API, code, et cetera, right? Or cloud code, whatever it is, right? OpenAI, yes, they have that business, Kodaks and API business, but really their majority of the revenue is consumer, right? And it's consumer subscriptions. But they have no way to upsell, you know, to make money off of all the free users, right? In any other application, consumer app, the free user still pays via ads, but this is not compatible with AI, right? Like it's a helpful assistant. You can't just make the users with result worse by injecting ads. Banner ads don't really work in AI either. So it's like, how do you now monetize them? And I think with the router, they're getting really close to figuring out how to monetize that user, right? With the new COV applications, if you saw her product that she launched at Shopify, I think it was Shopify, was an agent.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  5. I think, yeah, I mean, and they talked about how they've been able to dramatically increase their infrastructure capacity because I myself was just regularly using O three or 4.5, right? And now I'm forced to use auto, which sometimes gives me the 03 equivalent thinking model, but sometimes gives me just the regular base model, which sucks. But like, I think for the freezer, it's actually quite interesting, right? The free user was not getting thinking models pretty much ever or not using them, or in many cases they just opened the website and asked their query. And now sometimes their query gets routed there. So sometimes they get a way better model. But now sometimes the open AI can gracefully degrade them.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Isn't it even more interesting open AI can now control how much computer wants to allocate to you, right? If we're in a high load situation, maybe tune the router a little bit so it's less, right? Maybe I have no idea what they're doing behind the curtain, but there's this meme out there at the moment that basically all they did, which is a meme, right? It's not true, but all they did is take all three plus a couple of smaller models, put a router in front, and offer that at a lower blender price, essentially, right? I think there's a little bit of that. Cost suddenly matters and they figured out a way how they can steer that.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Again, like, this is something that OpenAI's first thinking models, you know, the first few generations of 01, 03, would think for a long time and waste a lot of tokens, if you will. And when you look at, for example, anthropics thinking models, even when you put them in thinking mode, they think a lot less, right? To get to the same results or better results, right, as OpenAI was. And so OpenAI, I think, optimized a lot of like, well, if I ask, like, I think the silliest one I had asked was like, oh, three once, is pork red meat or white meat? And it thought for like 48 seconds. I was like, what are you doing? Like this should just tell me the answer. And so the nice thing is that GPD5 will think a lot less even if you select thinking manually, but more importantly, they have the sort of auto functionality, the router, which lets them decide whether or not, hey, do I route to the regular model? Do I route to maybe mini if you're at a rate limits or do I route to thinking, right? And how much do I think? But in general, the thinking model will think less.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I think it depends on what tier reviews you are, right? If you're just using GPD 5 and before you were $20 or $200 a month subscriber, you no longer have access to 4.5, which in my opinion is still a better pre-trained model for certain things, or you no longer have access to O3, which would think for 30 seconds on average maybe, right? Whereas Jupyter 5, even when you're using thinking, only thinks for like five to ten seconds on average, right? Which is an interesting sort of phenomenon, right? But basically like GPD-5 is not spending more compute per se. The model did get a little bit better on a vanilla basis, right? 40 to 5 is actually quite a bit better. But when you think about what is this curve of intelligence, right? It's like the more compute you spend, the better the model gets. And that's whether it's a bigger model, which GPD-5 isn't, right? You can see it's not a bigger model. It's roughly the same size, or you think more, right?

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  9. I think, Dylan, you've done exceptional job in covering what's happening in the AI hardware space, AI semi space, and now more and more data center space as well. Just looking at it, currently the most valuable company on the planet is an AI semi-company, right? I think biggest IPO so far in AI was an AI cloud company. This is currently where it's happening, right? And any gold rush in the early days is the peaks and troubles that make money. And I think this is the stage that we're in. So super excited to have you here today.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source

  10. NVIDIA is going to have better networking than you. They're going to have better HBM. They're going to have better processed node. They're going to come to market faster. They're going to be able to ramp faster. They're going to have better negotiations with whether it's TSMC or SK Hynix and the memory and silicon side or all the rack people or like copper cables, everything, they're going to have better cost efficiency. So you can't just like do the same thing as NVIDIA. You have to really leap forward in some other way. You have to be like 5x better.

    2026-07-15 · a16z Podcast · From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure · IDENTIFIED FROM THE TRANSCRIPT · source