YouSaid · the spoken record

Nathan Labenz

lines on the record
112
first
2025-10-14
most recent
2025-10-14
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. One foot in front of the other, no quantum leaps, like it would probably feel pretty manageable, I would think, in terms of the pace of change. Hopefully society could absorb that and kind of adapt to it as we go without one day to the next, like, oh my God, all the drivers are getting replaced or that want to be a little slower because you do have to have the actual physical build out. But in some of these things, customer service could get ramped down real fast, right? A call center has something that they can just drop in and it's like this thing now answers the phones and talks like a human and has a higher success rate and scales up and down. One thing we've seen at Waymark, small company, right? We've always prided ourselves on customer service. We do a really good job with it. Our customers really love our customer success team. But I looked at our intercom data and it takes us like

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Fine tuning of things to just like capabilities we have it haven't been applied to particular problems yet. So, just going through the economy and just sitting with people and being like, why are you doing this? Let's document this. Let's get the model to learn your particular niche thing. That would be a real swag. And in some ways, I kind of wish that were the future that we were going to get because it would be unmethodical, you know, kind of

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Tech that we have could do so much for us, even as is. I think if progress stopped today, I still think we could get to 50 to 80 percent of work automated over the next like five to ten years. It would be a real slog. You'd have a lot of co-scientist type breakdowns of complicated tasks to do. You'd have a lot of work to do to go sit and watch people and say, why are you doing it this way? What's going on here? You'd handled this one differently. Why did you handle that one differently? All this tacit knowledge that people have and the kind of know-how, procedural, just instincts that they've developed over time. Those are not documented anywhere. They're not in the training data. So the AIs haven't had a chance to learn them. But again, when I say like no breakthroughs, I still am allowing there for like.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah, and it's so good. And the safety, you know, I think whatever people want to argue about jobs, it's going to be pretty hard to say 30,000 Americans should die every year so that people's incomes don't get disrupted. It seems like you have to be able to get over that hump and say like saving all these lives, if nothing else, is just really hard to argue against. But we'll see. without influence, obviously. So, yeah, I mean, I am very much on team abundance and My old matcha, I've been saying this less lately, but adoption accelerationist, hyperscaling pauser, the

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Yeah, and with the, you know, another thing that I would say has been slower to materialize than I would have expected are. AI culture wars or sort of the Ramping up of protectionism of various industries. We just saw Josh Hawley, I don't know if he introduced a bill or just said he intends to introduce a bill to ban self-driving cars nationwide. God help me. I've dreamed of self driving cars since I was a little kid, truly like sitting at red lights. I used to be like, there's got to be a way.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Mobile app type things spit out for you for far less and far faster than, and probably honestly with significantly higher quality and less back and forth with an AI system than with your kind of. Middle of the pack developer in that timeframe. One thing I do want to call out, there are definitely people have concerns about progress moving too fast, but there's also concern, and maybe it's rising about progress not moving fast enough in the sense that a third of the stock market is MAG 7. AI CapEx is over 1% of GDP. And so we are

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  7. I guess I think less. I think that's probably true even if we don't get like full-blown AGI that's better than humans at everything. I think you could easily imagine a situation where of however many million people are currently employed as professional software developers, some top tier of them that do the hardest things can't be replaced, but there's not that many of those. And the real rank and file, you know, the people that over the last 20 years were told, learn to code, you know, that'll be your thing. Like the people that are the really top top people didn't need to be told to learn to code, right? They just, it was their thing. They had a passion for it. They were amazing at it. We may not, it wouldn't shock me if we still can't replace those people in three, four, five years' time. But I would be very surprised if you can't get your nuts and bolts web app.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Spit out a lot more tokens, and so that you give back on a per token basis, it's dramatically cheaper, more tokens generated does eat back into some of that savings. But everybody seems to expect the trends will continue in terms of prices continuing to fall. And so, you know, how many more of these like price reductions do you have to then be able to do the power law thing a few more times?

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Incremental performance for 10x the cost is weird. It's definitely not the kind of thing that we're used to dealing with. But for many things it might be worth it and it still might be cheaper than the human alternative. You know, if it's like, well, cursor cost me Whatever, 40 bucks a month or something, would I pay 400 for however much better? Yeah, probably would I pay 4,000 for however much better? Well, it's still a lot less than a full-time human engineer. And the costs are obviously coming down dramatically too, right? That's another huge thing. GPT-4 was way more expensive. It's like 9. It's like a 95% discount from GPT-4 to GPT-5 No small thing, right? I mean, Apple stab was a little bit hard because the chain of thought

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  10. I think I'd probably rather have the models in a lot of cases. I mean, it obviously depends on the exact person you're talking about. But truly forced choice today, and then you've got cost adjustment as well, right? I'm not spending nearly as much on my cursor subscription as I would be on an actual human engineer. So even if they have some advantages. I also have not scaffolded, I haven't gone full co-scientist on my cursor problems. I think that's another interesting. You start to see why Folks like Sam Altman are so focused on questions like energy and the $7 trillion build out because these power law things are weird. And, you know, to get...

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  11. And once you have that, how's anybody who doesn't have that going to catch up with you? Obviously, some of that remains to be validated, but I do think They have been pretty intent on that for a long time. Five years from now, are there more engineers or fewer engineers? I tend to think less. You know, already, if I just think about my own. Life and work, I'm like, would I rather have a model or would I rather have like a junior marketer? I'm pretty sure I'd rather have the model. Would I rather have the models or a junior engineer

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  12. A few hundred research engineer people to having unlimited overnight and like what would that mean in terms of how much things could change and also just our ability to steer that overall process. I'm not super comfortable with the idea of the companies tipping into a recursive self-improvement regime, especially given the level of control and the level of unpredictability that we currently see in the models. But that does seem to be what they are going for. So in terms of like why I think this has been the plan for quite some time. Even you remember that leaked anthropic fundraising deck from maybe two years ago where they said that in 2025 and 2026 the companies that train the best models will get so far ahead that nobody else will be able to catch up. I think that's kind of what they meant. I think that they were projecting then that in the 25-26 timeframe they'd get this like automated researcher.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  13. At what point does that ratio really start to tip where the AI is like doing The bulk of the work. GBD5 notably wasn't a big update over 03 on that particular measure. And it also wasn't going back to the simple QA thing. GPT-5 is generally understood to not be a scale-up relative to 40 and 03. And you can see that in the simple QA measure, it basically scores the same on these long-tail trivia questions. It's not a bigger model that has absorbed like lots more world knowledge. It is Cal is right. I think his analysis that it's post-training, but that post-training is potentially entering the steep part of the S curve when it comes to the ability to do even the kind of hard problems that are happening at OpenAI on the research engineering front. And, you know, yikes. I'm a little worried about that, honestly. The idea that we could go from these companies having

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  14. They do want to create the automated AI researcher. That's another data point, by the way, from this was from the 03 system card. They showed a jump from like low to mid single digits to roughly 40%. PRs actually checked in by research engineers at OpenAI that the model could do. So prior to 03, not much at all, low to mid single digits. As of 2003, 40%. I'm sure those are the easier 40% or whatever. Again, there will be caveats to that. But that's, you're entering maybe the steep part of the S curve there. And that's presumably pretty high end. I don't know how many easy problems they have at OpenAI, but presumably not that many relative to the rest of us that are out here making generic web apps all the time. So at 40%, you got to be starting to, I would think, get into some pretty hard tasks, some pretty high value stuff.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Baton back to you, it now uses a browser and the vision aspect of the models to go try to do the QA itself. So it doesn't just say, okay, hey, I tried my best, wrote a bunch of code, like, let me know if it's working or not. It takes that first pass at figuring out if it's working. And again, that really improves the flywheel, just how much you can do, how much you can validate, how quickly you can validate it, the speed of that loop is really key to the pace of improvement. So it's a problem space that's pretty amenable to the sorts of rapid flywheel techniques. Second, of course, they're all coders, right, at these places, so they want to solve their own problems. That's like very natural. And third, I do think on the social vision competition, who knows where this is all going.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Utopia or dystopia is really the big question there, I think, right? I mean, is maybe one part technical, two parts social in terms of why code has been so focal? The technical part is that it's really easy to validate code. You generate it, you can run it. If you get a runtime error, you get the feedback immediately. It's somewhat harder to do functional testing. Replit recently, just in the last 48 hours, released their v3 of their agent. And it now, in addition to code, code, code, try to make your app work. V2 of the agent would do that. And it could go for minutes and, you know, in some cases generate dozens of files and had some magical experiences with that where I was like, wow, you just did that whole thing in one prompt and it worked amazing. Other times it will sort of code for a while and hand it off to you and say, okay, does it look good? Does it working? And you're like, no, it's not. I'm not sure why you get into a back and forth with it. But the difference between v2 and v3 is that instead of handing

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Humans, I think, are, you know, the leadership is maybe the bottleneck or the will in a lot of places might be the bottleneck, and software might be an interesting case where there is just so much pent-up demand, perhaps, that it may take a little longer to see those impacts because you really do want 10 or 100 times as much software. Let's talk about code because it's, you know, it's where Anthropic made a big bet early on, perhaps inspired by the sort of automated researcher, you know, recursive self-improvement, desired future. And then we saw open AI make moves there as well. When we flesh that out or talk a little about what inspired that and where you see that going.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Decision will be, but I do see a lot of these things where I'm just like, when you really put your mind to it and you identify what would create real leverage for us, can the AI do that? Can we make it work? You can take a pretty large chunk out of high volume tasks very reliably in today's world. And so the impacts, I think, are starting to be seen there on a lot of jobs.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Really gnarly stuff like scanned documents, you know, handwritten filling out of forms. And they've created this auditor AI agent that just won a state level contract to do the audits on like a million transactions a year of these, you know, these packets of documents, again, scanned, handwritten, all this kind of crap. And they just blew away the human workers that were doing the job before. So where are those workers going to go? Like, I don't know. They're not going to have 10 times as many transactions. You know, I can be pretty confident in that. Are there going to be a few still that are there to supervise the AIs and handle the weird cases? And, you know, answer the phones? Sure. Maybe they won't go anywhere. The state may do a strange thing and just have all those people sit around because they can't bear to fire them. Like, who knows what the ultimate discourse is.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  20. But maybe you want that. Maybe there is no limit or no, maybe the regime that we're in is such that if there's 10 times more productivity, that's also the good. And we still have just as many jobs because we want 10 times more software. I don't know how long that lasts. Again, the ratios start to get challenging at some point. But yeah, I think the bottom, you know, the old Tyler Common thing comes to mind you are a bottleneck, you are a bottleneck. I think more often it is are people really trying to get the most out of these things and are they using best practices and have they really put their minds to it or not? Often the real barrier is there. I've been working a little bit with a company that is doing Basically, government dock review. I'll obstruct a little bit away from the details.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  21. In theory, maybe you get some more tickets. Maybe you don't need to adjust headcount too much. But when you get to 90% ticket resolution, are you really going to have 10 times as many tickets or 10 times as many hard tickets that the people have to handle? It seems just really hard to imagine that. So I don't think these things go to zero probably in a lot of environments. But I do expect that you will see significant headcount reduction in a lot of these places. And the software one is really interesting because the elasticities are really unknown. You can potentially produce x times more software per user or per cursor user or per developer at your company, whatever.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Handled like, I don't know. I think there's kind of a relatively inelastic supply. Maybe you'll get somewhat more tickets if people expect that they're going to get better, faster answers, but I don't think we're going to see like three times more tickets. By the way, that number was like 55% three or four months ago. So, you know, as they ratchet that up, the ratios get really hard, right? At half ticket resolution.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  23. A spectrum of service offerings to your customers. I once coded up a pricing page for a set. I actually just vibe coded up a pricing page for a SaaS company that was like basic level with AI sales and service as one price. If you want to talk to human sales, that's a higher price. And if you want to talk to human sales and support, that's a third higher price. And so literally that might be what's going on, I think, in some of these cases. It could very well be a very sensible option for people. But I just, I do see the... Intercom, I've got an episode coming up with. They now have this Finn agent that is solving like 65 of customer service tickets that come in. So, you know, what's that going to do to jobs? Are there really like three times as many customer service tickets to be

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  24. In terms of, I guess, what What is the expectation for jobs? I mean, we're starting to see some of this, right? We are definitely seeing no less than like Mark Benioff has said that they've been able to cut a bunch of headcount because they've got AI agents now that are responding to every lead. Klarna, of course, has said very similar things for a while now. They also, I think, have been a little bit misreported in terms of like, oh, they're backtracking off of that because they're actually going to keep some customer service people, not none. And I think that's a bit of an overreaction. Like they may have some people who are just, you know, insistent on having a certain experience and maybe they want to provide that. And that makes sense. I think you can have a...

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Given the limitations And you could see that in terms of some of the instructions and the help that the meter team gave to people. One of the things that is in the paper that they would, if they noticed that you weren't using cursor super well, they would give you some feedback on how to use it better. One of the things that they were telling people to do is make sure you tag a particular file to bring that into context for the model so that the model has the right context. And that's literally like the most basic thing that you would do in cursor. That's like the thing you would learn on your first hour, your first day of using it. So it really does suggest that these were, you know, while very capable programmers like basically mostly novices when it came to using the AI tools. So I think the result is real, but I would be very cautious about generalizing too much there.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Very mature code bases with high standards for coding and developers who really know their code bases super well, who've made a lot of commits to these particular code bases. So I would say that's basically the hardest situation that you could set up for an AI because the people know their stuff really well, the AI doesn't, the context is huge. People have already absorbed that through working on it for a long time. The AI doesn't have that knowledge. And again, a couple generations ago models. And then a big thing too is that the people were not very well versed in the tools. Why? Because the tools weren't really able to help them yet. I think the sort of mindset of the people that came into the study in many cases was like, well, I haven't used this all that much because it hasn't really seemed to be super helpful. They weren't wrong in that assessment.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  27. What applications did they have open? Maybe they took a little longer with cursor than doing it on their own, but how much of the time was Cursor the active window and how much of it was some other random distraction while they were waiting? But I think a more fundamental issue with that study, which again wasn't really about the study design, but just in the sort of Interpretation and kind of digestion of it, some of these details got lost. They basically tested the models or the product cursor in the area where it was known to be least able to help. This study was done early this year, so it was done with, you know, kind of one, depending on how you want to count, right? A couple releases ago. With code bases that are large, which again strains the context window. And that's one of the frontiers that has been moving.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Was a little bit too easy for people who wanted to say that, oh, this is all nonsense to latch on to that. And, you know, again, there's something there that I would kind of put in the Cal Newport category too, where for me, maybe the most interesting thing was the users thought that they were faster when in fact they seemed to be slower. So that sort of misperception of oneself, I think, is really interesting. Personally, I think there's some explanations for that that include like hitting go on the agent, going to social media and scrolling around for a while and then coming back, the thing might have been done for quite a while by the time I get back. So honestly, one like really simple one that we're starting to see this in products, one really simple thing that the products can do to address those concerns is just provide notifications like the thing is done now. So, you know, stop scrolling and come back and check its work. That in terms of just clock time, you know, it would be interesting to know.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  29. For one thing, I think that meter paper is worth unpacking a little bit more because this was one of those things. And I'm a big fan of meter and I have no shade on them because I do think do science, publish your results, like that's good. You don't have to Make every experimental result and everything you put out conform to a narrative. But I do think it was a little bit

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Of these summer's developments. I don't hear as much about AI 2027 or situational awareness to the same degree. I do talk to some people who've just moved it a few years back to your point. But yeah, Darkesh had his whole thing around. He still believes in it, but sort of, you know, maybe because this gap in continual learning or something to the effect that maybe it's just going to be a bit slower to diffuse and meter's paper, as you mentioned, showed that engineers are less productive. And so maybe there's less of a sort of concern around people being replaced to the next few years in mass. I think when we spoke maybe a year ago about this, or I think you said something like 50% of 50% of jobs, I'm curious if that's still your litmus test or how you think about it.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Frame my own work so that I'm kind of preparing myself and helping other people prepare for what might be the most extreme scenarios and kind of one of these things where if we aim high and we miss a little bit and we have a little more time, great. I'm sure we'll have plenty of things to do to use that extra time to be ready for whatever powerful AI does come online. But yeah, I guess I don't. My worldview hasn't changed all that much as a result of these summers developments. Anecdotally, I...

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Dario says 2027, Demos says 2030. I'll take that as my range. So coming into GPT-5, I was kind of in that space. And now I'd say, well, I don't know. Dario's got, what cars does he have up his sleeve? They just put out 4.1 opus. And in that blog post, they said, we will be releasing more powerful updates to our models in the coming weeks. So they're due for something pretty soon. You know, maybe they'll be the ones to surprise on the upside this time, or maybe Google will be. I wouldn't say 2027 is out of the question, but yeah, I would say 2020 still looks just as likely as before. And again, from my standpoint, it's like. That's still really soon, you know. So if we're on track, whether it's 28, 29, 30, I don't really care. I try to.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  33. The whole distribution out super much. I think there may be more just kind of shrinking the, you know, it's getting a little tighter because it's maybe not happening quite as soon as it seemed like it might have been. But I don't think too many people, at least that I don't think are really plugged in on this, are pushing out too much past 2030 at all. And by the way, obviously there's a lot of disagreement. The way I kind of have always thought about this sort of stuff is

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  34. And if they had surprised on the upside, it might have narrowed and narrowed in toward the front end of the distribution. And if they surprised on the downside or even just were purely on trend, then you would take some of your distribution from the very short end of the timelines and kind of push them back toward the middle or the end. And so his answer was like, AI 2027 seems less likely, but AI 2030 seems basically no less likely, maybe even a little more likely because some of the probability mass from the early years is now sitting there. So it's not that I don't think people are moving

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Logarithmic scale graphs, it shouldn't really change your mind too much. It's still above the trend line. I talked to Z about this. Z Mashwith's legendary InfoVore and AI industry analyst on a recent podcast too and kind of asked him same question like why do you think the even some of the most Plugged in sharp mines in the space have seemingly pushed timelines out a bit as a result of this. And his answer was basically just, it resolved some amount of uncertainty. You had an open question of maybe they do have another breakthrough. You know, maybe it really is the Death Star if they surprise us on the upside, then all these short timelines, you know, we could have expected. I guess one way to think about it is

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Certainly the graphs that they showed basically showed the kind of with and without thinking. The problem at launch was that that router was broken. So all of the queries were going to the dumb model. And so a lot of people literally just got bad outputs, which were worse than 03 because they were getting non-thinking responses. And so the initial reaction of like, okay, this is dumb. And that sort of, you know, traveled really fast. I think that kind of set the tone. My sense now is that as the dust has settled, most people do think that it is the best model available. And things like the meter, the infamous meter task length chart, it is the best. We're now over two hours and it is still above the trend line. So if you just say, you know, do I believe in straight lines on graphs or not? And how should this latest data point influence whether I believe on these straight lines on

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  37. It needs to use, or maybe there's been a bunch of different research projects on like skipping layers of the model if the task is easy enough, you could skip a bunch of layers. So you might have hoped that you could genuinely on the back end merge all these different models into one model that would dynamically use the right amount of compute for the level of challenge that a given user query presented. It seems like they found that harder to do than they expected. And so the solution that they came up with instead was to have a router where the router's job is to pick, is this an easy query, in which case we'll send you to this model? Is it a medium? Is it a hard? And I think they just have two really models behind the scenes. So I think it's just really easy or hard.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  38. I think another way to understand what they're doing here is they're trying to own the consumer use case and to own that they need to simplify the product experience relative to what we had in the past, which was like, okay, you got GBT4 and 4.0 and 40 mini and O3 and O4 Mini and other things, you know, 4.5 within there at one point. You got all these different models. Which one should I use for which? It's like very confusing to most people who aren't obsessed with this. And so one of the big things they wanted to do was just shrink that down to just ask your question and you'll get a good answer and we'll take that complexity on our side as the product owners. To do that, interestingly, and I don't have a great account of this, but one thing you might want to do is kind of merge the models and figure out, just have the model itself decide how much to think, or maybe even have the model itself decide how many of its experts, if it's a mixture of experts, architecture.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Amount of it was They kind of fucked up the launch, you know, simply put, right? We're tweeting Death Star images, which is the same when Liter came back and said, no, you're the Death Star. I'm not the Death Star. But I think people thought that the Death Star was supposed to be the model that was generally the expectations were set extremely high. The actual launch itself was just technically broken. So a lot of people's first experiences of GPT-5, they've got this model router concept now where

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  40. I don't know. That's probably not a full appreciation. We could go on for a long time, but I would say in summary, GPT-4 was not able to push the actual frontier of human knowledge. To my knowledge, I don't know that ever discovered anything new. It's still not easy to get that kind of output from a GPT-5 or a Gemini 2.5 or, you know, a Claude Opus 4 or whatever. But it's starting to happen sometimes. And that in and of itself is a huge deal. Well, then how do we explain the bearishness or the kind of vibe shift around GPT-5 then? One potential contributor is this idea that if a lot of the improvements are at the frontier, not everyone is working with sort of advanced math and physics and a day-to-day. And so maybe they don't see the benefits in their day-to-day lives in the same way that sort of the jumps in ChatGPT were obvious and shaped the day-to-day. Yeah, I mean, I think a decent...

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  41. And these are things that literally nobody knew before. And GPD4 just wasn't doing that. These are qualitatively new capabilities. That thing, I think, ran for days. You know, it probably cost hundreds of dollars, maybe into the thousands of dollars to run the inference. That's not nothing, but it's also very, much cheaper than years of grad students. And if you can get to those caliber problems and actually get good solutions to them, like what would you be willing to pay, right, for that kind of thing?

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Breakdown of the scientific method, optimized prompts for each of those steps, and then gave this resulting system, which is scaling inference now kind of two ways. It's both the chain of thought, but it's also all these different angles of attack structured by the team. And they gave it legitimately unsolved problems in science, and it in one particularly famous kind of notorious case, it came up with a hypothesis which it wasn't able to verify because it doesn't have direct access to actually run the experiments in the lab, but it came up with a hypothesis to some open problem in virology that had stumped scientists for years, and it just so happened that they had also recently figured out the answer, but not yet published their results. And so there was this confluence where the scientists had experimentally verified and Gemini in the form of this AI co-scientist came up with exactly the right answer.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  43. So, yeah, I think that's really. That's really hard jumping capabilities to miss. I also think a lot about the Google AI co-scientist, which we did an episode with. You can check out the full store on that if you want to. But they basically just broke down the scientific method into a schematic. And this is a lot of what happens when people, there's one thing to say the model will respond with thinking and it'll go through reasoning process and the more tokens it spends at runtime, the better your answer will be. That's true. Then you can also build this scaffolding on top of that and say, okay, well, let me take something as broad and aspirational as the scientific method and let me break that down into parts. Okay, there's hypothesis generation. Then there's hypothesis evaluation. Then there's, you know, experiment design, there's literature review, there's all these parts to the scientific method. What the team at Google did is created a pretty elaborate schematic that represented their best.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  44. So there's a lot of weird stuff, right? The jagged capabilities frontier remains a real issue, and people are going to find peaks and valleys for sure. But GPT-4, when it first came out, couldn't do anything approaching IMO gold problems. It was still struggling on like high school math. And since then, we've seen this high school math progression all the way up through the IMO gold. Now we've got the frontier math benchmark that is, I think now like up to 25%. It was 2% about a year ago or even a little less than a year ago, I think. And we also just today saw something where, and I haven't absorbed this one yet, but somebody just came out and said that they had solved a canonical, super challenging problem that no less than Terrence Tau had put out. And it was like this thing happened in, I think, days or weeks of the model running versus it was 18 months that it took professional, not just any professional mathematicians, but like really.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Wrong move that is not optimal and thus allows the other player to force a win. And I ask the models if somebody can force a win from this position. Only very recently, only the last generation of models are starting to get that right some of the time, almost always before they were like tic-tac-toe is a solved game. You know, you can always get a draw. And they would wrongly assess my board position as player can still get a draw.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  46. I definitely don't think either one is dead. We haven't seen yet what 4.5 with all that post training would look like. Yeah. And so one of the things that you mentioned that Calysis missed was that it way underestimated the value of extended reasoning. And so what would it mean to fully sort of appreciate that? Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models with no access to tools from multiple companies. And that is night and day compared to what GPT-4 could do with math, right? And these things are really weird. Like it's nothing I say here should be intended to suggest that people won't be able to find weaknesses in the models. I still use a tic-tac-toe puzzle to this day where I take a picture of a tic-tac-toe board where one of the players has made a

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  47. If people take the time or go to the trouble of providing the necessary information, I can kind of access the same facts that way. So you have a kind of, do I want to push on this size and do I want to bake everything into the model? Or do I want to just try to get as much performance out of a smaller tighter model that I have? And it seems like they've gone that way. And I think basically just because they're seeing faster progress on that gradient, you know, in the same way that the models themselves are always kind of in the training process taking a little step toward improvement, you know, the outer loop of the model architecture and the nature of the training runs and where they're going to invest their compute is also kind of going that direction. And they're always looking at like, well, we could scale up over here, maybe get this kind of benefit a little bit, or we could do more post-training here and get this kind of benefit. And it just seems like we're getting more benefit from the post training and the reasoning paradigm than scaling. But I don't think either one is.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Would lose recall. They'd sort of unravel as they got into longer and longer contexts. Now you have obviously much longer context and the command of it is really, really good. So you can take dozens of papers on the longest context windows with Gemini, and it will not only accept them, but it will do pretty intensive reasoning over them and with really high fidelity to those inputs. So that skill, I think, does kind of substitute for the model knowing facts itself. You could say, geez, let's try to train all these facts into the model. We're going to need a trillion or who knows, five trillion, however many trillion parameters to fit all these super long tail facts. Or you could say, well, a smaller thing that's really good at working over provided context can

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Time, the current models are really smart, and you can also feed them a lot of context. That's one of the big things that has improved so much over the last generation when GPT-4 came out, at least the version that we had as public users was only 8,000 tokens of context, which is like 15 pages of text. So you were limited. You couldn't even put in like a couple papers you would be overflowing the context. And this is where prompt engineering initially kind of became a thing. It was like, man, I've really only got such a little bit of information that I can provide. I got to be really careful about what information to provide, lest I overflow the thing and it just can't handle it. There were also as context windows got extended. There were also versions of models where they could nominally accept a lot more, but they couldn't really functionally use them. You know, they sort of could fit them at the API call level, but the model.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Trained with the same power of post training that GBT5 has had. And so we don't really have an apples-to-apples comparison, but people did still find some utility in it. I think maybe the way to understand why they've taken that offline and gone all in on GPT-5 is just that model's really big. It's expensive to run. The price was like way higher. It was a full order of magnitude plus higher than GPT-5 is. And it's maybe just not worth it for them to consume all the compute that it would take to serve that. And maybe they just find that people are happy enough with the somewhat smaller models for now. I don't think that means that we will never see a bigger GPT 4.5 model with all that reasoning ability. And I would expect that that would deliver more value, especially if you're really going out and trying to do esoteric stuff that's pushing the frontier of science or what have you.

    2025-10-14 · a16z Podcast · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · IDENTIFIED FROM THE TRANSCRIPT · source