YouSaid · the spoken record
David Luan
- lines on the record
- 64
- first
- 2024-06-24
- most recent
- 2024-06-24
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Everything's an S curve, right? The giant model scaling, base model scaling S curve for the last couple years, we were here. We were at the sharpest point of improvement. You know, you could double the cost of your amount of $100 million to $200 million. And that would be the fastest and easiest way to deliver a smarter thing to the world. And now if you're getting to billion dollar training runs and two billion dollar training runs and $4 billion training runs, it's really freaking hard to go get more money to go make the base thing bigger. And so because of that, now the critical path for model improvement is shifting over to this broader sort of simulation slash synthetic data slash RL loop path. I think it's just a natural consequence of the fact that it's so expensive to just keep scaling.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Positive and negative data on how to solve math problems, and then that makes the model smarter. So the second way of improving model performance is just starting to be tapped now, and that's also going to absorb a boatload of compute. Because of that, I actually am not worried about the diminishing returns to the compute over time.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Whole new way of making models smarter is not just making the base model larger, but it's by having the base model collect data for itself to learn how to get smarter. So let me give you a concrete example of this, right? Concrete example of this is right now, let's say you want to train an LLM to get better at solving math problems. The way you do it is you collect lots and lots of positive solutions to hard math problems, and you throw it in the data set, right? And you're like, hey, like model, go get smarter at this thing. But a much better way to solve this problem is you give the model that you're training access to a theorem-proving and math environment, right? Like give it a Jupyter notebook, theorem proving library out there that a lot of people use as well. Give the model direct access to those tools and then say, hey, I want you to experiment. I want you to go try solving this problem. And then reflect on it. Like, did you do a good job? Like, is this problem solved? If no, try again. So now what you get to do is you get to have the model play with the simulated world basically to collect.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“So, the better way to go think about it, in my view, is like there's two parts to model scaling with compute. One part to it is you simply make the model bigger and then you throw more data and more GPUs at it. If we go look at CPUs and data centers, right? For a long time, we had Moore's law, right? Every year chips would get better at some predictable pace, and everybody's like, oh, Moore's law is going to die. You know, we're at three nanometers or whatever. There's like no more nanometers left. But what actually happened is you go look at the amount of compute available even for chips it's actually continued to trend up because now what we do is we build systems that have multiple chips in them so we have like both the scale up of a single chip and the scale out and as a result every year humanity has more and more compute available to it it's the same thing with giant model scaling the base model itself even if the base model itself stops scaling at some point as you throw more compute at it there's a whole new way to go make models smarter that is just being tapped right now and that”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't think so. And here's why I don't think so. It all depends on what axis you use, right? The way that giant model scaling works, historically, if you go look all the way back to GPT-2 to GPT-3 to 4, et cetera, the way it's worked is that let's say just to use a reductionist analogy, every incremental GPU you throw at the problem actually does have diminishing returns, but every doubling of GPUs you throw at the problem has very predictable, consistent returns. It's kind of like a logarithmic curve versus a straight line, right? Depending on what axis you used to go look at it. So put another way for just scaling up a base language model, you need to double the amount of compute for that language model for it to be predictably consistently smarter. Does that make sense?”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Exactly. So it's like the switch from like hiring a thousand people to go sit around thinking about how to put together small rockets versus creating like the Apollo project. It's a lot better to say, hey, our goal is to go to the moon and we're going to hire however many people it takes to go solve going in the moon. It's very different than just like a giant mass of people organically self-organizing to do that.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“While the first one actually is, and going back to the eras of AI discussion we were having, right, what OpenAI realized before basically everybody but DeepMind was that the next phase of AI after Transformer was not going to be about research paper writing. It was going to be about let's choose a major unsolved scientific problem and just try to solve it. And so like that led us to go build a culture of instead of like loose collections of federations of researchers, let's put a giant team around how to solve like robot hand control. Let's put a giant team around being humans at like one of the most popular video games on the planet, right? Let's put a giant team around scaling GPT until this thing is like a generalist reasoning and chat engine. That's just a totally different framework from like this very academic curiosity driven research. And I think that that's the right framework. And I think that's a big part of how we build adept now as well.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Make like super low level breakthroughs and modeling because it kind of just worked for everything. And then you got to go take that thing to solve really, really big problems.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“While I think what happened was in twenty seventeen the transformer came out. I was running engineering at OpenAI at the time and I was working really closely with ILIA. And Ilia and I were just sitting around and he was just like, look, this transformer thing is real. It's going to be the next most important thing. Let's get all of our teams looking at how we can use this thing. The thing that most people in the general public don't know about is we didn't invent transformer at OpenAI. It was invented at Google. But what Transformer did, though, it was the first time you had a model that was generally applicable to any machine learning task. Back in the day, if you wanted to understand images, you used a convolutional neural network. If you wanted to generate text, you used RNNs. If you wanted to beat humans at Go, you used a tree search or RL, right? So you have these all these different models that you would use to solve problems in AI. And then transformer kind of became like the universal model and base element of AI that came out at that time. And once Transformer came out, in a weird way, you kind of stopped needing to”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Problem in AI. Like, how do we create a model that better understands how to generate images? And just go work on that of their own curiosity and drive and maybe some interest in glory and fame through papers. And they do that for like six months or so, and then out pops out this research paper that gets posted to archive and goes to a journal that just solves the problem. That's huge, right? And so the reason I call it bottom up is because it's just driven by the natural interactions between all these researchers in a setting. They figure out what they want to do.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“So I have this worldview of AI progress as being a part of a couple different phases, right? And I like to think about pre-2012 as basically being prehistory. Of course, all of the OGs in the field would probably not like that if I characterized it that way. But before 2012, most of the things we tried just didn't really work, right? Like you had things like a sheep being identified as cats and dogs and chatbots that barely said anything coherent, et cetera. But I think like there was a period between 2012, like 2017 or 2018 where deep learning went from something that people didn't believe in to being like the dominant paradigm in the field. And so during that 2012-2018 era, the way people made progress, what I mean by bottom-up basic research is you hire the most brilliant scientists, they come to work every day with like no near-term objective they're being held accountable to. And they just work together and they think about, you know, hmm, like, I wonder what it'd be like if we could solve this open technical problem.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, Google Brain was and also now is part of DeepMind is a really magical place. I think during the peak days of AI progress on the research side, right, where every day there was a new paper that came out that just changed the world, that like 2012 to 2018 or so era, Google Brain was just incredibly dominant. They did an amazing job picking talent. Like the people who invented Transformer, the people who invented the diffusion model, people who did all of these new optimization techniques that we all take for granted today, they were all a brand at the same time, like truly the Bell Labs of the era. I learned a lot about how to make pure bottom up, like to see what good pure bottom-up basic research looks like at Google Brain.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, thanks, Harry, for having me on it. I've got to watch some of your cool previous episodes, so it's a real honor to be on here”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source
“What OpenAI realized before basically everybody but DeepMind was that the next phase of AI after Transformer was not going to be about research paper writing. It was going to be about let's choose a major unsolved scientific problem and just try to solve it. The second way of improving model performance is just starting to be tapped now and that's also going to absorb a boatload of compute. Because of that, I actually am not worried about the diminishing returns to the compute over time. I think every tier one cloud provider existentially needs to win here.”
2024-06-24 · The Twenty Minute VC · 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept · IDENTIFIED FROM THE TRANSCRIPT · source