YouSaid · the spoken record

Arvind Narayanan

lines on the record
58
first
2024-08-28
most recent
2024-08-28
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Right. So there are a lot of sources that haven't been mined yet. But when we start to look at the volume of that data, how many tokens is that? I think the picture is a little bit different. 150 billion hours of video sounds really impressive. But when you put that video through speech recognizer and actually extract the text tokens out of it and deduplicate it and so forth, it's actually not that much. It's an order of magnitude smaller than what some of the largest models today have already been trained with. Now, training on video itself instead of text extracted from the video, that could lead to some new capabilities, but not in the same fundamental way that we've had before where you have the emergence of new capabilities, right? Models being able to do things that people just weren't anticipating. So they're kind of shocked that the AI community had when I think back in the day, I think it was GPT-2.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  2. So, if we look at what's happened historically, the way in which compute has improved model performance is with companies building bigger models. In my view, at least, the biggest thing that changed between GPT 3.5 and GPT-4 was the size of the model. And it was also trained with more data, presumably, although they haven't made the details of that public and more compute and so forth. So I think that's running out. We're not going to have too many more cycles, possibly zero more cycles of a model that almost an order of magnitude bigger in terms of the number of parameters than what came before and thereby more powerful. And I think a reason for that is data becoming a bottleneck. These models are already trained on essentially all of the data that companies can get their hands on. So while data is becoming a bottleneck, I think more compute still helps, but maybe

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  3. So, when ChatGPT was released, people found a thousand new applications for it that OpenAI might not have anticipated. And that was great. But I think developers, AI developers took the wrong lesson from this. They thought that AI is so powerful and so special that you can just put these models out there and people will figure out what to do with them. They didn't think about actually building products, you know, making things that people want, finding product market fit, and all those things that are so basic in tech, but somehow AI companies diluted themselves into thinking that the normal rules don't apply here.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  4. I think that's possible, generative AI companies specifically. I made some serious mistakes in the last year or two about how they went about things.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  5. And they want to replace these institutions with a script. And that just didn't seem like the right approach to me. So both from a technical perspective and from a philosophical perspective, I really soured on it. While there are harms around AI, I think it has been a net positive for society. I can't say the same thing about Bitcoin.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  6. So I spent years of my time on this. I really believed that decentralization could have tremendous societal impacts. How is this going to make society better? It was not the money angle. But by around 2018, I had started to get really disillusioned. And that was because of a couple of main things. One is in a lot of cases where I had thought Kretor blockchain was going to be the solution, I realized that that was not the case. While there is potential for crypto to help the world's unbanked, the tech is not the real bottleneck there. And the other part, if it was just a philosophical aspect of this community, I do believe that many of our institutions are in need of reform or maybe decentralization, whatever it is. And that includes academia, by the way. So many reforms so badly needed. And in an ideal world, we would have this hardbird, important conversation about how do you fix our institutions. But instead, these students have been sold on block.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Sure. So I'm a professor of computer science, and I would say I do three things. One is technical AI research, and another is understanding the societal effects of AI. And the third is advising policymakers.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source

  8. We're not going to have too many motorcycles, possibly zero more cycles of a model that's almost in order of magnitude bigger in terms of the number of parameters than what came before and thereby more powerful. And I think a reason for that is data becoming a model. These models are already trained on essentially all of the data that companies can get their hands on. So while data is becoming a bottleneck, I think more compute still helps, but maybe not as much as it used to.

    2024-08-28 · The Twenty Minute VC · 20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton · IDENTIFIED FROM THE TRANSCRIPT · source