YouSaid · the spoken record

Aravind Srinivas

lines on the record
225
first
2024-06-19
most recent
2024-06-19
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Where any improvement in AI, semantic understanding, natural language processing improves the product and more data makes the embeddings better, things like that, or subdividing cars. Where more and more people drive A bit more data for you, and that makes the models better, the vision systems better, the behavior cloning.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  2. You can just call it a fancy autocomplete. It's fine. Except it actually worked. At a deeper level than before. One property I wanted for a company I started was it has to be AI complete. Was something I took from Larry Page, which is Want to identify a problem where if you worked on it, you would benefit from the advances made in AI. The product would get better Because the product gets better, more people use it. Therefore, that helps you to create more data for the AI to get better That makes the product better. That creates the wheel. It's not easy to Have this property for most companies don't have this property. That's why they're all struggling to identify where they can use AI. It should be obvious where you should be able to use AI. And there are two products that I feel truly nail this. One is Google search

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Yeah, so I got together, my co founders, Dennis and Johnny, and all we wanted to do was build cool products with LLMs. It was a time when it wasn't clear where the value would be created. Is it in the motto? Is it in the product? But one thing was clear. These generative models are transcended from just being research projects to actual user-facing applications. GitHub co-pilot was being used by a lot of people, and I was using it myself and I saw a lot of people around me using it. And Rick Carpathy was using it. People were paying for it. So, this was a moment unlike any other moment before where people were having AI companies where they would just keep collecting a lot of data, but then it would be a small part of something bigger. But for the first time, AI itself was the thing.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  4. You need to insert some element like random seeds where even though the core intelligence capabilities are the same level, they have different worldviews. And because of that, it forces some element of new signal to arrive at. Both are truth-seeking, but they have different worldviews or different perspectives because there's some ambiguity about the fundamental things. And that could ensure that both of them arrive at new truth. It's not clear how to do all this without hard coding these things yourself.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Be cool. I mean, yeah, that's a self play idea. I think that's where it gets interesting, where it could end up being an echo chamber, too, right? Saying the same things, and it's boring, or it could be like you could.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  6. A session, like even a half an hour chat with that AI Completely changed the way you thought about your current problem. That is so valuable.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Yeah, Elon's trying to figure out how to go to Mars, right? And obviously redesigned from Falcon to Starship. If an AI had given him that insight when he started the company itself, said, Look Elon, like I know you're going to work hard on Falcon, but you need to redesign it for higher payloads. And this is the way to go. That sort of thing will be way more valuable. It doesn't seem like it's. Easy to estimate when it will happen. All we can say for sure is it's likely to happen at some point. There's nothing fundamentally impossible about designing a system of this nature. And when it happens, it will have incredible, incredible impact.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Maybe that's just like one particular question. Let's assume a question that has nothing to do with how to solve Parkinson's or whether something is really correlated with something else, whether Ozampic has any side effects. These are the sort of things that I would want more insights from talking to an AI than the best human doctor. And to date doesn't seem like that's the case.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Our more in-depth understanding of an existing, like more in-depth understanding of the origins of COVID? What we have today That it's less about arguments and ideologies and debates and more about truth.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  10. And it's hard to measure because you don't really know there saying that on a front end like this. The timeline is best decided when we first see a sign of something like this. Not saying at the level of impact that PageRank or any of the fast bourier transform, something like that, but even just The level of a PhD student in an academic lab. Not talking about the greatest PhD students or greatest scientists. Like if we can get to that, then I think we can make a more accurate estimation of the timeline. Today's systems don't seem capable of doing anything of this nature.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  11. You may not, and that's okay, but at least it'll force you to think. That's something I didn't consider. Like you'll be like, okay, why should I, how's it gonna help? And then it's going to come and explain. No, no, no. Listen, if it just look at the text patterns, you're going to overfit on websites gaming you. But instead, you have an authority score now.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So, these are the sort of things that I feel like AIs are not there yet to truly come and tell us, he likes, listen, you're not supposed to look at text patterns alone. You have to look at the link structure. Like that sort of a truth.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  13. There are some beautiful algorithms humans have come up with you have electrical engineering background. So fast Fourier transform, discrete cosine transform, right? These are really cool algorithms that are so practical yet so simple in terms of core insight.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  14. I'm talking about like. Real truth to questions that we don't know Explain itself Helping us understand why it is If we see some signs of this, at least for some hard questions that puzzle us, I'm not talking about things like it has to go and solve the clay mathematics challenges. It's more like real practical questions that are less understood today. If it can arrive at a better sense of truth. I think Elon has to stop thinking, right? Like, can you build an AI that's like Galileo or Copernicus where it questions our current understanding and comes up with a new position, which will be contrarian and misunderstood, but might end up being true.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Can it truly create new knowledge What does it take to create new knowledge? At the level of a PhD student in an academic institution. The research paper was actually very, very impactful.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  16. I don't think it would be like one single moment. It doesn't feel like that to me. Maybe I'm wrong here. Nobody knows, right? But it seems like it's limited by a few clever breakthroughs on how to use iterative compute. It's clear that the more inference computed throw at an answer, like getting a good answer, you can get better answers. But I'm not seeing anything that's more like or take an answer. You don't even know if it's right and have some notion of algorithmic truth, some logical deductions. And let's say you're asking a question on the origins of COVID, very controversial topic. Evidence in conflict interactions. Sign of higher intelligence is something that can come and tell us that the world's experts today are not telling us because they don't even know themselves.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  17. So that is where we need to be clear about the regulation is not on the whole conversation around the weights are dangerous. That's all really. It's more about like Application who has access to all this.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And then the answer ended up being transformer. But instead was done by an AI instead of Google Brain researchers. Now, what is the value of that? The value of that is like trillion dollars, technically speaking. So would you be willing to pay $100 million for that one job? Yes. But how many people can afford $100 million for one job? Very few. Some high net worth individuals and some really well capitalized companies.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  19. We're calling it inference. Fluid intelligence, right? The facts, research papers, existing facts about the world, ability to take that, verify what is correct and right, ask the right questions, and do it in a chain. Do it for a long time. I'm not even talking about systems that come back to you after an hour, like a week. Or a month, you would pay, like, imagine if someone came and gave you a transformer-like paper. Let's say you're in 2016. You asked NAI, an EGI, hey, I want to make everything a lot more efficient. I want to be able to use the same amount of compute today, but end up with a model 100x better.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  20. It's not much about. I think at some point it's less about the pre-training or post-training. Once you crack this sort of iterative compute of the same weights.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Correct, or rather, who's even able to Controlling the compute might just be like cloud provider or something, but who's able to spin up a job? Just goes and says, Hey, go do this research and come back to me and give me a great answer

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Less about access to a model's weight, it's more access to compute that is. Putting the world in more concentration of power and few individuals because not everyone's going to be able to afford this much amount of compute Answer the hardest questions.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  23. You don't even need to ask the exact We suggested it's more a guidance for you. You could ask anything else. And if AIs can go and explore the world and ask their own questions, come back and come up with their own great answers. It almost feels like you got a whole GPU server that's just like, hey, you give the task go and explore. Drug design, like figure out how to take alpha fold 3 and make a drug that cures cancer and come back to me once you find something amazing. And then you pay like, say, $10 million for that job. But then the answer came up with you. It's a completely new way to do things. And what is the value of that one particular answer? If So that's the sort of world that I think we don't need to really worry about AIs going rogue and taking over the world.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Indication to me that we can mimic Feynman's curiosity. We could mimic Feynman's ability to thoroughly research something and come up with non-trivial answers to something. But can we mimic his natural curiosity about just his spirit of just being naturally curious about so many different things and endeavoring to try and understand the right question or seek explanations for the right question? It's not clear to me yet.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Think we have like some work in AIs have explored this curiosity driven exploration, like a Berkeley professor Aliosha froze, written some papers on this where in RRL, what happens if you just don't have any reward signal, an agent just explores based on prediction errors. And he showed that you can even complete a whole Mario game or like a level by literally just being curious because games are designed that way. by the designer to like keep leading you to new things. So I think, but that's just like works at the game level and like nothing has been done to really mimic real human curiosity. So I feel like even in a world where you call that an AGI if you can, you feel like you can have a conversation with an AI scientist at the level of Feynman. Even in such a world, I don't think there's any.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  26. I also think it's what kind of makes us really special. I know you talk a lot about this, you know, what makes human special is allowed beauty to how we live and things like that. I think another dimension is we're just deeply curious as a species

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  27. It's possible, right? Haven't cracked it, but nothing says we cannot ever crack it. What makes human special though is like our curiosity? Even if AI has cracked this. It's used asking them to go explore something. And one thing that I feel like AIs haven't cracked yet is being naturally curious and coming up with interesting questions to understand the world and going and digging deeper about them.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Come back and just blow your mind. I think that's if we can achieve that, that amount of inference compute where it leads to a dramatically better answer as you apply more inference compute. I think that would be the beginning of real reasoning breakthroughs.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Like, can you have a conversation with an AI where it feels like you talk to Einstein? Feynman, where you ask them a hard question, they're like, I don't know. And then after a week, they did a lot of research.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Bootstrap, interact, and improve. So maybe when recursive self improvement is cracked, yes, that's when intelligence explosion happens, where you've cracked it. You know that the same compute when applied iteratively keeps leading you to like a Know increasing IQ points or like reliability, and then you just decide, okay, I'm just going to buy a million GPUs and just scale this thing up. And then what would happen after that whole process is done, where there are some humans along the way providing yes and no buttons. That could be a pretty interesting experiment. We have not achieved anything of this nature yet. Know at least nothing I'm aware of unless it's happening in secret in some frontial lab. But so far it doesn't seem like we are anywhere close to this.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  31. For self play, go or chess, you know, who won the game that was Signal. And that's according to the rules of the game. In these AI tasks, like, of course, for math and coding, you can always verify if something is correct through traditional verifiers. But for more open-ended things, like say predict the stock market for Q3. What is correct? You don't even know. Maybe you can use historic data See if you predicted well for Q2 and you train on that signal, maybe that's useful, and then you still have to collect a bunch of tasks like that and create an RL suite for that. Or give agents tasks like a browser and ask them to do things and sandbox it and completion is based on whether the task was achieved, which will be verified by humans. So you do need to set up an RL sandbox for these agents to play and test and verify.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  32. It's hard to say it's not possible. Of course, there are some simple arguments you can make. Like, where is the new signal to this? The AI coming from? Like, how are you creating? New signal from nothing.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Yeah, but this is a good bet to make that if you have a model that's pretty good at math and reasoning, it's likely that it can handle all the corner cases when you're trying to prototype agents on top of them.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So, I think it's going to be pretty important. And the way it transcends just being good at math or coding is if getting better at math or getting better at coding translates to greater reasoning abilities on a wider array of tasks outside of two and could enable us to build agents using those kind of models. That's when I think it's going to be getting pretty interesting. It's not clear yet. Nobody's empirically shown this is the case.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Mathematically, you can prove that it's related to the variation lower bound with the latent. And I think it's a very interesting way to use natural language explanations as a latent. That way you can refine the model itself to be the reasoner for itself. And you can think of constantly collecting a new data set where you're going to be bad at trying to arrive at explanations that will help you be good at it, train on it, and then seek more harder data points, train on it. And if this can be done in a way where you can track a metric, you can start with something that's like, say, 30% on some math benchmark and get something like 75, 80%.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Correct. So the idea for the star paper is that you take a prompt, you take an output, you have a dataset like this, you come up with explanations for each of those outputs, and you train the model on that. Now, there are some impromptu where it's not going to get it right. Now, instead of just training on the right answer, you ask it to produce an explanation if you were given the right answer, what is the explanation you provided? You train on that. And for whatever you got to write, you just train on the whole string of prompt explanation and output. This way, even if you didn't arrive with the right answer, if you had been given the hint of the right answer, you're trying to reason what would have gotten me the right answer and then turning on that.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  37. It's not that weird that such tricks really help a small model compared to a larger model, which might be even better instruction tuned and more common sense. So these tricks matter less for the GPT-4 compared to 3.5. But the key insight is that there's always going to be prompts or tasks that your current model is not going to be good at. And how do you make it? Good at that by bootstrapping its own reasoning abilities. It's not that these models are unintelligent, but it's almost that we humans are only able to extract their intelligence by talking to them in natural language. But there's a lot of intelligence they've compressed in their parameters, which is like trillions of them. But the only way we get to extract it is through exploring them in natural language.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Chain of thought is a very simple idea where instead of just training on prompt and completion, what if you could force the model to go through a reasoning step where it comes up with an explanation and then arrives at an answer, almost like the intermediate steps before arriving at the final answer. And by forcing models to go through that reasoning pathway, you're ensuring that they don't overfit on extraneous patterns and can answer new questions they've not seen before, barely going through their reasoning chain.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Absolutely These are the kind of architectures we should explore more where small models, and this is also why I believe open source is important, because at least it gives you like a good base model to start with and try different experiments in the post-training phase to see if you can just specifically shape these models for being good reasoners.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Basic common sense stuff. But it's hard to know what tokens are needed for that. It's hard to know if there's an exhaustive set for that. But if we do manage to somehow get to a right data set mix that gives good reasoning skills for a small model, then that's like a breakthrough that disrupts the whole foundation model players because you no longer need That giant of cluster for training. And if this small model, which has good level of common sense, can be applied iteratively. It bootstraps its own reasoning and doesn't necessarily come up with one output answer, but things for a while bootstraps for a while. I think that can be like truly transformational.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  41. It memorizes everything You can ask the question why do you need to memorize every single fact to be good at reasoning? Somehow that seems like the more and more compute and data you throw at these models, they get better at reasoning. But is there a way to decouple reasoning from facts? And there are some interesting Where they're training small language models. They call it SLMs, but they're only training it on tokens that are important for reasoning. And they're distilling the intelligence from GPT-4 on it to see how far you can get if you just take the tokens of GPT-4 on data sets that require you to reason, and you train the model only on that. You don't need to train on all of regular internet pages. Just train it on like...

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Another rag architecture, the retrieval augmented architecture, I think there's an interesting thought experiment here that We've been spending a lot of compute in the pre training. To acquire general common sense. But that scene's brute force and inefficient. What do you want is a system that can learn like an open book exam? You've written exams like in undergrad or grad school where people allow you to come with your notes to the exam. Versus, no notes all Think not the same set of people end up scoring number one on both.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  43. That data, like, oh, for this coding query, make sure the answer is formatted with these markdown and syntax highlighting tool use and knows when to use what tools. Can decompose the query into pieces. These are all like stuff you do in the post training phase, and that's what allows you to build products that users can interact with, collect more data, create a flywheel, go and look at all the cases where it's failing, collect more human annotation on that. I think that's where a lot more breakthroughs will be made.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Not easy to make these systems controllable and well behaved without the RHF step. By the way, there's this terminology for this. It's not very used in papers, but people talk about it as pre-trained post-trained. An RLHF and supervised fine tuning are all in post training phase. And the pre training phase is the raw scaling on compute. And without good post-training, you're not going to have a good product. At the same time, without good pre training, there's not enough common sense to actually have, you know. The post training have any effect. You can only teach a Generally intelligent person, a lot of skills. And that's where the pre-training is important. That's why you make the model bigger, the same RLHF on the bigger model ends up like GPT-4 ends up making ChatGPT much better than 3.5.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  45. And then the GPT-3 happened, which is you just scale up even more data. You take Common Crawl, and instead of 1 billion, go all the way to 175 billion. That was done through analysis called the scaling loss, which is for a bigger model, you need to keep scaling the amount of tokens. And you train on 300 billion tokens. Now it feels small. These models are being trained on like tens of trillions of tokens and like trillions of parameters. But this is literally the evolution. Then the focus went more into pieces outside the architecture on data, what data you're training on, what are the tokens, how DDoop they are. And then the Shinshila insight, it's not just about making the model bigger, but you want to also make the data set bigger. You want to make sure that tokens are also big enough in quantity and high quality and do the right evals on like a lot of reasoning benchmarks. So I think that ended up being the breakthrough, right? It's not like attention alone was important. Attention parallel computation transformer.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Getting to the same performance, which means if you run the same job, you would get something that's way better. If you apply the same amount of compute. And so they just train Transformer, like all the books, like storybooks, children's storybooks. And that got really good. And then Google took that inside and did Bert, except they did bidirectional, but they trained on Wikipedia and books. And that got a lot better. And then OpenAI followed up and said, okay, great. So it looks like the secret sauce that we were missing was data and throwing more parameters. So we'll get GPT too, which is like a billion parameter model and trained on a lot of links from Reddit. And then that became amazing. Produced all these stories about a unicorn and things like that, if you remember.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  47. Ilya Sutskyr has been saying unsupervised learning is important, right? They wrote this paper called Sentiment Neuron. And then Alec Radford and him worked on this paper called GPT-1. It wasn't even called GPT-1. It was just called GPT. Little did they know that it would go on to be this big. But just said, hey, let's revisit the idea that you can just train a giant language model and learn natural language common sense. That was not scalable earlier because you were scaling up RNNs. But now you got this new transformer model that's 100x more efficient.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  48. It's a very clever insight that look, you want to learn causal dependencies, but you don't want to waste your hardware, your compute, and keep doing the back propagation sequentially. You want to do as much parallel compute as possible during training. That way, whatever job was earlier running in eight days would run like in a single day. I think that was the most important insight. And whether it's cons or attention, I guess attention and transformers make even better use of hardware than cons because they apply more compute per flop. Because in a transformer, the self-attention operator doesn't even have parameters. The QK transpose softmax times v has no parameter, but it's doing a lot of flops. And that's powerful. It learns multi-order dependencies. I think the inside than OpenAI took from that is, hey,

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Would say it's almost like the last answer that, like, nothing has changed since 2017, except maybe a few changes on what the nonlinearities are and how the square root of descaling should be done. Some of that has changed. And then people have tried mixture of experts having more parameters for the same flop and things like that. But the core transformer architecture has not changed.

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source

  50. And so they just said threw away the RNN. And that was powerful. And so then Google Brain, like Vaswani et al., the transformer paper, identified that, okay, let's take the good elements of both. Let's take attention. It's more powerful than cons. It learns more higher order dependencies because it applies more multiplicative compute. And let's take the inside and WaveNet that You can just have a all convolutional model that fully parallel matrix multiplies and combine the two together and they built a transformer. And that is the

    2024-06-19 · Lex Fridman Podcast · #434 – Aravind Srinivas: Perplexity CEO on Future of AI, Search & the Internet · IDENTIFIED FROM THE TRANSCRIPT · source