YouSaid · the spoken record

Edo Liberty

lines on the record
34
first
2024-02-22
most recent
2024-02-22
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Why? This can't be the right thing to do. So I'm very excited about us as a community truly understanding how the different components interact and how to build everything much more in some sense correctly. And I hope we get to build some exciting products. By we, I mean the community gets to build some exciting products this year. I think we're going to see a year of a lot of experimentation that went that people went through last year are going to take the production and to build cool products this year. And I can't see, I can't wait to see how that looks like. I have a feeling that this yield is going to be very, very exciting for consumers of AI.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  2. When we learn how to build the subsystems of AI correctly and for each one of them to do the roles optimally, either we're going to be able to do to achieve the same tasks much cheaper, faster, better, or we're going to still want to use the same amount of resources, but achieve much more. What happens today is that we have very crude tools and we try to use everything for everything. Delightfully or shockingly enough, depending on who you are, that kind of works. I mean, we found this like very, very efficient, very general purpose tools, right? But they're still very general purpose. They're still super blunt instruments. Again, as a technologist or somebody who cares deeply about how things are built, you kind of see the inefficiency and it hurts the brain to figure out that we take half the internet and cram it into GPU memory. I'm like, holy.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  3. It's hard to say. I really do want to see a A distillation in some sense of foundational models. And by distillation, I know it's a, I don't mean what usually people say, but does distillation of models? I don't mean that. I mean, the separation of reasoning and knowledge, foundational models get it fundamentally wrong.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  4. So, I mean, there's a ton. We're an infrastructure company. And so we obsess about ease of use and security and stability and cost and scale and performance also as an engineer at heart. I'm very excited about those things. And all of that is coming. Again, serverless is becoming faster, bigger, better, more secure, easier to use. And we're starting to really grapple with what very large companies and very kind of trailblazing tech companies are going through. I said that getting AI to be truly knowledgeable is still complex. I think we're starting to grapple with deeper issues that the entire information retrieval community has been dealing with for about 40, 50 years now. We're starting to see those come to the fore in rag and in AI in general.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  5. Again, I think it's necessary, especially when you build applications, you want the response to conform to some format or have some property. I think that's a given. You should do that. It runs its course after a while. I mean, it's in some sense, you get what you get. It's necessary, but there's a limit in what you can do with that. And RAG, I think, is incredibly powerful. But like I said before, when we talked about canopy, that's not simple either. I mean, it's simpler than the other ones, but silico is working understanding experimentation and so on. This is almost the hallmark of a nascent market when the simplest solution is still somewhat complex.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  6. I can answer both of the scientists and as a business owner, right? As a scientist, I'm all for fine tuning. We have all the evidence to show that done right, it helps tremendously. Business owner, I can tell you that it's actually extremely hard to do. I mean, this is something that unless you have the research team and the AI experts that know how to fine-tune, you might actually make things significantly worse. So there's nothing that says that more data is going to make your model do better. In fact, it oftentimes regresses to something significantly worse. With prompt engineering.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  7. Information will never be available to your foundational model again. So you don't even have to devise some complex mechanism for forgetting. You just don't know it anymore. What are the main reasons why people attach vector databases to a foundational model? It gives you this operational sanity that is almost completely impossible without it.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  8. What people do with vector databases is actually incredibly simple, right? You don't fine-tune your model on your own proprietary data, at which point you know for a fact it doesn't contain any proprietary data because it's never seen any of it. And then at retrieval time or at, you know, whenever you apply the chat or the agent, you retrieve the right information from the database, give it as context to the model, but only do inference. You don't actually retrain and you don't save that interaction, at which point that data doesn't exist anywhere. It's like an ephemeral thing. And the added benefit to that is, by the way, that you can be GDPR compliant. You can actually delete data. So if, you know, if you're a company like a legal company and somebody deletes a document, you can just delete it from the Vector database.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  9. Yeah. So that's a very common and reasonable thing to be concerned about. Data leakage can happen in two main ways. A, if you use a service for your foundational model that frankly retrains their models with your data or records it, right? Or saves it in some way that is opaque to you, right? That is a huge problem and I think a lot of people are struggling with that. The second is if you're building an application in-house, whatever it might be, and you fine-tune your models on added data, that itadata might end up popping where it shouldn't in anselves to other people's questions or whatever.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  10. Well, theoretically that may be possible, but clearly practically that's not feasible, right? So at some point the context window just becomes gigabytes and gigabytes and gigabytes of data like terabytes. I mean, where do you stop, right? And so already today, we have users who use not a very large models, you know, maybe a few billion parameters. And the vector database next to the model contains trillions of parameters, right? And they get much better performance that way, right? Attaching all the context to everything you do, I think runs its course very, very quickly. And it's also unnecessary, to be honest

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  11. I should take First of all, those companies sell their services by the token. So the fact that they allow you to use infinite context windows is not shocking. Okay, that's good for business. The second thing is there's plenty of evidence that increasing the context size doesn't actually improve results unless you do this very carefully, right? So just what's called concept stuffing is not helping. You just pay more and don't actually get much for it. And the last thing that even that, even if you kind of buy into the marketing, that runs its course, right? I don't need Google because I can, every time I query Google, I can send the internet along with my query, right? It's like, yeah.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  12. That's right. No, I mean, I agree 100%. I mean, this is exactly what we're experiencing. And in fact, we already see, even though new players in the Vector database space that basically started to try to take us down all took the open source angle, we already see them, even young as they might be, they are already struggling with the open source strategy. Serverless is the fourth almost complete rewrite of the entire database at Pinecone.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  13. Managed and multi tenant service and to be able to run that at scale and provide the cost scale trade-offs, we actually run a very, very complicated system. And in some sense, even if we gave it as open source to somebody, they wouldn't know what to do with it. It will be a Herculean effort to even run this thing. The right decision was basically that We should offer this as a service, we should manage end to end. And as long as you give people a fully reliable interface and you keep doing that year after year, you earn the trust and the ease of use, that open source becomes, I hope, not an issue.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  14. When we started pine cone, we asked the very basic question of why do people open source the platform, right? One of it was to earn trust. One of them was to get contribution from the community and one of them was a channel to users. And we figured we can earn trust by being excellent while we do and providing an amazing service. need external contribution. And in fact, if you look at statistics, even companies that are open source, 99% of the contributions are actually from the company itself. Not 99, but high 90s. And so that doesn't actually make a huge difference. And in terms of experience, we figured that we can actually provide a much better experience and much better access to the platform than what OpenSource does. And Pinecone is a food.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  15. I'll say that most databases started before cloud was really a fully mature. Product or market or platform. Okay. And so that was the precursor to PLG, essentially, or whatever. It was PLG, right? That was the only way to put a technically complex product at the hands of engineers was to open source it. And you see, I think all, I mean, maybe not all, but definitely the larger databases that are open source out there. I think that's the reason they did that.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  16. Again, I mean, we see this in the market a lot, people tell me, hey, I already use tool X or database Y and why not? And frankly, oftentimes when it's some tiny workload, you're just learning how to use embeddings for the first time and so on, it might actually work okay. It's when people try to actually do something in production, they're trying to scale up, they'll try to actually push the envelope, or they're trying to launch a product that needs to have some unit economics attached to it that makes sense for that product. That's what people run into huge problems. And so many of them just start with us to begin with. To be honest, a lot of them are enthusiasts and they actually kind of enjoy learning how to use a new kind of database. And are you user experience is smooth enough and so many tutorials and notebooks and examples that they actually find it exciting. But I guess.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  17. I mean, I think it'll be used for boosting and other types of levers to control your search. I think the mode of you baking keywords into that is going away. Yes.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  18. To look at it and improve by boosting and all sorts of other tricks that you can bake into spouse vectors, including keywords. My guess is that that's not going to be the dominant mode of search in the very near future.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  19. So it's interesting. Our research actually shows that when you do this well, you very rarely need keywords alongside embeddings. But getting embeddings to perform perfectly is actually, it could be quite intricate. And we find that it's very convenient to have keywords alongside embeddings and to score those things together. We call this hybrid search. And in fact, we made this even more general and we said, okay, why not, you know, keywords under the hood are actually represented as sparse vectors. That's true of any keyword search, by the way. This is just kind of the mathematically identical. And then we said, why don't we just make this more general and just say, hey, you can give idle sparse ODense vectors or both of them and kind of have the best of both worlds. And people find that very convenient. And so I'd highly encourage people.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  20. Actually, go to production with it, you understand the limitations. With other search technologies, this is, again, this is the wrong search mode. If you're searching with keywords, Just now finding relevant information because the embeddings, the contextual space in which these pieces of text documents or images live is in vector space in high-dimensional numeric space, not in keyword space. And like everyone that's ever searched the inbox for an email, you know for a fact you have and not find it knows that keyword search has a deeply flawed retrieval system.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  21. I'll just go back to the fundamentals about what are you trying to achieve, right? And what we're trying to achieve is to give as much context and as much knowledge to foundational models as possible to that easily at scale on a budget, get to a unit economics that actually works for your product, which is incredibly hard to do with AI with many discussions going on about that now. Those other products don't work. They don't work either because they don't scale in terms of the efficiency scale cost, the trade-offs that they can offer because they're not designed to do this. They're designed to do something else. They kind of thought about Vectorindex as a bolt-on, you know, retrofitted feature. And so, yes, it works at small scale, but when you try to...

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  22. Is not Jira tickets, you know, and JIRA tickets are not Slack messages and you might be building a different product, but at least you have some end-to-end starting point that already does something and you can start improving on.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  23. Yeah, so the Victor database is in PineCone specifically a very foundational model, are very foundational pieces of technology. We're very deep in the stack. And to build a A proper full end to end solution say Notion Q&A, there's quite a lot that you have to build on top of it. You have to ingest documents and what's called chunk them. You have to figure out how to break them into like factoids and pieces of information. You have to embed everything with models. You have to ingest them into the vector database. When you get a query, you have to figure out how to manipulate it and how to embed that. You have to search over it. You have to re-rank. There's a lot. There's a whole system you have to build around it. And a lot of people told us that this is actually quite complex and they're right, right? We put out canopy as really an example. It is an end-to-end kind of cookbook. If you just take this, it should work. You should probably once it works, you should figure out how to make it better for your own application, right? Because medical data.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  24. Basically resolve those main problems. It's incredibly easy to operate it scales massively. I mean, again, there's no theoretical limit to how much you can scale. We've tested it with tens and tens of billions with live customers in live traffic. And I'm not going to go into the architectural design, but it's actually designed to be incredibly efficient, like asymptotically better than what can be done with any other architecture. It's fundamentally about removing the all limits so people can actually have all the information they need ready for the foundational models.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  25. Play in that field. We very quickly figured out with our own customers and our own experimentation that something else is much more significant, which is the scale and cost. If you want to be able to answer correctly, you just have to know a lot. If you want to do that, you have to ingest hundreds of millions, billions, sometimes tens of billions of vectors into your vector database. And you want to query it efficiently in terms of cost. You just don't want that to explode in terms of, again, spend. And finally, you want to do that easily. So you don't want to spend weeks and months setting things up and getting it to work. And doing that in our old architecture and frankly with any other architecture today that's not serverless is very difficult. And serverless is here to...

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  26. So, canopy is actually an open source that we PineCon Serverless is just going pine. It's just pine cone, but serverless. What it does is basically removes the limits from what people used to experience before. When we started Pinecone, a lot of the applications had to do with recommendation engines and anomaly detection and other problems where usually the scale was actually fairly small and the requirements had to do with super low latencies and sometimes high throughput. And as a result, you still see a lot of databases kind of

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  27. Yeah, 100%. First text is probably most of what we see nowadays models are really good at images and so on, but text is still the predominant data type. Notion, Q&A now runs on PineCone and they serve essentially question answering with AI to tens of thousands and probably hundreds of thousands of their own customers. Gong does the same thing with sales calls again serves all of their use cases for all of their customers and so on. So one of the most common patterns is companies that themselves become trailblazers and innovators with AI and they themselves hold a lot of their own users or customers text and they want to search over it or generate information on top of it. That ends up being an incredibly common pattern.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  28. So, most companies do use their own company data. It could be whatever it is. Depends on the application they're building. It could be legal data, medical records, internal wiki information, sales calls, you name it. There's an infinite variety. I want to say that this is just rag. I mean, this is just semantic search. I mean, there are many other applications that we didn't talk about, but we can keep it focused on this application for this conversation.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  29. And you could see that if you augment all of them with rag, even on the internet, which is data that they were trained on, you can reduce hallucinations significantly up to 50% sometimes. Interestingly enough, many of them actually start behaving quite similarly in terms of level of accuracy, even though without rag, they actually have quite different behaviors. So it's sort of both like a uniform improvement and a little bit of leveling the playing field. Because we know we can do that very well now, now you can do that also with proprietary data, with data inside your company and so on, stuff that, of course, is not available on the internet and stuff that those models were never trained on. And interestingly enough, again, the quality ends up being incredibly high.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  30. LLMs are already very, very good at representing data in the way that they want to consume it, which is these embeddings. And so you can add question time in real time at the time of the interaction, go and find relevant information. And relevant might be associated with or call it with or something that is similar to whatever it is that you're being asked about. And once you bring that into the context, you can now give much more accurate answers, right? And as a side experiment, we actually loaded what's called common crawl, which is the top internet pages crawled fairly frequently. We load that into PineCone and saw what happens when you augment GPT 3.5 and 4 and llama and mixture and models. from coherent

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  31. The obvious way, in some sense, to add knowledge to your conversational agent, whether it's chat or what have you. We talk about it as generative AI now, but it's much more general than that, is to, again, not shockingly bring the relevant information into the context, right? So that you can actually arm the foundational model with the right pieces of content, with text, with images, with what have you, right? You want to be able to retrieve that from a very large corpus of knowledge that you have, whether it's your own company's data or whether it's the internet or what have you. So happens that

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  32. The tsunami wave of AI that we're going through right now didn't hit yet, but in 2019 the earthquake had already happened. Deep learning models and so on have already been grappled with large language models and transformer models like Bert and others started being used by the more mainstream engineering cohorts. You could already kind of connect the dots and see that where this is going. In fact, before starting pinecone, I myself had founder anxiety between are we already too late versus nobody knows what the hell this is and we're way too early. And it took me several months of like wild swings between those two things until I figured maybe the fact that I have those too early, too late mood swings maybe means exactly the right time.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  33. And that is pine cone when we started that category. People called me concerned and said, what is the vector and why are you starting a database? And now I think they know the answer.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT

  34. So, pinecone is a vector database, and what vector databases do very differently is that they deal with data that has been analyzed and vectorized. I'll explain a second what that means by machine learning models, by large language models, by foundational models, and so on. Actually any models, really understand data in a numeric way. Models are mathematical objects, right? And when they read a document or a paragraph or an image, they don't save the pixels or the words. They save an earic representation called an embedding or a vector. And that is the object that is manipulated, stored, retrieved, searched over and operated on by vector databases very efficiently at large scale.

    2024-02-22 · No Priors · Improving search with RAG architecture with Pinecone CEO Edo Liberty · IDENTIFIED FROM THE TRANSCRIPT