YouSaid · the spoken record
Matei Zaharia
- lines on the record
- 38
- first
- 2023-04-06
- most recent
- 2023-04-06
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Ways of doing them recipes or abstractions or whatever that are actually easy enough for everyone to do. And one analogy I like is when I was learning programming, which was sort of like mid late 90s, I got these books on web applications. And it was very complicated. There was a book on MySQL. There was a book on Apache web server like CGI bin, all these things you have to hook together. And now most developers can make a web application in like one function and even non programmers can make something like Google Forms or Salesforce or whatever that sort of, you know, basically it is a custom application. So I think we're far away from that in data NML, but it could sort of look like.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think we're still at the early stages of AI on unstructured data. So things like text and images and so on, really having an impact in applications. So I think ChatGPD related features that every application is going to add will change the way we work with computing. And they'll also change data analytics to some extent because you'll be able to use this data. And honestly, I also think that in terms of just basic data infrastructure and ML infrastructure, we're still pretty early also. It's still many different tools you have to hook together, a lot of complex integration, and you need a lot of sort of specialized people to do it. And I think over time, like I increasingly think that basically, especially because of the capabilities of these AI models, every software engineer will need to become an ML engineer and a data engineer also as they build the application and will figure out”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“About how can we quickly validate something, but at the same time, and even in research, I had the same thing. You try to pick topics that will matter. Like, for example, when I was doing my PhD, I didn't do a ton with machine learning. And I knew people who did it. I helped them out. I built infrastructure, but I didn't do ML research myself. And then later, I kind of decided like, yeah, I am going to do some things, especially around this connecting machine learning to external data sources like search engines. And I know it's going to take a while to really learn about it and get an intuition and stuff, but I think this is going to matter long term because I think the local parsing semantics of what the sentence means is kind of solved already. And the interesting thing will be like, you know, doing this in a bigger system.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“So, especially as a company goes, right? It's like actually it becomes slower to change direction super dramatically. So you really want to think about what will we do long term or our CEO Ali has this decision rule of with any decision I ask about like, hey, which one am I sort of more likely to regret like five years from now? Not five months from now, but like, you know, if I don't do this or whatever, like, what's going to happen? So you try to think about where things are going to go. Of course, you do want to collect data and sort of update your thoughts about it, test hypothesis. And that, I think, is something you can get from research too. Like in research, we always think when we have an idea, it's sort of a race to figure out, is it a good idea and can I publish it? Because the research community values novelty a lot, being the force to do something, you know, for better or worse. It's not amazing, but if you just reproduce a thing that someone else did, unfortunately, you don't get as much credit.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Say like seven or something, and then it tries to make up the explanation, and it's like wrong. But if you tell it, do the explanation for things step by step and then answer, it's more likely to be right. But you can imagine other versions of that, like if it had a scratch pad, if it had a way to backtrack to say, you know, this is kind of a dead end, it might become better. So I think stuff like that, that's kind of around the model. It's still an AI system, but it's not just one giant DNN can further improve its abilities.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“But the other thing that we're learning though from this is it does seem that the type of data you put in and the kind of fine tuning, essentially it's like weighing the data has a lot of impact. So this instruction tuning stuff is like really we have only a few examples of instruction following, but since we do fine-tune the model, it's as if we put a very high weight on it and had lots of examples of that in our training set. I mean, I think it's still an open question. Like, for example, if you made a lot of examples of logical puzzles, right? Like you just generate some problems and solutions, would you get a model that's better at logical reasoning? There are other things you can do. I also think a big problem with current models, Anka hinted at this before, is we're just calling them to generate one token at a time. So for example, you've probably seen this chain of thought. He's a ning thing. Like if you ask a model a math problem and it just tries to answer like how many sheep were there?”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“I just think a lot of things like scale sublinearly in general. Now, it's hard to tell for things like reasoning and so on, but certainly in classical machine learning, like for example, you're trying to learn a function that separates positive and negative examples. And as you add more data, your accuracy doesn't really improve linear.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Small models that you can train for a specific test that are good. Self driving cars is another example. They rapidly improved in quality up to a point. And then they kind of plateaued and they're still not really ready for primetime. Eventually you hit some limits.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“I think first of all, as a researcher, you think a lot about the long-term trends. What could things look like five or ten years from now? What's kind of the fundamental things here? So, for example, this thing about LLMs being commoditized, or honestly, the thing about them kind of maxing out at more parameters, I think many people hadn't really thought about that. But if you think back, like, you know, there is a lot of room to improve efficiency usually in hardware and software for an application. And this particular application is kind of simple because it is all basically like two or three different types of matrix operations. So like it's sort of the hardware designer's dream to do this stuff. And also there are usually diminishing returns from scale in terms of quality of models in general. And you can also kind of see it in other areas like in computer vision, for example. We don't have trillion parameter models. You get actually”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“A good one early on was deployment infrastructure for like how do we deploy and update our software across all the clouds and so on. And we soon realize it's better to go with really standard things like Kubernetes and tools like that than to try to do something custom because they're evolving very quickly. So yeah, that's kind of a good example where at the beginning you say how hard can it be? Let's just build something. But then you realize, wait, every month there's like new stuff coming out and maybe this isn't where we want to focus on.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Let's see, like a lot of research, at least in computer science, the kind of stuff that I've worked on, a lot of research has basically is mostly prototyping. It's like, can we showcase an idea? But it's not really software engineering of like, we'll build a thing that can be maintained and runs flawlessly in the future and supports problems. So I think you should kind of unlearn just the focus on short-term stuff and think about how is this going to go over time eventually, right? There is a phase of the company where you're just prototyping to get a good fit, but you should design things so they can evolve into something that's very reliable long term. The other thing is, I think unlearned trying to invent everything from scratch, you should really be careful about like, hey, where am I doing something unique? Or if I'm doing something different from others, like why is it, right? Don't do it just for gigs. So, because then research is very tempting to say, you know, I did this new thing. I'm going to try all the fanciest.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Are a lot of challenges along the way. I think just being able to learn about all the aspects of a business and how much complexity there is in each one, you know, starting out as a more technical person, I didn't really know what to expected, but there's a ton of depth in each one. And if you understand them, if you really try to understand them, get to know the culture of people there, like really get to know what they're thinking about, you can make much better decisions across multiple aspects of your company.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“No, it really wasn't. Yeah, I mean, as a graduate, I've always been interested in just doing things that help people that have an impact, help people do cool things. And I had seen these open source technologies out there for distributed data processing. I thought, okay, well, I'll try to start one and see how it goes. I wasn't sure that people would really pick it up and use it. But I wasn't looking to be a founder necessarily. I was just looking to do something useful in this emerging space. Honestly, I was at least considering to be a computer science professor. And I thought if I'm going to be a professor and all the most exciting computing is happening in data centers today, and like, I don't know how that works, how am I going to teach computer science to people? So I better learn about that stuff. But it turned out to be something, you know, more broadly interesting.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“You want your startup to have a long term defensible moat, ideally something that goes over time also. So anything around a unique data set, for example, or unique feedback, interaction you have is always good, right? Honestly, even something like adding ML features in your product that just kind of learn from your users and do better recommendation and so on could eventually become a mode where others just can't easily catch up. But I think anything that's around custom data sets is sort of safest.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think there are a lot of these. I think it's very early on. Probably one of the most obvious ones is just the domain or vertical specific models and tools. I actually think even a lot of the enterprises that have a lot of the data in various domains might turn more into data or model vendors of some form in the future as they use this to build something that no one else can. So I wouldn't be surprised at all if you see the next wave of companies for say security analytics or like biotech or analyzing financial data or stuff like that really built around LLM technology in there. And I also think in general in the app development space, like how do you develop apps that incorporate these tools? It's very open. It's not clear what the best way to do it is. And you might end up with really good programming tools that focus on this problem. I would say, you know, for people thinking about startups and so on, like.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“That data, like the questions and answers plus the data in the actual documentation to essentially automatically answer many”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“So, you know, it's always an arms race, and like you're trying to do the best you can because every percent of accuracy you do better in can translate into huge impact. With language models specifically, and especially with kind of conversational ones, I think the really exciting thing is interfaces to people. And I think customer support is a very obvious one. Maybe things like recommendations or asking questions on a product page in retail, things like search augmented with stuff is one. And we've also found that just internal apps in a company that have a lot of internal data can benefit from this kind of thing. So like one of the things we've built, for example, is inside Databricks, we have all these resources for engineers to understand how different parts of the product work, how to operate it, like all the APIs. And people used to just ask each other questions in these Slack channels for each team. And we could use.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, great question. So, traditional ML, we're seeing actually virtually all major enterprises, you know, in all industries are using it. It's changed a lot in the past decade, actually. So it's very good for forecasting things in general and for automating certain types of decisions. So basically like, for example, optimizing your supply chain, right? You don't have time to look at exactly everything that's going on and think about it and have a meeting, but if you do order the right amount of parts to meet your demand this week or if you minimize the amount of time, you know, an agricultural product sits in a warehouse and degrades in quality or stuff like that, it matters a lot. And it can have a huge impact on the profitability of a company. So we're seeing a lot of that people applying it to automate supply chain and to automate basically their operations in various ways. And then there are more classic use cases like FARD detection and stuff like that.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“There were lots of groups working on at least models integrated into search engines, sometimes into calling other tools as well, like calculators. I think it's still a little bit open-ended. There's one extreme where people say the model will figure out what tools to use on its own. I think for enterprise use cases, that's a little bit more than you really need. You can kind of give it some tools and feed it stuff, and it doesn't have to discover and read the manual to figure out which one to use. But yeah, I think that's another piece you'll need for really powerful applications. And then I do think Infrastructure, like just basic training and serving infrastructure is important too, when you start to care about performance, like about latency and speed. And you can see some of the new search engines using these models are not that fast, right? Like a little bit slow, you know, it would be nice to have it faster. And for automated analytics, it's even more important that it's efficient. So there could be, I think there'll be a lot of.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. So definitely the first piece is you need a data platform that could actually build reliable data. So we think that's like the bread and potatoes of getting anything. You need a basis to sort of build on. So we think that will become really important. And maybe data platforms will have to evolve a little bit to be better at supporting unstructured data like text and images and so on and to do quality assessment and stuff like that for it. That's one piece. I think another piece you need is you need the MLOps piece of like being able to experiment with things, deploy them, A, B test them and so on and see what does better and improve it incrementally. And I also think these models will need a good connection to operational systems inside the company to do really powerful things with like the latest data. So you saw probably the support for tools in Chad GPD before.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Kind of token by token generation we're doing now is not an amazing format for reasoning because you have to like linearly like say one thing at a time. So it's not really good for making plans or comparing versions. I think to get a really smart application, you'll need to combine today's language modeling with some other sort of framework around it that uses it multiple times or explores a planned space or whatever. And then you might get something good. And it's also possible that the very largest models are simply memorizing more stuff. So like they're impressive in terms of trivia. I can ask it about some random topic and it'll know, but they're not really smarter at solving even a basic word problem. So yeah, I'm not sure. Unfortunately, especially with training from the web, it's often very hard to tell apart like reasoning from memorization, essentially. They didn't see that thing before. So I think actually being able to do experiments where you train.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Are there results like this? Is it does seem that the core tech is getting commoditized very quickly. So if you just want to run something like today's chat GPD, it will be a lot cheaper because all these hardware manufacturers are building devices that are specialized and much cheaper. And another thing that's making it less expensive is we're figuring out ways to get a smaller model with less data, fewer parameters and stuff to get similar performance. That I think is happening faster than at least I would have thought, you know, a few months ago. So at least to get something with today's capabilities, I think it'll be very affordable and you might just be able to run it locally on your phone or something. The question of how large can if you make a much larger model, is it going to be a lot smarter? I think it's still a bit unknown. I mean, there are people who argue it's going to be very good at reasoning, but at the same time, this”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, great question. So to me, at least the relationship between scale of the model versus quality of the data and supervision you put in versus like design of an application around it and those things and like overall quality, I think the relationship is not 100% clear yet. Like to get a really reliable model that say, I don't know, can make a pharmacy prescription or something like that, maybe you need a trillion parameters. Maybe you actually need a really carefully designed data set and like supervision process, which is kind of traditional sort of ML engineering type work. Or maybe you actually need a clever application where like you're chaining together a couple of models and things and you're saying, well, does this make sense? Can I find a reference? Can I show this example to a human if it's really hard? So I think it's a little bit open. The thing I can say for sure, especially and Dolly and like,”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Models and with big ones to reduce hallucination from them. But I think it's still an open question. But I think if we can figure it out, then these become”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Sources and making them really produce reliable results. Because if you use ChatGPD or GPD4, the two big problems with it are number one, like the knowledge is not up to date. It only knows stuff that was strained on. And number two, a lot of the things that says are inaccurate and it's confident but like in various ways. And I think you can tackle both of these by combining some kind of language model with a system that pulls out like vetted data either from documents like a search engine or from APIs and tables and stuff like that inside your company. For example, when I talk to the chatbot in my bank, it should know my latest bank account balance and transactions and stuff. If I'm like, can you cancel the payment I made? Because I unsubscribed. You should just know what that means. So cracking, how exactly to do that isn't easy. It may actually be easier with small.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, Databricks definitely were using everything we learned from Dahlia and we're learning from our customers to just offer a great suite of tools for training and operating LLM applications. We already have a popular MLOps platform and we also have this open source project called MLFlow that integrates with a lot of tools out there that are offering us built around. So you can expect some nice integrations into that. Separately, we're also working on Databricks product features that use language models internally and learning a lot from developing those and feeding that into our product. So I think in the next few months, you can expect it. And we also have this big user conference data AI summit coming up in June that will probably have a lot of stuff about this. And I would say as a researcher and also kind of with my databricks hat on, the thing I'm most excited about is really connecting these models with reliable data.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so it's based on cloning this other model from Stanford called Alpaca by doing it with an open data set. And that itself was based on something that Meta released, I think, maybe three weeks ago or less called Lama, which is they took a modest size model, 7 billion parameters, and they trained it on a ton of data. I think they said 1.4 trillion tokens or something like that, which is, I don't know how many bytes of data it was, but it was multiple terabytes of data basically. And they said, hey, by just training this for longer, we got a small model that's actually producing pretty high quality content for its size. So there were all these kind of woolly sort of animals out there. And we thought it's just too perfect to clone it. And there are all these other things. I don't know. There are all these things. Yeah.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“We had an example because we've actually been building a slightly bigger version of this too, and we had this question with who is the author of Snow Crash, which is Neil Stevenson. And the initial DALI model said Neil Gaiman. So it's still a Neil, it's still an author, but it's the Han Neil.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think to me the most interesting thing is it's surprisingly good at just freeform like kind of fluent text generation. So you can tell it to create a story or create a tweet or create a scientific paper abstract. And it does a pretty good job at that. And before that, whenever I talk to my NLP researcher friends, they thought that that creativity was the thing that required a lot of parameters from something like GPT-3. Like they actually told me, oh, the knowledge intensive stuff, like remembering facts, tell me the capital of France and whatever, that's not surprising that a small model with a few parameters can do it, but the creativity, that's like really hard. So this one is actually pretty good at the creativity and generation. It's less good at remembering lots of facts, which kind of makes sense given the parameters. So if you ask it about common topics, you know, it'll be good. If you ask it like the author of a book, you know, it might give the wrong one.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“It's very interesting, and I think there's actually a lot of research still to be done here because these models have been mostly locked up in these very large companies for a while, and everyone thought it's too hard to reproduce them. So the interesting thing is, you know, language models had existed for a while. You basically trained them to complete words.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Became Dali. But yeah, we've been looking at this space for a while and seen incredible demand for these kinds of applications.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Listen to the things you're telling it to do and do those, as opposed to just completing text or just telling you a small amount of information, like this is a positive or negative sentiment. So we really wanted to see whether it's possible to democratize this and to let people build their own models with their own data without sending it to some centralized provider that's trying to sort of learn from everyone's data and kind of control their destiny in this space. We were exploring different ways of doing it. And in particular, like Dolly is partly based on this great result from some other faculty members at Stanford called Alpaca, where they tested a way to, you know, basically they use the model to generate a bunch of realistic conversations, and then they use this to train another model that can now kind of carry on conversation on its own. And so we tried essentially cloning that approach, but starting with an open source model. And it actually worked pretty well. And so that's our”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah. So we've had customers working with large language models of various forms, even before ChatGPT came out, but they were doing the more standard things like translation or sentiment analysis or things like that. A lot of them were tuning models for their specific domains. I think we had like almost a thousand customers that were using these in some form. But then when ChatGPD came out in November, it got people interested in using these for a lot more than just analyzing a bit of data and instead creating entire new interfaces or new types of computer applications, new XP answers in them. And so there was an intense interest in this, even at a time when the industry in general is being conscious about spending and which things are really required and so on. This was an exciting one. The really exciting thing about ChatGPD, as you both know, is the instruction following or basically the ability of it to kind of carry on a conversation and like.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so I'm a computer science professor there, so I split my time between that and databricks. And we work on a bunch of things. We usually looking farther ahead into the future. And we've worked a lot on scalable systems for machine learning, how to do efficient training on lots of GPUs and stuff like that, or how to do efficient serving. And then another thing I'm really excited about that we started about three years ago is looking at knowledge intensive applications where you combine a language model with something like a search engine or an API you call or something like that and you try to produce a correct result maybe for a complicated task like do a literature survey and then like tell me you know what you found about this thing with a bunch of references or counter arguments or whatever and i have a great group of phd students that are working on that and you know exploring different ways to do it”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, well, we definitely didn't anticipate necessarily to go to this site, right? A lot of things can go wrong, but we were excited about the confluence of a few trends. So, first of all, it's so easy to collect large amounts of data and people are doing it automatically in many industries. And second, cloud computing makes it possible to scale up very quickly, do experiments, scale down, and so on, which enables more companies to work with this kind of thing. And then the third one was machine learning. So we thought, you know, these are powerful trends. And the exciting thing for us as a company is we didn't invent cloud computing. We didn't necessarily invent big data or anything, but we were able to start at a point in time when many companies were thinking to move into this space and just provide a great platform for that. And there's this migration already happening. And if you provide the best platform as people are migrating to the cloud, they'll consider it.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Sure. Yeah. So Databricks offer is pretty comprehensive data NML platform in the cloud. It runs on top of the three major cloud providers, Amazon, Microsoft, and Google. And it includes support for data engineering, data warehousing, machine learning. And most interestingly, all this is integrated into one product. So for example, you can have one definition of your business metric that you use in your BI dashboards and the same exact definition is used as a feature in machine learning. And you don't have this drift or copying data and you can just kind of go back and forth between these worlds. The company has about 6,000 employees now. And last year we said that we crossed a billion dollars in ARR and we're continuing to go. It's a consumption-based cloud model where customers that are successful can go over time and bring in new use cases and so on.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“Sure, yeah. So Databrick started by Gup of 7 researchers at UC Berkeley back in 2013. And we were really excited about democratizing basically the use of large data sets and of machine learning. So we had seen the web companies at the time were very successful with these things, but most other companies, most other organizations, things like scientific labs and so on weren't. And we were really excited to look at making it easier to do computation on large amounts of data and also to do machine learning at scale with the latest algorithms. So we had started doing our research. We worked with some of the web companies. We also started open source projects, like most notably Apache Spark, which was essentially the first version of it was my PhD thesis. And we had seen a lot of interest in these. And we thought, you know, it would be great to start a company to really reach enterprises and make this type of thing much better.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source
“So, we really wanted to see whether it's possible to democratize this and to let people build their own models with their own data without sending it to some centralized provider that's trying to sort of learn from everyone's data and kind of control their destiny in this space.”
2023-04-06 · No Priors · The Future is Small Models, with Matei Zaharia, CTO of Databricks · IDENTIFIED FROM THE TRANSCRIPT · source