YouSaid · the spoken record
Jared Quincy Davis
- lines on the record
- 52
- first
- 2024-08-22
- most recent
- 2024-08-22
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Pretty soon. We think it's a fair leadership approach. We've seen a lot of interesting evidence that we'll speak more about pretty soon for things like code generation and agentic tasks and the code regime for things like design, chip design, for things like actual neural network design or network of network design even funny enough in recursive ways. So it's actually a really good turns out a lot of these problems that we care about have that property was verifiable and you can compose these systems and bootstrap your way to much higher performance than people might have imagined. So it seems pretty applicable downstream, but there's a lot of open questions, a lot of work to do further. And I think part of our hope is that the community will explore this more and that these types of workloads that are a bit more paralyzable will become more and more common. There'll be a lot more batch inference, a lot more synthetic data generation, and you won't necessarily need the big interconnected cluster that maybe only Open AI can afford to do kind of cutting-edge work in the future.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“In many ways, we're not originating these. I think some of these are largely into systems like Alpha Code, alpha geometry already. I was pretty inspired to see the alpha geometry results recently as well. Yeah, I think that what we'll see people doing is kind of composing, this sounds funny, but massive networks where maybe each stage in the network will basically be maybe some best of a best of K component with many, many calls to different language models, Claude, Gemini, GPT-4, each with their own spikes in terms of capabilities, kind of throw multiple of them at questions in many cases, and then kind of choose the best response. You might also ensemble that with other components like classical heuristic-based systems and simulators, et cetera, and kind of compose large networks that may make millions of calls to answer a question. I think that type of approach, it sounds kind of farcical right now, but I think it'll seem comes”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Jim and I, 1.5, and llama 3 put one, and things like that. So 2.8% or 3% is actually a pretty major gap on MML you. So pretty intriguing. And I think a lot of practitioners are hopefully going to explore this setting a lot more”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Here, and we were able to, in one case, a prime factorization, the performance go from 3.7% to 36.6%. On prime factorization, which is pretty hard, kind of factor it, you know, taking a number that's a composed composite of two primes, two, three digit primes, and factorizing it, factoring it in two, the constituent primes, have a classic problem that pops up a lot in cryptography. And then also we looked at subjects in the MMLU and found that for kind of the subjects you would expect, math, physics, electroengineering, this type of approach was really helpful. Now we use language models. It doesn't have to be a language model. It could be a simulator. These could be unit tests as your verifier, et cetera. But we think this type of approach kind of points towards maybe a very different paradigm for getting better performance than just kind of scaling them models and doing a whole new pre-training from scratch. The MMLU performance bump was about 3%. And to put that in perspective, the gap between some of the previous best models is often less than 1% between, for example,”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“And there are a lot of cases where this is true. Most software engineering and computing tasks kind of classically have this property. We looked at things like prime factorization or a lot of math tasks. Classically, it can take someone years of suffering to write a proof. And you can read the proof in a couple of hours, right? I think we've all had that experience with some training. So there are many examples like this. And so one thing you can do is you can have models, you can embarrass some horizontally scale out and generate many, many candidate responses and then relatively cheaply check those candidate responses and kind of do a best of K type of approach. And turns out the judge model or the verifier choosing the best candidate response might actually have a lot higher accuracy at selecting the best candidate response from the set is you can kind of repeat this as a procedure to actually bootstrap your way to really, really high performance in many cases. And so, you know, we did have some preliminary investigation.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“But we haven't yet elucidated the principles for how to construct networks of networks, so to speak, these combination systems, so to speak, where you have many, many calls, maybe external components. And so one principle that we To explore Be one thing you can prove to figure out how to compose these calls or whether compose these calls. Verifiable means that it's kind of easier to check an answer than it is to generate an answer.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“So, I think it's kind of in this regime that we're just talking about where more and more often to go beyond the capabilities on Frontier accessible to today's state-of-the-art models and kind of get GPT-5 or GPT-6 early practitioners are starting to do these things, oftentimes implicitly, where they'll call the current state of the art model many, many times. There may be scenarios where maybe you're willing to expend a bit of a higher budget. Maybe it's code or something. And if I said that I can give you a 10% better model, you know, for code, many developers might pay 10x for that access to that. Instead, $20 a month, they might be very willing to pay $200 a month, right? For obvious reasons. And so there's a question of what do you do in that setting? And so people are, you know, if you're willing to call the model many times, you can compose those mini calls into almost a network of network calls. And I guess one of the questions is how then should you compose these networks of networks? What principles should guide their architecture? We kind of know how to construct neural networks.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“One thing you can do that people are doing sometimes is you unroll the current state of the art model many, many times, basically doing chain of thought on the current state of the art model. And then you take what required six steps with the previous day of the art, and that becomes your training data for the next day of the art. That type of approach to generating data that then you'll filter that down to high quality examples and then training malls on it looks very different than just throwing more poor quality data into a massive supercomputer to get the next generation.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Train a small model to be smarter than it should be, wasting money, but then that small model is actually really cheap to inference because it's small and it's way smarter than it should be for its size. And so I think people are getting more sophisticated at thinking about cost in a more of a life cycle way. And that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference. Which is actually a really, really horizontally scalable workflow that you can parallelize. You don't need interconnecting the same way. You don't even need to save their systems in the same way. I think you're seeing that type of workload maybe grow in prominence. And so just to give one, just unpack one more statement on that.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Reframing of the scaling laws that we hold near and dear as an ecosystem. I think that the Chinchilla paper scaling laws that DeepMind uncovered have fuel a lot of the scaling effort. But actually, one kind of funny way of looking at those results is they show that if you want to make a monolithic model as smart as possible, there is an ideal way to distribute parameters, compute, and basically trade iterations. One of the funny thing I think some people have done, like no, is actually choose to maybe inefficiently.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“And so, this is what they did. They generated a million candidate responses and basically filter down to the best one as a way to solve coding, so to speak. So that's pretty interesting. And you're seeing that type of regime more and more. Same with alpha geometry, really a powerful system was just announced recently one the silk metal level and you know in the IMO more broadly not just geometry for a broader class of problems and you know these are kind of compound systems with major kind of synthetic data generation pieces and so I think you're seeing people kind of move computation around interpolate between training and inference for example to make the best use of the info they have and this is actually a kind of a funny I think”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“And they distilled that larger model into the 70B and 8B variant, right? And so they got a very, very high quality small variant. Another example is with alpha code 2 was able to achieve extremely high code proficiency in when competitions with really, really small language models. But what they did was they called the model a million times for every question. That's an adversely paralyzable workload. You can scale it horizontally infinitely. They called it a million times per query and then had a nice kind of pretty elegant regime to filter down to the top 10 responses, which they hadn't tried one by one.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“And the small model did not need the kind of big interconnected cluster. You can train it on a pretty small cluster. However, it was still a non-trivial endeavor because they had to curate and obtain this kind of high quality data. And so one of the things I think you're seeing is for these models like some of the Llama 3.1, 8B, and 70B variants, those models are really small, but they're extremely smart. They're smarter than much, much larger systems, like the prior generation from OpenAI. And the way that they trained Lama B and 70B looks a little bit different. So what they did was they did generate a ton of synthetic data Seems with Llama 3.145B.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Pretty different paradigm. I think myself and a number of my collaborators and Matte and others have kind of termed this compound AI systems. I think you actually see it with these most recent models like alpha geometry. A code and lama three. And so I think this actually points the way towards what the AI infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters. And it was a little bit more interesting. And so maybe I'll use Phi 3 as an example. With 5.3, they take a little bit of a different approach where they train a really high Microsoft trained a really high quality small model on high quality data.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“How will we continue to push the scaling loss? One fact about the scaling loss curves is that they're all plotted on logarithmic or house. And things get better predictably, but it requires a continued 2xing or 10xing to get that next bump in performance. But it's quite a bit harder to get the next 10xing. And so it starts to come kind of intractable pretty quickly. And so it's already prompted, I think, a lot of innovation. So Google, for example, has been doing a lot with training across facilities, for example, across data centers, interconnecting them for these models like Palm 2, something that previously would have seemed a bit inconceivable, or things like DePaco or DLOCO, these models they released that are now trained across facilities. And that's one innovation, but I think actually an even slightly more radical thing is we're starting to see a shift towards”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“I think there's a massive shortage of both power, space, and interconnect for the kind of largest of clusters. It's actually very hard to come by and to construct or to find a really large interconnected cluster. It kind of starts to vanish the larger the cluster gets. Like there are a lot more one K clusters than 10K clusters and 20k clusters and you can keep going. Now, I think one thing that, and it'll get harder and harder to keep the scaling going from there. You know, I think there's one question, which is,”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Think it's a little bit of all the above to be clear. And there'll be many workloads and use cases for which having state-of-the-art extremely large interconnected clusters is a really valuable thing. You know, part of something though that we're noticing and also trying to promulgate further is our basically paradigms that don't require this though as well. And so here's why I'd say things we true at once.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Particularly if you look at it more broadly in terms of total AI compute capacity, the percentage that's accessible, useful, and used is a pretty de minimis fraction of the total.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“for our own platform was actually something that we planned to make available to other people just to use in their own clusters as well. Actually, one of the reasons why we invested in Spot is because we reserve very aggressively healing buff for ourselves so that the GPU failure we can automatically swap in another GPU and a user won't perceive a disruption. And so we actually, we maintain buffer for that reason. And so actually being able to pack that buffer with preemptible nodes is actually a really useful thing. But now we're allowing other people to do this, including third-party partners who want to, for example, make their healing buffer available to others through foundry is really offsetting their economics and the cost of their cluster for them. So it's a really, really powerful thing. And so between Mars and Spot, you can kind of see how these things are really interconnected and in a nice way. But yeah, the number of GPUs available, there's quite a few.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Actually, an iPhone 15 Pro now is actually stronger than a V100 as a funny example. It has about 35 teraflops in F16, I believe, where V100 is around 30. And so there's actually quite a bit of compute in the world, broadly speaking. So I guess the point I'm making. Now, it's not all useful. It's not all interconnected. It's not accessible. It's not all secure. But there's just one point to make. There's a lot of compute in the world. And even for the high-end GPUs, there's a lot more than people think. By many measures, utilization of these even H100 systems will kind of state the art, the most viable, the most precious, et cetera. is in many cases 20% 25% or lower according to some pretty high quality data I've seen from some great sources here yeah so quite quite slow as I mentioned even during these pre-trading rounds it's often 80 lower because of the healing buffer partially that actually ties to another product that we've launched which is this product that we built actually larger for ourselves called Mars it's kind of a funny name but it's monitoring alerting resiliency and security it's basically a suite of tools that we've”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, Bitcoin, though, used a lot of ASICs in particular. Ethereum actually had a higher relative ratio of GPUs. And some of the larger GPU, sorry, Ethereum mining providers like Hive, for example, had less than 1% of global hash power. And so you can actually start to extrapolate. And, you know, they had had quite a few GPUs, like tens of thousands of NVIDIA GPUs. It starts to give you a number of the scale capacity, and then I'll have other mining equipment for way less than 1% of the total half power, about 0.1%. So quite a bit of hash power in Ethereum. And I think that's kind of one proxy. But to give you one more anecdotal along that line,”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“An aggressive guess. Yeah, that's an aggressive guess, and you're directly very correct. It was about 10 to 20 million. Which is quite a substantial scale. And you can, by the way, Sisch check this really easily by looking at the basically the hash power in Ethereum at the peak, which is around 900 tera falls per second, I believe, but a tera hash per second, about a petah hash per second, so quite a bit of fleet power. And a typical V100 will give you between, I believe, 45 to 120 mega hash per second if you really know what you're doing. So that's good one guess. There are tens of millions of Jute.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“The very peak, the tippy top of Ethereum, how many V100 equivalents were there, given there were 10,000 for two weeks for GPT-3? How many were there in Ethereum? Noting, by the way, that these were running 24-7 in Ethereum. So you can modulate your guess based on that.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“At the Pico Ethereum. And so I kind of asked people this question, and it's fun to see people guess. Can I solicit a guess actually a lot? Can you make a guess there? You might know. We have this conversation.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“This looked at least a couple of years ago. This is an evolving thing, but how this looked the example of GPT-3 in its training is kind of an interesting one. It's a bit dated, but I'll use it just because the numbers of the types of machines and numbers there are public, as is not the case for the case for some of these other systems. And so GPT-3 was trained on 10,000 B100 GPUs in an interconnected cluster in Azure for about 14.6 days. To put that in perspective, it was a state-of-the-art system. It was kind of by many estimates eight figures for a single run at the time. So it's a pretty substantial investment by OpenApp. And that tells you that 10,100s run continuously for 14.6 days is quite a bit of compute. I think one interesting kind of maybe trivia question then is how many equivalent GPUs normalized in terms of the number of flops, you know, they weren't fully interconnected, but just interesting proxy measure anyway. Were they in the Ethereum network?”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“One kind of funny bit of trivia that I've posed to a few people that I think reveals how off base a lot of our priors are is kind of what percentage of the world's GPU pedophile capacity or exafloc capacity is kind of owned by the major public clouds. And I've asked as many people and I typically have gotten guesses in the high tens of percent. And the only time I got a lower guess was from Sacey at Microsoft who guessed basis points, which actually is correct. It's a very small, small amount. Maybe as one anecdote to maybe illustrate”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“What each workload needs and cares about, and how this might evolve over time. Actually, that ties to a little bit to the compound AI systems concept. And also to the LLAMA 3.1 release in a funny, tangential way. But yeah, that's kind of one analogy for the product that we launched on the Foundry Cloud platform around Spot.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Now, yeah, so it's all relative convenient. It's kind of seamless for everyone involved and creates a ton more effective space, allows us to get much better economics out of the machines and is really, really helpful for companies. That's one really powerful, I think, thing that we've done. And this is kind of offered in the context of a spot product. I think people are somewhat familiar with spot usage in the typical cloud context. It's a lot more challenging to do with GPUs for a few different reasons, which is why there are not very many GPUs available on spot, definitely not at scale and definitely not with Interconnect. And so one of the things we want to do is enable this. And it makes a lot of other things possible that are pretty neat. And so it's a mechanism we're employing a few in a few different ways. So that's, I think, one thing, and I'll give some more analogies to explain why that's powerful. What we found that companies are using this hypechism quite a bit for everything from training, which is classically seen as a workload that's difficult for spot, but also especially for things like inference, batch inference especially. And that actually opens up another interesting conversation about the different classes of workloads and what they”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“I'm stressing this analogy, but you get the idea, it would be kind of inconvenient. And so, part of what we did was added more and more convenience features, which we could broadly call spot usability. And these are a lot of things that we continue to add. And so, you know, the scenario that we're now at is basically you letting someone else use your spot show up. The sensor kind of knows, okay, you're here. And then the car in your spot's automatically moved via conveyor to another spot where managing the spaces to ensure we can move it somewhere. And then when the person comes to pick up their car, it's kind of brought to them.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“And then the luck can also make a couple of bucks. And so it was kind of a win win win for everyone. And you're kind of double clicking the lot, and it's really, really efficient. Now, that sounds great, but it wouldn't quite work if you showed up to your reserve spot and there was someone parked there. That might be a little bit aggravating. It also might not work if you were forced to call two hours in advance and say, hey, I'm coming. And then the person who was parked in your spot had to leave their dinner reservation to move their car. That wouldn't be a fun model. And so I think one thing that we had to do was kind of create the analog of a system to make this all really convenient. And so, you know, maybe the V1 of this system was you came into the lot and the sensor was triggered saying you're here to pick up to go to your reserve spot and then some valet ran to the car that was parked there and moved it and they maybe moved it from the second floor to the tenth floor and then the person who had parked there previously now comes to the counter asks the valley where their car is and gets some ticket saying it's on the 10th floor and but there are no stairs so they have to walk up the stairs and get it”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“An hour, I'm still getting a massive discount, but it's $3K a month effectively, right? Which is actually pretty substantial. And if you're only using it 40 hours a week when you're in the office as a typical worker, it actually might be effectively $16 an hour as opposed to $12. So it's actually worse. And so I think one kind of funny analogy for one of things that we want to do with a couple of these products is kind of create enable the equivalent of allowing people to park pay as you go in someone else's reserve spot. And that sounds kind of funny, but you can imagine that, okay, that'd be actually an interesting thing. And if you could do that, then depending on how much, how the percentage of the spots that are typically reserved, you might have 10x the effective capacity and a lot. And then also it can kind of be a one-win. Instead of the pay-sy-go person paying 12, you know, they can pay something much lower. In this case, I'll say seven, but they actually be a lot lower. The person who owns the spot instead of paying for can make five equivalent.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“So, I think right now AI Cloud is the AI cloud business is very much like a parking lot business. And that sounds really funny because cloud is supposed to be high tech and you can hardly conceive of a less sophisticated business, at least on the surface than parking lots. And what do I mean by that? Well, there's fundamentally there are two models in the parking lot business. One is pay as you go. For pay-as-go, the racer luxurious, you know, you may or may not find a space. I'm sure many of us have the experience of driving through SF and seeing a lot full sign for lot after lot as we drive around trying to park. And if you do get a spot, you might pay the $12 an hour rate or something like that. I'm choosing that rate because it's the rate of AWS, you know, for on demand. On the other hand, if you want to kind of guarantee that you'll have a spot and also have a better rate, you can basically buy a spot reserved. So you can have your own reserved parking spot in your building. Maybe you pay four.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“In terms of where the market's at. And yeah, I think that's leading to a challenge state of affairs that's going to continue to bring a lot of pain for people. And so, you know, there are several things I think we've employed to do this better. Some are mixed as kind of business model innovations and technical innovations at the same time. But I think we're making a pretty substantial dint in this, but also in a way that's really viable economically for us. And does it involve buying all GPUs and taking undue risks on them per se? And so that's kind of a lot of what we've tried to do is can we do something a lot more efficient? Can we find some leverage, some points of leverage to address this problem?”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Right, yeah, and that's challenging. That means that you have to raise all the capital you may need up front, means that that's a challenging model. In perpetuity, it means also that you can't quite grow the product elastically as demand and interest in it grows. So you're kind of ball necked by the supply chains in a way that I think developers haven't experienced in quite a while. So it's a pretty challenging state, I think, of affairs. And it's also a very challenging from a risk management perspective for these companies because they're making these big commitments that are potentially, you know, if they don't work out or pre-catastrophic for them on this hardware, paying all upfront, paying for these long duration contracts, et cetera, it's very challenging thing. And there's no analog yet the market's not mature enough that there's any analog. So we have in other domains like in commodities markets like wheat, oil, et cetera, where you can buy options and futures and hedge and sell back and things like that. It's kind of still pretty in cool.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Utilize to make the economics kind of work. But yeah, if you could do those two things, it'd be a really, really big deal. That has to be free in the class. And I think that's one of the things that's really absent in AI cloud today. In AI Cloud today, you're kind of forced to get really long-term reservations, often three years for a fixed amount of capacity. No one really wants 64 GPUs for three years or 1,000 GPUs for three years. You want 10,000 for a couple of months, maybe nothing for a little while. Maybe don't know how much you need because you're launching a new product and you're not sure how much demand there will be for it and how much inference capacity you'll need, et cetera. It's very, very challenging if you have to reserve the total amount that you may need up front for a long duration.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Quotes from Will Silfanders to kind of illustrate this was the thing they were trying to think about what was unique about the cloud, what you could do in the cloud that you couldn't do anywhere else. And the killer idea that they converged on fundamentally, the cloud made fast free, quote unquote, fast was free in the cloud is how they put it explicitly. And the idea was if you have a workload that was designed to run for 10 days on 10 machines, in the cloud, you could theoretically run it on 100 machines for one day or 10,000 machines for 15 minutes. And that would cost the exact same. And so you could run it a thousand times faster for the same cost theoretically. And it's like, wow, that's kind of a big deal, right? You can run something a thousand times faster for the same cost and then give the compute back. Now, what that would require, though, and now him to make that actually work, one being that you'd have to be able to kind of reshape a workload that was designed to run on 10 machines to run on 10,000 machines, not trivial. Another is that you actually have to have the 10,000 machines worth capacity in the cloud and make sure”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Definitely not nobody. I think Christian people and a lot of builders really saw the value really early, particularly startups in the VC community. I think Berkeley and some of the researchers there really got it early, like 2009 and earlier. By 2015, I think it was already around 3 billion of run rate. So it was not small. It was very far from what it is today, but it was not small. They didn't break it out yet. But it was not small at all. It was a meaningful thing. I think though people didn't recognize it would become anything like it is today. And it still wasn't clear to people what the value proposition was. You may recall that there was drawbacks, for example, famously exited the cloud and went back on prim and that story they published why they did that and the economics of it and that caught a lot of attention and people were saying I'm not sure if cloud actually makes sense anymore, et cetera. And I think people kind of lost track of that. I think one of the insightful things that some of the early cloud systems companies like Databricks and Snowflake recognized was what are the value propositions to the cloud that's really special.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“And so in 2009, 2010, Mate, Sir, you mentioned wrote this paper called Above the Clouds with some collaborators at Berkeley of Berkeley-Bround Computing. And they talked about why COD would be a big deal. And I think it was not to say the least not appreciated that the points they were making were valid at the time. I think it was kind of not clear to people to say the least. Fast forward a few years, Bezos in 2015 said AWS was for all intents and purposes market-sized unconstrained. And he was kind of laughed by many. That was seen as a really ludicrous statement to make. Fast forward, no one's laughing. In 2019, I remember people saying it's not clear if cloud would be a big deal. Snowflake hadn't yet scaled.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“First, I guess that's a little bit of context for people the cloud as we currently know it is arguably one of the most important business categories in the world. That's, I think, pretty clear. The biggest companies in the world are either clouds, they're Azure AWS, GCP, core component of the biggest companies in the world, or NVIDIA, who sells to the clouds, obviously. So it's clearly an important category. AWS is arguably a trillion dollars plus of market cap if you burke it out from Amazon. So at one point, relatively recently, it was all of Amazon's profit and more. So it's an important category, to say the least. Cloud we know today really started in 2003, which is when Amazon decided 50 people initially, and I believe to start working on AWS. And they worked on this for three years with an Amazon, launched in 2006 in March with S3. Later that year in September, they launched EC2. And that kind of was the beginning of the cloud. It took quite a while for this model to catch on, and people were not quite clear on why this would be useful, even until quite recently.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“This opens up a pretty interesting conversation around what is cloud. I think we've kind of forgotten in this current moment what cloud was originally supposed to be, and it was value proposition was intended to be. I think current AI cloud is not cloud in the originally intended sense by any means. So we should pull on that thread. But I'd say right now, it's basically co-location. Yeah, it's basically co-location, right? It's not really cloud. Yeah. Maybe, yeah, it's definitely worth pulling on that thread a little bit.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Lead to some degradation or challenge downstream and stop the entire workload, right? You've probably heard people have talked a lot about Infiniband and the fact that NVIDIA, part of NVIDIA's advantage, comes from the fact that they do build systems that are also state-of-the-art from networking perspective, right? And their acquisition of Malinox was one of the better of all time, arguably, from market cap creation perspective. And the reason they did this is because they realized that it would be really valuable to connect many, many machines into a single, almost contiguous supercomputer that almost acts as one unit. Yeah, the challenge of that, though, is that there now are many, many more components and many more points of failure. And these things kind of, you know, the point of failure kind of multiplies, so to speak.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Key characteristic of the large regime is that you have to somehow orchestrate a cluster of GPUs. Perform a single synchronized calculation, right? And so it becomes a bit of a distributed systems problem. I think that's one way of characterizing the large regime. Now, a consequence of that is that you have many components. Are all kind of collaborating to perform a single calculation. And so any one of these components failing can actually potentially”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“I think it's more of the complexity has grown, right? And we're in a different regime now. I think that it's fair to say, so maybe stepping back again to definitions. We throw the term large around a lot in ecosystem. I guess one question is what does large mean? And one useful definition of large that I think roughly Is that a large language model, you enter the large regime when the Essentially, the amount of compute necessary to contain even just the model weights starts to exceed. Capacity of even a state of the art single GPU or single node. I think it's fair to say you're in the large regime when you need multiple state-of-the-art servers from NVIDIA or from someone else to even just contain the model. Just basically run e-training and definitely even just to contain the model. That's definitely the larger team.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Funnily enough, the newer, more advanced systems anecdotally fail a lot more than historical systems that were worse in some ways.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“I think Jensen and Vidy have done one of one of many is they've basically taken an entire data center's worth of infrastructure and compressed it down into a single box. When you look at it from that perspective, the fact that 70 pounds isn't quite as alarming, but it is, these are really gnarly systems. And when you end up composing these individual systems, these DGXs or HGXs into large supercomputers, what you're often doing is you're interconnecting thousands, tens of thousands, hundreds of thousands of them. And so the failure probability kind of multiplies. And so because you have millions perhaps of individual components in this supercomputer, the probability that it will run for weeks on in, and this is basically a verbatim quote from Jensen's keynote, is basically zero. That's a little bit of a challenge. And funnily enough, I think the AI infrastructure world is still somewhat incoherent. And so doesn't have perfect tooling, broadly speaking, to deal with these types of things. And so one of the compensatory measures I think people take is reserving this healing buffer, for example. I think that disconnect maybe helps explain why these things fail. And it's actually...”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“Depending on how bad of a batch they have and the frequency of failure in their cluster. And so even in that case, now there are often also large gaps in intermissions between training workloads. Even if you have the GPUs are dedicated to a specific entity. And so even in those most conservative cases, which I'll come back to the less conservative extreme cases, utilization is really can be really quite a bit lower than people would imagine. So we can pull on that case a bit more because I think it's actually quite counterintuitive and really interesting. I think there's a really fundamental disconnect between people's mental image of what GPR GPUs are today and what they actually are. I think that in most people's minds, GPUs are chips, right? And we talk about them as chips. But actually the H100 systems are truly systems. They're 70 to 80 pounds, 35,000 plus individual components. They're really kind of monstrosities in some sense. And the remarkable thing that”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“One of the most, I'd say, positive cases with the highest utilization, which is the case where you're running kind of an end-to-end pre-training job. And so that's the case where you've done a lot of work up front, you've designated a time that you're going to run this pre-training workload for, and you're really trying to get the most utilization out of it. And for a lot of companies, utilization, even during this phase, is sub 80%. So why? One reason is actually that these GPUs, particularly the newer ones, actually do fail a lot. As black practitioners would know. And so one of the consequences of that is that it's very common now to hold aside 10 to 20% minimum of the GPUs that a team has as buffer, as healing buffer in case of a failure, so you can slot something else in to keep the training workload running. Sometimes less than 50% extra.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“The AI offerings from the existing major public clouds and kind of some new GPU clouds haven't really re-envisioned things. And by thinking about a lot of these things, a bit anew, we've been able to improve the economics by and make cases 12 to 20x over lower tech GPU clouds and the existing public clouds. And we'll partially based on some of these products that we'll talk about today that we're releasing and a lot of new things that we're working on. We think we can push that quite a bit further as well. And so our pirate products are essentially infrastructure as a service. So our customers come to us for Elastic and really economically viable access to stay-the-card systems and also a lot of tools to make leveraging those systems really seamless and easy. And we've invested quite a bit in things like reliability, security, elasticity, and just the core price performance.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“So, Foundry, we're essentially a public cloud built specifically for AI. And what we've tried to do is really reimagine All of the systems under girding we call the cloud end to end from first principles for AI hour clothes. And we've showed you this in a bit of a new way. I think.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT
“That's a lot of what we worked on with Foundry saying, can we build a public cloud? Built specifically for AI workloads where we reimagine a lot of the components that constitute the cloud end to end from first principles. And in doing that, can we make things that Currently, it cost a billion dollars, cost $100 million, then $10 million. And over time, and that'd be a pretty massive contribution. I think it would increase the frequency of events like Alpha Pole 2 by 10x, 100x, or maybe even more super linearly. And we're already starting to see the early signs of that, but quite a lot of room left to push this agenda. So really exciting. So that's kind of maybe an initial introduction preamble to how we thought about it. And I can trace that line of reasoning a bit more. But that's kind of part of what we've done.”
2024-08-22 · No Priors · The marketplace for AI compute with Jared Quincy Davis from Foundry · IDENTIFIED FROM THE TRANSCRIPT