YouSaid · the spoken record
Dylan Patel
- lines on the record
- 170
- first
- 2024-10-02
- most recent
- 2024-10-02
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Why wouldn't like not a five nanometer, but like of two nanometer A16, A14, these are the nodes that will be in that time frame of 2028 used for AI. I could see like 60, 70, 80% of it. Like, yeah.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Again, this is like a question of how scale pilled you are and how much money will flow into this and how you think progress works. Like, will models continue to get better or does the line does the line slope over? I believe it'll continue to skyrocket in terms of capabilities. In that world...”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it'll be like. Multi hundred billion dollars and But then, like, it'll be, I truly believe people are going to be able to use their prior generation clusters alongside their new generation clusters. And obviously, smaller batch sizes or whatever, right? Or use that to generate and verify data, all these sorts of things.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“I think more reasonably it'll be like the total summation of the flops that you deliver to the model Across pre training.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“The other thing to say is like the way you count flops on a training run is really stupid, you can't just do active parameters times tokens times six, right? Like that's that's really dumb because like the paradigm, as you mentioned, right, is like, and you've had many great podcasts on this synthetic data and like RL stuff, post training, like verifying data and like all these things generating and throwing it away, like all sorts of stuff search, like inference time compute, all these things like aren't counted in the training flops. So you can't like say 1830 is a really stupid number to say because by then, you know, the actual flops of the pre-training may be X, but the data to generate the for the pre-training may be way bigger or like the search inference time may be like way, way bigger, right?”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“If you increase the globally produced flops by like 30x. You're on your or 10x, you're on your, and the cluster size grows, or the cluster size grows by, you know, 30 to 5 to 7x, and then you get multi-site going better and better and better. You can get to the point where multi-million chip clusters, i.e. even if they're like regionally not connected right next to each other, are right there.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“This is this, you know, like draw the line, right? Like, log log lines, let's number goes up, right? You know, if you do that, right? Like, you're going from 100k to 300 to 500k, where the equivalent is a million, you just 10x year on year. Do that again, do that again, or more, right? If you increase the pacing farm.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Not in the near term, it's more again back to the concentration versus decentralization point. Because the largest cluster is 100,000 GPUs. NVIDIA is manufactured close to 6 million hoppers, right? Across last year and this year. So like what that's fucking tiny, right?”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“This is like, you know, like Sam has a superpower, no? Like, it's like recruiting and raising money. That's like what he's like.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“One gigawatt one site In 2026, but then you have a number of five different locations, each with some with multiple sites, some with single site. You're easily north of two, three gigawatts. And then the question is, can you start using the old chips with the new chips? And like the scaling, I think, is like, you're going to continue to see flop scaling much faster than people expect. I think as long as the money pours in, right? Like that's the other thing is like there's no fucking way you can pay for the scale of clusters that are being planned to be built next year for OpenAI. unless they raise like 50 to $100 billion, which I think they will raise that like end of this year, early next year.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Great question. This is where, like, you know, this is where you need the secrets, right? Of like an anthropics got similar plans with Amazon and you go down the list, right? Like people.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“By the end of the year, yes. No, no, no. So one cluster is like the wishy washy definition, right? Multi-site, right? Can you do multi-site? What's the efficiency loss when you go multi-site? Is it possible at all? I truly believe so. Whether it's, what's the efficiency loss is a question, right?”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Next year, 300 to 500,000, depending on whether it's one side or many, right? 300 to like 700,000, I think, is the upper bound of that. But anyways, it's about like when they tier it on, when they can connect them, when the fibers connect it together. Anyways, 300 to like 500,000, let's say, but those GPUs are two to three x faster. Versus the 100K cluster. So on an H100 equivalent basis, you're at a million chips next year in one cluster.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“So then there's the definition game, right? Like Elon claims he has the largest cluster at 100k GPUs because they're all fully connected.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“How much power? And you only look at net ads since 2022 instead of like the total capacity at each data center, then you're still north of multi gigawatt, right? So they're spending $10 billion on these fiber deals with a few fiber companies, Lumen, ZO, like, you know, a couple other companies. And then they've got all these data centers that they're clearly building 100K clusters on, right? Like old crypto mining site with Coreweave in Texas or like this Oracle Crusoe in Texas and then like in Wisconsin and Arizona and a couple other places. There's a lot of data centers being built up and providers, QTS and Cooper and like, you know, you go down the list. There's like so many different providers and self-build, right? Data centers I'm building myself.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, each GPU is getting more power consumption too, right? Like it's like, you know, the rule of thumb is like H100 is like 700 watts, but then like total power per GPU all in is like 1,200, 1300 watts, 1400 watts, but next generation NVIDIA GPUs are, it's 1,200 watts for the GPU, but then it actually ends up being like 2,000 watts all in, right? Like, so there's a little bit of scaling of power per GPU, but you already have 100K cluster, right? OpenAI in Arizona, XAI in Memphis, and many others already building 100K clusters of H100s. You have multiple, at least five, I believe, GB200, 100K clusters being built by Microsoft slash OpenAI slash their partners for them. And then potentially even more, 500k GB200s, right, is a gigawatt, right? And that's like online next year, right? And like the year after that, if you aggregate all the data center sites and like.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Their actions, so Microsoft has signed deals north of $10 billion with fiber companies to connect their data centers together. There are some permits already filed to show people are digging, you know, between certain data centers. So we think with fairly high accuracy, we think that there's five data centers, massive, not just five data centers, sorry, five like regions that they're connecting together, which comprises of many data centers, right?”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Basically, Chinese allegiance, but I think OpenAI wanted to use the data center from them, but instead, like the US forced Microsoft to, I feel like this is what happened is forced Microsoft to do a deal with them so that G42 has a 100k GPU cluster, but Microsoft is like administering and operating for security reasons, right? And there's like Omniba in Kuwait. Like the Kuwait, like, super rich guy spending like five plus billion dollars on data centers, right? Like you just go down the list, like all these countries, Malaysia has, you know, you know, 10 plus billion dollars of like data center, you know, AI data center buildouts over the next couple years, right? Like, and go to every country, it's like this stuff is happening. But on the grand scheme of things, the vast majority of the compute is being built in the US and then China and then like Malaysia, Middle East, and like rest of the world.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“100k GB200 cluster going up in the Middle East, right? And the US, there's like clearly stuff the US is doing, right? G42 is the UAE data center company, cloud company. Their CEO is a Chinese national, or not a Chinese. He's Chinese.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Right? Like, there's a little bit more equipment required, but that's like not too hard. Why don't they? I think there's like true security risks, right? If you're China or if you're the US lab like to build a fucking data center with all your IP and fucking Ethiopia like you want AGI to be in Ethiopia like you want it to be that accessible like people you can't even monitor like like being the technicians in the fucking data center or whatever right or like powering the data center all these things like there's so many like you know things you could do to like you could just destroy every GPU in a data center if you want if you just like fuck with the grid right like pretty pretty like easily i think”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“2025, 2026. Before I talk about the US, I think it's important to note that there's like a gigawatt plus of data center capacity in Malaysia next year. Now, that's like mostly by dance, but like there's like, you know, in PowerWise, there's the humongous damming of the Nile in Ethiopia and the country uses like one-third of the power that that dam generates. So there's like a ton of power there to like.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Fuck you. Like, why are you not going to pay X dollars per kilowatt hour? Because to me, the marginal cost of power is irrelevant. Really, it's all about the GPU cost and the ability to get the power. I don't want to turn it off eight hours a day.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“And this is on a multi-year, including debt financing, including cost of operation, all that, right? Like when you do a TCO, total cost of ownership, like it's like 80% is the GPUs, 10% is the data center, 10% of the power, rough numbers, right? So it's like kind of irrelevant, right? whether or not you like like how expensive the power is right yeah you'd rather do what taiwan does right when like power like what do they do when when there's droughts right they like like force people to not shower”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“For these activities is why it's really difficult. And the economic regulation around that. But the real thing is, if you look at the cost of ownership of an H100, let's just say you gave me a billion dollars and I already have a data center, I already have all this stuff. I'm paying regular rates for the data centers. I'm not paying through the nose or anything. Paying regular rates for power, not paying through the nose. Power is sub 15% of the cost. And it's sub 10% of the cost, actually, right? The biggest 75 to 80% of the cost is just the servers.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“It's so expensive. Yeah, so when you look at the power cost of a large cluster, it's trivial to some extent, right? Like, you know, like the meme that like, oh, you know, you can't build a data center in Europe or East Asia because the power is expensive, that's not really relevant. Or power's so cheap in China and the US, that's where the only places you can build data centers. That's not really the real reason. It's the ability to generate new power.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Cooling and all these other things. This core scientific one is going to be like 1.5, 1.6, i.e., even though I have 300 megawatts of generation on site, I only deliver like 180, 200 megawatts to the chips”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, yeah. So this XAI is doing, right, is like, it's not that polluting on the scheme of things, but it's like you have 14 mobile generators and you're just burning natural gas on site on these like mobile generators that sit on trucks, right? And then you have power directly two miles down the road. There's no unequivocal way to say any of the power is because two miles down the road is a natural gas plant as well, right? There's no way to say this is like green. You go to the core weave thing is a natural gas plant is literally on site from core scientific and all that, right? And then the data centers around it are horrendously inefficient, right? There's this metric called POE, which is basically how much power is brought in versus how much gets delivered to the chips, right? And like the hyperscalers, because they're so efficient or whatever, right? Their POE is like 1.1 or lower, right? I.e. if you get a gigawatt in 900 megawatts or more gets delivered to chips, right? Not wasted on your...”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, they're going through partners, right? Because this is the other interesting thing the big tech companies can't do crazy shit like Elon did. ESG.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Powering these crypto mining data centers from the company called Core Scientific. And so they're just converting that. There's a lot of conversion, but like the power is already there, the power infrastructure is already there. So it's really about converting it, getting it ready to be water cooled, all that sort of stuff, and convert it to 100,000 GB200 cluster. And they have a number of those going up across the country. But that's also like tapped out to some extent because Nvidia's doing the same thing in Plano, Texas for a 32,000 GPU cluster that they're building.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“And that cost about $5 billion, right? $4 billion, right? Not one billion. So the scale that SSI has is much smaller, by the way, right? So their size of cluster will be maybe one-third or one-fourth of the size, right? So now you're talking about 25 to 32k cluster, right? You still don't have that, right? No one is willing to rent you a 32k cluster today, no matter how much money you have, right? Even if you had more than a billion dollars. So you now, it makes the most sense to build your own cluster one instead of renting it or get a very close relationship like a OpenAI Microsoft with CoreWeave or OpenAI Microsoft with Oracle slash Crusoe. The next step is Bitcoin, right? So OpenAI has a data center in Texas, right? Or it's going to be their data center. It's like they kind of contracted all that. Coreweave, there is a 300 megawatt natural gas plant on site.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“There's obviously no free data center lunch, right? In terms of, and you can just, you know, take that based on the data. We have shows that there's no free lunch per se. Like immediately today, you need the compute for a large cluster size or even six months out, right? There's some, but like not a huge amount because of what X did, right? XAI is like, oh, shit, we're going to go like, we're going to go buy a Memphis factory, put a bunch of like generators outside, like mobile generators usually reserved for like natural disasters. A Tesla battery pack drives much power as we can from the grid. Tap the natural gas line that's going to the natural gas plant like two miles away, the gigawatt natural gas plant, like just like send it and like get a cluster built as fast as possible. Now you're running 100k GPUs, right?”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Every six months or three months, the cost of intelligence, like OpenAI and GBD GBD4, what, February 2023, right? $120 per million tokens or something like that was roughly the cost. And now it's like $10. It's like the cost of intelligence is tanking, partially because of compute, partially because the model's compute efficiency wins, right? I think that's a trend we'll see. And then that's going to drive adoption as you scale up and make it cheaper and scale up and make it cheaper.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Hour right at first shorter term or midterm deals, right? Now it's like if you want a six month deal, you could get like $2.15 or less, right? Like, and like the natural cost, if I have a data center, right, and I'm paying like standard data center pricing to purchase the GPUs and deploy them is like $1.40 and then you add on the debt because I probably took debt to buy the GPUs or cost equity, cost of capital, gets up to like $1.70 or something, right? And so you see deals that are like the good deals, right? Like Microsoft renting from Core Weaver, like $1.90 to $2, right? So people are getting closer and closer to like there's still a lot of profit, right? Because the natural rate even after debt and all this is like $1.70. So like there's still a lot of profit when people are selling in the low twos like GPU companies that people were deploying them, but it is a buyer's market in a sense that it's gotten a lot cheaper, but cost of compute is going to continue to tank, right? Because it's like sort of like, I don't remember the exact name of the law, but...”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“So I think for the Ilya question, it was like a cluster of 256 GPUs or even 4K GPUs is kind of cope, right? It's not enough, right? Yes, you're going to make compute efficiency wins, but with a billion dollars, you probably just want the biggest cluster in one individual spot. And so small amounts of GPUs. Probably not possible to use, right? Like for them, right? And that's what most of the sales are, right? Like you go and look at GPU list or VAST or like Foundry or a hundred different GPU resellers, the cluster sizes are small. Now, is it a buyer's market? Yeah, last year you would buy H100s for like $4 or $3.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Obviously, though, like the whole pitch is you're going to make some research breakthrough that's like compute efficiency win, data efficiency win, whatever it is, you're gonna make some breakthrough, but you need compute to get there, right? Because your GPUs per researcher is your research velocity, right? Obviously data centers are very tapped out, right? Not in terms of tapped out, but like every new data center that's coming up, most of them have been sold, which has led people like Elon to go through this insane thing in Memphis, right? I'm just trying to, I'm just trying to square the circle, yeah.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Okay, so the constraints were a US Israeli firm because that's what SSI is, right? And your researchers are in the US and Israel, you probably can't build data centers in Israel because power is expensive as hell. And it's probably risky, maybe. I don't know. So still in the US, most likely. Most of the researchers are here, or a lot of them are in the US, right? Like Paltor or whatever. So I guess you need a significant chunk of compute.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Not going to say there are converging, but they are getting a little bit closer in terms of how big is the Mattmole unit size and like some of the topology and world size of the scale up versus scale out network. Like there is some convergence slightly. Like not saying they're similar yet, but like already they're starting to, but then there's different architectures that people could go down and paths. So you see stuff like from all these startups that are trying to go down different tech trees because maybe that'll work. But there's a self-fulfilling prophecy here too, right? All the research is in transformers that are very high arithmetic intensity because the hardware we have is very high arithmetic intensity and transformers run really well on GPUs and TPUs and like you sort of have a self-fulfilling prophecy if all of a sudden you have an architecture which is theoretically it's way better, but you can get only like half of the like usable fops out of your chip, it's worthless because even if it's 30% you know compute efficiency win it took twice it's half as fast on the chip right”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, that's the quiet part out loud, right? But there's already divergence in capabilities there, right? If you look at image recognition, China destroys American companies, right? On that, right? Because the surveillance. You have this divergence in tech tree and like people can like.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Like here, but that's because state space models potentially have a huge advantage in video and image and audio, which is like stuff that China does more of and is further along and has better capabilities in. So it's like there's...”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“And it applies to use cases, it applies to data, right? Like American models are very important about, let me learn from you, right? Let me be able to use you directly as a random consumer, right? That is not the case for Chinese model, I assume, right? Because there's probably very different use cases for them. China crushes the West at video and image recognition, right? ICML, like Albert Gu at, you know, of Cartesia, like state space models, like every single Chinese person was like, can I take a selfie with you? Man was harassed in the U.S. Like you see Albert and he's like, it's awesome. He invented state space models, but it's not like state space models are like.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Directly coupled, or is the compute cell, right? So these are like things that China's investing hugely, and you go to conferences like, oh, there's 20 papers from Chinese companies slash universities about computed memory. Or like, you know, hey, like, because the flop limitation is here, maybe Nvidia pumps up the on-ship memory and like changes the architecture because they still stand to benefit tens of billions of dollars by selling chips to China, right? Today it's just like neutered American chips, right? Neuter chips that go to the US, but like it'll start to diverge more and more architecturally because they'd be stupid not to make chips for China, right? And Huawei, obviously, again, like has like their constraints, right? Like, where are they limited on memory? Oh, they have a lot of networking capabilities, and they could move to certain optical networking technologies directly onto the chip much sooner than we could, right? Because that is what's optimal for them within their search space of solutions, right? Because this whole area is like blocked off.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“So, you can take this to a very large conclusion leap, and it's like, oh, you know, neuromorphic computing or whatever is the optimal path. And that looks very different than what a transformer does, right? Or you could take it to a simple thing, which is the level of sparsity, like coarse grain sparsity, i.e. like experts and all this sort of stuff, the arrangement of what exactly the detention mechanism is because there are a lot of tweaks. It's not just like pure transformer attention, right? Or like, hey, how wide versus tall the model is, right? That's like very important, like demod versus number of layers, right? These are all like things that would be different. And like, I know they're different between like, say, a Google and an open AI and what is optimal. But what really it really starts to get like, hey, if you were limited on a number of different things like, like China invests humongously in compute and memory, you know, which is basically the memory cell is.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“X amount of compute of TPU versus GPU, compute optimally what is the best thing, you'll diverge in what the architecture is. And I think that's important to note, right?”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“I think that there's a few vectors to go here, right? One is, you mentioned, and I think it's important to note, is that hardware has a huge influence on the model architecture that's optimal. And so it's not a one-way street that better chip equals the optimal model for Google to run on TPUs given architecturally than what it is for OpenAI with Nvidiaged stuff, right? It is like absolutely different. And then even down to networking decisions that different companies do in data center design decisions that people do, the optimal, like if you were to say.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Simultaneously to where NVIDIA Gen to Gender Gen will do more than 2x performance per dollar. I think that's very clear. And then hyperscalers are probably going to try and shoot above that, but we'll see if they can execute.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“What has happened is the output per individual has soared because of EDA electronic design assistance tooling, right? Now this is all still like classical tooling. There's just a little bit of inkling of AI in there yet, right? What happens when we bring this in is the question and how you can solve this search space somehow with humans and AI working together to optimize this so it's not most of the power data movement and then the logic the compute is actually very small to flip side the compute is first of all compute can get like 100x more efficient just with like design changes and then you could minimize that data movement massively right so you can get a humongous gain in efficiency just from architecture itself and then process node helps you innovate that there right and uh power delivery helps you innovate that system design chip to chip networking helps you innovate that right like memory technologies there's so much innovation there and there's so many different vectors of innovation that people are pursuing”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“The question is how much can we advance the architecture, right? The challenge, the other challenge is like the number of people designing chips has not necessarily grown in a long time, right? Yeah, like company to company, it shifts, but like within like the semiconductor industry in the US and the US makes the vast majority of leading edge chips, the number of people designing chips has not grown much”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Compute ALUs or Ethiopologic Unit designs, right? But even then, the vast majority of the power doesn't go there, right? The vast majority of the power goes to moving data around, right? And then when you look at what is the movement of data? It's either networking or memory, you know, you have a humongous amount of movement relative to compute and a humongous amount of power consumption relative to compute. And so how can you minimize that data movement and then maximize the compute there are 100x gains from architecture, even if we like literally stopped shrinking, I think we could have 100x gains from architectural advancements over what”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, useful compute, right? What is, you know, if the goal is optimize intelligence per picajoule, right? And intelligence is some nebulous nature of what the model architecture is, but then picajoule is like a unit of energy, right? How do you optimize that? So there's humongous innovations possible in architecture, right? Because vast majority of the power on a h100 does not”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source
“I think it's important to state that semiconductor manufacturing and design is the largest search base of any problem that humans do because it is the most complicated industry that anything that humans do. And so, you know, when you think about it, right, there's 2010, 2011, right? 100 billion transistors on leading edge chips, right? Blackwell has 220 billion transistors or something like that. So what is, and those are just on-off switches. And then think about every permutation of putting those together, contact ground, et cetera, drain source, blah, blah, blah, with wires, right? There's 15 metal layers, right? Connecting every single transistor in every possible arrangement. This is a search space that is literally almost infinite, right? You could like the search space is much larger than any other search space that humans know.”
2024-10-02 · Dwarkesh Podcast · @Asianometry & Dylan Patel — How the semiconductor industry actually works · IDENTIFIED FROM THE TRANSCRIPT · source