YouSaid · the spoken record

Steeve Morin

lines on the record
88
first
2025-02-24
most recent
2025-02-24
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. No, you want HBM to be clear. No, SRAM, this will not deliver. It's a dead end in terms of scaling SRAM means scaling the surface means you get depreciating problems. It explodes everywhere, right? So you need some SRAM, right? So we'll have bigger amounts of SRAM into chips. And of course, bigger what's called external memory into chips. The issue with HBM is that it's still slow. And yes, maybe Nvidia has a stronghold and they can prevent you from getting some. So that would be like, I call it the Nutella situation in which 80% of the hazelnuts market, right? So yes, you can do a competitor, but who with you buy the nuts from, right? So there will be a need for HBM. There will be a need for Esram. I would say better, more dedicated architecture will be able to deliver these things. And then there's like the next.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Because the access to external memory prevents it. So HBM is all the rage, right? But HBM compared to SRAM is absolutely dark slow. So this is the problem you get. So HBM is like the best we can do, but it's still slow versus SRAM.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  3. So the way models reason today is they reason in tokens. So it's as if you think to yourself, you would say out loud what you're thinking. So, yes, it works, but it is a bit inefficient, right? And you lose information doing this latent space reasoning is this without going, I would say, to English or whatever, right? So staying in what's called the latent space, which is where all the information of an LLM, let's say an LLM, an LLM lives, right? So this is very much how we work as humans and we move toward what Jan Lecan calls energy-based model in which we have different types of longer or shorter, I would say, thinking times, if you will, right? So that fundamentally, GPUs cannot deliver this. Plain and simple at scale.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Pushed by reasoning. So reasoning not in the sense that you see on deep seek and whatever, right? Reasoning and what's called latent space reasoning. Latent space reasoning and agents will push the market towards different types of compute.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  5. It's hard to say. I mean, you need some Esram, but if you can have a smaller process node, but if you can hook yourself with external memory, then yes, you can do that a lot better. But the thing is, if you go full-blown ESRAM, then there's no magic. You will have to pay the price.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  6. To what they're called their wafer scale engine, which is a chip the size of a wafer. I mean, most likely it's interconnected, but it's huge, right? And it has to be water cool. They have a copper, you know, I would say needles that touch the chip. It's crazy stuff. Very, very impressive technology, mind you, but very, very expensive. So my bet is, I think there will be chips on the market that do that at much lower price. And there's two companies I see going in that direction. One is called Etched, and the other one is called VSOR. That's the two I see. Because if you can deliver this, I would say the price that is comparable to GPUs, you've won.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So there's this trick. Because here's the thing there's no magic. This little trick is called Esram. Esram is memory on the chip directly. So that is very, very fast memory. But here's the problem with Esram, is that SRAM consumes surface on the chip, which makes it a bigger chip, which is very hard in terms of yield, right? Because the chances of problems are higher and so on. SRAM is, I would say, very, very, very fast memory, which gives you a lot of advantage when you do very, very high inference, but it's terribly expensive. And if you look at, for instance, Grok, they have on their generation, this generation, they have 230 megabytes of SRAM per chip. A 70B model is 140 gigabytes. So you do the math, right? Cerebras has a 44 gigabytes of Eshram in

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Let's say, for instance, cerebrus. Incredible technology. Incredibly expensive. So how will the market value the premium of having single stream very high tokens per second? There is value into that, right? As we saw with Mistral and perplexity. But I think that was none of the loss. I don't know. I don't have the details. But I think it was done at the loss that cerebrus put it out. Today there's three actors on the market that can deliver this. I think this will be, I would say, the pushing force for change in the inference landscape, agents and reasoning. So that is, you know, very high tokens per second only for you.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Technically speaking, he is right, but realistically speaking, I don't sure I agree. The thing is, these chips are on the market. They're here. Our tab on Chrome and get one. That is something that I don't take lightly. Availability, that is, right? I think Nvidia is just to stay, at least if not for the H100 bubble bust, because these chips are going to be on the market and people will buy them and do inference with them. Remains to see the OPEX and the electricity, et cetera. But the thing is, the only chips that are really frontier on that sense are probability. And then the upcoming chips. But the thing is, they're great chips, but they're not on the market. Or there are outrageous prices, like millions of dollars to run a model.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So, I might tell you that I think he already started. I'm getting cold emails for discounts from services I never heard about. And I started getting these emails probably around October, November. Some people are left with a lot of CapEx that they don't know what to do with. You know, it's a different thing to build a cluster and run a training and do a training run that it is to build literally a cloud provider or hyperscaler or whatever you want to call it. So there are a lot of people who do their training runs on the regular providers but then move to regular hyperscaler when they do production. So I very much worry there will be an oversupply of these chips. The problem is that remember the chips are the collateral. So somewhere in the US or whatever there's going to be a data center with like a thousand GPUs that

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  11. It might well be, yes, we are being spared a bit because Blackwell is late and others are getting cancelled. So H series, I would say, are still in the active. But yes, absolutely. But, you know, what choice do you have? This is the thing

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  12. That's one example. Yeah. Another one is choosing the right compute. It's like kind of, I would say, a virtual circle because provisioning compute is very hard. So if you lose compute, it's very bad. You are essentially incentivized to overbuy in the case of Amazon or Google that would be buying reserved compute, which you're not going to use because if you buy it on demand, you will get tremendously ripped off. So that creates this like face scarcity of compute because that people buy preemptively because they rate shit ton of money and they're not using it. So this is a major problem too.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  13. At least in that, I would say regular back engineering, this is a problem everybody knows, right? Everybody is doing it because the savings are so huge, but on AI, nobody really had the problem. So now they're coming up to it. So this is one example.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So, for instance, probably the number one thing, depending on how you deploy, but if you're deploying inference, the number one thing that will get you is what's called auto scaling. So as your systems get more and more loaded, you want to provision because these things are tremendously expensive. You want to provision them as you scale, right? So you want to say, I have a thousand GPUs, you know, 24 hours, even if there's nobody under production, I will pay for them, which is, mind you, what people are doing today. This is crazy. So what you want to do is you want to provision, compute as you grow your needs, right? And you want to do it up and you want to do down. Probably the number one thing that gives you a lot of efficiency in terms of spend. Like we're talking multiples, like, you know, five, you know, sometimes 10x, you know, improvement. The thing is, this is a problem.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  15. There's a lot of duct tape. Here's also probably one of the problems that training on first principle is actually two passes, forward and backward, right? It's called forward pass and backward pass, right? Inference is running only the forward pass. So that's how things are today. There are people who are trying to specialize a bit because, you know, at some point duct tape doesn't really work out. And when you run big scales, that makes a problem. And it's a problem that's growing because a lot of people are coming on the market with the needs for inference. That wasn't the case a year and a half ago or a year ago. OpenAI had this problem, right? Maybe anthropic had this problem. But it wasn't a universal problem yet. And now it's becoming a universal problem.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Of it like doing a painting and doing a million painting. The tools you will use, the process you will do. If you do one painting, what you favor is the speed at which you can do a stroke and do some iteration. If you do a million, what you want is a process, a process that is reliable that can deliver you efficiently, a million paintings. So that is the same for training versus inference. If you want to run millions of instances of a model, you cannot hack your way to do that. By the way, people do hack their way today, but this is probably the fundamental difference.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Will not go, you will avoid that, right? If you can. And this is why models have the sizes they have, is so that people can run them without the need to connect multiple machines together. It's very constraining in terms of the environment. So that is probably the fundamental difference. The need for interconnect. And number two is ultimately do you really care about what your model is running on as long as it's outputting whatever you want it to output?

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  18. So these two obey fundamentally different, I would say, tectonic forces. So in training, more is better. You want more of everything, essentially. And the recipe for success is the speed of iteration. You change stuff, you see how it works, and you do it again. Hopefully it converges. And it's like, you know, changing the wheel of a moving car, so to speak. So that is training. On inference, this is a complete reverse. Less is better. You want less headaches. You don't want to be woken up at night because inference is production. You could say that training is research and inference is production. And it's fundamentally different. In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, if you can avoid to have interconnect between, let's say, a cluster of GPUs, of course, you will

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Absolutely, yes. Yeah, yeah, yeah. Here's the problem though. Let's say you are on Google Cloud and you're on TPUs. Suddenly you just remove that 90% chunk on the spend. The problem is that for multiple software reasons which we are solving DML is that they're not really, I would say, a commercial success. They are very much successful inside of Google, but not much outside of Google. Amazon, same is pushing very, very hard for their trainium chips. So the future I see is that you use whatever your provider has because you don't want to pay 90% outrageous margin and try to make a profit out of that.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  20. That takes, let's say, a 30% margin. So you are a very thin crust on a very big cake. It's a bit of a losing game if you go all in on one provider. You want optionality.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  21. So, I would divide it in two categories, well, three categories. The GPUs you can buy or rent. The TPUs you can rent and the TPUs you can buy. This is how the market is structured today, right? Right now, if you want to go dedicated, at least in the cloud, there's two options, TPUs and Traniums on Google, Trainium on Amazon. So these are available chips. You can rent them today. If you want to buy GPUs or rent GPUs, their GPUs, we know it all the time. And there's this new wave of computing, which are dedicated chips you can actually buy. The 10 storant, the etched, the Visora. So I think it will be a mix of, you know, for instance, let's say you are in Google Cloud. Of course, you don't want to do NVIDIA. You get ripped off. Here's the dirty secret is that NVIDIA, like a TSMC sells you at 60% margin, NVIDIA sells you at 90% margin. And on top of that, there's...

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  22. The matrix is read like hundreds of times during the multiplication. So there's a lot of transfers going on. And so far, Milanux with Infiniband had the best technology. So that's why a lot of people... When you do inference, not so much. You don't care when you do inference.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  23. To crack because they're so big and there's a lot of memory transfers and so on. Actually, that's why Grock achieves not Grok, but Lagrok, cerebral, and all these folks, they achieve very high performance single stream. It's because the data is right in the chip. I don't have to get it from memory, which is slow, which GPU has to do. So there's a lot of these things that ultimately make it a good trick, but not, I would say, dedicated solution per se. That said though, the reason probably NVIDIA won, at least in the training space, is because of Melanox, right? Not because of the raw compute, because you need to run lots of these GPUs in parallel. So the interconnect between them is ultimately what matters, right? So how fast can they exchange data? Because remember when you do a matrix multiplication, let's say, you read

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  24. So, the way it worked is that you can think of a screen as a matrix. And if you have to render pixels on a screen, there's a lot of pixels and everything has to happen in parallel, right? So that you don't waste time. Turns out, you know, matrices are very important thing in AI. So there was this cool trick in which we essentially tricked the GPU into back that was like probably 20 years ago. We would trick the GPU into believing it was doing graphics rendering where actually we would making it do parallel work, right? It was called GPGPU at the time, right? So it was always a cool trick, but it was not dedicated for this. The pioneers probably were, of course, Google with TPU, which are very much more advanced on the architectural level, but essentially the way they work, it kind of works for AI, but for LLMs, that starts.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Hard to say because Nvidia has a very, very vertical approach. They do more of more, right? Like if you look at Blackwell, it's actually crazy what they did for Blackwell, they assembled two chips, but the surface was so big that the chip started to bend a bit, which further perpetuated the problem because it then didn't make contact with the heat sink and so on. So they are very much, and the power envelope, they push it to a thousand watts. It requires liquid cooling and so on. So they are very much in a very vertical foot to the pedal in terms of GPU scaling. But the thing is, GPUs are a good trick for AI, but they're not built for AI. It's not a specialized chip. It is a specialization of a GPU, but it is not an AI chip.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Because you don't want to wait 50 seconds for whatever thinking, right? And agents, it's the same. So these two, I think, are the shot that might make NVIDIA change its course with respect to chips. I mean, they're not idiots, right?

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  27. I think this is where NVIDIA can be attacked. I mean, why agents and why reasoning? The difference is for agents and reasoning, you need to wait until the end of the request to get whatever it is you came for. You don't really care about the speed at which the text outputs, which is what you want in a chat, right? You only care about how much time does it take between the beginning of my request and the end. And so that fundamentally changes the incentives from throughput bound to latency bound. And so GPUs, let's say you're running a GPUs at, let's say, 10,000 tokens per second. You very much like to do it 100 times 100, right? And they can do that, but they cannot do, they cannot give you 10,000 tokens per second only on you per stream, what we say. But in terms of, you know, agents are reasoning, this is exactly what you want.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Not much ultimately. The two things that could really very much shake the industry, the chip industry, in my opinion, is our agents and reasoning

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Legitimate because it was built on the A100, I would say, financial model, which was at Generation Zero. We do training. But when it's last generation, we do inference. And it worked beautifully, right? For A100. Then H100 comes along and inference is it's worth five times the price and it maybe runs twice in terms of performance on inference that is. On training, it's a lot better. But on inference, it's like maybe twice as fast. Actually, when it came out, it ran at the same speed than the A100. So there's a money gap that's going to have to be bridged sometime, right? And the part that worries me is that I see amortization plans in like, you know, six, seven years, right? With the GPUs at the collateral. And I'm like, well, I'm not sure how it's going to work because at least when they came out, they were five times the price and they're just two times faster.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Because the chips out there. There's a lot of things. But in my opinion, there's going to be a need for inference. Very hard to say whether it will be worth everybody's money to do it on H100. That is a bubble that I think will blow some time. I'm kind of afraid of that, to be honest.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Want to resell and people just use NVIDIA because it's there, right? But it's by far not the most efficient platform. And arguably, even in terms of software, it's not the best software platform. So that is probably true of the most, I would wager the most important reasons.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  32. So, of course, it runs on AMD, it runs on even Apple and so on. But there was always the tens of little details that not exactly run like you would expect and there's work involved, but then also there's supply. So probably this is the number one thing. The second thing is there's a lot of GPUs on the market pretty much all of them are NVIDIA. The reason being that if you think, you know, in layers and you say, all right, I'm going to buy, let's say, GPUs and I'm going to sell them to folks to maybe not even do training, right? Just do inference. Then most likely, if you look at it that way, you'll end up buying NVIDIA because everybody will want to run an NVIDIA because nobody knows really how to do whatever. And they've trained on NVIDIA. So they're like, I can just reuse my code and so on. So there's like this self-perpetuating circle of people just buy NVIDIA because they

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  33. Oh, yes, absolutely. Yeah, yeah. PyTorch is the ML framework that people use to build, actually train models, right? You can do inference with it, but by far the most successful framework for training is PyTorch. And PyTorch was very much built on top of CUDA, which is NVIDIA software, right? Let's just say the strings of PyTorch make it

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So there's a few reasons. Probably the most important one is the PyTorch CUDA, I would say duo, and that's very, very hard to break. These two are very much intertwined.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  35. No, absolutely. You can get order, like probably an order of magnitude, more efficiency depending on the hardware you run on. That is substantial. Not a lot of people have that problem at the moment. Things are getting built as we speak. But a simple example is if you switch from NVIDIA to AMD on a 70B model, you can get four times better efficiency. In terms of spend, right? So that is substantial. That is very much substantial. Now the problem is getting some AMD GPUs, right?

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Yes, you actually can see it. It's been happening for a while. Models now are not the right abstractions, at least if you look at closed source models. They're not really models. They're more like backend. And there are a lot of tricks that you feel like you're talking to one model, but ultimately you're talking to a constellation, an assembly of backends that produces a response. Probably the number one, you know, I would say obvious thing would be that if you ask a model to generate an image, then it will, you know, switch to a diffusion model, right? Not the LLM. And there's many, many more tricks, the Turbo models and OpenAI do that. There's a lot of tricks. So definitely models in the sense of getting weights and running them is something that is ultimately going away because, you know, in favor of full-blown backends, right? You feel like you're talking to a model, but ultimately you're talking to an API. The thing is, that API will...

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  37. So at the very bottom of things, ZML is an MML framework that runs any models on any hardware. We sit ultimately at the infrastructure layer. We enable anybody to run their model better, faster, more reliably. But on any compute whatsoever. Doesn't really matter. It could be NVIDIA. It could be AMD, could be TPU, and whatnot. And we do all that without compromise. That's the key point. Because if there's a compromise, then it's not really agnostic, right?

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source

  38. The thing with NVIDIA is that they spend a lot of energy making you care about stuff you shouldn't care about. And they were very successful. Like, who gives a shit about CUDA? OpenAI is amazing, but it's not their compute. Ultimately, if you don't own your compute, you're starting with something at your ankle. In five years, I would say 95% inference, 5% training. You have the products, the data, and the compute. Who has all three? Google has like Android, Google Docs. They have everything. They can sprinkle everywhere. This is the sleeping giant in my mind.

    2025-02-24 · The Twenty Minute VC · 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML · IDENTIFIED FROM THE TRANSCRIPT · source