YouSaid · the spoken record

Gavin Uberti

lines on the record
84
first
2023-12-12
most recent
2023-12-12
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. You'd have to go to the room, turn the machine on, wait for the vacuum to be drawn, turn the diffusion pump on, wait for that to be pumped out. It had to be done totally manually because the automated box had been broken. Gunnar would go sit there and supervise us while we messed around with that thing until like nine PM at night. Almost every night. And that got me interested in engineering, got me to Harvard in some ways, and got me to where I am right now.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  2. My school needed an electron microscope, and we didn't have the money to buy one, but he saw a $1,000 broken Jewel 87 machine that was a user for parks being sold in Everett. He, some of my friends and myself, drove over in a U ha haul, picked up the machine, weighed several thousand pounds, drove it back to the school, and spent the next year and a half fixing it. We got it working, actually. It was a very old microscope. It had a polarite camera, if you can believe that. It would flash pixels one at a time as it scanned across the image, because each pixel would be exposed, they'd slowly be turned into the full image, and we ripped that out and placed it with an Arduino. We tried to go and rip out the high voltage box with a newer one because it kind of broke sometimes. We got that working. We never figured out how to keep it from leaking, but that's what we had. The whole microscope would take about 30 minutes to get run.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Kindest thing anyone's ever done for me this guy Gunar mine when I was a kid in middle and high school that was very interested in engineering but young it's hard to go turn that into an impactful results Gunar dropped out of college in Germany went to work for Microsoft for a couple decades and then decided that instead of working for Microsoft he wanted to go teach high school which is quite a pivot he went and taught programming classes at my high school which I token they were good but really he stayed very very late to run the robotics program to run Zanier stuff Gunnar thought that a bunch of high schoolers could build a nuclear fuser to go turn a deuterium into a helium real nuclear fusion it wouldn't generate power these things haven't built before but the thought that much the high schoolers could do it was a little bit crazy that

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  4. That's just too big to run on a device. And if you're going to run in a data center, it might as well use a smarter model. And you might say, well, hundred billion is going to be cheaper, but I think there are enormous economies of scale here because the cost of loading into those weights is so much higher than the cost of doing marginal extra computation. It is really easy to kind of get added to that batch. So I think we're going to see a promising future for Edge and a promising future for the biggest of the big and these hyperscale data centers and nothing in between.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Is there a market for smaller models? I think absolutely. I don't think there's a market for medium-sized models, though. I think there's a lot of domain for a thing that's super, super low latency, going to the server and then going back is so slow that it causes problems, much like, for example, speech recognition today. It used to be the case that when you said hello Google and then asked it a question, it would send the audio to the data center, it would transcribe it, and it would come back, and the latency there was annoying. So Google built AS, CB recognition model on the device for better latency. Now, that model is very small. And I do think we're going to see other very small transformers for tasks, like selecting the next autocomplete word, doing a little bit of drafting to help with your texts. These things are not going to be the crazy smart agents we're going to have in data centers. However, I don't think there's a future for medium-sized models. Things in the hundred billiamp parameter range.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I believe this is going to be the future there's not a whole lot of advantage training your own thing from scratch Remember how we talked earlier about the pre-training phase where you do things that are tendentially related but easy to grade then you do the RLHF on top of that to really make those capabilities come out pre-training phase is way more expensive than the RLHF stuff on top and the pre-training is going to look the same regardless of whether you're an AI for finance an AI for law an AI for therapy There's no sense in a large company doing their own thing from scratch doing the expensive thing again when it's going to look the same you just fine tune

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Think this data is very, very valuable, but it can be learned later. And the reason I'm so confident that it's okay to learn this stuff later, because humans do it all the time. If you go to work for, say, a Bloomberg or company that has valuable financial data, you come in there, having a good understanding of the world around you, but not having the Bloomberg secrets, and you're able to get caught up pretty fast. I think the same thing will happen for these AI models. OpenAI will train GPT X. It'll cost ten billion dollars, and then on top of that, Bloomberg will say, hey, I'm going to fine-tune that. With my Bloomberg proprietary data, I'm going to pay OpenAI 10 million for this, and I'm going to use that fine-tuned version. Other companies say semiconductor company that wants to build a semiconductor-based AI for coding chips, or some investment bank, you name it. They will find tune on top. In the second piece of wise

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Is so incredibly good that it will outweigh the hardware disadvantage it will have biggest transformers are favored. It will outweigh the software advantage that has because transformer libraries are so common. It will outweigh the end user application because end users are used to working with these transformer models. That is just such a heavy lift. And again to go back to this neural networks analogy, random forests or SVNs could have been the dominant AI paradigm, I think. But now to go cash that up to the state of neural networks is totally infeasible.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  9. The world was fair, there could be. But the world is not fair. The transformer, by virtue of being the dominant architecture, is suddenly the thing that has way, way more humanists thinking about trying to improve, trying to optimize and out of people, hardware too. Nvidia is building a little bit of hardware support for the transformers to optimize certain kinds of memory loads into their next GPUs. That is going to preferentially favor the transformer over any replacement that comes out. For in software, NVIDIA has three very highly automatic transformer libraries. That means if you want to run a transformer model, it'll be super easy. If you want to run, say, some other kind of thing, state space diffusion, you have to make the model so much better that it outweighs the factor at two three performance improvement you get by optimizing it so heavily. So for a new architecture to go and beat the transformer.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  10. And when I look out what's happening in this landscape, it is the folks who are building on top of it, who are building the most successful, coolest products, AI games based on having this memory, real conversational things, speech to speech, although with the higher latency than I would like today, or a primitive AI lawyers of the future. So again, the transformer is not just the next fad. This is a building block.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  11. Take it back to this idea of transformer is becoming the way neural networks are, where it's not going to be replaced. It's going to be a building block. Going to build, say, retrieval augmented generation on top of that to make the model smarter. She read Microsoft Paper where they have agents living in a simulation. You're going to build memories on top of these transformers to let them remember things. So a real key storage cache. I think that folks have talked and talked and talked about how the transformers going away. A much more productive line of reasoning as how do I build on top of this.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  12. I think there's a lot of value in a leader being technical, understanding your team's problems, understanding why they're doing the things that they're doing is a requirement for bringing in the right people. There aren't that many levers you can push at the top. It's very coarse things. To figure out the right guy to bring in, you have to understand that problem. And that's what being technical gets you. This is why Gray Brockman OpenAI despite being the CTO, you still see that guy coding. He can't understand how far along the model is unless he himself has seen the internals. Even though I love the folks who work here and I trust them completely, I can't understand how far along design of eradication, how far along is this block honored chip unless I'm talking to that guy, and I understand why he's doing what he's doing. And also, to be quite frank, it's fun.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Google, Microsoft, Amazon, you name it. What does a good leader do? Kind of tie this all back? It looks from afar like Saint Alman himself as pulling all the strings of open eye and mathemating the whole thing. But I think it's the folks under him who are doing the magic and the folk under them who are making that into a real product. The role of a good leader is to set the vision, get the right people in those chairs, and not running out of money. That's it.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  14. You guys are college students, he says. This is semiconductors. Just to get started, you're going to need two to three million dollars and a hell of a lot more after that. This isn't dating, perhaps. You guys have never built the chip before. And then we raised five and a half million. We went back to Mark and asked if he knew anybody to be the chip architect. And after a serious conversation, he used to be a chip architect, and we talked him out of retirement. So Mark Ross is now our chief hardware architect. And from there, Canada, the rest of the dominoes fell. We got Ajat, our VP of engineering, fantastic chip guy, Intel VP before this, who had been overbuilding these very large chips on an advanced nose with teams, just like a dozen people. That is not how that marmali happens. One of the few good men there who can still build the chip on a budget, though we got a Zapadik to go founder Oradine, we got Reynold the first ship guy at Cruise, and dozens other vets from

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  15. So we go to Mark. We show him our tech. And he says, well, Looks like this isn't going to work, but it's too early to tell. You guys should go build a functional simulation. You guys should go write a white paper, and then he come back to me, and we can talk about, hey, here's the problem that you're going to hit. Here's why this isn't going to work. And you guys are going to learn a lot by doing it. And then maybe from that we can figure out how to pivot and build the chip that is viable. Over here, a lot of long nights. Chris and I do just that. We write our technical white paper. We do the original architecture for the chip. We build our functional simulation and we go back to Mark. And Mark says shit, this works.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Let me tell you a bit of the story of how we got here, and then I'll bring into what I think a good leader has to do. When I was on my gap year working for Octo ML, living as a digital nomad versus often as dorm room, we had this idea for, hey, let's build a transformer specialized ASIC. I bet we get a huge amount of performance improvement. We went to a big industry vet, who told us that nope, not gonna work. But also, he said that, despite being a veteran, the industry vet, he wasn't the guy to know for sure, but he knew just the guy to tell us why our ship would not work. Mark Ross, his resident dower. Kind of person who makes a little bit scared. He was a CTO of Cypress Semiconductor for a time, which the soldiers were nine billion dollars twenty nineteen, and Marquez shipped, I think five chips that have done more than a billion in revenue.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  17. We will not fall into this trap and that we will be better. I already know the metrics on our chip. I already have PPA figures. We will just have so much more rock compute, substantially more than an order of magnitude. That's not just for compute, but also for latency. People talk a big game about, hey, if the constant compute goes down, you can do so much more stuff. And I do think that's true. But it's not what gets me excited. With chips that have 20 times the better latencies than the state of the art GPUs because they're able to run a huge prompt in total parallel, that lets you build products you could not otherwise build. Lets you build this real speech-to-speech machine. Lets you build a machine that ingests the whole internet in a seconds instead of hours. That is what gets me excited.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  18. There's got to be more than 10 times. I don't want to share exact numbers right now. We got to have some secret sauce. But because we're able to specialize, you can get a huge amount more compute on these chips. And let me just give a concrete example here. It takes about ten thousand transistors to build a feasible add unit. That's the building block of any measure multiplier. And Nvidia has not that many on their chip, only about four percent of their chip is the matrix multipliers because to feed them requires so much more circuitry. That's what the cost of being so flexible. Or as us, we will have more than an order of magnitude, more of these multipliers, an order of magnitude more raw flops, and of course turning that into a useful throughput, that is still not trivial. You need a very good memory bandwidth there as well. But the reason I can confidently

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  19. They weren't better. If you look at, say, the amount of memory bandwidth that they have versus the GPUs they competed against, it's the same if not worse. If you look at the amount of flops compute they have. Again, it's the same, if not worse And that, I think, is a much bigger problem than just the software itself.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Even say AMD. It's no contest. Uda is a great because it is production great. And the only way to solve this problem besides being old, I think, is to be less ambitious. If you're doing a simpler thing, you can get to real production great quality much faster, and luckily for us and for specialized ships, because there is so much less programmability on the thing, because you're only taking transformers to transform or compile to code, the software stack is much, much easier. Now, that's not to say I want to be complacent about it. We already have a prototype of this running internally, running with our chip and simulation, and because of this, we can begin iterating on that software, getting it stable, reliable production grade, so that when the chips come back, the software works the way that you expect it to. The second piece of why these companies didn't succeed is that

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  21. To the question of why did other AI companies fail? People have tried to build general purpose AI chips in the past and have generally lost out to GPUs. GraphCore, Sabanova, Cerebrus, and Grok have had some modest success, have not taken over the world the way that folks thought that they would. And I think this is two reasons. First, the software, and second, the chips generally weren't better. We'll do software first. It is extremely hard to do what NVIDIA does. CUDA has been developed over, I think it came out in two thousand seven, more than fifteen years, and because it is so old, it has become stable, reliable, the workers of people can depend upon. A startup trying to get into the space won't have that.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  22. Value that more than $50 billion. So because when a special allotted ship move into a space, general purpose devices become no longer competitive, the person who gets there first is almost guaranteed to win. I think that to go back to us, this means we have to move as fast as possible. Back when we began this company last year, four transformers were big before this was obvious. Then yes, we could afford to be a little more relaxed. But today, when people are saying why is there no transformer ASIC, we must move as fast as possible, ship on time, and ship silicon that works for us to win.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  23. But specialized ships like ours are much easier to program. In a world where you burn the transformer architecture into the silicon, you don't need CUDA in the same way. If the world goes to specialized ships, an avidia kind of loses that moat. So while I do believe they will compete, they're not going to be first. The only people who can really innovate are startups like ourselves. So I see a world where we get to market first, and NVIDIA gets to market a year later or so. Now, what happens in that year? I think again going to the crypto analogy makes a lot of sense. A lot of companies, thirty, I think, tried to build Bitcoin mining ASICs. They're much simpler than GPUs. You can tape them out much faster. But the companies that got to the space first ended up winning. There it's now this duopoly between Micro BT and Bitmain. Bitmain, when they're testing the waters for IPO a couple years back

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Be clear, NVIDIA is a fantastic company. And I believe they will eventually build specialized transformer chips much late we're doing. It just makes sense. But right now they're stuck in the innovator's dilemma. The H one hundred is selling like hot cakes and has an enormous margin on it. And if they were to build, not even a transport specific ASIC, but say eight hundred for inference, or eight hundred with a little bit of the graphic stuff ripped out, that would reduce the margin they have on the H one hundred and would be bad for them entirely. They can't innovate or the vantage goes away. And this is especially true for specialized chips. At one of the biggest motes for NVIDIA is CUDA, their software stack that is very, very flexible. It's able to run graphics able to run AI, crypto. That is a great moat in the world where general purpose ships, general purpose AI ships are dominant.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  25. And now, what I did not mention here is the end application. What does the user see? I don't think this will be built by OpenAI and bianthropic and by these model companies. It'll be folks building on top here, this fifth level of a stack. And some of those folks will be incumbents, big established firms that want to build some new AI product. A lot of it will be startups, who say, hey, chatbots are boring, I have this new interface, speech to speech. I have this new interface. Enormous amounts of document ingestion and querying. I have this new interface. AI powered search of the whole internet. So I think there's the most opportunity. I think there's opportunity at all these levels, building a high products, innovating AI models, building new kinds of data centers that are designed to operate at these crazy high scales, building specialized chips and building them fast for running these models.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Maybe microns. Two of those are in South Korea and Microns Americans. So these people are other, really, really critical folks in that supply chain. And folks talk all the time about NVIDIA, TSMC, ASML relationship. They don't talk enough about NVIDIA to Samsung, NVIDIA to Eskinish relationship, because these memory modules are so critical for building the state-of-the-art GPUs. But then these companies themselves, where do they get their machines from? And now we get to a little bit of a diminishing returns. ASML is a much smaller company than TSMC, and there are so many others that help make these chips as well. But I think that this supply chain of model provider on top, then hyperscale data center below that, then chip designer below that, and chip fab below this. This is the fundamental stack of the future.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Is almost entirely NVIDIA. But I think that's going to change. AMD is coming out with a new MA300X next year. I think it's going to be a viable competitor to the H100. And companies like us are trying to replace this idea of general purpose AI chips entirely and build specialized ones just for the transformers. So I think that us, along with some of the general purpose guys, become this layer underneath it. But us, NVIDIA and AMD cannot make the tips ourselves. We have to go to a fat house. TSMC is the one of choice, but the chip is just one piece here. On a chip like the MF or on the H one hundred, the memory actually cost more than the silicon itself by more than a factor or two. Now, these memory modules are made by only two or three companies as well, Eskienix, and

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  28. I know that it looks from afar like Microsoft and Google and Amazon are dominating everything. And they do dominate, but at one layer of this, enormously complex supply chain. Let's think about it from the top. At the very top here we have the state of the AI models being trained by folks like OpenAI and Anthropic. These folks have enormously good talent. I think that's the big moat, but they don't really run the compute themselves. They get it from one of the hyperscalers. So we have the models at the top. Below this, we have the hyperscale tech giants. Those are the ones people are scared of. Google, Microsoft, and Amazon are the big three in this domain. Buying those chips, racking and stacking them, supplying data center with power, and then selling that to these companies on top. Where do those companies get to their chips from? Right now,

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Well, one of the beauties of language models, I think, is that they learn these concepts themselves. You can test this. You can go ask GPT-4, here's a scenario, is the human behaving ethically in this scenario. The AI knows what is good, what is ethical, and it got the understanding by reading all the text we've generated as a species. So when we ask it to go say, hey, pick the response that is more ethical, pick the response that's better. It is using that concept that developed from reading all there is to read.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  30. Generating your own training data is the final frontier here. One good analogy is the machines that played Go. They try training AIs on human games to make them better than humans at the game of Go. And it worked, but you ran out of those games very fast. So what Google did with the Alpha Go is they made the model play against itself and learn from those games. And the same trick has begun to show some promise in natural language processing where you have an AI that writes out some output, then you ask it to analyze that output and say, was this good or was this bad? As a human, you can do that. You can get better. That's why I revise the things that I write. And I think that technique of the models generating their own training data will allow us to scale even once we run out of all the images and video available on the web.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  31. I learned to pick up objects. I didn't do it from a book. I did it by looking with my own eyes and seeing, hey, does this behave the way that I think it does? And we've begun to kind of dip our toes to it. People have built vision transformer models. And usually what they do is they train a model on text. And then at the very end, they add in the image part. It just adds a little side feature. And I think that's wrong. I think transformers tomorrow will be trained from day one with enormous amounts of video data. So that solves the problem in the medium term. But eventually, these models will get hundred times bigger, but require a hundred times more data.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Trees to fix us for it. And that didn't really work, didn't scale up. But the thing that did scale up is making neural networks very big and feeding them huge amounts of data. So the lesson of the bitter lesson is that the only techniques that will really advance the field of artificial intelligence are those that are able to leverage the cost of compute getting cheaper, that are able to leverage these enormous data centers that must be built. And you mentioned data as well. A key part of these learning algorithms is data. There's only so much data available in the world. Some people have raised the alarm that we're about to run out. The models aren't going to be able to go and get the next trillion tokens. It doesn't exist. And yes, I do think that'll be a problem eventually. But we haven't even begun to tap into the most valuable source of data video. I don't know about you, Patrick, but when I learned to walk.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  33. A bishop is worth three. I'm going to add a rule to my machine that says, hey, if you can go trade a bishop for a rook, that's a good trade to make. So you should go make that move. And again, that works a little bit, but it doesn't really generalize. So the strategy that eventually worked for chess that enabled the blue to become the World Chess Champion was not encoding human knowledge, not hard coding rules, but by just leveraging a massive amount of computation. Deep blue did search, much like this tree search we were talking about earlier. It would test hundreds of thousands of positions. I'm the one that it liked best and then plain news to get there. And the same thing is happening for artificial intelligence. In the very old days, the image recognition, people were saying hey, there's a block of white pixels and this kind of shape it might be a bird. It might be a light bowl. They had these big.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  34. So, the bitter lesson is an essay put out by Rich Sutton, great pioneer AI in 2019, before the whole transformer craze. Rich looked out at the field and saw that what would happen all the time if somebody would take a model that worked, and they'd hard code some knowledge and do it. For example, somebody would say, hey, here's a great image recognition model, but it gets this one weird case wrong on occasion. I'm going to hard code a hacky fix on it. And inevitably, doing this whackamal fix strategy works in the short term. You get a little bit better performance. You can go publish your paper all is well, but it never scales. The only techniques that have routinely worked are those that leverage the fact that compute is getting cheaper. For example, I mentioned chess engines before. In the very old days, people would write chess engines by saying, well, a work is worth five points.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Too. But we cannot be a passive about this. We're in a good state now, but that could change. So I think that export bans for state of the art AI chips are really in the national interest here.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  36. If I said that we were going to have threats of war over a ship fab 10 years ago, that sounded like a sci-fi novel too. This has to be the conclusion here. Some countries are going to have these hyperscale data centers. And if you don't have one, there's going to be really hard to catch up. China is famously trying to catch up on semiconductor development. They realize how critical that is for AI, but they did not get on the EUV train when it was leaving the station. They're now on this older immersion technology, and it's really holding them back. And while they have produced some great quote unquote seven nanometer chips, I do think they'll continue getting better technology. China is not state of the art for semiconductors today and does not seem like it's not a path to be either. And I think for data centers as well, the United States is very lucky to have all the big hyperscalers, very lucky to have the big AI companies.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  37. I think we'll see a similar thing for these AIs. You'll have a few smaller data centers. Don't try to compete on being as purely smart as the big megaliths. But they'll find some niche.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  38. These companies build. But they'll be much more like TSMC, the peer play AI generating house, and they will be some full stack integrator. I just think the world cannot support a hundred different ten billion dollar models, but it can support three. So I see a world where you have these few enormous leading edge state of the art development houses, possibly at the current hyperscalers, possibly at some new company that specializes on just this kind of thing. And there'll be a tale, much like there are for semiconductors. Not every application needs bleeding edge chips. For example, the US government in US military care much more about chips being produced in a reliable American factory than the chips being state of the art. So if fabs like global foundries have accepted, they're not going to be as good as TSMC. But they're still a market for them to persist by being American.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Again, there's a great analogy in the world of chipmaking. For the world of chipmaking used to be pretty cheap, anyone designing their chips probably had their own little chip fab because it cost a couple million dollars. And as these things became more and more expensive, all but the biggest players became priced out. Now, only two or three firms, Samsung, and maybe Intel, are able to make the world state of the art ships. I think the same thing will happen for a generative AI and these megascale data centers. Today, many companies are trying to explore probing the waters because it is so cheap to do so. But as having a state of the AI model becomes a hundred million billion ten billion dollar endeavor, it'll be much more common to outsource that to one of the few firms who will have one of these megascale data centers. And don't get me wrong, you'll still be able to fine-tune on top of these.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Almost got to do it. That's the only way that we get to our $10 billion training runs. That's the only way that we get to GPT-6. That's the only way that we get these AI generated books and movies and TV shows that would make the artist of today be impressed. Someone has to do it, and might as well be me.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  41. People don't like rocking the boat because of how expensive it is to build a ship, only for it to come back dead. And once a few people do it, it'll be the obvious solution, but nobody wants to be first.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Linear in the number of transistors going from 100,000 to 88 billion, like you have on the NVIDIA GPUs, is more than 80 billion over 100,000 times harder. So instead of being able to do it on one beefy server, you suddenly a huge, huge data center working on this optimization problem to get it working. So I am optimistic there's a future where, just like for a hobbyist building a chip on an old node today, we'll be able to get companies building chips on new nodes by spending up huge data centers to do this back end process and training these language models to help them right the front end. I think the upper limit really is a couple of months instead of the four or five years of today. A lot of work's going to have to happen to bring that down. It's also a problem where the first ships that are made with some faster process probably aren't going to work.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  43. I really want to go call out a lot of the Google Silicon ecosystem. I want to call out Oba Foundries, the Skywater Foundry for open sourcing their PDKs to help make this process fast. Now, if you're a hobbyist who wants to build a chip on one of these older technologies, you can do it from Conception to putting it together a few transistors, taping it out weeks, months, even if you're starting from absolute scratch. That is so cool. They're able to do for pretty cheap because they take all the hobbyist projects, put them together on one mask set. You make a one photo mask that has all the chips on it. And I'm really optimistic that this can be scaled up. Now, it's definitely not easy. Hundreds of thousands of transistors on them. And the amount of effort it takes to computer to decide where to put these things is super.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Into your model as training data because it's all proprietary. And I'm very optimistic that somebody will be able to go to synopsis and go to all the old chip vendors and say, hey, your code for CPUs from like two thousand five. That's probably not very useful to you anymore. Can I use that to train my AI model? And if enough companies are even used to say yes, ideally to open source it, but probably not. Maybe just to go train that model, then you can get a language model that was good at Verilog, as good at Verilog as it is at say C and Python, and use that to speed up writing some of this artillery development.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Instead of nine months by using one of these automated tools. And we haven't gotten to this point with enormous modern skilled chips yet. It's possible in theory, and I'm very hopeful that more smart folks will begin working on this problem. And the same thing for actually writing the chips, doing this RTL coding. Today it's all done by hand. But for example, C and Python, we have seen language models themselves begin to help out with it. GPT four is very good at both those languages, as is GitUp Copilot, because there's so much code in C and Python available online. So having access to these AI tools can make a developer much faster for a Verilog in these hardware languages, that's not true yet. While software loves open source, hardware really doesn't.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Is said and done. You have parts of the wafer that are hardened in just the right pattern for you to put your transistors down on them. So this is a very long process. Real production samples coming off an assembly line for modern semiconductors. And if we want to build the next generation of transformer specific ASICs for training and for inference, this has to shrink. Now there's been a huge amount of progress to make this faster, especially on older nodes. A Google in particular has driven a lot of this. They have a series of open source efforts to automate parts of this back end process because really, it's not that hard. It's just checking rules, and it should be done by a computer. And if you're building on a node like Ostrom, if you're building for old technology, old semiconductor detecting like two thousand five, you can have the back end done automatically in a day

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  47. They check a huge number of problems, and eventually after this back end process is all finished, they'll give it to a masked house, mask shop, who will take these designs of where your transistor goes, and they'll say all right, to make the transistors, we need to go and have silicon in these parts of the chip, but not these parts of the chip. If you ever heard a screen printing, the process kind of works like that. From this file, this GDS two file, which says where there should be and where there shouldn't be a transistor, they make a photo mask that will shine light on the part where our transistor and not on the part where there aren't. And then those photo masks are sent to fab that actually makes the chip, who then uses them to project an image onto those silicon wafers after coating them with a chemical called photoresist that hardens when exposed to light. So that means that

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Other. But now you have to put on an actual silicon wafer and lay them out. That's a hard problem, so you'll give it to a back-end engineer who will make sure that while both beyond the logic being correct, there's no issue where hey, this block takes too long and is not going to be done by the time this other block has to take its results. They call that time enclosure, and they'll make sure that on a chip, especially a large chip, process issues, manufacturing problems in the TSMC line, can make it so that one part of the chip behaves differently than another part of the chip, even things moving at different speeds. So they'll make sure that it works in a matter what kind of wafer is made, whether it's fast or slow, what temperature it is. They call this getting all the corners closed, and a number of other checks too. There's not going to be some weird radio frequency interference thing, and there's not going to be some pool, some vacuum in the ship or other photoresistor will flow in.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  49. To block how to behave. Then this block is verified. Because it costs so much to manufacture semiconductors, you have to make sure there is not a mistake in the RTL block, for you go pull the trigger and tell TSMC or whoever it is to begin cranking out wafers off their line. So a verification engineer will very carefully write a test bench to check that that block behaves the way that it should. And they'll also check that their test bench actually runs every single line in that ferlog file. They call that line coverage. And they'll go a step further. That ferlog, that code, will be compiled to a net list. This is how each group of transistors talks to each other, and then you'll run toggle coverage on this. You make sure every one of those transistors will flip to either state. That's what they call the front end of the chip. Now you have a netlist that comes out of that, a set of all the transistors on how they connect to each other.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source

  50. Well, there are a few steps involved in bringing a new chip to market. No matter what year you're building in, first have to say, okay, what's the problem we're trying to solve? And then from there, you'll devise an architecture and a microarchitecture to sort of a more detailed variant of the architecture of how your chip works at a high level to solve this problem. And as part of this architecture work, you're going to get a number of blocks. You'll say, okay, in order for me to do this giant matrix multiplication, I'll need a matrix multiplication block. And that matrix multiplication block that we made above many sub-blocks that each do a multiply add operation. It'll be made up of a control block that sends a signal to the full matrix block, and so on. Then you'll take these blocks, number of blocks which it has varies widely, but it can be as high as hundreds, and you'll give it to an RTL engineer who will write code in a hardware description language kind of programming language.

    2023-12-12 · Invest Like the Best · Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356] · IDENTIFIED FROM THE TRANSCRIPT · source