YouSaid · the spoken record

Chris Lattner

lines on the record
301
first
2023-06-02
most recent
2023-06-02
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. And so C is way more complicated because of C in the legacy than it would have been if they would have theoretically designed a from scratch thing. There's lots of people right now that are trying to make C better in recent text C. It's going to be great. We'll just change all the syntax. But if you do that, now suddenly you have zero packages. Don't have compatib

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Well, it paints you into corners. So again, I'm very happy with Python. So joking, all joking aside, I think that the indentation thing is not the actual important part of the problem. Yes. Right, but the fact that Python has amazing dynamic metaprogramming features and they translate to Beautiful static metaprogramming features, I think, is profound. I think that's huge, right? And so Python, I've talked with Guido about this. It was not designed to do what we're doing. That was not the reason they built it this way, but because they really cared and they were very thoughtful about how they designed the language, it scales very elegantly in the space. But if you look at other languages, for example, C and C++, if you're building a superset, you get stuck with the design decisions of the subset.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Being a superset of Python. And that's a really hard technical problem, but it's, in my opinion, worth it, right? And it's worth it because it's not about any one package. It's about this ecosystem. It's about what Python means for the world. And it also means we don't want to repeat that Python 2 to Python 3 transition. Like we want people to be able to adopt this stuff quickly. And so by doing that work, we can help lift people.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  4. And integers in Python can be an arbitrary size integer. If you care about fitting in a going fast in a register, in a computer, that's really annoying. And so you can choose two paths on that, right? You can say, well, people don't really use big integers that often, therefore, I'm going to just not do it and it will be fine. Not a Python superset. You can do the hard thing and say, okay, this is Python. You can't be a superset of Python without.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  5. And then exactly. Once you're in the New World, then you can build all kinds of cool tools to say, hey, should you adopt this feature? And we haven't built those tools yet, but I fully expect those tools will exist. And you can quote unquote modernize your code or however you want to look at it. So, I mean, one of the things that I think is really interesting about Mojo is that there have been a lot of projects to improve Python over the years. Everything from getting Python to run on the Java virtual machine, PyPy, which is a JIT compiler, there's tons of these projects out there that have been working on improving Python in various ways. They fall into one of two camps. So PyPy is a great example of a camp that is trying to be compatible with Python. Even there, not really, it doesn't work with all the C packages and stuff like that, but they're trying to be compatible with Python. There's also another category of these things where they're saying, well, Python is too complicated. And I'm going to cheat on the edge.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  6. Yeah, right. And you can move C code to C. Then you can adopt classes. You can adopt templates. You can adopt other references or whatever C++ features you want. After you move C code to C, you can't use templates in C. And so if you leave it a C, fine, you can't use the cool features, but it still works. And C and C code work together. And so that's the analogy, right? Now, here, Python is bad and a Mojo is good, right? Mojo just gives you superpowers. And so if you want to stay with Python, that's cool. But the tooling should be actually very beautiful and simple. Because we're doing the hard work of defining a superset.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So, and this is why, among other reasons why we use tabs. So, first of all, by being a superset. It's like C versus C Can you move C code to C?

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Not yet, so we're missing some basic features right now. And so we're continuing to drop out new features on a weekly basis. But at the fullness of time, give us a year and a half maybe two years.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  9. So they're definitely, I'm a huge fan of that work, by the way, and it composes well with what we're doing. And so it's not like we're fighting or anything like that. It's actually just goodness for the world. But it's just a different path, right? And again, we're not working forwards from making Python a little bit better. We're working backwards from what is the limit of physics.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Yeah. Well, I mean, you could argue Mojo is redesigning CPython, but why not make CPython faster and better and other things like that? There's lots of people working on it. So, actually, there's a team at Microsoft that is really improving, I think CPython 3.11 came out in October or something like that, and it was 15% faster, 20% faster across the board, which is pretty huge, given how mature Python is and things like this. That's awesome. I love it. Doesn't run on GPU. It doesn't do AI stuff, like it doesn't do vectors, doesn't do things. I'm 20% is good, 35,000 times better.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  11. And so by pulling this all into Mojo, what you get is you get one world. You get the ability to say, cool, I have untyped very dynamic, beautiful, simple code. Okay, I care about performance for whatever reason, right? There's lots of reasons you might care. And so then you add types, you can parallelize things, you can vectorize things, you can use these techniques, which are general techniques to solve a problem. And then you can do that by staying in the system. And if you have that one Python package that's really important to you, you can move it to Mojo. You get massive performance benefits on that. Another advantage is, if you like, stack types, it's nice. It's nice if they're enforced. Some people like that, right? Rather than being hints. So there's other advantages too. And then you can do that incrementally as you go.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  12. But the interpreter and the wrappers and the implementation details and the conventions, and it's just this really complicated mess. And when you do that, now suddenly you have a debugger that debugs Python. They can't step into C code. So you have this two-world problem

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  13. It's being used in memory. You'd have to know a lot about how the CPython interpreter works. It has, for example, reference counting, but also different rules on how to pass pointers around and things like this. Super low-level fiddly, and it's not like Python, it's like how the interpreter works. And so that gets all exposed out, and then you have to define wrappers around the low-level C code. This means you have to know not only C, which is a different world from Python, obviously, not only Python.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Yes, it's complicated. I mean, this is what we do. So, I mean, we make it look easy, but it is complicated. But what we do is we use. The CPython existing interpreter, so it's running its own bytecodes, and that's how it provides full compatibility. It gives us CPython objects. We use those objects as is. And so that way we're fully compatible with all the CPython objects and all it's not just the Python part, it's also the C packages, the C libraries underneath them because they're often hybrid. And so we can fully run and we're fully compatible with all that. And the way we do that is that we have to play by the rules. And so we keep objects in that representation when they're coming from that world.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  15. And so the run as fast as they do in the traditional CPython way. But what that does is that gives you an incremental migration path. And so if you say, hey, cool, well, here's the Python ecosystem is vast. I want all of it to just work. But there's certain things that are really important. And so if I'm doing weather forecasting or something, well, I want to be able to load all the data. I want to be able to work with it. And then I have my own crazy algorithm inside of it. Well, normally I'd write that in C++. If I can write in Mojo and have one system that scales, well, that's way easier to work with.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Exactly, but we're not willing to wait for that. Python is too important, the ecosystem is too broad. We want to both be able to build Mojo out. We also want to do it the right way without intense time pressure. We're obviously moving fast. And so what we do is we say, okay, well, let's make it so you can import an arbitrary existing package. Arbitrary, including like you write your own on your local disk or whatever. It's not like a standard package. And import that using CPython. Because CPython already runs all the packages, right? And so what we do is we built an integration layer where we can actually use CPython. Again, I'm practical to actually just load and use all the existing packages as they are. The downside of that is you don't get the benefits of Mojo for those packages.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Without CPython someday. Not today, but someday. That'll be a beautiful day because then you'll get a whole bunch of advantages and you'll get massive speedups and things like this.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  18. Yep. So in the fullness of time, Mojo will solve for all the problems and you'll be able to move Python packages over and run them in Mojo.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Somebody was just tweeting about that yesterday, which is pretty cool, right? And again, interpreters, compilers, right? And so without changing any code, without also jit compiling or doing anything fancy, this is just basic stuff. Move it straight over. Now, Mojo will continue to grow out, and as it grows out, it will have more and more and more features. And our North Star is to be a full superset of Python. And so you can bring over basically arbitrary Python code and have it just work. And it may not always be 12x faster, but it should be at least as fast and way faster in many cases as the goal. Now I'll take time to do that. And Python is a complicated language. There's not just the obvious things, but there's also non-obvious things that are complicated, like we have to be able to talk to CPython packages that talk to the CAPI. And there's a bunch of pieces to this.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Yes. And this has two sides of it. Mojo's not done yet, so I'll give you a disclaimer. Mojo's not done yet. But already we see people that take Small pieces of Python code, move it over. They don't change it. And you can get 12x speedups.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  21. If you get it wrong, it's not 5% or 1%. It could be 2x or 10x. If you think about it, you really want to make use of the full memory you have, the cache, for example. But if you use too much space, it doesn't fit in the cache. Now you're going to be thrashing all the way back out to main memory. And these can be 2x, 10x major performance differences. And so this is where getting these magic numbers and these things right is really actually quite important.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  22. It's optimized. Yeah. Well, so all of these, the details matter so much to get good performance. This is another funny thing about machine learning and high performance computing that is very different than C compilers we all grew up with, where if you get a new version of GCC or new version of Clang or something like that, maybe something will go 1% faster. And so compiler engineers will work really, really. But when you're talking about an accelerator or an AI application, or you're talking about these kinds of algorithms, these are things people used to write in Fortran, for example, right?

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  23. So, tiling is a memory optimization. It says, okay, let's make sure that we're keeping the data close to the compute part of the problem instead of sending it all back and forth through memory every time I load a block.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  24. So think about this in levels, hierarchical levels of abstraction. If you zoom all the way into a compute problem, you have one floating point number. And so then you say, okay, I can do things one at a time in an interpreter. It's pretty slow. So I can get to doing one at a time in a compiler. I can see. Then I can get to doing four or eight or 16 at a time with vectors. That's called vectorization. Then you can say, hey, I have a whole bunch of different, what a multi-core computer is. It's basically a bunch of computers. So they're all independent computers that can talk to each other and they share memory. And so now what Parallelize does is it says, okay, run multiple instances of this on different computers. And now they can all work together on problem, right? And so what you're doing is you're saying keep going out to the next level out. And as you do that, how do I take advantage of this?

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  25. And that's something that a lot of machine learning infrastructure and tools and technologies don't have. Typical state of the art today is you walk up, particularly if you're deploying. If you walk up with a new model, you try to push it through the converter, and the converter crashes That's crazy. The state of ML tooling today is not anything that a C programmer would ever accept. And it's always been this kind of flaky set of tooling that's never been integrated well and it's been never worked together because it's not designed together. It's built by different teams. It's built by different hardware vendors. It's built by different systems. It's built by different internet companies. They're trying to solve their problems. And so that means that we get this fragmented, terrible mess of complexity.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  26. So to me, I kind of joke, what is a compiler? There's many ways to explain that. You convert thing A into thing B and you convert source code to machine code. You can talk about many things that compilers do. But to me, it's about a bag of tricks. It's about a system and a framework that you can hang complexity. It's a system that can then generalize and it can work on problems that are bigger than fit in one human's head. And so what that means, what a good stack and what the modular stack provides is the ability to walk up to it with a new problem. And it'll generally work quite well.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Modular stack provides it provides this infrastructure and the system for factoring all this complexity and then allowing people to express algorithms. You talk about auto-tuning, for example, express algorithms in a more portable way so that when a new chip comes out, you don't have to rewrite it all.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  28. What's happened with these accelerators now is you get multiple levels of memory, like in a GPU, for example, you'll have global memory and local memory and all these things. If you zoom way into how hardware works, the register file is actually a memory. So the registers are like an L0 cache. And so a lot of taking advantage of the hardware ends up being fully utilizing the full power in all of its capability. And this has a number of problems, right? One of which is, again, the complexity disaster, right? There's too much hardware. Even if you just say, let's look at the chips from one line of vendor like Apple or Intel or whatever it is, each version of the chip comes out with new features and they change things so that it takes more time or less time to do different things. And you can't rewrite all the software whenever a new chip comes in. And so this is where you need a much more scalable approach. And this is what Mojo and what the

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  29. So, again, what happened is we went through a phase of many years where people took the special case and hand tuned it, and tweaked it and tricked it out and they knew exactly how the hardware worked and they knew the model and they made it fast, didn't generalize. And so you can make Resonant 50 or AlexNet or something in Section V1. You can do that because the models are small. They fit in your head. But as the models get bigger, more complicated, as the machines get more complicated, it stops working, right? And so this is where things like kernel fusion come in. So what is kernel fusion? This is this idea of saying let's avoid going to memory and let's do that by building a new hybrid kernel, a numerical algorithm that actually keeps things in the accelerator instead of having to write all the way out to memory.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  30. I mean, maybe I'm a special kind of nerd, but you look at that. What is the limit of physics? How fast can these things go, right? When you start looking at that, typically it ends up being a memory problem. And so today, particularly with these specialized accelerators, the problem is that you can do a lot of math within them, but you get bottleneck sending data back and forth to memory, whether it be local memory or distant memory or disk or whatever it is. And that bottleneck, particularly as the training sizes get large, as you start doing tons of inferences all over the place, that becomes a huge bottleneck for people, right?

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  31. So, I mean, if you just look at the Python problem, right, you can say, how do I make Python faster? And there's been a lot of people that have been working on the Python 2x faster, 10x faster, or something like that, right? And there have been a ton of projects in that vein. Mojo started from the, what can the hardware do? What is the limit of physics? What is the speed of light? How fast can the sun go? And then how do I express that? And so it wasn't anchored relatively on MakePython a little bit faster. It's saying cool, I know what the hardware can do. Let's unlock that, right? Now...

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  32. So you start from simple and predictable models, and so you can have full control, and you can have coarse grain knobs that nudge the system so you don't have to do this. But if you really care about getting the best, you know, the last ounce out of a problem, then you can use additional tools. And there the cool thing is you don't want to do this every time you run a model. You want to figure out the right answer and then cache it. And once you do that, you can say, okay, cool. I can get up and running very quickly. I can get good execution out of my system. I can decide if something's important and if it's important, I can go through a bunch of machines at it and do a big expensive search over the space using whatever technique I feel like it's fairly up to the problem. And then when I get the right answer, cool, I can just start using it. And so you can get out of this trade-off between, okay, am I going to spend forever doing a thing or do I get up and running quickly? And as a quality result, these are actually not in...

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  33. You're going to turn So then you turn into a machine learning problem, and then you have a space of genetic algorithms and reinforcement learning and all these.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  34. If it's a matrix multiplication, do it over there. Or if it's floating point, do it on the GPU if it's integer, do it on the CPU, like something like that, right? And then you then get into this mode of like people care more and more and more. And you say, okay, well, let's actually make the heuristic better. Let's get into auto-tuning. Let's actually do search of the space to decide, well, what is actually better? Well, then you get into this problem where you realize this is not a small space. This is a many-dimensional hyperdimensional space that you cannot exhaustively search. Do you know of any algorithms that are good at searching very complicated spaces?

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Absolutely. So, I mean, in my opinion, this is an opinion. This is not everybody would agree with this, but in my opinion, the world benefits from simple and predictable systems at the bottom that you can control. But then once you have a predictable execution layer, you can build lots of different policies on top of it, right? And so one policy can be that the human programmer says, do that here, do that here, do that here, do that here, and like fully manually controls everything. And the system should just do it. Then you quickly get in the mode of like, I don't want to have to tell it to do it. And so the next logical step that people typically take is they write some terrible heuristic.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  36. And so the simplest thing is just saying do matrix operations over there. But then you realize you get a little bit more complicated because you can do matrix multiplications on a GPU. You can do it on. Neural net accelerator, you can do it on CPU, and they'll have different trade-offs and costs, and it's not just matrix multiplication. And so what you actually look at is you look at, I have generally a graph of compute. I want to do a partitioning. I want to look at the communication, the bisection bandwidth and like the overhead and the sending of all these different things and build a model for this and then decide, okay, it's an optimization problem of where do I want to place this compute?

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Yeah, so there's a pretty well known algorithm, and what you're doing is you're looking at two factors you're looking at the factor of sending data from one thing to another. Because it takes time to get it from that side of the chip to that side of the chip and things like this. And then you're looking at what the time it takes to do an operation on a particular block. So take CPUs. CPUs are fully general. They can do anything. But then you have a neural net accelerator that's really good at matrix multiplications. And so you say, okay, well, if my workload is all matrix multiplications, I start up, I send the data over the neural net thing, it goes and does matrix multiplications. When it's done, it sends me back the result. All is good.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  38. In a data center, you now have multiple different machines, sometimes very specialized, sometimes with GPUs or TPUs in one node, and sometimes with disks in another node. And so you get a much larger scale heterogeneous computer. And so what ends up happening is you have this multi-layer abstraction of hierarchical parallelism, hierarchical asynchronous communication, and making that, again, my enemy is complexity by getting that away from being different specialized systems at every different part of the stack and having more consistency and uniformity. I think we can help lift the world and make it much simpler and actually get used.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Choreographed, right? And so, again, one of the cool things about machine learning is it's moving things to like data flow graphs and higher level of abstractions and tensors and these things that it doesn't specify here's how to do the algorithm. It gives the system a lot more flexibility in terms of how to translate or map it or compile it onto the system that you have. And so what you need, the bottomest part of the layer there is a way for all these devices to talk to each other. And so this is one thing that I'm very passionate about. I mean, you know, I'm a nerd But all these machines and all these systems are effectively parallel computers running at the same time, sending messages to each other. And so they're all fully asynchronous. Well, this is actually a small version of the same problem you have in a data center.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Modern cell phone, you've got CFUs, and they're not just CPUs, there's like big.littl CPUs, and so there's multiple different kinds of CPUs that are kind of working together that are multi-core. You've got GPUs, you've got neural network accelerators, you've got Dedicated hardware blocks for media, so for video decode and JPEG decode and things like this. And so you've got this massively complicated system. And this isn't just cell phones. Every laptop these days is doing the same thing. And all these blocks can run at the same time and need to be

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Yeah, so what is heterogeneous, right? So heterogeneous just means many different kinds of things together. And so the simplest example you might come up with is a CPU and a GPU. And so it's a simple heterogeneous computer to say I will run my data loading and pre-processing and other algorithms on the CPU. And then once I get it into the right shape, I shove it into the GPU. I do a lot of matrix multiplications and convolutions and things like this. And I get it back out and I do some reductions and summaries and I shove it across the wire across the network to another machine. And so you've got now what are effectively two computers a CPU and a GPU talking to each other, working together in a heterogeneous system. But that was 10 years ago. You look at a modern cell phone.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Yeah, so again, go back to the simplest example of int. And so what Bull Swift and Mojo and other things like this did is we said, okay, pull magic out of the compiler and put it in standard library. And so, what modular is doing with the engine that we're providing and like this very deep technology stack, which goes into heterogeneous runtimes and a whole bunch of really cool things, this whole stack And hacked and changed by researchers By hardware innovators and by people who know things that we don't know because modular are some smart people, but we don't have all the smart people, it turns out.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  43. That no analog computers really well, or they know some GPU internal architecture thing really well, or they know some crazy sparse numeric, interesting algorithm that is the cusp of research, but they're not compiler people. And so one of the challenges with this new wave of technology trying to turn everything into a compiler is, again, it's excluded a ton of people. And so you look at what does Mojo do, what does the modular stack do, is it brings programmability back into this world. Like it enables I wouldn't say normal people, but like a different kind of delightful nerd that cares about numerics or cares about hardware or cares about things like this to be able to express that in the stack and extend the stack without having to actually go hack the compiler itself.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  44. And so with the compiler people said, they said, hey, cool. Well, I will go enumerate all the algorithms and I will enumerate all the pairs and I will actually generate a kernel for you. And I think that this has been very, very useful for the industry. This is one of the things that powers Google TPUs, PyTorch 2 is rolling out really cool compiler stuff with Triton, other technology and things like this. And so the compiler people are kind of coming into their fore and saying like, awesome, this is a compiler problem, we'll compiler it. Here's the problem Not everybody's a compiler person. I love compiler people, trust me, right? But not everybody can or should be a compiler person. It turns out that people.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  45. And so within machine learning algorithms, for example, people figured out that, hey, if I do a matrix multiplication, I do a relu, right, the classic activation function, it is way faster to do one pass over the data and then do the relu on the output where I'm writing out the data, because rel is just a maximum operation, right? Max was zero. It's an amazing optimization. Take MapMulu, squish it together in one operation. Now we have MattMull reluct. Well, wait a second. If I do that, now I just went from having two operators to three. But now I figure out, okay, well, there's a lot of activation functions. What about leaky relue? What about a million things that are out there, right? And so as I start fusing these in, now I get permutations of all these algorithms.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  46. New exotic crazy accelerators is people have been trying to turn this from a let's go write lots of special kernels problem into a compiler problem And so we and I contributed to this as well as an industry went into a like let's go make this compiler problem phase let's call it and much of the industry is still in this phase by the way so I wouldn't say this phase is over and so the idea is to say look okay what a compiler does is it provides a much more general extensible hackable interface for dealing with the general case

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  47. What that would mean in terms of the daily impact on the products we use, that would be huge. Now, if you're building an analog computer, you may not be a compiler specialist. These are different skill sets, right? And so you can hire some compiler people if you're running a big company, maybe. But it turns out these are really exotic new generation of compilers. This is a different thing, right? And so if you take a step back out and come back to what is the status quo, status quo is that if you're Intel or you're NVIDIA, you keep up with the industry and you chase and, okay, there's 1,900 now, there's 2,000 now, there's 2100, and you have a huge team of people that are like trying to keep up and tune and optimize. And even when one of the big guys comes out with a new generation of their chip, they have to go back and rewrite all these things. So really it's only powered by having hundreds of people. They're all like, frantically trying to keep up. And what that does is that keeps out the little guys. And sometimes they're not so little guys. The big guys are also just not in those dominant positions. And so what has been happening, and so a lot of you talk about the rise of.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Building a phone, you want something that looks very different. If you have something like a laptop, you want something that looks maybe similar but a different scale. AI ends up touching all of our lives, robotics, and lots of different things. And so as you look into this, these have different power envelopes. There's different trade-offs in terms of the algorithms. There's new innovations in sparsity and other data formats and things like that. And so hardware innovation, I think, is a really good thing. And what I'm interested in is unlocking that innovation. There's also analog and quantum. And although the really weird stuff, right? And so if somebody can come up with a chip that uses analog computing and it's 100x more power efficient.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  49. So, Chris's philosophy on hardware Right, so my philosophy is that there isn't one right solution. Specialization happens. If you're building, if you're training GPT-5, you want some crazy supercomputer data center thingy. If you're making a smart camera that runs on batteries, you want something that looks very different.

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  50. And so, what happened with both TensorFlow and PyTorch is that the explosion of innovation in AI has led to, it's not just about matrix multiplication and convolution. These things have now like 2,000 different operators. On the other hand, you have, I don't know how many pieces of hardware there are there. It's a lot, it's not even hundreds, it's probably thousands across all of Edge and across all the different things that

    2023-06-02 · Lex Fridman Podcast · #381 – Chris Lattner: Future of Programming and AI · IDENTIFIED FROM THE TRANSCRIPT · source