YouSaid · the spoken record
Mihir Bellare
- lines on the record
- 62
- first
- 2025-10-13
- most recent
- 2025-10-13
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“No, there is something, I don't know if I should say profound, but there is something about it which tells what these LLMs can or cannot do. So, one of the examples that I often tell is suppose I ask you what is 769 times 1025?”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“So you're reducing it to a very specific state space when it comes to confidence in an answer. And this is kind of a manifold that you can go on. And then, I mean, do you You have kind of a conclusion of what that means for systems or what that means for reasoning, or is it just a nice way to articulate the bounds of LLMs?”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“But when I say I'm going to dinner with Martin Cas Know the LLM now, this is information rich. This is sort of a rare phrase. And now the sort of realm of possibilities reduces because Martin is only going to take me to Michelin Star restaurants. I'm not going to go McDonald's. You did what I'm saying. The moment you add more context. You make the prompt information rich, the prediction entropy reduces”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“And the prompts also you can categorize into two kinds of prompts. One prompt is you can say high information entropy. And one prompt is low information entropy. So, the way these manifolds work, the LLMs start paying attention to prompts that have high Information entropy And low prediction entropy So, what do I mean by that? So, when I say I'm going out for dinner. So, when I say, I'm going out for dinner, that phrase the LLMs have been trained. Have seen it a lot, and there are many different directions I can go with it. I can say, I'm going for dinner tonight, I'm going for dinner to McDonald's or I'm going to dinner blah, blah, blah. There are many different.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“Not thermodynamic entrop So suppose you have a vocabulary of, let's say, 50,000 different tokens and you have a distribution next token distribution over these 50,000 tokens. So let's say the cat sat on the if that is a prompt, then the distribution will have a high probability for Mac. Or hat or table, and a very low probability of, let's say, ship or whale. Or something like that, right? So, because of the way it's trained, it has these distributions. Their distributions can be low entropy or high entropy. A high entropy distribution means that there are many different ways that the LLM can go. With a high enough probability for all those paths. Low entropy means that there are only a small set of choices for the next token.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“It can produce something which makes sense. The moment it sort of veers away from the manifold, then it starts hallucinating and starts spotting nonsense. Confident nonsense, but nonsense. So it creates these manifolds. And the trick is the distribution that is produced. You can measure the entropy of the distribution. But entropy the way Shanna just said it. Shannon entropy.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“BC with an edge index of what 60? Yeah, so ultimately what all these LLMs are doing, whether the early LLMs or the LLMs that we have today with all sorts of post training RLHF, whatever you do, at the end of the day, what they do is they create a distribution for the next token. So given a prompt, these LLMs create a distribution for the next token or the next word, and then they pick something from that distribution using some kind of algorithm to predict the next token, pick it, and then keep going. Now, what happens because of the way we train these LLMs, the architecture of the transformers, and the loss function, the way you put it is right, it sort of reduces the world into these Bayesian manifolds. And as long as the LLM is going in sort of traversing through these manifolds, it is confident.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“Roughly. So you've reduced the dimensionality of the probe to a geometric manifold, and then you can actually formally specify kind of how far you can reason within that manifold. And the articulation is that we, or what are the intuitions that we as humans do the same thing is we take this very complex, heavy-tailed stochastic universe, and we reduce it to kind of this geometric manifold. And then when we reason, we just move along that manifold.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, okay, so can I just try to take a rough sketch at it and then you just tell me how wrong I am. You're trying to describe how LLMs work. And one thing that you found is that they reduce a very, very complex multidimensional space into basically a geometric manifold that's a reduced state space. So it's reduced degrees of freedom, but you can actually predict where in the manifolds the reasoning can move to.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“Beyond the black box. Actually, we should put this in the notes for this, but the single best talk I've ever seen on trying to understand how LLM's work is one that Fishall did at MIT, which Hari Bala Krishan pointed me to, and I watched that. So he did that work. And then he's doing more recent work that's actually trying to scope out not only how LLMs reason, but like it has some reflections on humans reason too. And so I just think he's doing some of the more profound work and trying to understand and come up with models, formal models for how LLM's reason.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“And so we actually view the world in an information theoretic way. It is actually part of networking. And with all this AI stuff, there's so much work trying to create models that can help us understand how these LLMs work. And in my experience over the last three years, the ones that have most impacted my understanding, and I think have been the most predictive, are the ones that Vishal has come up with. He did a previous one that we're going to talk about called Matrix, is it?”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source
“Any LLM that was trained on pre 1915 physics would never have come up with the theory of relativity. Einstein had to sort of reject the Newtonian physics and come up with this space-time continuum. He completely rewrote the rules. AGI will be when we are able to create new science, new results, new math. When an AGI comes up with a theory of relativity, it has to go beyond what it has been trained on to come up with new paradigms, new science. That's my definition of AGI.”
2025-10-13 · a16z Podcast · Columbia CS Professor: Why LLMs Can’t Discover New Science · IDENTIFIED FROM THE TRANSCRIPT · source