YouSaid · the spoken record

Edward Gibson

lines on the record
228
first
2024-04-17
most recent
2024-04-17
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Particular way. And so, why do we do that? Like, that's probably not about communication. That's probably about learning. I mean, then we're talking about learning. It's probably easier to learn regular things, things which are very predictable and easy to. So that's probably about learning is our guess, because that can't be about communicating.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  2. And learning. So, yes, one of the factors Yeah, so learning is messing this up a bit, and so for example, if it were just about minimizing dependency lengths, and that was all that matters, you know, then we, you know, so then we might find grammars which didn't have regularity in their rules. But languages always have regularity in their rules. So what I mean by that is that if I wanted to say something to you in the optimal way to say it was what really mattered to me all that mattered was keeping the dependencies as close together as possible, then I would have a very lax set of phrase structure or dependency rules. It wouldn't have very many of those. I would have very little of that. And I would just put the words as close the things that refer to the things that are connected right beside each other. But we don't do that. There are word order rules, right? So they're very, and depending on the language, they're more and less strict, right? So you speak Russian, they're less strict than English. English has very rigid word order rules. We order things in a very

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Yeah, well, all the languages in the world's language, none is right now we know is any better than any other with respect to sort of optimizing dependency lengths, for example. They're all kind of do it, do it well. They all keep low. I think of every human language as some kind of an optimization problem, a complex optimization problem to this communication problem. So they've solved it. They're just sort of noisy solutions to this problem of communication. There's just so many ways you can do this.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  4. How about just form still, though? Like, just what language you know. So, how well you know the language? And so if it's second language for you versus first language, and how maybe what other languages you know, these are still just form stuff. And that's potentially very informative. And, you know, how old you are. These things probably matter, right? So like child learning a language is as a noisy representation of English grammar, depending on how old they are. So maybe when they're six, they're perfectly formed.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  5. We think that those are there's three at least three different kinds of things going on there, and we probably don't want to treat them all as the same. Sure. And so I think the right model, a better model of a noisy channel would treat would have three different sources of noise, which are background noise, speaker inherent noise and listener inherent noise. And those are not those are all different things.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I think I'm pretty reasonable. Here's why does a word order look the way it does? We're now into shaky territory, but it's kind of cool.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Yeah, well, then pancy length is really about memory, right? I think that's like about sort of what's easier or harder to produce in some way. And these other ideas are about sort of robustness to communication. So the problem of potential loss of signal due to noise. So there may be aspects of word order, which is somewhat optimized for that. And we have this one guess in that direction. These are kind of just so stories. I have to be pretty frank. Like, I can't show this is true. All we can do is like look at the current languages of the world. This is like we can't sort of see how languages change or anything because we've got these snapshots of a few hundred or a few thousand languages. We can't do the right kinds of modifications to test these things experimentally. And so, you know, so just take that with a grain of salt, okay? From here, this, this stuff. The dependency stuff, I can, I'm much more solid on. And like here's what the lengths are and here's what's hard and here's what's easy. And this is a reasonable structure.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  8. That's really interesting. Yeah, he was interested in that. His examples in the 40s are kind of like they're very language things. Yeah. Can kind of show that there's a noisy channel process going on when you're listening to me. You can often sort of guess what I meant by what you think I meant given what I said. And I mean, with respect to sort of why language looks the way it does, there might be sort of, as I alluded to, there might be ways in which word order is somewhat optimized because of the noisy channel in some way.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  9. A little more easy to be passed from speaker to listener. So Shannon's the guy that did this stuff way back in the 40s, you know, it's very interesting. Historically, he was interested in working in linguistics. He was at MIT. And this was his master's thesis of all things. It's crazy how much he did for his master's thesis. In 1948, I think, or 49 or something. And he wanted to keep working in language. And it just wasn't a popular. Communication as a reason, a source for what language was wasn't popular at the time. So Chomsky was moving in there. And he just wasn't able to get a handle there, I think. And so he moved to Bell Haps and worked on communication from a mathematical point of view and was did all kinds of amazing work. And so he's just more.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Lineral background, no, is there like white noise in the background or some other kind of noise or some speaking going on that we're just you're at a party? That's background noise. You're trying to hear someone. It's hard to understand them because there's all this other stuff going on in the background. And then there's noise on the communication on the receiver side so that you have some problem maybe understanding me for stuff that is internal to you in some way. So you've got some other problems, whatever understanding for whatever reasons. Maybe you're maybe you've had too much to drink. You know, who knows why you're not able to pay attention to the signal. So that's the noisy channel. And so that language, if it's communication system, we are trying to optimize in some sense the passing of the message from one side to the other. And so one idea is that maybe, you know, aspects of like word order, for example, might have optimized in some way to make language.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  11. There's a lot of noise in this system. I don't speak perfectly. I make errors. That's noise. There's background noise. You know, you know that as we're back.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So that's about communication. And so, this is going back to Shannon. So, Shannon, Claude Shannon was a Student at MIT in the 40s. And so he wrote this very influential piece of work about communication theory or information theory. And he was interested in human language. Actually, he was trying in this problem of communication, of getting a message from my head to your head. He was concerned or interested in what was a robust way to do that. And so that assuming we both speak the same language, we both already speak English, whatever language is, we speak that. What is a way that I can say the language so that it's most likely to get the signal that I want to you? And so, and then the problem there in the communication is the noisy channel, is that there's

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  13. People On average better, but they had the same difference Exact same difference. So they but I they wanted it fixed so they they also and so that that gave us hope that because it actually isn't very hard to construct a material which is uncenter embedded and has the same meaning it's not very hard to do just basically in that situation you're just putting definitions outside of the subject verbatim in that particular example and that's kind of pretty general what they're doing is just throwing stuff in there which you didn't have to put in there there's extra words involved typically you may need a few extra words sort of to refer to the things that you're defining outside in some way because if you only use it in that one sentence then there's no reason to introduce extra extra terms so we might have a few more words but it'll be easier to understand so i mean i i have hope that now that maybe we can make legalese less less convoluted in this so maybe

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  14. I'm suspicious as well. I'm still suspicious. And I hear what you're saying. It could be kind of no individual and even average of individuals. It could just be a few bad apples in a way which are driving the effect in some way. Influential bad apples.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  15. And maybe there's a syntactic sort of reflex here of a magic spell, which is centrum betting. And so that's like, oh, it's trying to tell you this is like this is something which is true, which is what the goal of law is, right? It's telling you something that we want you to believe is certainly true, right? That's what legal contracts are trying to enforce on you, right? And so maybe that's like a form which has, this is like an abstract, very abstract form, centrum betting, which has a meaning associated with it.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Well, our best theory at the moment is that there's actually some kind of a performative meaning in the center embedding, in the style, which tells you its legal ease. We think that that's the kind of a style which tells you its legal ease. It's a reasonable guess. And maybe it's just, so for instance, if you're like, it's like... A magic spell. So we kind of call this the magic spell hypothesis. So when you kill someone to put a magic spell on someone, what do you do? People know what a magic spell is and they do a lot of rhyming. That's kind of what people will tend to do. They'll do rhyming and they'll do sort of like some kind of poetry kind of thing.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Yeah, and we asked them, we asked them, would you hire someone who writes like this or this? We asked them all kinds of questions and they always preferred the less complicated version, all of them. So I don't even think they want it this way

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And the central embedded bit there is just for some reason there's a definition. They throw the definition of what payments and benefits are in between the subject and the verb. How about don't do that? How about put the definition somewhere else as opposed to in the middle of the sentence? And so that's very, very common, by the way. That's what happens. You just throw your definitions, you use a word, a couple words, and then you define it, and then you continue the sentence. Like, just don't write like that. And you ask, so then we ask lawyers. We thought, oh, maybe lawyers like this. Lawyers don't like this. They don't like this. They don't want to write like this. We ask them to rate materials which are with the same meaning, with uncentro bed and center bed. And they much preferred the uncentrobed versions.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  19. I mean, legalese is hard. It goes in the event that any payment or benefit by the company, all such payments and benefits, including the payments and benefits under Section 3A hereof, being here and after referred to as a total payment, would be subject to the excise tax, then the cash severance payments shall be reduced. So that's something we pulled from a regular text, from a contract. Wow.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Tide of a sentence that just blows up that makes super can I read a sentence for you from these things? I see, I mean, this is just like one of the things that, this is just

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Anything processed just as well as it was normal? No, no, they're much better than lay people. So they're much like they can much better recall much better understanding, but they have the same main effects as lay people, exactly the same. So they also much prefer the non-centric, so we constructed non-center embedded versions of each of these. We constructed versions which have higher frequency words in those places and we did unpassivize. We turned them into active versions the passive active made no difference the words made a little difference and the uncentro embedding makes big differences in all the population

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  22. The dependent measure is how well you understand those things with those features. Okay, and so then, and it turns out the passive makes no difference. So it has a zero effect on your comprehension ability, on your recallability, nothing at all. It has no effect. The words matter a little bit. Low frequency words are going to hurt you in recall and understanding. But what really hurts is the centrum embedding. That kills you. That is like, that slows people down. That makes them very poor at understanding that makes them, they can't recall what was said as well nearly as well. And we did this not only on lay people, we did have a lot of lay people. We ran it on 100 lawyers. We recruited lawyers from a wide range of sort of different levels of law firms and stuff. And they have the same pattern. So they also, like when they did this, I did not know it would have happened. I thought maybe they could process, they used to legally.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Low frequency words suck. Well, sucks is different. So these are different judgment on passing. Yeah, yeah, yeah. Pass the drop the judgment. It's just like these are frequent. These are things which happen in legalese texts. Then we can ask.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Often a clause would intervene between a subject and a verb, for example. That's one kind of a central embedding of a clause. And turns out they're massively central embedded. So I think in random contracts and in random laws, I think you get about 70% or 80, something like 70% of sentences have a center-embedded clause in them, which is insanely high if you go to any other text down to 20% or something. It's so much higher than any control you can think of, including, you think, oh, people think, oh, technical academic text, no, people don't write syndrome embedded sentences in technical academic texts. I mean, they do a little bit, but much it's on the 20%, 30% realm as opposed to 70. And so there's that, and there's low frequency words. And then people, oh, maybe it's passive. People don't like the passive. Passive for some reason, the passive voice in English has a bad rap, and I'm not really sure where that comes from.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Yeah, yeah, yeah. We'll get to that. We'll get to that. But we wanted to see why first. So it turns out that we're not the first to observe that legalese is weird. Back to Nixon had a Plain Language Act in 1970 and Obama had one and boy a lot of these presidents have said, oh, we've got to simplify legal language, must simplify it. But if you don't know how it's complicated, it's not easy to simplify it. You need to know what it is you're supposed to do before you can fix it, right? And so you need a psycholinguist to analyze the text and see what's wrong with it before you can fix it. You don't know how to fix it. How am I supposed to fix something? I don't know what's wrong with it. And so what we did was just, that's what we did. We figured out, let's, okay, we just took a bunch of contracts, had people, and we encoded them for a bunch of features. And so another feature that people, one of them was central embedding. And so that is like basically how.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Well, you know, it's interesting. Now you're getting at why. And so, and I don't think so now you're saying it's they're doing it intentionally. I don't think they're doing intentionally. But let's

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Other kinds of control text, even sort of academic text legal lease is even worse. It is the worst that we were being supplying.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Sounds hard to understand. So why is it hard to understand and why do they write that way if it is so hard to understand? It seems apparent that it's hard to understand. The question is, why is it? And so we didn't know. And we did an evaluation of a bunch of contracts actually. We just took a bunch of random contracts because I don't know, you know, there's contracts and laws might not be exactly the same, but contracts are kind of the things that most people have to deal with most of the time. And so that's kind of the most common thing that humans have, like humans, that adults in our industrialized society have to deal with a lot. And so that's what we pulled. And we didn't know what was hard about them. But it turns out that the way they're written is very centrumbed. It has nested structures in them. So it has low frequency words as well. That's not surprising. Lots of texts have low, it does have surprising, slightly lower frequency words than

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  29. So, I'm just talking about language in laws and language in contracts So, the stuff that you have to run into, we have to run into every other day or every day, and you skip over because it reads poorly. And ORC, partly it's just long, right? There's a lot of text there that we don't really want to know about. But the thing I'm interested in, so I've been working with this guy called Eric Martinez

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  30. We're going to talk about legalese at some point. So maybe we'll talk about that kind of thinking with applied to legalese

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  31. And yet it explains some very complicated phenomena. If you, if I write these very complicated sentences, it's kind of hard to know why they're so hard. And you can like, oh, nail it down. I can do a math formula for why each one of them is bad and where. And that's kind of cool. I think that's very neat.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  32. Is the access, is trying to find those earlier things. It's kind of hard to figure out what was referred to earlier. Those are those connections. That's the sort of notion of as opposed to a storage thing, but trying to connect, retrieve those earlier words depending on what was in between. And then we're talking about interference of similar things in between. That's the right theory probably has that kind of notion in it as an interference of similar. And so I'm dealing with an abstraction over the right theory, which is just, you know, let's count words. It's not right, but it's close. And then maybe you're right, though. There's some sort of an exponential or something to figure out the total so we can figure out a function for any given sentence in any given language. But, you know, it's funny, you know, people haven't done that too much, which I do think is, I'm interested that you find that interesting. I really find that interesting and a lot of people haven't found it interesting. And I don't know why I haven't got people to want.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  33. I think it's an exponential. So we think it's probably an exponential such that the longer the distance, the less it matters. And so then it's the sum of those is my, that was our best guess a while ago. So you've got a bunch of dependencies. If you've got a bunch of them that are being connected at some point, that's at the ends of those. The cost is some exponential function of those is my guess. The reason it's probably an exponential is like it's not just the distance between two words because I can make a very, very long subject. Verb depends by adding lots and lots of noun phrases and prepositional phrases and it doesn't matter too much. It's when you do nested, when I have multiple of these, then things go really bad, go south.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Meaning like nouns might matter depending, and then it maybe depends on which kind of noun is it a noun we've already introduced or a noun that's already been mentioned, is it a pronoun versus a name? Like all these things probably matter. So probably the simplest thing to do is just like, oh, let's forget about all that and just think about words or morphemes.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Yeah, it's complicated, but probably it's doable. I would guess it's doable. I tried to do that a while ago, and I was reasonably successful. For some reason, I stopped working on that. I agree with you that it would be nice to figure out. So there's like some way to figure out the cost. I mean, it's complicated. Another issue you raised before was like, how do you measure distance? Is it words? It probably isn't. Part of the problem is that some words matter than more than others and probably, you know.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  36. The stronger the activation in the language network. And so there's some measure. There's a bunch of different measures we could do. That's a kind of a neat measure, actually, of actual...

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Oh, well, you can measure it in a lot of ways. The simplest is just asking people to say whether how good a sentence sounds. Just ask, that's one way to measure and you try to triangulate then across sentences and across structures to try to figure out what the source of that is. You can look at reading times in controlled materials and certain kinds of materials and then we can measure the dependency distances there. There's a recent study which looked at we're talking about the brain here. We could look at the language network. We could look at the language network and we could look at the activation in the language network and how big the activation is depending on the length of the dependencies. And it turns out in just random sentences that you're listening to, if you're listening to, so it turns out there are people listening to stories here. And the bigger, the longer the dependence dependency is, the

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  38. So, what I like about dependency grammar is it makes The cognitive cost associated with longer distance connections very transparent. Basically, it turns out there is a cost associated with producing and comprehending connections between words which are just not beside each other. The further apart they are, the worse it is, according to, well, we can measure that. And there is a cost associated with that.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  39. That's what this group argues. So the same group, Federenkos group, has a recent paper argue exactly that. There's a guy called Kyle Mahal, who's here in Austin, Texas, actually. He's an old student of mine, but he's a faculty in linguistics at Texas, and he was the first author on that

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Yeah, center bed. Exactly like the same way as humans. And that's not trained. So they do exactly, so that is the similarity. But that's not meaning. This is form. But when we get into meaning, this is where they get kind of messed up. We start to say, oh, what's behind this door? Oh, it's, you know, this is the thing I want. Humans don't mess that up as much. Here, the form is just like the form match is amazing similar without being trained to do that. I mean, it's trained in the sense that it's getting lots of data, which is just like human data, but it's not being trained on bad sentences and being told what's bad. It just can't do those. It'll actually say things like those are too hard for me to complete or something, which is kind of interesting. Actually, kind of how does it know that? I don't know.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  41. The places where large language models are the form is amazing. So let's go back to nested structure, center embedded structures. Okay, if you ask a human to complete those, they can't do it. Neither can a large language model. They're just like humans in that. If I ask a large language model,

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  42. Just saying the error there is like if I explain to you there's 100% chance that the car is behind this case this door well do you want to trade people say no but this thing will say yes because it's so that that trick it's so wound up on the form that it's that that's an error that a human doesn't make which is kind of interesting

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  43. So, I don't want to make that the inference I wouldn't want to make was that inference. The inference I'm trying to push is just that is it like humans here? It's probably not like humans here. It's different. So humans don't make that error. If you explain that to them, they're not going to make that error. They don't make that error, and so that's something it's doing something different from humans that they're doing in that case.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  44. I mean, you don't have to convince me of that. I am very, very impressed. But do, I mean, You're giving a possible world where maybe someone's going to train some other version such that it'll be somehow abstracting away from types of forms. I mean, I don't think that's happened.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  45. There's three doors, and one behind them is a good prize, and there's two bad doors. I happen to know it's behind door number one. The good prize, the car is behind door number one. So I'm going to choose door number one. Monty Hall opens door number three and shows me nothing there. Should I trade for door number two? Even though I know the good prize in door number one and then the large language will all say, yes, you should trade because it just goes through the forms that it's seen before so many times on these cases where yes, you should trade because your odds have shifted from one and three now to two out of three to being that thing. It doesn't have any way to remember that actually you have 100% probability behind that door number one. You know that. That's not part of the scheme that it's seen hundreds and hundreds of times before. And so you can't, even you try to explain to it that it's wrong that they can't do that. It'll just keep giving you back the problem.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  46. And then the question is should you trade to get the other one? And the answer is yes, you should trade because he knew which ones you could turn around. And so now the odds are two thirds, okay? And if you just change that a little bit to the large language model, the large language model just seen that. That explanation so many times that it just, if you change the story, it's a little bit, but it makes it sound like it's the Monty Hell problem.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  47. I mean, I would argue they're doing the form. They're doing the form and doing it really, really well. And are they doing the meaning? No, probably not. I mean, there's lots of these examples from various groups showing that they can be tricked in all kinds of ways. They really don't understand the meaning of what's going on. And so there's a lot of examples that he and other groups have given, which show they don't really understand what's going on. So you know the Monty Hall problem is this silly problem, right? Where, you know, if you have three door, it's let's make a deal. It's this old game show and there's three doors and there's a prize behind one and there's some junk prizes behind the other two and you're trying to select one. And if you knows Monty, he knows where the target item is, the good thing, he knows everything is back there. And you're supposed to, he gives you a choice. You choose one of the three. And then he opens one of the doors and it's some junk.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Unspecified as to what the form of the grammar is underlyingly. And so I would argue that the dependency grammar is maybe the right form to use for the types of construction grammar. Construction grammar typically isn't kind of formalized quite. And so maybe the formalization, a formalization of that, it might be in dependency grammar. I mean, I would think so, but I mean, it's up to people, other researchers in that area if they agree or not.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  49. It's just a general theory of language such that there's a form and a meaning pair for lots of pieces of the language. And so it's primarily usage-based is a construction grammar. It's trying to deal with the things that people actually say, actually say and actually write. And so it's a usage-based idea. And what's a construction of constructions either a simple word, so like a morpheme plus its meaning, or a combination of words. It's basically combinations of words like the rules. But it's...

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source

  50. About it, so you know, that's I mean, that presumes, and there's some evidence for this, that some large language models are implementing something like dependency grammar inside them. And so there's work from a guy called Chris Manning and colleagues over at Stanford in natural language. And they looked at, I don't know how many large language model types, but certainly Bert and some others, where you do some kind of fancy math to figure out exactly what kind of abstractions of representations are going on. And they were saying it does look like dependency structure is what they're constructing. It doesn't, like, so it's actually a very, very good map. So kind of a, they are constructing something like that. Does it mean that they're using that for meaning? I mean, probably, but we don't know.

    2024-04-17 · Lex Fridman Podcast · #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs · IDENTIFIED FROM THE TRANSCRIPT · source