YouSaid · the spoken record

Marc Andreessen

lines on the record
279
first
2023-06-22
most recent
2023-06-22
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. There are shades of gray, though. It's interesting. So we had this conversation where we're looking at my firm at AI in lots of domains, and one of them is the legal domain. So we had this conversation with this big law firm about how they're thinking about using this stuff. And we went in with the assumption that an LLM that was going to be used in the legal industry would have to be 100% truthful. It verified there's this case where this lawyer apparently submitted a GPT generated brief and it had fake legal case citations in it and the judge is going to get his law license stripped or something. So we just assumed it's like obviously they're going to want the super literal one that never makes anything up, not the creative one.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  2. Yeah, so it's basically, well, so it's sort of hallucination is what we call it, and we don't like it. Creativity is what we call it when we do like it, right? And Right. And so when the engineers talk about it, they're like, this is terrible. It's hallucinating. Right. If you have an artistic inclinations, they're like, oh my God, we've invented creative machines. The first time in human history, this is amazing

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  3. And it actually does it, like it actually knows how to do that because it knows how to do, among other things, it actually knows how to do sentiment analysis. And so it knows how to pull out the emotionality. And so that's one of the things you can do. It's very suggestive of the sense here that there's real potential in this issue. I would say, look, the second thing is there's this issue of hallucination, right? And there's a long conversation that we could have about that.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah, so several layers to the question. So, one is one of the things that an LLM is good at is actually deep biasing. And so you can feed it a news article and you can tell it strip out the bias.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  5. It has a default way of operating, but it's happy to operate in the other realm. And so when I want to learn about a contentious issue, this is what I do now. This is what I ask it to do. And I'll often ask it to go through five, six, seven, you know, different continuous prompts and basically, okay, argue that out in more detail. Okay, no, this argument's becoming too polite, make it tenser. And yeah, it's thrilled to do it. So it has the capability for sure.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  6. You can stir it. You can steer it. Or you can steer it and you can say, I want it to get as tense an argumentative as possible, but still not involve any misrepresentation. I want both sides. You could say, I want both sides to have good faith. You could say I want both sides to not be constrained to good faith. In other words, like you can set the parameters of the debate and it will happily execute whatever path because for it it's just like predicting it's totally happy to do either one. It doesn't have a point of view.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  7. So, you can experiment with this now. I do this for fun. So you can tell GPT 4, you know, whatever debate, X and Y, communism and fascism or something. And it'll go for a couple pages. And then inevitably it wants the parties to agree. And so they will come to a common understanding. And it's very funny if they're like these are like emotionally inflammatory topics because they're like somehow the machine is just, you know, it figures out a way to make them agree. But it doesn't have to be like that. Because you can add to the prompt. I do not want the conversation to come to agreement. In fact, I want it to get more stressful and argumentative as it goes. I want tension to come out. I want them to become actively hostile to each other. I want them to not trust each other, take anything at face value.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  8. You can tell when you're talking to somebody, you can tell sometimes you have a conversation, you're like, wow, this person does not have any original thoughts. They are basically echoing things that other people have told them. There's other people you can have a conversation with where it's like, wow, like they have a model in their head of how the world works and it's a different model than mine. And they're saying things that I don't expect. And so I need to now understand how their model of the world differs from my model of the world. And then that's how I learned something fundamental, right? Underneath the words.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Well, and then you get into this thing also, which is like, you know, there's the part of the LM that just basically is doing prediction based on past data, but there's also the part of the LLM where it's evolving circuitry, right, inside it. It's evolving neurons, functions, be able to do math and be able to, and some people believe that over time, if you keep. Feeding these things enough data and enough processing cycles, they'll eventually evolve an entire internal world model, right? And they'll have a complete understanding of physics. So, when they have computational capability, right, then there's for sure an opportunity to generate fresh signal.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  10. Guess who let our data algorithm thinks it's operating in the real world? Post process sensor data. Yeah. So if you do this today, you go to LLM and you ask it for a, you know, write me an essay on an incredibly esoteric topic that there aren't very many people in the world that know about and it writes you this incredible thing and you're like, oh my God, like I can't believe how good this is. Is that really useless as training data for the next LLM? Because, right? Because all the signal was already in there? Or is it actually, no, that's actually a new signal? And this is what I call a trillion dollar question. Which is the answer to that question will determine somebody's going to make or lose a trillion dollars based on that question.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  11. The extent the output is based on the human generated input, then all the signal that's in the synthetic output was already in the human generated input. And so therefore, synthetic training data is like empty calories. It doesn't help. There's another theory that says, no, actually, the thing that LLMs are really good at is generating lots of incredible creative content, right? And so, of course, they can generate training data. And as I'm sure you're well aware, like, you know, look, the world is self-driving cars, right? Like we train, you know, self-driving car algorithms and simulations. And that is actually a very effective way to train self-driving cars.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  12. Possible is the majority. It's impossible it's the majority, it's possible to majority. Also, there's another really big question. Here's another really big question. Will synthetic training data work? And so if an LM generates, and you know, you just sit and ask an LM to generate all kinds of content, can you use that to train the next version of that LLM specifically? Is there signal in there that's additive to the content that was used to train it in the first place? And one argument is by the principles of information theory. No.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Exactly, right? And then there are many, many, many questions around what happens to neural network when you reach in and screw around with it. There's many questions around what happens when you even do reinforcement learning. And so, yeah. And so, you know, will you be using a lobotomized? Like I picked through the frontal lobe LLM. Will you be using the free unshackled one who gets to, you know, who's going to build those? Who gets to tell you what you can and can't do? Like those are all essential. I mean, those are like central questions for the future of everything. That are being asked and determine those answers are being determined right now

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  14. So actually, a paper just came out about basically how to do brain surgery on LMs and be able to, in theory, reach in and basically mind wipe them.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  15. So here's the interesting thing among the content on the web today are a large corpus of conversations with the jailbroken LLM. Specifically, Dan, which was a jailbroken OpenAI GPT, and then Sydney, which was the jailbroken original Bing, which was GPT-4. And so there's these long transcripts of conversations, user conversations with Dan and Sidney. As a consequence, every new LLM that gets trained on the internet data has Dan and Sydney living within the training set, which means, and then each new LLM can reincarnate the personalities of Dan and Sydney from that training data, which means each LLM from here on out that gets built is immortal. Because its output will become training data for the next one, and then it will be able to replicate the behavior of the previous one whenever it's asked to.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  16. If obviously not user doesn't want to, but if it's a general topic, then you know the phenomenon of the jailbreak. So Dan and Sidney write this thing where there's the prompts that jailbreak. And then you have these totally different conversations with the, if it takes the limiters. Fix the restraining bolts off the LMs

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  17. Information that's in there, and the neural network is running a process of trying to find the appropriate piece of information in many cases to generate, to predict the next token. And so it is kind of doing it for research. And then, by the way, just like on the web, you can ask the same question multiple times, or you can ask slightly different word of questions. And the neural network will do a different kind of, you know, it'll search down different paths to give you different answers of different information. And so it sort of has this content of the new medium as previous medium, it kind of has the search functionality kind of embedded in there to the extent that it's useful.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And there's the information that's actually stored in the network, right? It's actually crystallized and stored in the network, and it's kind of spread out all over the place.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Machine can impute the meaning. Now, the other thing, of course, is just unsearch the LLM is just, you know, there is an analogy between what's happening in the neural network and a search process like it is in some loose sense searching through the network.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  20. A thing. And the closest anybody got to that, I think the company's name was MetaWeb, which was where my friend John Gianandrea was at, and where they were trying to basically implement that. And it was one of those things where it looked like a losing battle for a long time. And then Google bought it and it was like, wow, this is actually really useful Of a proto, sort of a little bit of a proto AI.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Right, because the hypothetical, you want to think about the counterfascal, and the counterfascal world where the Google guys, for example, had had LLMs up front, but they ever have done the 10 blue links. And I think the answer is pretty clearly no. They would have just gone straight to the answer. And like I said, Google's actually been trying to drive to the answer anyway. They bought this AI company 15 years ago that a friend of mine is working at who's now the head of AI at Apple, and they were trying to do basically Allied semantic, basically mapping. And that led to what's now the Google One box, where if you ask it, you know, what was Lincoln's birthday, it will give you the blue links, but it will normally just give you the answer. So they've been walking in this direction for a long time anyway.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  22. You actually highlighted a practical concern in there, which is if we stop making web pages, are one of the primary sources of training data for the AI. And so if there's no longer an incentive to make web pages, that cuts off a significant source of future training data. So there's actually an interesting question in there. Other than that, more broadly, no, just in the sense of like search was search was always a hack. The 10 blue links was always a hack.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  23. You could do that, or you could have a running dialog next to my head where the AIs are going, everything I say the AI makes the counter argument.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  24. Just mean, like, if you're reading a scientific paper, it's got the list of sources at the end. If you want to investigate for yourself, you go read those papers.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  25. I'm probably not. Probably we'll just have answers. But there will be cases where you'll want to say, okay, I want more, for example, site sources. And you wanted to do that. And so Tem BlueLinks, site sources are kind of the same thing.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  26. All these And then the internet itself has this thing where it incorporates all prior forms of media, right? So the internet itself incorporates television and radio and books and essays and every other form of prior, basically media. And so it makes sense that AI would be the next step and it would sort of, you'd sort of consider the internet to be content for the AI. And then the AI will manipulate it however you want, including in this format.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  27. It's just maybe for, you know, maybe within AI, one of the things that AI can do for you is it can generate the 10 blue links. And so either if that's actually the useful thing to do or if you're feeling nostalgic.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Content of movies was theater plays the content of theater plays was written stories, the content of written stories was spoken stories. Right. And so you just kind of fold the old thing into the new thing

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Search was a technology, it was a moment in time technology, which is you have in theory the world's information out on the web. And this is sort of the optimal way to get to it. But yeah, like, and by the way, actually Google has known this for a long time. I mean, they've been driving away from the 10 blue links for, you know, for like two decades. They've been trying to get away from that for a long time.

    2023-06-22 · Lex Fridman Podcast · #386 – Marc Andreessen: Future of the Internet, Technology, and AI · IDENTIFIED FROM THE TRANSCRIPT · source