YouSaid · the spoken record
Nat Friedman
- lines on the record
- 118
- first
- 2023-03-22
- most recent
- 2023-03-22
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“And then, yeah, I think that it became quite obvious that it was good because we had hundreds of internal users who were GitHub engineers. And I remember the first time I looked at the retention numbers, they were extremely high. It was like, I remember 660 plus percent after 30 days from first install. Like if you installed it, the chance that you were still using it after 30 days is like over 60%. And it's very intrusive product. I mean, it's sort of always popping UI up. And so if you don't like it, you will disable it. Indeed, 40-something percent of people did disable it. But those are very high retention numbers for an alpha first version of a product that you're using all day. And so then I was just incredibly excited to launch it. And now it's improved dramatically since then.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I think Alex had the idea of saying we can use the cursor position in the AST to figure out heuristically whether you're at the beginning of a block in the code or not. And if it's not the beginning of a block, just complete a line. If it's the beginning of a block, show inline a full block completion. So the number of tokens you request and when you stop gets altered automatically with no user interaction. And then the idea of using this sort of gray text like Gmail had done in the editor. And so we got that implemented. And it was really only kind of once all those pieces came together and we started using a model that was small enough to be low latency, but big enough to be accurate, that we reached the point where the median new user loved copilot and wouldn't stop using it. And that took four months, five months of just tinkering and sort of exploring the other dead ends that we had along the way.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“We still have that, but the UI for that wasn't good. And we had it, I think, so that to get a function body synthesized, you would hit a key. And then I don't know why this was the idea everyone had at the time, but several people had this idea that it should display multiple options for the function body, and then the user would read them and pick the right one. And I think the idea was that we would use that human feedback to improve the model. But that turned out to be a bad experience because first you had to hit a key and explicitly request it. Then you had to wait for it. And then you had to read three different versions of a block of code, reading one version of a block of code takes some cognitive effort, doing it three times, takes more cognitive effort. And then most often the result of that was like none of them were good or you didn't know, which one to pick. So that was also like you're putting a lot of energy and you're not getting a lot out, sort of frustrating. So once we had that sort of single line completion work,”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Stack overflow type of thing. So, like that turns out to be maybe not a sufficient condition for a product to be good because it was just at the time the models were just not reliable enough, they were not good enough. You know, I ask you a question. 25% of the time you give me an incredible answer that I love, 75% of the time your answer is useless or wrong. It's not a great product experience. And so then we started thinking about code synthesis. And our first attempts at this were actually large chunks of code synthesis, like synthesizing whole function bodies. And we built some tools to do that and put them in the editor. And that also was not really that satisfying. And so the next thing that we tried was to just do simple small scale autocomplete with the large models. And we used the kind of intellience drop down UI to do that. And that was better, like definitely pretty good, but the UI was not quite right. And we lost the ability to do this large scale synthesis.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Off to build large scale AI and eventually AGI. And so I think it was some combination of those three people kind of coming together to make it happen. But I still think it was a very prescient bet. I've said that to people and they've said, well, a billion dollars is not a lot for Microsoft. Yeah, but like there were a lot of other companies that could have spent a billion dollars to do that and did not. And so I still think that deserves a lot of credit. Okay, so GBD3 comes out. I pinged Sam and Greg, I think Brockman at OpenAI. And they were like, yeah, like we've already been experimenting with GPT-3 and derivative models encoding context. Let's definitely work on something. And to me, at least, and a few other people, it was not incredibly obvious what the product would be. Now I think it's trivially obvious. You know, autocomplete, my gosh, isn't that what the models do? But at the time, actually, my first thought was that it was probably going to be like a Q&A chat.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't know. Actually, I've never asked him. But I'm not sure. It's a good question. And I think if you're Satya and you're running this multi-trillion dollar company, you're trying to execute well and serve your customers, but you're always looking for the next gigantic wave that is going to upend the technology industry. It's not just about trying to win cloud. It's like, okay, what comes after cloud? And so you have to make some big bets. And I think he thought AI could be one. And I think Kevin Scott deserves a lot of credit for really advocating for that aggressively. And I think Sam Altman did a good job of building that partnership because he knew that he needed access to the resources of a company like Microsoft.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I mean, I've talked about this a little bit. So look, GPT-3 came out in May, I think, of 2020, and I saw it, and it really blew my mind. I thought it was amazing. And I was CEO of GitHub at that time. I thought, like, I don't know what, but we've got to build some product with this. This is, you know, we've got to build something. So Sanatia had, at I think, Kevin Scott's urging already invested in OpenAI like a year before GPT-3 came out. Like this is quite amazing. And he invested like a billion dollars.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Leave, and so the people who really care about building great products and serving the customers, maybe they don't want to work for the acquirer. And the set of people that are really load-bearing around kind of the long-term success is small. And when they leave or get disempowered, you get very different behaviors.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, it is true most acquisitions are destructive of value. What is the value of a company? In an innovative industry, the value of the company, a lot of it boils down to its ability culturally to produce new innovations and is some sensitive harmonic of cultural elements that sets that up, that makes that possible. And it's quite fragile, I think. And so if you take a culture that has achieved some productive harmonic and you put it inside of another culture that's really different, the kind of mismatch of that can destroy the productivity of the company. So I think that maybe one way to think about it is companies are a little bit fragile. And so when you acquire them, it's relatively easy to break them up. I mean, they're also more durable than people think in many cases too. I would say another version of it is the people who really care.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Then people in the team feel the excitement of shipping stuff. I think GitHub was a company that had a little bit of stage fright about shipping previously and sort of break that static friction and ship a little bit more. I think felt good. And then the other one is just the learning loop. By trying to do lots of small things, I got exposed to like, okay, this team is really good, you know, or this part of the code has a lot of tech debt. Or, hey, we shipped that and it was actually kind of bad. How come that design got out? And so whereas if the project had been some six month thing, I'm not sure my learning would have been quite as quick about the company. There's still things I missed and mistakes I made for sure. But that was part of how I think no one knows counterfactually whether that made a big difference or not, but I do think that earned some trust.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And we eventually found something that we could fix by the end of the day. And what I'm thinking is, I hope I said was, well, we need to show the world is that GitHub cares about developers. Not that it cares about Microsoft. Like if the first thing we did after the acquisition was to add Skype integration, developers would have said, oh, we're not your priority. Like you have new priorities now. And so the idea was just to find ways to make it better for the people who use it and have them see that we cared about that immediately. And so I said, we're going to do this today, and then we're going to do it every day for the next 100 days. And it was cool because I think it created some really good feedback loops, at least for me. One was, you know, you ship things and then people are like, oh, hey, I've been wanting to see this fixed for years and now it's fixed. It's a relatively simple thing. So you get this sort of nice dopaminergic feedback loop going there.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Where he'd just been allowing people to give GitHub feedback, and people had been voting on this stuff for years. And I kind of shared my screen and put that up, sorted by votes, and said, like, we're going to pick one thing from this list and fix it by the end of the day and ship that. Like just one thing. And, you know, I think people were like, this is the new CEO's strategy. And they were like, I don't know. We can't, you know, do database migrations. They can't do that in a day. And like, and then someone's like, well, maybe we can do this. We actually have a half implementation of this.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And then it said, Nat Friedman's going to be CEO. And I had, I don't want to overstate or whatever. But I think a couple people were like, oh, Nat comes from open source. He spent some time in open source. So Nat's going to be run independently. So I don't think they were really that calmed down, but at least a few people thought like, oh, maybe I'll give this a few months and just see what happens before I migrate off. And then my first day as CEO after we got the deal closed, like 9 a.m. the first day I was in this room and we got on Zoom and all the heads of engineering and product. And I think maybe, I don't know what people were expecting, but I think maybe they were expecting some kind of longer-term strategy or something. But I came in and I said there was this GitHub at no official feedback mechanism that was publicly available, but there was several GitHub repos that community members had started. Isaac from NPM had started Juan.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, I was really paranoid about it. And I really cared about what developers thought. I think there's always this question of who are you performing for? Like, who do you actually really care about? Sort of who's the audience that's in your head that you're trying to do a good job for, impress, or in the respect of whatever it is. And though I love Microsoft and care a lot about Satya and everyone there, I really cared about the developers. You know, I'd grown up in this open source world. And so for me to do a bad job with this central institution and open source would have been a devastating feeling for me. It was very important to me not to. So that was sort of the first thing is just that I cared. The second thing is that the deal leaked. It was going to be announced, I think, on a Monday, leaked on a Friday. And Microsoft's buying GitHub. And the whole weekend, there were like terrible posts online, you know, people saying we got to evacuate GitHub as quickly as possible.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Which we could try Three weeks later, he said, okay, go do this. Scott will support you on this. Three weeks later, we had a signed term sheet and an announced deal. And then it was an amazing experience for me. I'd been there less than two years. And Microsoft was made up of and run by a lot of people who'd been there for many years. And they trusted me with this really big project and made me feel really good to be trusted and empowered. I had grown up in the open source world. And so for me to get an opportunity to run GitHub, it's like, I don't know, getting appointed mayor of your hometown or something like that. It felt cool. And I really wanted to do a good job for developers. And so...”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I said, look, the challenge that we have is there's an entire new generation of developers who have no affinity with Microsoft, and the largest collection of them is at GitHub. And if we acquire this and we do a merely competent job of running it, we can earn the right to be considered by these developers for all the other products that we do. And to my surprise, Satya replied in like six or seven minutes and said, I think this is very good thinking. Let's meet next week or so and talk about it. And I ended up at this conference room with him and Amy Hood and Scott Guthrie and Kevin Scott and several other people. And they said, okay, tell us what you're thinking. And I kind of did a little 20 minute ramble on it. And Satya said, I think we should do it. And why don't we run it independently like LinkedIn, Nat, you'll be the CEO. And he said, do you think we can get it for $2 billion? And I said.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I wrote him an email, just a memo. It sort of said, I think it's time to do this. There was some noise that Google was sniffing around. I think that may have been manufactured by the GitHub team. But it was a good catalyst because it was something I thought made a lot of sense for Microsoft to do anyway. And so I wrote an email to Satya sort of a little memo saying, you know, hey, I think we should buy GitHub. Here's why. Here's what we should do with it. And the basic argument was developers are making IT purchasing decisions now. It used to be the sort of IT thing, you know, and now developers are leading that purchase. And this sort of major shift in how software products are acquired. And Microsoft really was an IT company. It was not a developer company in the way most of its purchases were made. But it was founded as a developer company, right? And so Microsoft's first product was a programming language.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“This was like my first week. It was like March or April of 2016 And then he said, yeah, it's a good idea. We thought about it. I'm not sure we can get away with it or something like that. And then it was about a year later, a little more than a year later.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, so I had started a company called Xamarin together with Miguel de Casa and Joseph Hill, and we had built kind of mobile tools and platforms. Microsoft acquired the company in 2016. And I was excited about that. I thought it was great. But to be honest, I didn't actually expect or plan to spend more than a kind of a year or so there. But when I got in there, I got exposed to what Satya was doing and just the quality of his leadership team. I was really impressed. And actually, I think I saw him in the first week or so I was there and he asked me, what do you think we should do at Microsoft? And I said, I think we should buy GitHub.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“In a way, they're all looking forward. My Herculaneum project is looking backwards. I think it's extremely exciting and cool, but it is sort of a funny contrast.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, first, there's so much data on the internet. I mean, the two kind of primitives that you need to build models are you need lots of data. We have that in the form of the internet. We digitize the whole world into the internet. And then you have you need these GPUs, which we have because of video games. So you take the internet and video game hardware and you smash them together and you get machine learning models. And they're both commodities. And so I think the data, I don't think anyone in the open source world is really going to be data limited for a long time. There's so much that's out there. Probably people who have proprietary data sets that are readily scrapable have been shutting those down. So get your scraping in now if you need to do it. But that's just on the margin. I still think there's quite a lot. So I think, look, there's going to be a time, this is the Euro proliferation. This is a week of proliferation. Like we're going to see four or five major AI announcements this week, you know, new models, new APIs, new platforms, new tools from all the different vendors.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Bad things happen. The probability mass of my belief, though, is that it's probably better in the whole for more people to get to tinker with and use these models, at least in their current state. And so, for example, when Georgie Gerganoff this weekend did a 4-bit quantization of the llama model and got it in francing on a M1 or M2, I was very excited and I got that running and it's like fun to play with. Now I've got a model that is very good. It's almost GPT-3 quality runs on my laptop. I've sort of grown up in this world of the tinkerers and open source folks and the more access you have, the more things you can try. I think I do find myself very attracted to that.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, I don't know. My opinion's been changing. I have increasing worries about kind of safety issues like not the hijacked version of safety, but But in the long term, I do think there are worlds that we should be a little bit concerned about, although I don't know what to do about where, yeah, like...”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I mean, Igor's really, really great guy and brilliant, but he also happens to have trained state-of-the-art models at DeepMind and OpenAI. And so, you know, the set of people who have that set of, you know, I don't know whether that's a consideration or how big of an effect that is, but it's the kind of thing that it would make sense to value if you think there are sort of valuable secrets that have not yet. Proliferated. So I think they're going to try to slow it down. Publishing has certainly slowed down dramatically already, but I think there's just a long way to go before you're anywhere in hedge fund or Manhattan Project territory and probably secrets will still have a relatively short half-life.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I think ML engineer salaries and compensation packages will probably be adjusted to try to address this because you don't want your secrets walking out the door. There are engineers, you know, Igor Babushkin, for example, who has just, I believe, joined Twitter, I think, is that right. Elon hired him to train, I think that's public. I think it is.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I was just wondering about that some more today. I mean, it seems to be sort of trying to close the valves. But I think there's a lot of things working against them in this regard. So, one is again that the secrets are relatively simple. Two is that you're coming off this academic norm of publishing and really, like the entire culture is based on sort of sharing and publishing. Three is, as you said, they all live in group houses, some are in polycules. There's just a lot of intermixing. And then it's all in California. And California's non-compete state. We don't have non-competes. And so you'd have to change the culture. Everybody their own house and move to Connecticut and then maybe it would work.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“It's two hours in a year. And to my shock, they told me eighty percent reduction in their profits. Like it would have a huge impact. And then I asked, okay, so how long would the lookback window have to be before it would have like a relatively small effect on your business? And they said 10 years. So that I think is just quite strong evidence that the world's not perfectly efficient because these folks make billions of dollars using secrets that could be related in like an hour or something like that. And yet others don't have them or they're secrets wouldn't work. And so I think there are different levels of efficiency in the world, but on the whole, our default estimate of how efficient the world is is far too charitable.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And I asked them, I was trying to understand the role secrets play in the success of a hedge fund. And the reason I was interested in that is because I think the AI labs are going to enter a new similar dynamic where their secrets are very valuable. Like if you have a 50% training efficiency improvement in your training runs cost $100 million, that is a 50 million dollar secret that you have that you want to keep. And hedge funds do that kind of thing routinely. And so I asked some traders at a very successful hedge fund if you had maybe your smartest trader get on Twitch for 10 minutes once a month and on that Twitch stream describe their 30 day old trading strategies, right? So not your current ones, but the ones that are a month old. What would that affect your business after 12 months of doing that? 12 months, 10 minutes a month, 30-day lookback.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“No, I mean, look, certainly parts of the world are more efficient than others. And you can't assume equal levels of inefficiency everywhere. But I'm constantly surprised by how even in areas you expect to be very efficient, there are things that are sort of in plain sight that know, and it's not that I see them and others don't. There's lots of stuff I don't see too. I was talking to some traders at a hedge fund recently.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, I mean, it turned out to be fractal anyway, just like all the little bits that you have to get right to do a thing and have it work. I hope we got all of them. So I think that's part of it is just not believing the world's efficient than just allowing your enthusiasm to cause you to commit to something that turns out to be a lot of work and really hard. And then you just are like stubborn and don't want to fail. And so you keep at it. I don't know. I think that's it.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And so, yeah, like that's frequently what happens is the commitment to do it is impulsive, and it's done that of enthusiasm. And then you get into it and you're like, oh my God, this is really much harder than we expected. But then you're sort of committed and you're stuck and you're going to have to get it done. Like I thought this project would be relatively straightforward. We're just going to take the data and put it up. But of course, everything is and truly 99% of the work has already been done by Dr. Seals and his team at the University of Kentucky. I am a kind of carpetbagger. I've shown up at the end here, you know, to like try to do a new piece of it.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Let myself be impulsive. And so frequently, what you do, there's this great image that I found and tweeted which said, we do these things not because they are easy, but because we thought they would be easy.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“They do that now. But yeah, I mean, I don't know. I think part of it is I just fundamentally don't believe the world is efficient. And so if I see an opportunity to do something, I don't have it. I used to, but I no longer have a reflexive reaction that says, oh, that must not be a good idea if it were a good idea someone would already be doing it. Like someone must be taking care of housing policy in California, right? Or somebody must be taking care of this or that. And so I think first I like I don't have that filter that says the world's efficient and don't bother someone's probably got it covered. And then the second thing is I kind of have learned to trust my enthusiasm. You know, this gets me in trouble too. But if I get really enthusiastic about something and that enthusiasm kind of persists, I just indulge it and just think, oh, yeah, I'm going to go. I like doing the things I'm enthusiastic about. And so I just kind of...”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I'm a little bit mystified by why people don't do more things too. Like I think First of all, I don't know. Maybe you can tell me why are more people doing things? I think most rich people are boring and they should do more cool things. So I'm hoping that”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“If you can scan one scroll and you know it works and you can generalize the technique out and it's going to work on these other scrolls, then the money, which is probably low millions, maybe only $1 million to scan the remaining scrolls will just arrive. It's too sweet of a prize for that not to happen. and kind of return on excavating the rest of the villa will be incredibly obvious too because if there are thousands more papyrus schools in there and we now have the techniques to read them then there's gold in that mud and you know it's got to be dug out and you know it's amazing how little money there is for archaeology it's you know literally for decades no one's been digging there so that's my hope is that like this this is the catalyst that you know it works somebody reads it they get a lot of glory we all get to feel great and then the diggers arrive”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“That's my personal hope for this. I always like to look for these cheap leveraged hacks, these moments where you can do a relatively small thing and it creates, you kick a pebble and you get an avalanche. The theory is, and Grant shares this theory, but the theory is that if you can read one scroll, just one scroll, we only have two scanned scrolls. There's hundreds of surviving scrolls. It's relatively expensive to book a particle accelerator.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I think they will. I think if we just let them do it, they'll get it solved. It might take a little bit longer because it's not. Huge number of people, and there is a big search space here, but I mean, yeah, if we didn't launch this contest, I'd still think this would get solved. But it might take several years, and I think this way it's likely to happen this year.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Or voxels, I guess, per across an entire layer of papyrus. That's probably enough. And we've also seen with the machine learning models, Dr. Seals has got some PhD students who have actually demonstrated this at 8 microns. So I think that the ink recognitional work, I think the data is in it. The data is clearly physically in the scrolls, right? The ink was carbonized, the papyrus was carbonized, but not as like a lot of data actually physically survived. And then the question is, did the data make it into the scans? And I think that's very likely based on the results that have seen so far. And so I think it's just about a smart person solving this or a smart group of people or just a dogged group of people who do a lot of manual work that could also, you know, be true. You may have to be smart and dogged.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Can be done I think it can be done because we've recognized ink from a CT scan on the fragments. And I think everything else is probably geometry and computer vision. The scans are very high resolution. So they're eight microns, eight micrometers, and they're taken, if you kind of stood a scroll on N like this, they're taken in these slices through it, right? Like this. So it's like in the z-axis from bottom to top, there are these slices. And the way they're represented on disk is each slice is a TIFF file. And for the full scrolls, each slice is like 100, 100 something megabytes. So they're quite high resolution. And then if you stack, for example, 100 of these, they're eight microns, right? So 100 of these is 0.8 millimeters. Millimeter is pretty small. So we think the resolution is good enough or at least right on the edge of good enough that it should be possible. There's sort of like seem to be six or eight picks.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“I think this is a very legitimate concern. They're different. When you have a broken off fragment, there's air above the ink. So when you CT scan it, you have kind of ink next to air. Inside of a rapt scroll the ink might be next to papyrus, right? Because it's pushing up against the next layer and your model has to, your model may not know what to do with that. So, yeah, I think this is one of the challenges and sort of how you take these models that were trained on fragments and translate them to the slightly different environment. But maybe there's parts of the scroll where there is air on the inside. We know that to be true. You can sort of see that here. And so I think it should at least partly work if and clever people can probably figure out how to make it completely work.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“We talked about that, but I think the What we basically concluded was the search space of different ways you could solve this is pretty big. And we just wanted to get it done as quickly as possible. So having a contest means lots of people are going to try lots of things and someone's going to figure it out quickly. Mini's may make it shallow as a task. And so I think that's the main thing. Like probably someone could do it, but I think this will just be a lot more efficient. And it's fun too. I think this is fun. I think it's interesting to do a contest. Who knows who will solve it or how? People may, they may not even use machine learning. We think that's the most likely approach for recognizing the ink, but they may find some other approach that we haven't thought of.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Now you've got some ground truth. Then you do a CT scan of that broken off fragment. And that turned out to work in the case of the fragments. Okay, so now I think this is sort of the why now. This is why I think launching this challenge now is the right time because we have a lot of reasons to believe it can work. Like, and the core techniques, the core pieces have been demonstrated. It just all has to be put together at the scale of these really complicated scrolls. And so, yeah, I think if you can do the segmentation, which is probably a lot of work, maybe there's some way to automate it. And then you can figure out how to apply these models inside the body of a scroll and not just to these fragments, then it seems like you could probably read lots of text.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Rays. And, you know, you look, I look at the x-ray scans and I cannot, at least in any of the renderings that we've seen, I can't see the ink, but the machine learning model can pick up on sort of very subtle patterns in the x-ray absorption at high resolution inside these volumes in order to identify ink. And we've seen that. And so you might ask, okay, how do you train a model to do that? Because you need some kind of ground truth data to train the model. So the big insight that they had was to train on broken off fragments of the papyrus. So as people tried to open these over the years, you know, in Italy, they destroyed many of them, but they saved some of the pieces that broke off. And on some of those pieces, you can kind of see lettering. And if you take an infrared image of the fragment, then you can really see the lettering pretty well in some cases. And so they think it's 930 nanometers. They take this little infrared image.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“The scrolls are just real messed up. They were long and tightly wound, highly distorted by the volcanic mud, which not only heated them but deformed, partly deformed them. And so just the segmentation problem of identifying each of these layers throughout the school is, you know, it's doable, but it's hard. Those are a couple of challenges. And then the other challenge, of course, is just getting access to scrolls and taking them to a particle accelerator. So you have to have scroll access and particle accelerator access and time on those. It's expensive and difficult. Dr. Seals did the hard work of making all that happen. And so the good news is, very recently, just in the last couple of months, his lab has demonstrated with a convolutional neural network the ability to actually recognize ink inside these extra.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Of techniques, and then read the contents of it. And it turned out to be, I think, an early part of the book of Leviticus, you know, something of the Old Testament or the Torah. And that was like a landmark achievement. And so then the next idea was to apply those same techniques to this case. And so, okay, this is proven hard. I think there's a couple things that make it difficult. One is that the primary one is that the ink used on the Herculaneum papyri, it is not very absorbent of X-ray. It basically seems to be equally absorbent of X-ray as the papyrus, or very close, certainly not perfectly. And so you don't have this nice bright lettering that shows up kind of on your tomographic 3D X-ray. So you have to somehow develop new techniques for finding the ink in there. So that sort of problem one and been a major challenge. And then the second problem is...”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“It's called the Engetty Scroll, and it was carbonized actually under slightly similar circumstances. I think there was a temple that was burned. The papyrus scroll was in a box, so it kind of, it's like a Dutch oven. It kind of carbonized in the same way. And so it was not openable. It just fall apart. And so the question was, could you non-destructively read the contents of it? And so he did this 3D x-ray, the CT scan of the scroll, and then was able to do two things. First, the ink gave a great x-ray signature. And so it looked very different from the papyrus. There was a high contrast. And then second, he was able to segment the wines of the scroll, you know, throughout the entire body of the scroll and identify each layer and then just geometrically unroll it using fairly normal flattening computer vision.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Dr. Seals actually pioneered this field of what he calls and is now widely called virtual unwrapping. And he did it actually not with these Herculaneum scrolls. These things are like expert mode. They're so difficult. I'll tell you why soon. But he did it initially with a scroll that was found in the Dead Sea in Israel.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“My understanding. And they can tell from the style of writing, they can date, you know, they can date some of these scrolls. And so there is some old stuff in there and the library of Alexandria was burned 80 or 90 years prior. And so again, maybe wishful thinking, but there's some rumors that some of those scrolls were evacuated and maybe some of them would have ended up at this substantial, prominent Mediterranean villa, God knows what would be in there. That would be really cool. I think it'd be great to find literature. Personally, I think that would be exciting, like beautiful new poems or stories. We just don't have a ton because so little survives. And so I think that would be fun. I think you had the best crazy idea for what could be in there, which was text, which was GPT watermarked. That would be a creepy feeling.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so what could be in there? I don't know. You and I have. Speculated about this. Well, I think it would be extremely exciting not to just get more Epicurean philosophy, although that's fine too, but almost anything would be interesting and additive. Dreams are, I think it would maybe have a big impact to find something about early Christianity, like a contemporaneous mention of early Christianity. Maybe there'd be something that the church wouldn't want. That would be exciting to me. Maybe there'd be something some color or detail from someone commenting on Christianity or Jesus. I think that would be a very big deal. We have no such things as far as I know. Other things that would be cool would be old stuff, like even older stuff. So there were several scrolls already found in there that they know were hundreds of years old when the villa was buried. So the villa was probably constructed about 100 years prior.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source
“And so we actually tried to replicate many of the heartbreaking 1700s, 18th century unrolling techniques like they used rosewater, for example, or they tried to use different oils to soften it and unroll it. And most of them are just very destructive. They poured mercury into it because they thought mercury would slip between the layers potentially. So yeah, this is sort of what they look like. They shrink and they turn to ash.”
2023-03-22 · Dwarkesh Podcast · Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI · IDENTIFIED FROM THE TRANSCRIPT · source