YouSaid · the spoken record
John Schulman
- lines on the record
- 95
- first
- 2024-05-15
- most recent
- 2024-05-15
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“And it knows everything I've done, it's proactively suggesting things for me to try or it's going and doing work in the background.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I would expect something like the mental model of a helpful assistant or helpful colleague to become more real where you can share more of your everyday work or have it like instead of just giving it one-off queries you would have a whole project that you're doing and it knows about everything you've done on that project so far. You can tell it can like even proactively make suggestions like maybe you can tell it oh yeah like remember to ask me about this and if I've made any progress on it so I think like proactivity is one thing that's been missing Yeah, I'd really love to see better, like moving away from sort of one-off queries, like using the model kind of like a search engine as part of search engine and more towards having a whole project that I'm doing in collaboration with the model.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Would definitely expect things to move in that direction. It's unclear what's going to be the best form factor, whether it's like something that's like a clippy that's on your computer and helping you with something or if it's more like helpful colleague in the cloud. So we'll see which kinds of form factors work the best. And I would expect people to try all of them out. Yeah, I would expect more like.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“You don't exactly need that because you can get quite a bit out of generalization. So if you like the base model has already been trained on tons of documentation, tons of code with shell scripts and so forth. So it's already seen all the FFM peg man pages and lots of bash scripts and everything. So like the base, even just giving the base model a good FewShot prompt, you can get it to answer queries like this and just training a preference model like for helpfulness will Even if you don't train it on, probably even if you don't train it on any STEM, it'll somewhat generalize to STEM. And So, not only do you not need examples of how to use FFMPEG, you might not even need anything with programming to get some reasonable behavior in the programming domain.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, yeah, and I'd say there have been times when we needed to hire different experts for some of our campaigns. Some of the people are very, some of them are very talented and like we even find that they're like at least as good as us, the researchers at doing these tasks and they're like much more careful than us. So I would say I would say the people we have now are quite skilled and conscientious.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I would say it varies a lot. So we've definitely hired Raiders with different skills. for different kinds of tasks or projects. So I would say a decent mental model is just look at people who are on Upwork and other platforms like that, like who's doing sort of odd jobs with remote work. Yeah, it's a pretty international group, though. There's a decent number of people in the US. We hire different people of people for different types of labeling, like whether we're more focused on writing or STEM tasks. So people doing STEM tasks are more likely to be in India or other sort of middle or lower middle income countries, whereas people doing more like English writing and composition tend more to be like US-based.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“It is somewhat possible to copy or to spin up more of these efforts. There's also one force that sort of makes it less of a mode is that you can distill the models or you can take someone else's model and clone the outputs or you can use someone else's model as a judge to like do comparisons. So I think the more big league people probably aren't doing that because it goes against terms of service policies. And it would also be sort of hit to their pride. But I would expect some of the smaller players are doing that to get off the ground.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Think there's something of a moat because it's just a very complex operation and there's, so it takes, you have to have a lot of skilled people doing it. And so there's a lot of tacit knowledge. There's a lot of organizational knowledge that's required. So I think post training, like to create a model that actually has all the functionality people care about is a pretty complicated, requires a pretty complicated effort. And this requires a lot of, this is basically an accumulation of a lot of R&D. So I would say that makes it somewhat of a moat that it's not trivial to spin this up immediately. Does seem like the same companies that are putting together the most serious pre-training efforts are also putting together the serious post-training efforts. So it seems like”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“They'll just learn from all the pre training data what people would find useful and helpful and what they'll have like some, there'll be some complex moral theories that they can have and they can, but of course there's still a lot of room to latch onto a different style or a different morality. So I think when we have. Like, if we were to write a doc, or if we're going to align these models, what we're doing is latching onto a specific style, a specific morality. And there's still, you still need a decently long document to capture exactly what you want.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, that's a good question. So I think these preference models do learn a lot of subtleties of subtleties about what people prefer would be hard to articulate in an instruction manual. Maybe if you. Obviously, you can write an instruction manual that has lots of examples of comparisons. And that's what the model spec has. It has a lot of examples with some explanation. So it's not clear what the optimal format is for describing preferences. I would guess that whatever you can get out of like a big data set that captures fuzzy preferences, you can distill it down to a smaller, a shorter document that mostly captures the ideas. And I would think that the bigger models are like they do learn a lot of these concepts automatically of what people might find. Like they'll have some”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“There might be some biases in the labeling that lead to verbosity, like the fact that we tend to train for one message at a time rather than the full interaction. If you only see one message, then something that just has like a clarifying question or maybe a short response with an invitation to follow up is going to be less complete than something that covers all possibilities. There's also a question of what people, whether people's preferences would change depending on how fast the model is streaming its output. Like clearly if you're sitting there waiting for it to waiting for the tokens to come out, you're going to prefer that it gets to the point. But if it just gives you a”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“That might account for some of the convergence, but also I think some of the things we're seeing are just what people like. I mean, I think people do like bullet points. They like the structured responses. People do often like the big info dumps that they get from the models. So yeah, I think there's, so it's not completely clear how much is just a quirk of the particular like choices and design of the post training processes and how much is actually intrinsic to what people actually want.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I don't know if it rubbed off on me from the model or what. But actually, I think there's also, there might be some funny effects going on where there's like unintentional distillation happening between the language model providers where like if you hire someone to go do a labeling task they might just be feeding feeding it into a model they might just be pulling up their favorite chat pod and feeding it in and having the model do the task and then copying pasting it back so there might be uh”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I would say there's a decent amount of room for variation in exactly how you do the training process. And I think we have a lot of, I'd say we're actively trying to improve this and make the writing more lively and more fun. And I think we've made some progress like improving the personality of ChatGPT. So it is more fun. And it's better when you're trying to chit chat with it and so forth. It's less robotic. I would say, yes, it's a kind of interesting question how some of the ticks came about, like the word delve. I've actually caught myself using the word a bit recently.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, well, definitely there's always progress in improving the efficiency. Whenever you have a 1D performance metric, you're going to find that different improvements can kind of substitute for each other. So you might find that post-training and pre-training both improve the metrics or like improve they'll have a different slightly different profile of which metrics they improve. But if at the end of the day you have a single number, they're both going to substitute for each. Somewhat. So I would say for something like a human evaluation, like what do humans prefer, we've definitely made a lot of progress on both sides, like pre-training and post-training and improving that.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Be really excited to see more research using base models to do simulated social science because these models have a probabilistic model of the whole world and you can set up like a simulated questionnaire or like a conversation and you can look at how anything is correlated, like any traits that you might imagine. You can see how they might be correlated with other traits. So it'd be pretty cool to see if people could replicate some of the more notable results in the social science like moral foundations and that sort of thing. Just like prompting base models in different ways and seeing what's correlated”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Work a lot. I guess there's also various incentives that there's various unfavorable incentives like, yeah, people are incentivized to make the baseline methods, like the methods they're comparing to worse. And there are other mild pathologies like trying to make your method seem sophisticated mathematically. But I would say overall, I feel like the field makes progress. And I would probably like to see a little bit more science and trying to understand things rather than more like hill climbing on benchmarks. Trying to propose new methods. And there's been a decent amount of that recently. But yeah, I think we could use more of that. And I think that's a good thing for academics to work on. Oh yeah, on the social sciences, on a slightly different note, I think actually.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Everyone has their complaints about the ML literature, but I would say overall, I think it's a relatively healthy field compared to some other ones like in the social sciences, just because, well, it's grounded, it's largely grounded in practicality and getting things to work. If you publish something that can't be replicated easily, then people will just forget about it. And it's accepted that often you don't just report someone's number from their paper. You also try to re-implement their method and compare it to your method on the same, say, on the same training data set. So I think if you publish methods that are really hard to implement,”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“We wanted it to be very actionable so that it wasn't just a bunch of nice sounding principles, but it was like each example kind of tells you something about some non-obvious situation and reasons through that situation.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, we have these different stakeholders. Sometimes they have conflicting demands and we have to make some call on how to resolve those conflicts. And it's not always obvious how to do that. So I would say we had to think through we just had to think through the trade-offs and basically the like the rough heuristic is that we mostly want the models to follow your instructions and be helpful to the user and the developer. But when this impinges on other people's happiness or way of life, this becomes a problem and we have to block certain kinds of usage. But we don't want to be too, we mostly want the models to just be an extension of people's will and do what they say. We don't want to be too paternalistic. We want to be kind of neutral.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“We don't want the models to expose us to legal risk and so forth. And then the rest of humanity, including people not part of the who might not be users or customers or anything. So obviously the user might ask. Ask the model to do something that we think is actively harmful to other people. And so we might have to refuse that. By the way, this isn't the order of priority necessarily. So this is just like we have these four or so classes of stakeholder. Actually, you could also say maybe in the future we'll say the model itself. So I would say we're not going there yet. But anyway, they.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Say we would need to make compromises between the needs of the different stakeholders involved. So we have this document that we're releasing called the Model Spec. And it's about how we want our models to behave in the API and in ChatGPT. And we sort of try to talk about this issue where there are different stakeholders involved. And sometimes there are conflicts between what they might want. like the In our case we were thinking of the stakeholders as the user or the end user that means like someone sitting in front of ChatGPT or some other app the developer so this is like someone using the API who might be serving other end users with their app like the platform which is OpenAI like”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“If the models are being used for these higher stakes use cases, then we would have to think about RLHF in a much different way than we are right now. So I would say we're not quite ready for that or the current methods might not be completely sufficient, but I would say.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“We can like they're better at being accountable to people than people are, then I would say maybe it's okay having the AIs run the firms. But I think that might be pretty far out. And I think we're more likely to be in a situation where they look better, like in the short term, but they still have some problem like the AI run entities still have some serious problems. And it's actually practical considerations that push you more towards having humans in the loop, at least for the near future.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, for example, there's some questions, are we actually confident that AI run companies are better in every way? Or do we think they're better most of the time, but occasionally they malfunction because AIs are still like, there's still less sample efficient in certain ways, like dealing with very wacky situations. So actually AI run firms have higher tail risk because they're more likely to malfunction in a big way. So I guess there might be some question, practical questions like that that would also determine how things play off, like play out, like maybe if you just require people to be accountable for various liability, this would also change the incentives a bit. So if it turned out that AIs are better at running everything and they're also completely benevolent and we've like totally solved alignment.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“How do Right Yeah, you would either have to have every country agree to this regulatory regime or you would need every model infrastructure or the model providers to agree to this kind of requirement. So it's definitely going to be non-trivial. I guess. Yeah, this is looking a ways ahead. So it's a little hard to imagine this world before seeing anything like it.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Right. So I think if we wanted to keep humans in the loop, which seems reasonable, and it turned out that firms with any humans in the loop were out-competed by firms that didn't have any humans, then I think then you would obviously need some kind of regulation that disallowed having no humans in the loop for running a whole company.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I think, well, we might not want to jump to having AI's rent hole firms immediately. I mean, we might want to have people overseeing important decisions and calling the shots. So even if the models are good enough to actually run a successful business themselves. Yeah, to some extent there might be choices there. And I think people will still have different interests and what they want, different ideas for what kind of interesting pursuits they want to direct their AIs at and like they can people could like. Do a lot of AI doesn't necessarily have an intrinsic, any kind of intrinsic desire of its own. We put it in the system. So I think. So people can still end up being, even if AIs become extremely capable, I would hope that people are still the drivers of what the AIs end up doing.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I would expect it to be used for more kind of technically sophisticated tasks. Like I gave the programming example earlier of doing like longer projects, but also helping with various kinds of research. So I would hope that we can use AI to accelerate science in various ways. because you can potentially have the models like understand all of the literature in a given field and be able to like Able to sift through tons of data more than a person would have patients to do. So I would hope that we can basically like. Yeah. Well, I hope the form factor would basically be that people are still driving all of this and you have your helpful assistance that you can use, you can sort of direct and point to lots of different problems that are useful to you. And everyone sort of has all these AIs helping them do more get more done.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, I would expect Expect things like okay, new modalities to be added over time or pretty soon. Yeah, I would expect the capabilities to generally keep getting better through a combination of pre-training and post-training. And that'll open up new use cases. So right now AI is still not a huge part of the economy. Like there's a pretty small fraction of jobs that it can help with at all. So I'd expect that to be higher over time. And not just from the models improving, also from people just figuring out how to integrate them into different processes. So even if we just froze the models at their current state, I think you would still see a lot of growth in how they're being used. So I would expect there to be a lot of like I would expect AI to be used much more widely.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Actually, anyway, I said something slightly wrong. But anyway, yeah, you can imagine something like that and just having a bigger model gives you more chances to get the right function. So that would. And then, of course, it's not just like you have a bunch of totally disjoint functions that have, you're taking a linear combination of. It's more like a library where you might chain the functions together in some way. So there's some composability. So, yeah, so I would just say there's like the bigger model has a bigger library of different computations, including lots of stuff that's kind of dormant and only being used some of the time. But those things, but it has like more space to look for the circuits to do something useful.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“You could argue that you're learning all these things in parallel, you're learning all these different computations in parallel and you just have more of them with a bigger model. So you have more chance that one of them is lucky and ends up having high winning, guessing correctly a lot and getting upweighted. So that's kind of like a. What would be the, yeah, there's some algorithms that work this way, like mixture, what is it, mixture, some kind of mixture model or multiplicative weight update algorithm. Yeah, there's some algorithms that kind of work like this. where you have like a some kind of mixture of i don't want to say mixture of experts because it means something different but uh like basically a weighted combination of experts with some learned gating uh and uh”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Like you could say that the model is sort of an ensemble of a bunch of different circuits that do the computation. So it has like, you could imagine that it's a bunch of computations that it's doing in parallel and it's doing some like the output is a weighted combination of them. And if you have more just width of, I mean, actually width is somewhat similar to depth because like with residual networks, you end up like the depth can do something similar to width in terms of like updating what's in the residual stream.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't think anyone has a good answer for a good explanation of the scaling law with parameter count. I mean, there are some. I don't even know what the best sort of mental model is for this. Clearly you have more capacity if you have a bigger model. But so you should be able to eventually get lower loss. But I guess why are bigger models more sample efficient? I guess I can give you some very sketchy explanations. Like they have.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Right, you might not be able to conclude that if transfer fails at GPD2 size, then it's also going to fail at a higher scale. So it might be that like for the smaller models, you learn these better shared representations, or the smaller models have to lean too much on memorization, whereas the larger models can learn how to do the right computation. So I would expect this to be true to some extent.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I would say it's pretty hard to do science on this type of question because you can't do that create that many pre-trained models. So maybe you can't train a GPD four sized model. You can't do ablation studies at GPD four scale. Maybe you can do like train a ton of GPD two size models or maybe even a GPD three-size model with different data blends and see what you get. So I'm not like aware of any results or public like public results on like ablations involving code data and reasoning performance and so forth. So that would be very interested to know about those results.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, okay. Yeah, I'll try to respond to all of that. So first, are we about to hit the data wall? I mean, I wouldn't draw too much from the... Since GPD4 was released because I mean, it does a while to train these models and to get all the, do all the prep to train a new model generation of models. So yeah, I wouldn't draw too much from that fact. I would say there are definitely some challenges from the limited amount of data. I wouldn't expect us to immediately hit the data wall, but I would expect the nature of pre-training to somewhat change over time as we get closer to it. In terms of generalization from different types of pre-training data,”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I'd say I just have a decent amount of experience at this point from the different parts of the stack from RL algorithms, obviously since I've worked on those since grad school to the data collection, the annotation process to Like language playing with language models. So, I mean, I'd say I just dabbled with these things. I'd say the people who do well at this kind of research have some view of the whole stack and have a lot of curiosity about the different parts of it and also sort of think about. You want to be both empirical and use-led experiments, update your views, but you also want to think from first principles somewhat like. Like assuming that learning works, like what would be the ideal type of data to collect and that sort of thing”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“So, there are a lot of different separate axes for improvement. Like you can, yeah, so we think about data quality, data quantity, just doing more iterations of the whole process of deploying and collecting new data and like changing what kind of annotations you're collecting. So there's a lot of things that stack up. But together they give you a pretty good effective compute increase.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, there are some arguments for that. I mean, right now it's a pretty lopsided ratio, but you could argue that the output generated by the model is like high quality compared to or higher quality than most of what's on the web. So it sort of makes more sense for the model to think by itself instead of just like training to imitate what's on the web. So I think there's a first principles argument for that. I would say we found a lot of gains through post training. So I'm not sure. So I would expect us to keep. Like pushing this methodology and probably increasing the amount of compute we put into it.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“I would say faster than I would have expected since GPT 2 I was pretty bought into scaling and pre-training and so forth being a good idea. But when GBD2 was done, I would say I wasn't completely sold on it being revolutionizing everything. Like I only really pivoted what I was working on and what my team was working on after GPD3. So after that, we kind of got together and said, oh yeah, let's This language model stuff works really well. Let's see what we can do here. But yeah, after GPD2, I wasn't quite sure yet.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“We also had another instruction following model trained with RL that was released a little before ChatGPT. So I think if you put a chat like wrapper on that, you would get something decently close. But that model, like if you just prompted it with chat. But that model had some differences in strengths. That model was pretty good at writing and poetry and so forth, but it wasn't. It wasn't as good at knowing its limitations and at factuality and so forth.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“You'd want to do some kind of iterative supervised fine tuning where you have humans edit the model generated outputs because it's really hard to get people to, like if you train on human generated data, even if it's really high quality, it's just hard for a model to fit that data perfectly because it might not be like it might not be something a model is capable of outputting. So you need to do something iterative that looks a little bit more like RL. So I think if you had done that, you could have gotten something pretty close, but that would have been kind of non trivial.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Not exactly. I mean, they could have, I don't remember the status of which models were available for fine-tuning. Assuming we had 3.5 available for fine-tuning at the time, you could have made something pretty decently close. But I'm not sure you would have, I don't think you would have been able to do just one iteration of fine-tuning where you have purely human written data and you fine-tune that. I think you would want to do several iterations. Like if you're not going to do RL, which we did.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“We're clearly more like it was easier to use. It was sort of more, it sort of like automatically had much more sensible behavior in terms of the model knowing its own limitations. That was actually one of the things that I got excited about as we were developing it. I realized a lot of the things that people thought were flaws in language models, like just blatantly hallucinating could be not completely fixed, but you could make a lot of progress with pretty straightforward methods. Oh, yeah, and also the other thing about chat was that when we had these instruct models, like the For people to get the idea of what the model was supposed to do. And so that, so as a result, I think the model had a much more coherent personality. And it was much easier to get pretty sensible behavior robustly.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, those models were really good and everyone got really excited about that after seeing the instruct fine-tune GPD4s. But so they were really, really good. They would occasionally give you amazing outputs, but they were also like a little bit, the model was clearly like pretty unreliable. Like you would sometimes hallucinate it a lot. And it was sometimes give you pretty unhinged outputs. So it was clearly not quite ready for primetime, but it was like obviously very good. And yeah, so I guess that people forgot about chat for a little while after that, about this alternative branch. But then we ended up, we pushed it further and we ended up mixing together all the data sets, like the instruct and the chat data and to try to get something that was the best of both worlds. And I think the models we, the chat models were like.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Like deemphasizing that later on because the model's internal knowledge was so good that the browsing wasn't the most interesting thing about it. And then we were thinking about, we had it out for beta testing to friends and family for a while and we were thinking about doing a public release. But at that time, actually GPD4 finished training in August or August that year. Actually, the flagship RL effort at OpenAI was the instruction following effort because that was the models that were being deployed into production. So like the first fine tunes of GPD4 used that whole stack. And that was.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“When you do question answering, it really wants to be in a chat because you always want to ask follow-up questions or sometimes you need a clarified, the model should ask a clarifying question because the question is ambiguous. So it was kind of clear after we did the first version of that that we should, the next version should be conversational. So anyway, we started working on a conversational chat assistant. This was built on top of GPD 3.5, which was done training at the beginning of 2022. And that model was quite good at language and code. So we quickly realized that it was actually quite good at coding help. And that was one of the things we were excited about. So yeah, we worked on that. We worked on that for most of the year. And we had browsing as another feature. And though we ended up”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Generation of models. Then at the same time, there were definitely a lot of people thinking about chat. So Google had some papers, like they had Lambda and earlier MENA. So they had these chat bots and it was more like you was more like a base model that was really specialized to the task of chat, really good at chat and like I think at least looking at the examples from the paper, it was more used for sort of fun applications like where the model would take on some persona and pretend to be that persona. It was not so functional, like help me refactor my code. So, yeah, there are definitely people thinking about chat. I had worked on a project before looking at chat called WebGPT, which was more about doing question answering with the help of web browsing and retrieval.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source
“Yeah, so early ChatGPT, we had OpenAI had these instruction following models. And the idea there was we had base models and people can prompt them in elaborate ways, but they're also kind of hard to prompt. You had to, they basically do autocomplete. So you have to set up a very good prompt with some examples. People at OpenAI were working on just taking the base models and making them easier to prompt so that if you just wrote a question, it would answer the question instead of giving you more questions or something. So we had these instruction following models, which were kind of like base models, but a little easier to use. And those are the original ones deployed in the API, or after GPD-3, those were the next.”
2024-05-15 · Dwarkesh Podcast · John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI · IDENTIFIED FROM THE TRANSCRIPT · source