YouSaid · the spoken record

Mark Chen

lines on the record
77
first
2025-09-25
most recent
2025-09-25
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. Maybe sometimes for completeness, you would go and actually do all of the mechanics of coding it from scratch yourself, but that's just a strange concept to them. Like, why would you do that? You just vibe code by default. And so, yeah, I mean, I do think, you know, the future hopefully will be vibe researching.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  2. You kind of imagine or like to think that you feel a set of the feelings that Lisa all felt too, right? It's like, wow, this is really crazy, right? And what are the possibilities? And this is something that I took decades to do. And I took a lot of hard work to get to the forefront of. So you really do feel an implication of that is these models, what can't they do, right? And I do feel like already it's kind of transformed the default for coding. This past weekend, I was talking to some high schoolers and they're saying, oh, you know, actually the default way to code is vibe coding. Like, you know, I think they would consider, oh, it's like.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  3. To kind of like speak to the recital moment, I think Alfa go for both of us was a very formative milestone in AI development. And at least for me, it was the reason I started working on this in the first place. And maybe partly because of our backgrounds in competitive programming, like I had this affinity to building these models, which could do very, very well in these forms of contests. And going from solving eighth grade math problems to a year later hitting our level of performance in these coding contests, it's crazy to see that progression.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  4. Yeah. Yeah, eventually, I think, especially with this latest coding tools like GPT-5, I've really kind of felt like, okay, this is no longer the way. You can do a 30 file refactor pretty much perfectly in 15 minutes. You kind of have to use it. Yeah. And so I've been kind of like learning this new way of coding, which definitely feels a little bit different. I think it is a little bit of an uncanny valley still right now where like you kind of have to use it because it is just like exerting so many things, but it's still a little bit not quite as good as a co-worker. So I think our priority is getting out of that in Cancune Valley. But yeah, it's definitely an interesting time.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  5. In terms of Cutting models being better. I mean, I think, yeah, I think it is extremely exciting to see this progress. I think that programming competitions have a nice kind of encapsulated test of ability to come up with some new ideas in this boxed environment and time frame. I do think if you look at things like, well, I guess the IMO problem six or maybe some very hardest programming competitions problems like I think there's still a little bit of headway to go for the models, but I wouldn't expect that to last very long. I do go a little bit. Historically, I've been like... Historically, I've actually been extremely reluctant to use any sort of tools. I just used Vim pretty much.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  6. What we found is the previous generation of the Codex models, they were spending too little time solving the hardest problems and too much time solving the easy problems. And I think that is actually just probably out of the box what you might get out of 03.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Yeah, so I think one of the big focuses of the Codex team is to just take the raw intelligence that we have from our reasoning models and make it very useful for real world coding. So a lot of the work they've done is kind of consistent with this. They are working on kind of having the model be able to handle more difficult environments. We know that real world coding is very messy. So they're trying to handle all of the intricacies there. There's a lot of coding that has to do with style, with just like kind of softer things like how proactive the model is, how lazy it is, and just being able to define in some sense like a spec for how a coding model should behave. They do a lot of very strong work there. And as you seem to like, they're also working on a lot better presets. Coders, they have some kind of notion of this is how long I'm waiting. I'm willing to wait for a particular.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  8. I expect this will evolve quite rapidly. I expect it will become simpler, right? Like I think, you know, maybe like two years ago, we would have been talking about what is the right way to craft my fine-tuning data set. And I don't think we are at the end of that evolution yet. And I think we will be inching towards more and more human-like learning, which RL is still not quite. So I think maybe the most important part of the mindset is to not assume that what is now will be forever.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  9. These different ideas and objectives in this extremely robust, rich environment given by pre-training. And so, yes, I think it's been perhaps the most exciting period in our research over the last few years where we've really found so many new directions and promising ideas that all seem to be working out. And we're trying to understand how to compare.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  10. RL is a very versatile method, right? And there are a lot of ideas you can explore once you have NRL system working a long time at OpenAI. We started from this before language models, right? Like we were thinking about, okay, like RL is extremely powerful thing. Of course, like on top of deep learning, which is that's incredible general learning method. But the thing that we struggled with for a very long time is like, what is the environment? Like, how do we actually anchor these models to the real world? Or like, should we simulate some island where they all learn to collaborate and compete? And then, you know, of course, came the language modeling breakthrough, right? And we saw that, oh, yeah, if we scale deep learning on modeling natural language, we can create models with this incredibly new understanding of human language. And so since then, we've been seeking how to combine these paradigms and how to get our all to work on natural language. And once you do, right? Like, then you kind of have the, well, you have the ability to actually execute.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  11. And I think it also makes sense to think about what the limits of open ended means. I think a while back, Sam tweeted about some of the improvements that we were making in having our models write more creatively. And we do consider the extremes here as well.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  12. But even if you want to solve a very well defined problem that is on a much longer scale, right? Prove this Millennium Prize problem. Well, that suddenly requires you to think about, okay, what are the fields of mathematics or other sciences that might possibly be relevant? Are there inspiration from physics that I must take? What is kind of the entire program that I want to develop around this? Now this becomes very open-ended questions. And it's actually hard to, you know, for our own research, like if all we cared about is reduce the modeling clause on a given data set, right? Like measuring the progress on that, like, are we kind of actually asking the right questions in research actually becomes like a fairly open-ended affair?

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  13. Oh, yeah, this is a question. I really like it. I think if you actually truly want to extend Research and discovering ideas that meaningfully advance technology on the scale of months and years. I think these questions stop being so different, right? Like it is one thing to solve like a very well posed constrained problem on the scale of an hour, right? And there's kind of a finite amount of ideas you need to look through. And that might feel extremely different from solving something very open-ended.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  14. Yeah, and I think reasoning is core to this ability to operate over a long horizon because you imagine kind of yourself solving a math problem. You try and approach it doesn't work. And you have to think about what's the next approach I'm going to take? What are the mistakes in the first approach? And then you try another thing. And the world gives you some hard feedback, right? And then you keep trying different approaches. And the ability to do that over a long period of time is reasoning and gives agents that robustness

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Actually, the ability to maintain depth is a lot of it being consistent over long horizons. So I think there are very related problems. And in fact, I think with the reasoning models we have seen the models greatly Extend the length over which they are able to reason and work reliably without going off track. Yeah, I think this remains a big area of focus for us.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  16. And actually, maybe on that topic, there's been this huge move toward agency and model development. But I think at least the state that it's in currently, users have sort of observed this trade-off between too many tools or planning hops can result in quality regressions versus something that maybe has a little bit less agency. The quality is at least observed today to be a bit higher. How do you guys think about the trade-off between stability and depth? The more steps that the model is undertaking maybe the less likely the tenth step is to be accurate versus you ask it to do one thing, it can do it very, very well. And to have it keep doing that one thing better and better, but more complex things, there's sort of that trade-off. But of course, to get to full autonomy, you are taking multiple steps. You're using multiple tools.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  17. So, the big thing that we are targeting with our research is producing an automated researcher. So automating the discovery of new ideas and of course a particular thing we think about a lot is automating our own work, automating ML research, but that can get a little bit self-referential. So we're also thinking about automating progress in other sciences. And I think one good way to measure progress there is looking at what is the time horizon on which these models actually can reason and make progress. And so now as we get to a level of near mastery of this high school competitions, let's say, I would say like we get to like maybe on the order of one to five hours of reasoning. And so we are focused on extending that horizon, both in terms of like the models, capability to plan over very long horizons and actually able to retain ability to retain memory.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  18. GPT5 is a definite improvement on me, O3 was definitely like that moment where the reasoning models became actually very useful on a daily basis, I think, especially for working through MF formula or a derivation, like it actually got to a level where it is fairly trustworthy, I can actually use it as a tool for my work. And yeah, I think it is very exciting to get to that moment, but I expect that, well, now as we're seeing these models actually able to automate, well, yes, like we're saying solving context problems over longer time horizons, I expect that that was quite small compared to what's coming over the next year. What is coming in the next one to five years? It would be just at whatever level you're comfortable sharing, what does the research roadmap look like?

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  19. I think one big thing for me was just how much it moved the frontier in very hard sciences. We would try the models with some of our friends who are professional physicists or professional mathematicians. And you already saw kind of some instances of this on Twitter where you can take a problem and have it discover maybe not like very complicated new mathematics, but some non-trivial new mathematics We see physicists, mathematicians kind of repeating this experience over and over where they're trying Kubernetes Pro and saying, wow, this is something that previous version of the models couldn't do. And it is a little bit of a light bulb moment for them. It's like able to automate maybe like what could take one of their students months of time.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  20. Which capability from GPD 5 before the release Surprise you the most when you were working through the eval bench or using it internally were there any moments where you felt like this was starting to get good enough to release because it was useful in your daily usage?

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  21. I mean, I think it is important to note that these evals, like IOI, at Coder, IMO, are actually real world markers for success in future research. I think a lot of the best researchers in the world have gone through these competitions, have gotten very good results. And yeah, I think we are kind of preparing for this frontier where we're trying to get our models to discover new things.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  22. To other things. So the way we think about it in this world, we definitely think we are in a little bit of a deficit of great evaluations. And I think the big things that we look at are actual marks of the model being able to discover new things. I think for me, the most exciting tread and actual sign of progress this year has been our model's performance in math and programming competitions. Although I think like they are also becoming saturated in a sense. And the next set of evals and milestones that we're looking at will involve actual discovery and actual movement on things that are economically relevant.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  23. One thing is that indeed for these evals that we've been using for the last few years, they're indeed pretty close to saturated. And so, yeah, for a lot of them, like, you know, inching from like 96 to 98% is not necessarily the most important thing in the world. I think another thing that's maybe even more important, but a little bit subtler. When we were in this GPT-2 GPT-3 GPT-4 era, there was kind of one recipe you just pre-train a model on a lot of data and you kind of use these evals as just kind of a yard seg of how this generalizes to different tasks. Now we have these different ways of training, in particular reinforcement learning on serious reasoning where we can pick a domain and we can really train a model to become an expert in this domain to reason very hard about it, which lets us target particular kinds of tasks, which will mean that we can get extremely good performance on some evils, but it doesn't indicate as great journalization.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  24. There's also a number of improvements across the board in this model relative to all three and the previous models. But our primary feasit for this launch was indeed bringing the reasoning mode to more people.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  25. So, I think GPT 5 was really our attempt to bring reasoning into the mainstream. And prior to GPT-5, right, we have two different series of models. You had the GPT kind of two, three, four series, which were kind of these instant response models. And then we had an O series, which essentially thought for a very long time and then gave you the best answer that it could give. So tactically, we don't want our users to be puzzled by which remote should I use. And involves a lot of research in kind of identifying what the right amount of thinking for any particular prompt looks like and taking that pain away from the user. So we think the future is about reasoning, more and more about reasoning, more and more about agents. And we think GPT-5 is this step towards delivering reasoning and more agentic behavior by default.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Thanks for coming, Jakob and Mark. Jakob, you are the chief scientist at OpenAI. Mark, you are the chief research officer at OpenAI, and you guys have both the privilege and the stress of running probably one of the most high profile research teams in AI. And so we're just really stoked to talk with you about a whole bunch of things we've been curious about, including GPD-5, which was one of the most exciting updates to come out of OpenAI in recent times. And then stepping back how you build a research team that can do not just GPD5, but Codex and ChatGPT and an API business and can weave all of the many different bets you guys have across modalities, across product form factors into one coherent research culture and story. And so to kick things off, why don't we start with GPD-5? Just tell us a little bit about the GPT-5 launch from your perspective. How did it go?

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source

  27. The big thing that we are targeting is producing an automated researcher, so automating the discovery of new ideas. The next set of evals and milestones that we're looking at will involve actual movement on things that are economically relevant.

    2025-09-25 · a16z Podcast · From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki · IDENTIFIED FROM THE TRANSCRIPT · source