← Return to Index Archived September 9, 2026
The Lead — Sep 9
THE PRAGMATIC ENGINEER · GERGELY OROSZ

Building Codex with Tibo Sottiaux

Codex lead Thibault traces the coding agent’s path from OpenAI’s infrastructure work to an open-source Rust harness, arguing that agentic software development shifts attention from code review toward intent, contracts and architecture. He also describes how cheap maintenance and re-architecture are remaking the pace of engineering inside OpenAI.

1h 13m / September 9, 2026 /aitechnologyproduct / Transcript sourced from openai
All episodes from The Pragmatic Engineer →·Listen on Apple Podcasts →

Overview

Thibault traces Codex from an internal effort to improve OpenAI's research infrastructure into an open-source coding agent built to serve external developers. The conversation focuses on how the Codex team works: why the core agent is written in Rust, how the harness and model improve together, and how agents are changing planning, reviews, maintenance, and software architecture.

He also describes the longer-term direction: coding agents that can work across local and cloud machines, understand a company's internal context, and take on more of the operational work that used to consume engineering time.

Key Takeaways

  • Codex was built with a deliberate separation between the agent core and the product interface. Thibault says Rust helped enforce that boundary while providing efficiency, correctness checks at compile time, and a strong base for a system expected to operate at large scale.

  • The team chose to open-source the CLI, SDK, and app server because a coding agent should participate in the developer community it serves. Open source also makes hiring easier: new engineers can inspect the repository and past pull requests before they join. The tradeoff is real, though. The team has to manage low-quality contributions and accepts that competitors can copy work before it ships.

  • Codex is not meant to be locked to OpenAI models. Thibault argues that supporting other model providers avoids pointless forks and gives developers room to switch models as capabilities change.

  • The harness is often "ahead" of the model. It supplies tools, instructions, safety controls, and behavior scaffolding that make a model useful before the model can reliably handle those behaviors on its own. As training improves, the system prompt and harness can shrink because the model needs fewer reminders and workarounds.

  • Product and research decisions are made jointly. When Codex struggles with a task, the team decides whether to fix it in the harness now or wait for a model improvement expected soon. They use agents to analyze feedback across coding and other work domains, then identify recurring failure modes.

  • Code review is moving toward review of intent, interfaces, and invariants rather than line-by-line inspection. If a component has clear guarantees around security, data access, resource use, and behavior, engineers can spend less attention on its internal implementation.

  • Maintenance and re-architecture are becoming far cheaper. Dependency upgrades, security patches, and broad codebase changes can increasingly be automated. That makes good boundaries and abstractions more valuable, since teams can revise what is inside a component without destabilizing everything around it.

Practical Steps

  • Keep a clear boundary between your agent logic, UI, and product-specific integrations. Choose interfaces that let each part evolve independently.

  • Treat agent behavior problems as a diagnosis exercise. Ask whether a failure needs better tools, clearer instructions, stronger guardrails, or a better model. Do not build a large workaround if an upcoming model capability is likely to remove the need for it.

  • Shift review earlier in the process. Before generating a large change, agree on the component's purpose, inputs, outputs, security constraints, and resource limits. Let automated checks handle more of the implementation detail.

  • Use agents for maintenance work you would otherwise postpone: dependency updates, changelog review, security patches, test additions, codebase exploration, and migration planning.

  • Make internal knowledge searchable by default. Broad access to project documents, decisions, public team channels, and code history gives agents enough context to answer onboarding and operational questions.

  • For engineers aiming at AI product teams, practice getting oriented quickly in unfamiliar systems. Ask good questions, follow the "five whys," and stay close to the people who will use what you build.

Notable Quotes

  • "Have you asked Codex?" - Thibault

  • "The harness, in a sense, is always a little bit ahead of the model." - Thibault

  • "What you need to agree on is, like, what does the box actually do and what are the invariants that must be satisfied." - Thibault

The cost of mistakes, I would say, is going down, but the good old rules of software engineering, like having good abstractions, really help. — From the episode

Full Transcript

Source: openai 1h 13m runtime

Codex is one of the most popular AI coding harnesses today, but how did it all start? Many of you will know today's guest, Thibault, from his generous and pretty frequent codex usage resets. He was also there when Codex as a product started and has led the broader codex team since. Today we cover how Codex started and why it was built in Rust and made open source. How code reviews are changing inside the Codex team and OpenAI. What it means when maintenance and re-architecting are getting ridiculously cheap. What the merge of Codex into ChantGPT looked like and the many underappreciated engineering challenges of this project. If you want to understand how teams inside of OpenAI plan, review and ship software, this episode is for you. This episode is presented by TurboBuffer, a ridiculously scalable, fast and cheap hybrid search engine built on top of object storage by an engineering team that I've really grown to like after spending time with them. TurboBuffer is the tool that companies like Entropic, Notion, Cognition and Harvey all use to connect their AI products to massive amounts of unstructured data. When I've talked with engineers who use TurboBuffer, the theme that always comes up is reliability and performance at scale. The reasons for this have everything to do with TurboBuffer's architecture. TurboBuffer uses only object storage for state and NVMe SSDs with memory cache for compute. Data in TurboBuffer is organized into namespaces. You can think of a namespace as a database table or a search index or an S3 prefix depending on the world you come from. When a namespace is not being queried, it stays on cheap object storage with no associated compute cost. When a namespace is active, TurboBuffer pulls it up into hot caching tiers so queries are very fast. This design fundamentally makes it effortless to scale to hundreds of millions of namespaces. If you're building a multi-tenant AI product, every user and their agent can have their own dedicated search index without any overhead. And each namespace can hold hundreds of millions of documents without any special configuration. You can scale TurboBuffer virtually without limit, and the performance, reliability, and operating model all stay the same. If you need to connect AI to lots of data, TurboBuffer should be your first choice. Check it out at TurboBuffer.com slash pragmatic. Thibaut, welcome to the podcast. So good to have you here. Thank you for having me. It's so good to see you again. It's good to do this last time we did it in person, now we're doing it in video. And I wanted to ask you, how did you get into tech? When did you first know that you want to work with computers? I think it's a good question. It was a long, long time ago. My parents actually decided to move out of Brussels, where I was born, and just thought it was great to just buy a small house and refurbish it, but it was in the middle of a village with not much going on. I think there was like roughly 200 people living there. Not many that I felt like I wanted to talk to or could make friends with. And so I kind of got stuck. This is like very early, like eight years old, I kind of got stuck as like, you know, computers and, you know, it's like early days of like the, for me, the internet. And you know, that was my way to learn about things. And so just the rest is just like, you know, came from that. Sort of like I owe it to my parents to, you know, have moved into the middle of nowhere. And then, you know, I had no choice but to get interested in computers. Once you finished high school, like you went on and went to university, right? Actually studying it properly. Yes. I studied mathematics, applied mathematics at university. I went there quite early. And so I graduated early as well. I thought for a long time that I would actually not make it and that I would drop out. I had just like small companies and small consulting business, like while I was studying. I was like working for banks, I was working for, I was like very interested in supply chain and applied mathematics problems. And I guess sort of like selling that and learning a lot through that. Eventually ended up in the startup world in Belgium. I did that for a little while and then moved to London to work initially at Google and then DeepMind and then now, you know, moved to be here at OpenAI, like this is California. I love the California weather. We can talk about that. It's been very good. Right after university, you started, you founded a startup, right? You had the startup bug in you or the entrepreneur, entrepreneurial bug. Yeah. So this startup was all about pharmaceutical supply chain, looking at the supply chain for clinical trials and like try to optimize and decide like, hey, you know, should you produce more medicine? Where should you send it? Where should you dispatch it? How do you avoid waste? And through that, making clinical trials more efficient. And this was using traditional, like non-ML techniques, traditionally more optimization solving, Monte Carlo simulations, these kinds of things, stochastic multistage optimization problem really. And we also applied it on steel industry and we applied it to electrical grid as well in Europe. It was like anything that sort of had the shape of like an optimization problem we sort of like get interested in. And, you know, to this day, like this, this, this company still exists and I think they do some of the most interesting work still, but it's changing a lot, you know, with modern AI for sure. But it's interesting because you kind of said like, oh yeah, that wasn't ML. It was just the traditional stuff. And then you go into like Monte Carlo simulation and optimization and this algorithm. I get a sense that you kind of just went deep, right? It was like, okay, like here's a problem space. Like how can I use mathematics, stuff that I learned, stuff that I didn't learn to just go deeper and deeper. Do I sense that correctly? Yeah, that's, that's why I was obsessed with applied mathematics is just really this idea of you have theoretical mathematics or you have theoretical science and physics and like there you just, you do it because there's something to be discovered and something beautiful about it. So it's all about patterns and pushing the frontier, but you don't necessarily always know like how you're going to apply it. And then there was like the real world, right? It's like, you know, there's all these cool problems that just lie around. And I was like very interested in seeing like, you know, how can I make the world better? And so like, how do I apply like, you know, sophisticated mathematics, you know, to just optimize the world around me. And that was like a lot of the thesis behind that startup. Yeah. And then after startup, you ended up at Google and first at Google London. It was in 2015. And I remember in 2015, Google was a really, really competitive place to get into, like, maybe as competitive as open AIs today in terms of the industry or in terms of prestige. You worked on Maps initially, and then you moved over to DeepMind. Can you talk a little bit what you worked on? And then why did you move on from an already really interesting space that you clearly loved, you know, like the optimization, logistics and all these things? Yes, I, I didn't, I didn't start on Google Maps. I started on a project that was meant to make the web faster and make to meant to make websites faster, especially on mobile. At the time, you know, Google was saying, so like seeing the transition from desktop to mobile and like more and more traffic going to like, you know, mobile phones. And so like wanting to get ahead of that. So funded like a number of a number of initiatives and projects. I was working on one of them. This was like really, really fun because it was a small group within actually the ads organization. It was meant to sort of like, you know, offset the loss for the ad revenue loss because of this shift of traffic to mobile and worked on it for roughly two years. And then it was canceled. And although it was like the most fun I've had on, you know, solving hard technical challenges, I learned a lot from not having product market fit, not having the right users, not having the right feedback loop, not trusting your product manager when they say the project is going well, when in fact it's not going well at all. And then, you know, one day it's just like this VP flew in from California and then it was just like, oh, yeah, it's like, you know, we're canceling this project. You know, unfortunately, you only have, you know, hundreds of users and this is clearly not Google scale. And then it's unbelievable, but people were surprised. And I think there's a lesson there that I carry with me, of course, is, you know, just always, always question, always go to like always, you know, deeply think about the impact that you're having, but also like the importance of the overall project that you're contributing. And then I moved into Google Maps, Google Maps was super fun, worked on reviews. And then after roughly a year, I couldn't ignore like DeepMind. It was just, it was this special place, headquartered in London. So many great things were happening. This was like really the early days, you know, with rumblings of things like AlphaGo. And they just seemed to be doing extraordinary things and, you know, just really tackling the very, very hardest problems that you can tackle. And like, with my background, I was obviously drawn to that. I started there, it's like I worked on a lot of the research infrastructure, research tooling. This is a theme that I carried on for almost a decade. And it's like, this is very much also the, like how I approach things is, how can I build tooling and products that help make others more efficient and bring a lot of utility to them? Initially, I was doing this for research. And then like over time, you know, I got like into thinking about things in a much more like more general and general and general way, you know, eventually like, you know, ending up where I'm now. And a fun story that you recently shared on X as well is how you were part of the team that built this internal Google bot that was, you know, if you want to say, similar to ChatGPT, but a year before ChatGPT. Can you can you talk about that? That is a new story. I haven't heard it before. This was part of DeepMind. There were like multiple efforts as well. There was like Braid as well that was separate at the time. They had their own efforts on large language models, but it was definitely something that was being explored. It was not the main thrust of DeepMind. DeepMind was like very much worried and busy, like thinking about grand challenges and games and, you know, thinking about RL, not in the language sense. And so there was like this group that was pushing on large language models and, you know, thinking about, you know, what if what if large text corpuses are everything? What if you just pushed. Pushed out and simplified to its core, you're making a lot of, you know, we found that you could make a lot of progress very quickly and then learn very quickly. And then Greg and Sam are, you know, people who were immensely supportive. And also Greg was very adamant that, you know, we would not just focus on ourselves, but we would also focus on benefiting the world. And so he just sort of encouraged that we would be thinking about this not just as a tool for OpenAI itself, but also as something that we would actually make into a product. And this is when we merged this research effort with this API effort, and we started building one thing. And then that led to a sprint, which was like the initial Cloud Codex that we launched, which didn't really have PMF because it was like a little bit too high friction. And we also launched the Codex CLI, and we continued to push. But it was always this idea of, hey, how do we get models to really help here? You mentioned that first you started to build this model to train on the Python code and actually help me build Infra better. But then you made this interesting decision where for Codex, you built it in Rust. And at the time, the model was not on distribution for Rust, right? It wasn't as good as Rust as it was in Python or TypeScript. Why did you make that kind of a decision? Was it kind of like, did you expect that it will catch up, or you figured that performance is more important? Because it was very counterintuitive. Most of the other harnesses built were actually not built in Rust. They were built on distribution on TypeScript or Python or something else. Yes. From first principles, like we very early on, we were thinking about the product interface and the agent as different things. So it was very important to build. The core of the agent in a way that was robust, that was secure as well, that was, you know, engineered for efficiency and scale. And having worked through projects over the years that go from, hey, this is a fun thing to, like, hey, we need to scale this to the scale of, like, the largest data center, the decisions early on are, like, really turn out to be quite important, as long as you don't sacrifice too much of the velocity. And so it's like, it's a trade-off, but we had very prolific and amazing Rust developers. Our internal models were not bad at Rust. And then you get a lot of validation as well at compile time. It's like, you know, statically verified and all these things. And that is great for agents too. So turns out, you know, it was quite clear that, you know, Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it. But primarily we were focused on correctness and we were focused on efficiency as well. Interesting. So you're saying, you know, it's worth, in your case, it was worth thinking ahead of where you want this thing to be. And for example, things like a language choice. Obviously with agents, you can rewrite a bunch of stuff and easier than in the past, but it's still like you can save yourself reworking by putting in the right, I guess, scaffolding or, well, the, you know, the baseline of what you're building on, right? I think we could have been successful if we had written it in TypeScript or, you know, maybe even Python and then it would have been fine. And then, you know, we would have rewritten it at some point. But having a very clean separation between the agent itself, which can exist irrespective of the product, it was a very important principle. And if you write everything in the same codebase, in the same language, it's like inevitably you're going to be a little bit sloppy and you're going to intertwine things more than you should. And then it's going to prevent further innovation after that. And so that was very important. And, like, the Rust boundary, in a sense, like, was very useful for that. One interesting decision that you made, which is unique across all of the major labs, is having this built in open source, right? The CLI is open source, the SDK and the app server are all open source. When and why did you decide that? It's not a given, especially, you know, there used to be jokes about OpenAI having things closed, but this is actually the opposite, where, like, this is open, whereas, like, some competitors would ship closed-source harnesses, which, again, I think it's very easy to understand why you want something closed source. Why did you want it open source? There was something really cool about the idea of having the code open source because fundamentally what you're building is you're building a coding agent. And so we were sort of, like, thinking about, well, if you have that, you know, you're obviously going to point it at itself and, you know, maybe, you know, you can build a community of, you know, contributors that use it to improve it, and then, you know, you can learn a lot from that. Also, it felt at the time, it's, like, you know, very clear to us that if we were going to be successful, open source itself would change, and the role of code itself would change. And so being part of that community seemed important instead of divorced from it. I think, you know, it's hard to solve problems if you don't sort of, like, witness them yourself. And then the other thing was just it still feels like early, but it was very early at the time. It felt like we would have some ideas for how to solve things well, and we were co-designing these, you know, with the training and the research, and it's all about expressing, like, the capabilities model in, like, the most flexible and the best way. But also we didn't have all the answers. Sort of being very open about, hey, this is what a good harness looks like. This is how we think about it. We did, like, a couple of, like, very technical, like, deep dives and blog posts, and we talked about it a lot. And we thought, you know, hey, it's just like the world is vast out there. There's, like, you know, crazy smart people. It's like, you know, we're going to get inspired by other open source projects as well. And so let's just make this a level playing field and sort of, like, encourage a lot of tinkering and exploration at this stage. Now, this has been now, you know, like a year later, a year and a half later, which is a very long time right now in this AI time frame. But looking back or taking the experience, what are the benefits you've seen, the kind of engineering benefits, the engineering team's benefits from being open source? And just honestly, what are things that are kind of hard about being open source, right? Like, there must be downsides. Like, just trying to get an honest take on both sides. Yeah, there are definitely downsides. It comes at a cost, right? The benefits are it's awesome to build in the open. It's awesome to have, like, a small repo as well. Like, whenever we hire someone and they join the Codex team, it's like they've seen the repo before, they've looked at PRs. The onboarding is done. Yeah, it's done. Yeah, it's like onboarding is just like you use Codex to look at the repo, you know, with you, and you ask some questions. But it's like, it's not a secret issue that you can get productive right away. We get a lot of good contributions, although we get, like, you know, a tsunami of, like, random stuff as well. Obviously, you and everyone else, right? Open source is changing. I think this is one of the examples. That's right. And then to me, it just, and to a lot of the team, it just brings a lot of energy to just be part of the community and, like, be directly contributing, not just saying that we care about the community, but actually doing things that, you know, you can see it's costing us effort, right? We don't have to do it. The downsides are, you know, it's separate from the rest of our code. So, you know, sometimes we have to draw, like, artificial boundaries and, you know, work across multiple repos. When we're working on something particularly exciting and, you know, we're building it in the open, then, you know, at times we find that, you know, others copy it, you know, before we have the time to release it. And it's like, it's just a little bit sad. But also, it's like it's part of the game. You know, it's like you're building in the open. It's like, you know, that's sort of like the contract that you signed. It's like, you know, you can copy it. We have a very permissive license as well. But it does sting a little bit when you're working on something and you're like, you know. And then the third thing is just like everyone else. It's like, you know, we are overwhelmed with, you know, random contributions and, you know, we have to deal with that additional tax. But then that pushes us to, you know, also, like, try and solve for it, right? Which I think is good. And on top of the open source, one thing that surprised me about Codex, and I didn't even know about it until recently, it's not tied to the OpenAI models. You can use other models with Codex. You know, like putting myself in a vendor's shoe, it might not be very obvious because, again, all the other vendors I look at, when they do a CLI, it's kind of use it with our models. Again, what made you decide to be this permissive about, you know, using or allowing to use your harness with other models? It felt quite natural if you are part of this community and building an excellent coding harness, it's like, why would you couple it to your model? That felt like sort of like quite disappointing to make that decision. So it didn't feel right. And in general, it's like I think, you know, it's like I kind of try to make decisions that I'm like, yes, you know, just like I can just sort of like explain it, you know, it is correct. It's the same reasoning with, you know, it is open source in the first place. It would have been trivial for anyone to fork it and then add support for another thing. But then you're just encouraging people to just, like, you know, go and use that fork, and then now suddenly you have overhead. And the only reason you have a fork is because, you know, you wanted to change, like, ten lines of code to add support for, like, another model provider. That feels very silly. So, like, you know, why not just support it in the first place? The other thing is, we benefit a lot from, like, being able to just give optionality. So, you know, it's like maybe today, you know, you love using OpenAI models, and, you know, you're super productive with them. But, like, tomorrow there's a new model that comes out, you want to try that. Why force you to go and completely change? More, and this is like just a step. It's going to be much more seamless in the future to, like, you know, use cloud machines and then, you know, maybe have a combination of, like, partial execution on your laptop, partial execution on cloud machines. And really the thing that we're thinking about that is very natural is as models just get better and more capable, they can leverage so much more compute and many more resources than are available on your local machine. And so it would be a constraint at some point to just limit execution on your local machine. But one thing that is great about it running locally, and I think the reason I love it when it runs locally, of course it's a pain because if I'm doing some work, it's like, you know, I have several agents, it's eating CPU. If I want to close my laptop, I cannot kind of leave it like half open, right? When I was in one of the offices of an AI company, I had it half open and they're like, Are you running agents? I'm like, Yeah, I have one running. I get it. But the reason I do it is because I have my local tools. I have my local Postgres database. I have my this, this, and that. How are you thinking about—the cloud is amazing, but it doesn't have this setup, or it's just a pain to set it up. Are you thinking or are you experimenting with, you know, making these setups? And I'm kind of reminded of a topic that we talked about pre-AI, which is cloud development environments. In like 2022, 2023, they were hot, and then we talked about AI more. But I think outside of large tech companies, like cloud dev boxes really never took off because there's a very big upfront cost, and then you need to pay, like, a maintenance cost as well, and you just, you know, don't benefit from it as, like, a solo developer or, like, a small team. With the level of capabilities that we have in agents now, it's like the setup almost is free, right? So, like, the setup cost and this maintenance cost is like, if your agent is capable of doing it, you know, it should just do it for you. So, for example, if you're saying like, Hey, you know, I have, like, I have my local... SQLite, or I have a local server and MCPs and whatnot. It's like, how hard is it to actually configure exactly the same setup and keep it in sync on a cloud DevBox? Well, maybe it's not that hard if the model just does it for you. And so I think we're going to see a resurgence of, you know, fully cloud-orchestrated machines, which then frees you from your laptop, right? It's like one thing that we've seen a ton of success with with ChatGPT Work is like it's just available on your mobile. I start my day just dictating a bunch of tasks into it next to the coffee, and it just does it. It has access to my calendar, it has access to my email, it has access to Slack, and it's just so awesome to just be able to walk around and, you know, get stuff done without having to, you know, carry my laptop everywhere. And I think it's the same thing, like, you know, we ship, like, Codex Remote, where, you know, execution is, like, still happening on your laptop, but it would be wonderful if, you know, you didn't have to keep your laptop open. Can you tell me a bit on how, in the past, how did you improve Codex? Because I remember when I first used Codex, this was one of the early versions. You know, like, you could talk to it, it did stuff, but for example, I said, like, All right, make this change, and it did that change, and I had unit tests, and it didn't run it. And then later, a few months later, I don't know exactly when, it just started to run it automatically. Were these things, did you improve the, you know, the script that runs, you know, the instructions? I'm not sure how exactly you call the, you know, the bootstrapping script or whatever that is. Is it improving the model? Like, as a dev, how can I imagine you making each version better between the harness and then between the model, and, like, what's the connection between the two? Yeah, this is a good question. So the— The harness, in a sense, is always a little bit ahead of the model. Oh, really? How so? What I mean by that is that you have the model. It's capable of certain things, but then you set it up with, like, a couple of crutches so that it can actually do the thing to a level of reliability and in a way that is, like, efficient and also with the behavior that you expect as a user. And sort of, like, that's the role of the harness, right? It's like, you know, provide guardrails, like safety, make it more efficient, make it more, like, steerable, controllable. And then the harness usually is also responsible for, you know, what we call, like, the developer message, which is sort of, like, injected in the context at the start of each turn. And so that affects, obviously, like, the purpose is to affect, like, the behavior of the agent throughout the turn. A lot of what you have is, like, the result of the harness and the model. Initially, like, maybe you're like, Oh, it doesn't run tests, so, you know, you have to remind it to run tests. And then, you know, we train a better model that is just, you know, capable of, like, better reflecting on what is it that you really want when you ask for something, and then, you know, you don't actually have to tell it anymore. So over time, what we see is, like, the system, the developer message shrinks, and then the harness also shrinks. Inside of the Codex team, do you have specific goals? Do you say, like, All right, right now the Codex as a harness and model combined is not very good at this, or it's kind of doing silly mistakes, or how can I imagine, like, how— As the engineering team, how you're working on the next, you know, version of Codex is, because the thing that I don't really get as a dev is like, okay, there's a model, which to me is this magical thing which will get better. Of course, I'm sure you have some feedback channels, but you also have the harness, which is the tools that you're building. Like, that's probably what the team is responsible for. How do you set even your goals, right? Like in traditional software, you'd be like, we will build this feature, and you build that feature because you know how to do it. But it feels a bit more fuzzy to me, this development process. Yeah, it is, and it's why we co-design most things. And it's a process where it's a collaboration between research and the engineering team, like primarily building the core agent harness. It's always a question of like, okay, we see today that, you know, we are very good at this, but we're not very good at this, and, you know, we have a desire to do, like, another thing because it would be a very cool product feature. And then, you know, we always sort of like look at it. It's like, okay, should this be like a harness change or should this be a model change? And if it's a model change, like how soon can we have it? Can we have it in a month? Can we have it in, you know, three months, six months? And we sort of like work through that. And then depending on, you know, how soon we can just fix it in the model at which level of training, then we might decide to not even do something in the harness at all and just wait for the model to solve it. You know, it's agents all the way, right? So we use agents to analyze, like, a lot of the feedback, to, like, you know, come up with themes, you know, to just help us have these conversations and decide on priorities. But we analyze it across all of coding. We analyze it across, like, all of, like, you know, the other domains like finance, comms, marketing, you know, all the things where our users are using these agents nowadays. And there's, like, you know, subcategories within those. And then we roughly know, like, you know, how well we perform, and then we're always pushing the frontier. And there's a thing that is interesting is, like, as we make, you know, as our pre-training model gets better, as we make the overall model better, like, the whole thing lifts up. But then there are sometimes things that we pay a little bit more attention to. You mentioned you analyze agents all the way. Can we talk about the software development lifecycle on Codex in the sense of whenever a new engineer joins a team, any team, it's like, okay, how are things done here? And, you know, back pre-AI, it would have been you join the company like Uber or Google, and they would tell you, like, cool, the way it works is we have an idea or the PM has an idea, we make a plan, we get together, we do some estimations, we break up the work, we code the work, we do tests, we do code reviews, we release, we do feature flags, and then, you know, we're on call. You know, that's how it used to be. When someone joins the Codex team, you know, they've clearly been contributing to the open source part, but what do you tell them? How do things get done here if they're, like, a total newbie? I introduce them to great people. And then the thing that they hear the most about, like, when they have a question, is like, Have you asked Codex? And Codex is, like, just by default at OpenAI is, like, plugged into everything. So it has access to Slack, it has access to all the documents, access to all the code. And it still surprises new starters that you can basically ask it anything, and it will very often just, like, come up with, like, a really good response. And so the easiest way to understand the state of a project, or who's working on something, or why a decision was made, is, like, critics knows about it all internally. And so you just use all of that. We do a lot of work in, you know, for that reason, we do a lot of work in public channels. We open up documents with, like, you know, fairly broad permissions, and so that, you know, everyone has access to this information as well, and so that, you know, your agent can go through things and, like, you know, reason through things. And then, you know, we have a couple of other things that are just really very helpful for team productivity and team collaboration that we haven't released yet, but are going to come, like some of it at Dev Day. All of that just sort of, like, makes you very grounded and in tune with the rest of the team, and allows you to, like, you know, just very, very quickly, like, understand the state of things and produce things yourself. The general recommendation is just like, hey, care about the user, care about the coherence of the product, care about the models and where they're going. If you're doing something and, you know, you're building, like, this 10,000 lines of code crutch to work around a model flaw, it's like, you know, you're probably doing the wrong thing. So we have a set of principles, but it's just really sort of like a team culture and ethos at this point, and, you know, it's just very much sort of carries on, you know, when people join. It's just like through the rest of the team, just like, you know, sort of like teaching the ropes. And then when I have an idea, I think— I think to attempt to do. I think you can have that discussion outside of the pull request. It doesn't have to be around code. So maybe this helps crystallize the, you know, like where a discussion needs to happen versus where we did it because maybe we didn't have the type of tool that we have right now. Yeah, I think this is going to change, and it was like a forcing function because, you know, you have to have that discussion, or like it's good to have that discussion before you merge it and it becomes production code. But I think there are other ways to have these discussions and, you know, design things together and make sure that the intent is good, and then the code doesn't matter as much. And it's interesting because when I think back of all my code reviews, like, of course I have, like, memories where, like, it was great, we had a good discussion, or I learned something really interesting. But a bunch of times, honestly, it was such a pain in the ass. Like, I was trying to get my stuff. You're pinged, Hey, could you review my code? And like, No, right now I'm busy. No, I really need this to unblock me. And then you context switch, and then I feel it's always been, like, good and bad, right? So I feel whatever we do, there will be always upsides and downside, but now they're just moving. So I guess one upside is, as an engineer, you might have to not give your attention to just kind of basic stuff that doesn't need your input per se. Yes, it saves time. And progressively what we're going to see is also, like, you have an agreement on, you know, the box and the overall contract of what it's supposed to do, and then, you know, what is inside the box. As long as you have, like, strict guarantees in terms of resource utilization, data access, security, these kinds of things, it's like what happens inside the box is, you know, it could be literally anything. It's like, don't really need to care. And, like, really what you need to agree on is, like, what does the box actually do and what are the invariants that must be satisfied. And I think that is then worthy, you know, having, like, a really good conversation on, you know, maybe assisted by your favorite agent. But then once you have that and you have that understanding, it's just like changing anything within the box is, like, you know, doesn't require further discussion, and it's like, you know, just really preserves your attention. The cost of maintenance has gone down. You know, maintenance is always such a hot topic whenever we build something inside of all these companies like Google, Uber, even startups. Like building was the fun part, but then maintenance was the painful, and that's when we learned, like, okay, it was not worth building it, et cetera. Inside of Codex and OpenAI, what do you see maintenance becoming cheaper, changing in terms of instead of what you're building, what the ambition is, the, I guess, custom tooling, those kind of things? Maintenance is really like sort of like a tax that you pay over time just to keep things running, and it's always been necessary, will continue to be necessary. But where I think it changes is, like, a lot of it is just going to be automated. So, you know, it's like, okay, you have this third-party dependency, it's like you need to upgrade the version. It's like, oh yeah, you can fully automate this. You know, if you have good changelog and, you know, and the code is well documented and, like, you know, and the model can just, like, reason through it, it's like, you know, it can just, like, blast through your codebase, do it in a couple of hours, and, you know, previously you would have, like, sort of punted on it because it's not the most fun thing to do. But it's actually really important for your business, so it's, like, really important for your project, like, you know, especially for security vulnerabilities. You want to stay up to date, right? You want to apply, you know, all these patches. I think that's just going to be fully automated. So a large part of, like, maintenance, it just kind of comes for free, right? And then I think it's awesome to also think about before, like, you know, when you wanted to just completely re- you have to do, like, a new architecture because you're trying to make space for, like, a new, you know, different kind of trade-offs, or you have a new understanding of, like, the workload, or you're trying to fit a new feature, and, like, suddenly you realize, like, your current system is just very limiting and you need to completely re-architect it. That was, like, a really, really costly endeavor, right? So, you know, sometimes, like, multiple years. And I think this is also, like, super, super accelerated now. So, like, the cost of mistakes, you know, I would say, like, you know, is going down. But then at the same time, the good old rules, I would say, of software engineering, of, like, you know, having good abstractions, like, really help. Like, you know, going back to this, like, having the box with invariants, like, you know, if you sort of, like, draw the right shape, you're going to be able to change things much more quickly within the box and, like, not affect the rest of the services or the rest of your infrastructure. And I think it's important. It's important to design for very quick iteration and change. I remember when I talked with Peter Steinberger, that was before he joined OpenAI, but about OpenCLAW and how he thinks about it. Like, you know, he told me that he doesn't read the code, but he kept thinking about, like, I could see that he's holding the architecture in his head, and he was telling me how he re-architects a lot, and he thinks about how to make it modular, how to allow a hundred contributors to each build their thing without stepping on each other's toes. So I'm hearing what you're saying, that this care, this planning, this structuring has become maybe just a lot more important too, which was something back in the day. You know, it was like the architect or the staff engineer or experienced folks were doing this thing, and other engineers around them were kind of building this, you know, smaller parts. But it sounds like now all engineers need to be aware of when you're building your software, right, and plan for it. Yeah, and the GPT models are getting better and better at this as well, of, like, you know, thinking about long-term maintenance and, like, good architecture, and, like, this is, like, a natural sort of, like, next step, right? It's, like, not just about code quality in the sense of, like, oh, is this code clean within this file, but, like, you know, is the architecture actually correct to reduce maintenance burden over time and, like, you know, make space for, like, future product or feature extensions or changes? And just really this act of, like, you know, engineering over time, that's kind of, like, something that models are starting to become capable of, like, thinking about very well. I think it's just kind of fascinating to understand that the software that we're building is just going through the life cycle much, much faster, right? You know, before you had, you know, you were scaling it, you were starting it, you know, maybe as, like, a small team of, you know, yourself, maybe a couple of engineers, and then you would add engineers, like, slowly, and then, you know, maybe after a year, you know, it's like if it's very, very successful, you would have 50 engineers on it or, like, 100 engineers on it. You would have time to see it coming. You would have time to see, like, you know, the humans on board, and, like, you know, you can think about the documentation and all of that stuff. But now it's just sort of like that explosion of, like, you know, suddenly you have, like, 100 agents contributing to this thing. It's like, you know, that can happen, like, you know, in a weekend. And so, you know, you're just going through it at, you know, major, major speed compared to before. Okay, but how do you and the folks at OpenAI, like, deal with this? Like, does it not mess with your mind? Like, you know what I mean, in the sense of, like, you've been in this business for quite some time now, like decades or well over, and there was a pace that we kind of got used to. And obviously it's now a lot faster. But how do you get your head around the fact that, A, it's faster; B, the stuff that you've been doing a year ago, right now you're not doing because now the model is good at it. And, you know, like, how do you kind of reconcile that? Because I'm sure there's stuff that you've been really good at related to software that now you can hand off to the agent. Do you not get a little bit of sting? You know, we talked about it stinging for your features to be implemented open source, but it can also sting that I've been really good at, like, I don't know, refactoring or right now it might be architecture, but maybe the model will be really good at that. And now I'm like, oh, okay, damn. Like, I'm glad, but also like, it would have been nice for me to do that. Yeah, I think there's like a craft aspect to it, which occasionally I still, you know, pull up an editor and, like, write some code, and it's just like, it feels nice. And it's sort of like I have fond memories of, like, late nights sitting in Vim and, you know, just like cranking it out, you know, drinking Coke Zero and, yeah, just not having to think about anything else other than, like, the problem in front of me. But really, I think it's all about being in the flow and solving problems. And what I find is, like, you know, folks here and also, like, everyone I talk to is just, like, adapting very quickly. And I think if you have a mindset where it's all about— Code is a tool to solve problems, and you can solve so many more problems. It's like before you wanted to benchmark something and you weren't quite sure where you were going to net at. It's like you can just do it. It's going to take you like no more than 30 seconds, you know, to launch something in the background and, you know, get proper numbers and be able to do like a better trade-off. It should make you a better engineer if you just really care about, you know, the outcome and the system working well. And so what it allows us to do at OpenAI, it allows us to run, you know, our inference much more efficiently. It allows us to, you know, get like much more effective compute and, you know, deploy that to the world. And so, like, everyone is just like very focused on that and solving important problems at the speed that was not possible before. And like, I haven't yet, you know, encountered someone who's like, Oh, that's not good. That's not fun. Do I understand correctly that it sounds like if you have ambitious problems, if you have way more problems than what you can solve today or tomorrow or the next week, sounds like this is not really a problem because when, you know, you get more efficient somewhere, you keep going, which is a lot of startups, right? Like startups are always way more ambitious than what they're able to do. We're not out of problems for sure, right? So, and I don't think we will be for a while. We have a long, long road ahead of us in terms of like mathematical breakthroughs, scientific breakthroughs, you know, making the world a better place, like just really building for humans and solving the most important problems that everyone is facing and just doing it in a deeply human way. That's what we're here for. Also, just going back to coding and like, you know, these late nights, it's like I think there's like, it's also like maybe like a glamorous version of it. It's just like I also had a lot of late nights where I was trying to refactor something. And, you know, it's just like it would be like three hours deep into the refactor and then realize, like. There were, like, many different permutations considered. And so there's, like, a very fun journalistic element to it, where we have a full recounting that Codex did over time. And, yeah, it's just, it's kind of become known as well as, like, the toggle arc of OpenAI, where, you know, we introduced, like, the work toggle, which there was also a lot of debate around of, like, you know, whether this was, like, the right thing. And then, you know, it's just, like, we kind of grew to just really like it. But over time, we're going to merge things further. So it's like we're really headed into this direction of, like, full unification. And, you know, we kind of view this as, like, a temporary state where, you know, you have, like, you have better, stronger capabilities when you're in work mode. But over time, we're bringing this, you know, all the way to, like, you know, everyone that uses ChatGPT. And how do you personally use Codex? Like, what's your working setup in terms of agents, in terms of tasks, in terms of what you manage with it? And related to this, I asked Peter Steinberger what I should ask about you, and he said, like, you need to ask him, how do you deal with the fact that you're involved with all these projects, your calendar is like Tetris, but usually you show up pretty cheerful. My calendar is fine. And it's just, I am capable of doing so many more things nowadays because I have a technology like Codex. And I actually shifted a lot of, like, my work on mobile using ChatGPT work, where whenever I have something that I want to take note of, I just, like, fire that off. I use dictation a lot. Whenever I have a question, instead of, like, writing it down to look into later or delegating to someone, I just, like, fire it off in ChatGPT work and I get, like, a report. It has, like, a whole bunch of, like, custom skills and custom instructions where it's now, like, very, very tailored to, like, you know, produce the kinds of reports and slide decks and code explorations, you know, in the style that I can consume effectively. And so every time I'm, like, between meetings or, like, you know, you'll kind of, like, see me, like, you know, I was just, like, dictating into my phone. As I said before, it's just like, we do a lot of work in public channels. We have, like, a lot in Slack. We have a lot in Notion and Google Docs as well. And so there's pretty much, like, there's no question really that I feel I cannot ask that, you know, Codex will be able to sort of, like, do at least a first pass of thinking through, whether it is, like, public sentiment on a feature, looking at production logs for, you know, how much usage we have on a certain thing, making a list of things that we should deprecate because they're not getting traction, understanding what a certain team is up to. It's like any question I have, I can get an answer to, like, you know, within 30 minutes. And so that's how I use it. I use it for everything. It's like my personal agent in, like, all the ways. And then oftentimes on weekends as well, I do some, like, code explorations or, like, I build some prototypes and I have fun, like, sort of, like, imagining the future of the product in some ways. And I do that with others on the teams. It's not always the same team. And it's just like, in one day, I can build things that I sort of, like, I had it in my system, right? It's like, I woke up one day, I was just like, we should explore what it means to build this. And then I can just sort of express all of that and get, like, something in front of people in a day so that they can think through it and criticize it and hopefully get inspired by it. It's like, by no means, you know, we need to ship it, but it's more like, okay, I flush it out of my system and then, you know, I go on and, like, you know, do other things. So it's just, like, so, I don't know, it's such a magical time and it's, like, so empowering. And as closing, what would your advice be for a software engineer slash AI engineer, someone who builds software, who would want to get the skill set and the experience to have the opportunity to work at a place like the Codex team, like OpenAI or, like, an AI startup? So, like, you know, just become this really great builder with these tools. Because the question that comes up often, like, should I start with the theory? How important are the basics? Should I just get really good at using the tools? Yeah, I think there are two things that are important, is a deep, deep curiosity for how things work and an ability to, like, you know, train yourself to understand things very quickly. And so it is the case that things will continue to change, but people that do extremely well at OpenAI are, like, you know, people that just sort of, like, are able to, like, grok a system quickly and, like, you know, also dive into, like, a new codebase and sort of, like, you know, make sense of it. But obviously, like, all of that is helped with agents nowadays, right? So just, like, there's so much information that you need to absorb and, like, you know, being able to understand and reason through it. I think a lot of that is asking good questions, really, about, you know, how do things work, and just, like, going into, like, the five whys, which I think, you know, you can just kind of keep digging and digging and digging, and, you know, you're learning very, very fast through that. The other thing is being in tune with the community or, you know, the people that you're trying to solve a problem for. It's like not everything is, like, solving a direct problem. Sometimes you're solving a problem that will be useful, you know, to, like, another group of people in the pursuit of, like, solving a problem for humans. But just being crisp about the taste or the needs or the requirements and being able to think clearly and, like, you know, exercising through this clarity of thought feels really important to me. Like, if you can't explain what you're trying to achieve, if you can't explain your intent, if you don't have a tie to a community, if you don't have the taste, it's going to be much harder to do great work. Awesome, Tibo. Well, thanks a bunch for this conversation. This was awesome. Thanks for having me. I've always wanted to get together with Tibo, and I'm glad that we finally made it happen. I appreciated how Tibo talked about not just the upsides of open source, but also the downsides, most notably how competitors can copy features you are just working on in the open right now and then ship it right before release, and just how much this stings. Plus, you get a lot of low-quality contributions that you still need to somehow deal with. Another interesting one was Tibo saying how the hardness is always a step ahead of the model. From the inside, the Codex team see their job as building clutches for the model with the hardness, the tools, and the setup instruction. And then the next version of the model will be trained to need fewer of these clutches. I'll be honest, as a dev, this sounds a little demotivating that the stuff I build in the next version of the model, it'll just know and we can get rid of it. Plus, I do suspect that it's not just about building these clutches, but also building tools that models will use. And it's not like the next version of the model will reinvent an MCP protocol or scales or plugins. At least I hope not. I also enjoyed hearing what the merge, merging ChatGPT and Codex, looked like from the inside. It was merging a previously fully local coding agent, Codex, into a managed cloud-based stack, and doing it efficient enough so that it can be included in OpenAI's $20 per month plan when $20 is not all that much in terms of compute purchase. It was pretty amusing to hear how Codex itself acted as a journalist of the whole project, as it was present in all the Slack conversations and all the documents, and so it could capture all the important debates and decisions. I'm not gonna lie, this part felt a little bit of a Big Brother feel to it, where the AI is always watching, but it could well become the new normal in startups in the future. I've not yet decided how I feel about this. And finally, I appreciated Tibo's advice for engineers to succeed: be curious, understand systems quickly, and be in tune with the group you are building for. It's reassuring to hear from Tibo as well how much the fundamentals still matter. Do check out the show notes below for deep dives on how Codex, Claude Code, and Cursor were built, and other related topics. If you like what you heard, please hit a rating on the podcast player that you're using. It means a lot to me and to the show. Thanks, and I'll see you in the next one.