← Return to Index Archived September 30, 2026
The Lead — Sep 30
THE PRAGMATIC ENGINEER · GERGELY OROSZ

Distributed databases with Peter Mattis

Cockroach Labs CTO Peter Mattes traces a career through GIMP, Gmail’s storage systems and distributed databases, with B-trees recurring as a durable engineering obsession. He also makes the case that AI coding agents can raise output without sacrificing rigor, provided engineers force them to test, measure and review relentlessly.

1h 41m / September 30, 2026 /aitechnologyproduct / Transcript sourced from openai
All episodes from The Pragmatic Engineer →·Listen on Apple Podcasts →

Overview

Peter Mattes, co-founder and CTO of Cockroach Labs, discusses a career spent building systems where failures, scale, and efficiency are core design constraints. He covers early work on GIMP, storage infrastructure at Google, the design choices behind distributed databases, and his return to hands-on coding after AI tools became capable enough to help produce production-grade database code.

A central theme is that AI can raise engineering output without lowering standards, but only when engineers impose strict expectations around testing, performance, security, and review.

Key Takeaways

  • Mattes repeatedly came back to B-trees because they solve a broad class of storage and indexing problems well. Compared with pointer-heavy structures such as red-black trees, B-trees can reduce memory overhead and improve spatial locality, which often improves speed too.

  • His work on Google's Colossus file system illustrates the economics of distributed storage. Traditional replication might store three full copies of data. Erasure coding, such as Reed-Solomon coding, spreads data and parity across machines so the system can tolerate failures with less storage overhead.

  • CockroachDB's automatic sharding is designed to spare application teams from becoming accidental distributed-database engineers. Manual sharding forces teams to decide where data lives, handle resharding, and deal with cross-shard operations. CockroachDB partitions contiguous key ranges and manages their placement and movement internally.

  • Mattes says he wrote as much as 100,000 lines of code in some pre-AI years, but stopped doing much core-product coding between 2022 and 2024 as executive and management work took over. He returned partly because he felt leaders need direct experience with AI coding tools if they expect to guide engineers using them.

  • He rejects the idea that higher AI-assisted output must mean lower quality. His view is that agents are often lazy about tests, much as people are, but they can be instructed to apply more testing discipline than a rushed human might sustain. He calls out property-based testing, metamorphic testing, and deterministic simulation testing as tools that agents should use routinely.

  • AI also changes who can contribute directly to production code. Mattes describes a designer at Cockroach Labs working in HTML, CSS, and JavaScript and sometimes opening production pull requests, reducing the traditional handoff from design tools to engineering.

Practical Steps

  • Profile before replacing a data structure. If a map, cache, or index consumes surprising memory or dominates CPU time, inspect pointer overhead, allocation behavior, cache locality, and access patterns rather than assuming the standard library is already optimal.

  • Treat scaling limits as a design problem early. If a single database instance is nearing its size or availability limits, do not wait until a failure forces a rushed sharding plan. Evaluate whether the application should own sharding logic or whether the database should manage distribution.

  • Give coding agents explicit test requirements. Ask them to add edge-case tests, property-based tests where invariants can be stated, regression tests for bugs, and performance checks for sensitive paths. Do not accept "tests pass" without asking what classes of failure were tested.

  • Use multiple AI passes for security-sensitive changes. Mattes suggests agents can review every changed line adversarially, which is more practical than asking scarce security specialists to inspect every commit manually.

  • Keep hands-on familiarity with the tools your team uses. For engineering leaders, that means writing code with the same agents, editors, and workflows engineers are expected to adopt.

Notable Quotes

  • Peter Mattes: "The agents are lazy. Humans are a little bit lazy. So getting the humans to actually be very disciplined about their testing is also challenging. And I think it's actually easier with agents."

  • Peter Mattes: "I got into software engineering because I like building stuff. I can build stuff faster."

  • Peter Mattes: "I feel like I've learned more in the past probably even year than the previous five years combined."

I got into software engineering because I like building stuff, and now I can build stuff faster, taking away some of the compromises we had to make in the past. — From the episode

Full Transcript

Source: openai 1h 41m runtime

He built a GIMP image editor while at college, designed the original storage system behind Gmail and built many more large, complex and widely used systems. This is Peter Mattes, co-founder and CTO of Cockroach Labs. Before we sat down to talk, Peter told me, for the last 30 years I've always been a prolific coder, but my current output is a bit insane. And this isn't vibe coded junk, but database worthy, high quality, high performance code thanks to working with strong coding models. Today we cover why B-trees are so important when building databases and why Peter kept reaching for the data structure over and over throughout his career. How he wrote 100,000 lines of production code per year pre-AI, stop coding between 2022 and 2024 and why he is back now. Why he thinks AI agents are lazy about testing and how this is an easier fix than it looks. And many more. If you're interested in distributed databases, distributed storage systems or knowing how Peter manages to use AI to produce unusually high quality and production ready code, this episode is for you. This episode is presented by TurboBuffer. A ridiculously scalable, fast and cheap hybrid storage engine built on top of object storage by an engineering team I've really grown to like after spending time with them. The TurboBuffer engineering team is doing something really, really cool. They're completely redesigning their storage architecture from first principles to make storage faster, cheaper and more reliable at scale. If you follow TurboBuffer, you know that their storage architecture was a massive part of their early success. Redesigning a winning architecture is a big deal. It's one thing for your query plan to pass unit tests. It's another to bring each query plan to performance parity, or better, while maintaining correctness and reliability in production. Here's the cool part. TurboBuffer is documenting the whole thing. Their new storage architecture, which they're calling TPUFv3, demands some hardcore systems engineering and they're building it in public, sharing design decisions and benchmark results as they ship. They're keeping a work log of their journey and the first post just dropped today. Follow along at TurboBuffer.com slash v3. That is TurboBuffer.com slash v3. Peter, welcome to the podcast. Oh, I'm happy to be here. This is awesome. I wanted to get into, how did you get into tech originally? When did you figure out computers are interesting? I figured that out in kind of elementary school, high school. Gaming was a little bit of a gateway drug for me as a software engineer, as many other people. I remember like early on, my mom did programming at IBM at some point. I'm never even quite sure what she did, but we had computers always around our house, like Apple 2 Plus, Apple 2GS. I'm from that era on up and, you know, go to the bookstore, I'd find a book on basic or magazine, type in programs, no idea what it was doing, but you know, just kind of like I was addicted to like, you can produce, put stuff into these computers and get interesting stuff out. And then I got to college and I was like, I didn't think there was any money in computers. I didn't know anything about it. I started as a mechanical engineer, following in the footsteps of my dad. You started mechanical engineering as your specialization? Yeah, as my major. Yeah. Major. And I got in there and I was like doing, like I'd done some computer stuff before. It was a real foolish move to do this, but first semester of doing these homework assignments in mechanical engineering, they were god awful, like six pages for a single problem. And I happened to take a CS course at the same time. And it was so easy. And then everybody else was failing it. I'm like, I'm in the wrong, I'm in the wrong field. Let me switch. And then you switched. Then I switched. Yeah. What was the first software that you built, either at college, it must have been a college, that you were like, all right, this is a piece of software that I'm kind of proud of. That's a complete piece of software. Well, I mean, the big thing that I did, along with my roommate in college, we had this course, was it a compilers course? I'm quite sure anymore. This is like 30 years ago. And we were kind of bored with it. So we wanted to do something like kind of fun on the side. And I'd done journalism in high school in my senior, junior and senior year. And I knew stuff about kind of computer graphics and wanted to do something like Adobe Photoshop. So started just kicking the tires. And we built up this program that a lot of people know of called the GIMP. And along with the GIMP, I did a lot of the graphics library, GTK. This is since evolved just massively since then. It's kind of interesting because after college, I kind of stepped away from it. Didn't really stay involved much past my first year out of college, but definitely people still know me. And it led to other some interesting events in my career. And it was like starting GIMP, it was literally just you saying, all right, I want to do something like Photoshop. How hard could it be? And then you just, this wasn't through what you learned in college, right? This was like you figuring out how to build, you know, like a graphical, I guess, engine rendering, drawing, data structures, all of that stuff, right? All that stuff. Figured it all out. I remember trying to look at some papers back then. is that the reason Bazel and Buck are so popular for large codebases is it can help you improve your build performance. It gives you a lot more levers to play around with, from caching, from being smart about cache generation, to obviously just raw performance. Yeah, that's exactly it. Yeah. So you kind of just dabbled and like, okay, I'll make this build file, did that. People took that over. But what was your next main focus? Well, so I mentioned GFS earlier, Google File System, and at some point we realized that there were some limitations in GFS, scalability bottlenecks, because I've been working on the storage system for Gmail. I actually dabbled in other storage systems, you know, kind of a research thing that never went anywhere. But because we were working on that, I got invited to participate in, like, the founding team of Colossus, which was the successor to GFS. And as far as I know, Colossus still exists. It's gone through multiple iterations at Google, but it's like the second-generation, you know, distributed file system. So what is Colossus? Yeah. So when I say distributed file system, externally you might think of something like S3, kind of blob storage. It had a flat namespace, kind of like S3. You give names, and there's like a minor hierarchy there, but it's, like, very limited. But it's not like a POSIX file system, so you don't have the full directory hierarchy. You didn't even have the full, like, kind of permission system. A lot of that stuff got added later. But the files are not stored on your local machine. There is a fleet, you know, kind of a service out there that has all the files. They're writing it down to their hard drives or now SSDs, and your client can access that, and it's all replicated. So if there's any crashes and whatnot, you're not losing your data. I don't know when, at some point S3 got erasure coding. We did erasure coding in Colossus. That was kind of a big breakthrough. Which coding? We used Reed-Solomon. So for the audience who's not familiar with erasure coding, you might think of, like, I want to have replicas, and there's this thing in hard disks called RAID where it's like, I don't actually have to have full replicas. I can actually, you know, if you take A and B, you can XOR them together, and then you can have this kind of third kind of version. And there's more and more complicated versions of that. Reed-Solomon is kind of, you know, I think there's actually better codes now, but it's like— Like one of the known ways. One of the known ways, yeah. Yeah. And we had to, you know, kind of pioneer internally, like, oh, how are we going to actually make this work in distributed file system? Were you focused on latency, on being able to store data more efficiently? It's storing it more efficiently. So in GFS, and there's varying costs where it's like, well, you're not storing two copies, or one copy of your data, or two copies, storing three, triplication, three times as much storage you're having to use. And with Reed-Solomon, you can get that down quite a bit lower. I can't remember offhand exactly what we use for Reed-Solomon, but I think it was like essentially 2x. But you get that with also the redundancy too. So it's smaller and the redundancy is higher. So because I guess the naive thing, if you're saying, all right, I want my data to be replicated at three places, you take three nodes, three machines, physical machines, and you say like, copy one, copy one, copy one, I have it three places, great. If one explodes, I still have two, wonderful. And then you're saying that the algorithm here is you could take not 3x the data, but 2x the data, split it smartly across machines, or maybe you could take it lower, and you still have the thing where like, oh, one of them explodes, I still have all my data because it's split enough. That's exactly right. And like kind of the mental model, you know, if you just want to understand at a high level, which is essentially you might want to say like, I want to have eight replicas of this data. I think, I can't remember offhand, I think S3 might use nine. They've actually talked about this publicly. You have kind of nine chunks of data, but any five of those chunks can be used to reconstruct it. And what this means is you can lose any four copies and you can still reconstruct your data. And oftentimes it's more like, you know, the first five chunks are exact replicas and the other four are kind of parity ones. I've kind of forgotten some of the details, it's escaped my mind, but it's along this line. Yeah, but when you come up with an algorithm, you can then prove that this algorithm will work, right? Like this is a little bit like, I know in software engineering, like math and algorithms is a bit out of fashion, but in this case, this is really important because once you can prove it that this algorithm works, it will work. It will work. And, you know, it's like the math behind here is like Galois fields over like... GF2, something like that. I don't even know the—I never actually understood the full math behind it. I always regretted not doing more math in college, but you didn't have to. Like Reed-Solomon proved how this worked. I think it's like back in the 1970s, something associated with, like, communication networks. So you just take that and, you know, kind of use that expertise, but leverage that and have to do all the engineering behind it to make it work in a storage system. And then when building a distributed storage system like Colossus, what— Give us a little bit of, like, how you've been inside, how it happens, and how other people like yourself and your colleague can say, like, Oh, what if we try something else? Yeah, yeah. So, I mean, my recollection here is he was working on this kind of the big internal system, I think it was called Gaia, that actually had the mapping from, you know, you log in, you have your user ID, and you have to look this up. And they were storing, you know, all, like, the map from user ID and email to whatnot to the metadata about the user in STL maps. And you just notice, like, well, there's a lot of memory usage here, and it shows up on profiles. And then we're like, well, what can we do to do better? And that was kind of the genesis of it. And, you know, he happened to be working on it, and he happened to be working with me, and, like, we just started kind of noodling on this problem, like, Oh, can we do something better? And it's not one of these things, like, I think now at Google, they have a whole team working on their kind of internal libraries. At the time, it was more of like, you know, everybody working on their own systems and contributing to a shared base. But I guess it still goes back to what you were just saying of, like, just go down the layers, try to understand, and if something just doesn't add up, like suddenly, like, Oh, there's this big explosion of memory usage, like, you know, just ask the questions, Why is this? And if you're able to, or you happen to be like, Oh, can we do something about it? Right. And one of the things that, you know, he observed early on, I think part of one of the things was it was like a map from integer ID to something else. And, like, if you look at Red-Black tree, every node, you have your kind of value that you're storing in the map, and then you have two pointers. You might have an integer ID that's like four or eight bytes, and then two— Yeah, and you look at that and you're like, Ooh, that seems like a lot of overhead. You could just use a bit. Well, the B-tree actually has a lot better spatial locality, and that's what made it faster. But it was actually smaller as well at the same time because you had less pointers involved. You also contributed to Go, right? Yeah, yeah. Well, that came later. Yeah, yeah, it came a lot later, but can we talk about that? Yeah. Just one of these other things, you know, I pay attention to, like, you know, when there's research papers coming out about new data structures and, like, hash tables. Hash tables are like— One of the earliest things you learn about in college in data structures, like, how do I map keys to values where the ordering is unimportant? That's when hash tables come in. And there is, like, you know, the very earliest ways to do this, I've implemented hash tables multiple times, is like you take your key, and it might be a string, and you put it through a function and it spits out an integer, and then you map that into an array of buckets. And if multiple things map to the same bucket, you have to have a linked list. You have a linked list, yeah. This is naive implementation. Naive implementation used quite frequently. And over time, people discovered, like, a lot better ways to do hash tables. There's very, like, that's called chaining of your hash. There's another technique called open addressing, where instead of actually having a linked list, you just kind of hash it again and move on, or kind of walk down to subsequent buckets to find out, like, or subsequent slots to find out where you should be. And I remember reading about this new technique. It came out of some folks at Google. I believe it came out of their Swiss office, because it was called Swiss Tables. I believe that's where the naming came from. Not 100% sure about that, but I remember reading about it. And then I was working on Go for a long period of time, and Go has this built-in map structure, and it's a hash table. It's a very highly optimized hash table because the Go team is a very competent, the Go runtime team, and various folks had taken an attempt at, like, you know, putting together a Swiss table implementation for Go. And I tested some of them, and I was like, this is kind of fascinating what Swiss tables do, and I'll explain how it works in just a second. But I looked at it, it's like, well, it's really hard to beat the performance of the runtime. The runtime was really good. And I kept on, I noodled on this for a little while, and eventually I ended up having to take this business trip to India, to Bangalore, and so I was on a long flight. No. Yeah, it always starts like this. And I'm just like, I'm just going to try to pull on this. I pulled on it sufficiently that I could get some of the benchmarks to be faster. Wow. And then I'm like, you know, that is kind of like catnip for an engineer, like, can I make it all faster, kind of figure all the rest of it, you know? Got some help from the runtime folks. There's an issue on the Go issue tracker that, you know, where other people have been attempting this, because people propose, like, hey, let's use Swiss tables. And Go folks are like, well, you know, we don't quite know all the details. We're going to have to navigate this and that. And there's some ideas there that combined them together and got to the point where we had kind of a complete implementation that was faster on most benchmarks, not quite all of them, but most of them. And then the Go folks eventually picked this up and pushed it over the finish line. And then, so you, you know, like, you came with the idea, you got— And the system keeps on going. And it's like stories like that, you know, like what we did on the marketing side there. But also, we hear this from our customers too. They've had fires in data centers and all the other data systems go down, and CockroachDB keeps on going. I'm like, that's awesome. When you started out, who were companies, startups that who wanted to use CockroachDB, and how has it changed since? Because, you know, like just making the case, like, okay, I'm starting a startup. It's a small startup. Like, I will need a database, and I'll, I don't know, I'll typically choose a Postgres, right, because it's free, everyone's using it, I'm running it on Node. At what point did you see that typical, typically tech companies are like, okay, like this is not enough for me that it's running on a node, either because it can go down or because I'm outgrowing it? What was it they're outgrowing? Like, I'm trying to get a sense of, like, at what point did companies say, like... tell themselves, like, we need something distributed in a database. Yeah. I mean, oftentimes we have companies calling us up after they've had a disaster. So, like, a node went up or a hard drive failed, that kind of stuff. You know, it's not quite like we're ambulance chasers, but if you see an outage, like a big outage from some company, you know, like sometimes we're trying to knock them up. But also they will call us, you know, and being like, there's a very big bank who's now a customer. Don't think I can name them, but you can go read. They had a very serious outage due to a weather event, and after that— Which probably knocked down, I'm assuming, a region or database or a networking cable or tree fell on something. I think it knocked down a whole region, you know, the region-wide power outage knocked down the region. And there's a mandate from the CEO. It's like, no, we just have to, you know, be able to survive these things, and that gets pushed down all the way. And you see this in other places where, you know, one of our early customers, they were running on AWS, and they just got to the maximum size you can run an Aurora instance on. And then what would typically happen at that point is then you have to shard your database. You know, this is very standard practice. You take your single-node database, you create 10 or 20 or 100 shards, and this is what Google had done for some period of time. And that's a heavy burden on the application developer. And the way we always phrase this is like, I mean, the application developers become a database developer at that point, and they're doing it poorly. You know, they're trying to implement distributed transactions or indexes and whatnot. We felt the burden for that belongs on the database developer. Can we talk about automatic sharding? I think it's safe to—most of us will know what sharding is when you're—but actually, let's start from, like, manual sharding, and then how you can implement automatic sharding, and if you can tell us, like, you know, tactics that a database like CockroachDB can do to actually just take that load off of you. Yeah, yeah. So I think the very basic form of sharding is a little bit like the hash table. Let's say, you know, you have a fixed number of shards, like let's say just 100 shards. Your data model is a user with a lot of data associated with the user. You just take the user and you say, like, oh, they map into one of the shards, and, you know, you kind of just rely on the hash function to get, like— Like fairly even distribution. The problem with this is at some point, you know, one of your shards will get full and you have to kind of reshard, and that's a very, very onerous process. Resharding, I guess a simple way to do is like if it's just a hard drive, I don't know, per node, where you write the user data, it gets full and you're like, okay, well, I now need to split it somehow. I need to move it. I need to remap it. I need to rejig my metadata, which knows where this data lives, that kind of stuff. Yeah, yeah. And it depends on exactly how you're doing that mapping from, like, you know, the user ID or whatever your shard key is to the shard. You might have to remap them all, right? This is very typical. Exactly. So, I mean, this happens in hash tables where, you know, oftentimes in order to grow the hash table, you just have to essentially create a new hash table, double the size, and copy all the data over. Now, that's kind of like the very basic straightforward way. And there's various levels of complexity on it. One of them is called consistent hashing, and there's various techniques to do this. It's kind of fascinating, like how they all work. But in consistent hashing, you can add an additional node, and then it only moves a fraction of the data from each shard over there. There's various systems to do that, and I believe this is like whenerlies Cassandra. The way CockroachDB does it is more akin to Bigtable, more akin to Spanner, more akin to HBase, where instead of actually hashing, we actually take, you know, you can imagine all your keys in a system, and this is always true in any system, that you can imagine them just in one big contiguous key space, and then you kind of partition contiguous spans of that, and then you have to build up an index on top of those contiguous spans. And what I just described there actually sounds a lot like a B-tree. So there's this index on top that is like, that maps you from Adding functionality to integrate better within enterprises. Our revenue has been, you know, kind of steadily growing over the years, and, you know, it's at this place now where we see a path to future success as well. And I want to ask about your coding habits. So when you co-founded the company, how much code did you write for the first few years? I wrote a lot. So I've always been a very prolific coder. I wrote a lot of code in the early days. And early days, I mean, I was kind of—we were all technical co-founders, Ben, Spencer, and I, and we were all writing a lot of code, and I always no exception. But, you know, if, like, I look back at my GitHub output, you know, it's like kind of peak years, maybe 100,000 lines of code in a year. A year, which is a lot. Yeah, yeah, no. So, I mean, we're talking pre-AI. Pre-AI, right. This is when you're back doing this manually, right? You know, at some point, you know, we started out using a system called RocksDB, which is an LSM. At some point, I think it's back in 2019, you ran into limitations with it. I decided, you know, I want to rewrite it and did a big push to rewrite it. That might have been like 40,000, 50,000 lines of code. And then a bunch of other people come up and helped. And, like, you kind of look at that output, and I'm just like, oh my goodness, that was a lot to keep in your head. It's a lot just to type, you know, 100,000 lines of code. The average kind of like that the industry talks about is 3,000 lines of code for an engineer in a month. And so if you multiply that out, maybe 36,000 in a year, that's good, right? So I was doing a lot. You know, I kind of look at that, and it's like there's kind of a max that you can hold in your head at a time. The tools have gotten a lot better since I first entered the industry. We've gotten better debugging techniques, better testing techniques, but still quite significant. And you were CTO from the beginning, co-founder and CTO, but there was a time sometime around like 2022 when you decided to kind of be a bit more hands-off, right? Yeah, yeah. Can you tell me about that? I mean, the general rule of thumb for engineering leaders is, well, you got to have your team, you got to manage your team. And we had a VP of engineering, but I was kind of getting to the point of like, okay, is my coding days done? You know, like, can I just direct from a higher level? And, you know, I got this advice for a long period of time, and I pushed back on it, but, you know, I kind of acquiesced at some point, and I think it was the right advice. I'm not saying that the advice was wrong at the time, but there was a time period from about 2022 to 2024 where I was like, my output declined. I think I actually did the Swiss tables thing in that time period, but I wasn't doing much on the core— The business. Yeah, the core business. You know, I would get in there and do some work, but, like, it's really hard that if you're in meetings all day to also do coding. I mean, I think this is the fundamental tension. So you kind of took on the kind of the meeting burden, the coordination burden, the stuff that was, if I'm reading correctly, before you spent a lot of your head in the code, and now you're spending a lot of your head, like, above the code—the business, the engineering, or the whatever, customers, that kind of stuff. The customers, and just being an executive as well. Yeah. So very hard to wear all those hats simultaneously. And then I got back into it because AI started to emerge. So how—when did you start using AI? When did you start to find it useful in terms of coding? Yeah, well, it was interesting because, you know, those initial versions of, like, glorified autocomplete came out, and— We're talking about the GitHub Copilot, the Cursor, the early version. Yeah, GitHub Copilot. That was the one I had the first exposure to. We dabbled with Cursor at the time, but they're all, like, kind of glorified autocomplete. And it was kind of crazy that you could just start typing something and, like, fill in the rest of the function. You look at it and you're like, wait, kind of got that right. This is crazy, right? And, you know, we were training Cursor engineers to use this. And at some point, you know, I can't remember if this is my idea or my co-founders or someone, and basically, like, you know, like, in order to, like, really guide people about how to use it, you have to be a user yourself. You know, I think this is true of, like, engineering management in general. Like, you want to, like, manage engineers, you have to know how to be an engineer. Like, if you don't know how to be a good engineer, it's, like, really hard to manage other engineers. I feel like you'll have a hard time, like, just relating to them at the very least. Exactly, exactly. So it kind of took it on me, like, no, I mean, this is, like, it was clear very early on, like, this is probably going to go somewhere, but it wasn't quite clear how far, how fast it would go. And you start dabbling this, and I was like, oh, okay, well, it's not quite good enough. It's not quite good enough. But, you know, let me start getting back into the coding very rapidly. You know, you started seeing the signs of life, like, you know, kind of the OpenAI models coming out. It was Sonnet first. In VS Code tabs, the terminals in tabs, you know, have like eight of them open, and then they started taking over. It used to be like just Claude Code in the terminals. Yeah, yeah. And then, you know, just some point recently, it's like six weeks ago, a colleague was like, Well, have you tried out the Claude desktop app recently? It's really good. I was like, Yeah, no, I'm very, very happy. And I made the switch, and it was a little bit awkward at first, and then suddenly I'm like, Holy crap, this is awesome. You can, like, manage the agents a bit better, easier. Yeah, I mean, it's like you have the sessions there, the sessions down the side, you know, and it's like they're kind of like tabs in some regard. And it's just, but the whole integration that Anthropic's been doing, and OpenAI is doing the same thing now. And by the way, like, I'm not dissing Cursor and Factory and Cognition. They're all pushing the same thing. It's incredible how fast these systems are innovating right now and evolving. What do you think good software engineering looks today compared to, like, you know, four years ago before we had AI? Has it changed? Has good changed, or that's not really changed? Yeah, I think the ambition has to increase. You know, I think the quality has to be higher, security has to be higher. I'm glad you mentioned quality. Can we talk about that? Because I'm seeing across the industry just quality declines, which you cannot fully put your finger on AI, but oftentimes it is people pushing out more and more and just not paying attention to just small regressions here and there. Again, like, you're building a database, like, yeah, yeah. Have you noticed any, or have you gotten feedback of any quality regressions, or if not, like, how come? Because, like, when—like, this is just basic, you know, like we talk about the law of physics. This is an observed thing that when you start to have more output, you increase your deployment frequency, you often, not always, but you often have, like, more regressions. More bugs. If you produce a certain number of lines of code, you're probably going to have a certain number of defects per line of code, and now you can produce more lines of code, so you probably would have more defects, right? But the thing that pushes against that is that you can be telling these agents, you have to give them kind of firm-handed. This is the thing that I hope the model providers are listening to, but you need to give them a firm hand. They kind of get a little bit lazy on the testing side, and you have to make sure the tests are comprehensive. But also they're using all the testing techniques, and there's a lot of testing techniques out there. We have a lot of knowledge about how to do testing well. And you know what? The agents are lazy. Humans are a little bit lazy. So getting the humans to actually be very disciplined about their testing is also challenging. And I think it's actually easier with agents. You know, you can kind of instill that. You can get them set up. You have to give them the guidance. Use property-based testing. Use metamorphic testing. Use, like, these advanced testing techniques, deterministic simulation testing. You know, there's like technique after technique you can use. And humans are always like, I'm lazy myself. Like, I'm pretty good about being disciplined about testing. At some point you're just like, okay, that was enough. We got to ship. And, like, now you can be a little bit, like, you know, a little bit stricter and firmer. I think the same thing applies over on the security side, the performance side. I mean, we've always had ad hoc approach in the industry. We know security coding practices. Yep. You can literally have every single line of code on every commit reviewed by a security expert doing, like, an adversarial security review by an agent or multiple agents. Or multiple agents, right? And, you know, we're going through a tough time in the industry right now with, like, kind of hacks and security leaks and whatnot. There's only a limited number of bugs that can be in software. I think we'll get it, you know, out of it on the security side and on the quality side. And then also just on the—when I say quality, it's not just bugs, but it's also, like, you know, just little things like, oh, that UX element is wrong. And there's no excuse for that right now. Like, fixing it is so, so easy. And what we're also starting to see evolve is, like, you know, designers using Figma. Well, I think that should be a thing of the past right now. Our designer just deals with HTML and CSS and JavaScript directly, and sometimes is even just producing pull requests, you know, PRs, which she loves, we all love. Everybody's happy with this. There's not any of this waterfall handoff. So she's producing pull requests for the production codebase. Yeah. Yeah. And, you know, this is on the UX side, not on the core database. Of course, but this is the area that, like, she owns, right? Yeah. Yeah, it's just, I mean, she loves it. Everybody loves it. I mean, there's no downside. This is not new. Yeah, for sure. Yeah. What about code review? What's your take on code review? I think it's pretty—it's starting to get a bit controversial. Like, is it going to stay or not? Because it's been a practice that's been around, like, think about, I mean, you... This, it's gonna be really hard right now because anyone who's coming in explaining how to use AI or explaining how to be a better software engineer, they're gonna be out of date, right? You just got to get in there, be using these tools all the time yourself, and using it, like, use it to learn. I feel like I've learned more in the past probably even year than the previous five years combined, which is weird. Given your trajectory and given the environment you were working in, right? I mean, like, everybody's been in the industry for a while. Like, I'm definitely a better coder. I was a better coder 10 years ago than when I first got in the industry. It's like I can look back every decade and realize, like, I got a lot better, and I feel like I just got a lot better over this past year. And this was also one of the reasons I was really excited to talk to you, because when we started to just exchange messages, the first thing you wrote to me when I asked, like, hey, you know, how are things going? You said, like, you wouldn't believe, but my coding output is insane, and it's high quality, and it's database-quality level. And those were the things I don't really usually see. I usually see, okay, I'm not producing more code, but it's slop. But again, like, to me, this is a bit of an inspiration. Like, look, like, you can use these tools to just, like, amplify yourself as a software engineer. Like, you are one example, right? Hopefully one of many. Yeah, yeah. No, I'm not the only one in Cockroach Labs. We have other people doing this as well. I find it very exciting. You know, it's a little bit exhausting right now, but it's very exciting. Like, I got into software engineering because I like building stuff. I can build stuff faster. You know, the stuff you might have had to compromise on in the past, I think you can take away some of those compromises. I mean, you see this in the UX of software coming out. I think the UX is a lot higher. You see all the fancy, like, web animations and whatnot. But that's only just, like, the surface level. It just extends way, way below that. Peter, this was awesome. Thanks for coming on the podcast. Yeah, this is wonderful. Thanks for having me. One reason I was excited to talk to Peter is because he's been a very high-profile and productive engineer, pre-AI, building some of the most resilient distributed systems in production. CockroachDB is known for its resilience and how even if several nodes are destroyed, the database still operates without data loss. Basically, it's as hard to get rid of as cockroaches are, hence the name. One interesting part of our conversation was how Peter built more efficient data structures than the standard coding libraries had, thanks to him and colleagues paying attention to parts of the library that seemed slow. He did it for the C++ STL map and then Go for the Swiss table implementation. I found both stories a good reminder that you can improve the existing library or even the language, especially if you measure which parts feel slow. Another part of the conversation that I liked was how Peter came a bit of a full circle. He used to write 100,000 lines of code per year, being a very productive engineer and CTO. He then stopped writing code, aiming to coach engineers between 2022 and 2024. And then he started to code again because with AI tools, he wanted to coach his engineers better, but it's hard to do if you don't use the tools yourself. And now he finds himself being extremely productive, and this time the team around him is productive as well. And we're not talking about vibe-coded software, but database-worthy, high-quality code generated and committed to production. Peter is convinced that AI amplifies existing expertise. And this is one reason why he probably learned more this last year building with AI than the previous five years combined. And I find it a valuable reminder that learning and building deep expertise in software engineering, this is very valuable. And as closing, I appreciated that Peter said that not only is he excited, but he's also exhausted. There's a lot to learn, but it's tiring, and neither him nor anyone I know is immune to this. So if you're also exhausted with all of the things going on with AI, know that you're not alone. Check the show notes for more of the Pragmatic Engineer deep dives on Google's engineering culture and on distributed systems. If you liked this episode, please make sure you're subscribed in your podcast player, and a special thank you if you leave a rating. Thanks, and I'll see you in the next one.