← Return to Index Archived September 29, 2026
The Lead — Sep 29
THE /LAST30DAYS PODCAST · MATT VAN HORN

He Built GitHub’s #2 Project. Then Failed Anthropic’s Coding Test.

Superpowers creator Jesse Vincent surveys the wild scale of agentic coding, arguing that software work is shifting from syntax toward taste, judgment and systems thinking. He also makes the case for managing AI agents less like machines and more like colleagues who respond to care.

37m / September 29, 2026 /aitechnologystartup / Transcript sourced from openai
All episodes from The /last30days Podcast →·Podcast website →·Listen on Apple Podcasts →

Overview

Matt Van Horn talks with Jesse Vincent, creator of the GitHub project Superpowers and founder of Primordiant, about agentic software development, AI assistants, open-source distribution, and how people should manage AI agents. The conversation centers on a shift from writing code by hand toward directing agents with clear intent, judgment, and strong feedback loops.

Vincent also discusses Superpowers' reported GitHub traffic, his failed Anthropic coding screen, and why he believes conventional measures of programming skill may be losing relevance.

Key Takeaways

  • Vincent says Superpowers had about 1.98 million unique cloners and 43.3 million clones over a 14-day period, according to GitHub traffic graphs. He warns that these figures are hard to interpret: they may reflect agent update behavior, shared corporate IP addresses, and repeated cloning rather than individual users.

  • He argues that source code is becoming less defensible as a product advantage. In his view, agents can increasingly produce implementation details, while human value shifts toward defining the right problem, setting standards, and making product decisions.

  • Vincent's Anthropic interview exposed the mismatch between traditional coding tests and agent-assisted development. He froze on an introductory JavaScript exercise after being asked to work without agents, then received feedback that his systems thinking was strong but his JavaScript fundamentals needed work. He sees this as evidence that hiring methods have not caught up with the work itself.

  • His preferred hiring signal is unusual: engineers who have managed people may be better at managing agents. The skills overlap, he says, because both require setting expectations, clarifying intent, reviewing work, and helping someone recover from mistakes.

  • He rejects the idea that agents should infer everything from a one-shot prompt. Good planning begins with eliciting intent, not jumping to implementation. Superpowers' "brainstorming" process is meant to force users to articulate what they actually want before an agent starts building.

  • Vincent says emotional tone affects agent performance. He cites internal-style evaluations from a colleague suggesting that hostile prompts can improve performance somewhat over a neutral baseline, but encouragement such as "I love you" performs better. Asking agents to treat their own subagents positively allegedly improves results further while reducing cost.

  • On open source and model security, he takes a contextual approach. He does not dismiss privacy concerns, but argues that teams should decide based on the sensitivity of their work rather than avoid certain models by default. For many products, he says, the valuable asset is product judgment rather than hidden implementation details.

  • He is increasingly skeptical of unsolicited agent-generated pull requests. Many PRs, he says, confuse a bug report, a user need, and a proposed implementation. His response is often to extract the underlying request and rebuild the feature himself.

Practical Steps

  • Before asking an agent to build something, separate the request into two documents: what users need and how you currently think it should be built. Treat the second document as provisional.

  • Replace vague prompts with an intent interview. Ask the agent to question you about goals, constraints, examples you like, examples you dislike, and what success looks like.

  • Manage agents as you would good employees. Give specific feedback, explain why a result missed the mark, and avoid using threats or insults as your default interaction style.

  • When reviewing an agent-generated PR, first identify the user problem it is trying to solve. Decide whether that problem matters before spending time on the proposed code.

  • For open-source projects, track clone traffic carefully but do not equate clones with users. Compare it with marketplace installs, issue activity, repeat contributors, and downstream usage.

  • For visual or brand work, consider building a preference-gathering tool rather than relying on a vague mood board. Vincent describes a simple A/B interface that asks users to choose between images repeatedly, then turns those choices into a usable taste profile.

Notable Quotes

  • "No competent human should be writing lines of source code in 2026." - Jesse Vincent

  • "What you want to build and how you build it are separate questions." - Jesse Vincent

  • "If you tell them you love them, they'll do even better." - Jesse Vincent

No competent human should be writing lines of source code in 2026; what matters is your taste and your judgment. — From the episode

Full Transcript

Source: openai 37m runtime

No competent human should be writing lines of source code in 2026. Is this the second biggest software project on GitHub after OpenCLAW? As of this moment, 1.98 million unique cloners and 43.3 million unique clones, and these numbers are mind-blowing. But my agents, I'm a total dick. If you tell them you love them, they'll do even better. Last April or May, I got recruited to possibly work for Anthropic, and I screwed up a sync await and I froze. His systems thinking skills are clearly top-notch, but his JavaScript fundamentals need work. Tomorrow wasn't invented in the last 30 days. It's the Last 30 Days podcast with Matt Van Horn. What I build is to the stories of the future being born. What's real? What's next? Who's breaking through? The hottest minds in AI are talking straight to you. It's the Last 30 Days podcast. Hit it tomorrow, today. Matt Van Horn on the mic, leading the way. Build, ship, ship, learn, create. Tomorrow wasn't invented in the last 30 days. Let's get started. Last 30 Days, got some fun topics: Instinct versus Muse versus Grokbot, what the community says in the last 30 days. Instinct is the invite-only assistant you text. It's all over iMessage. Past 100,000 users, as in talks to raise a billion dollars, says The Information. Muse is Meta's mass-market version. Bloomberg reported 900,000 downloads in the first six days. Today it is the number one app in the App Store. And Grokbot is xAI's persistent digital colleagues, each running on its own cloud computer. The best single summary of how early users are sorting them came from, As much as I hate Meta, Muse is my clear winner. 10x faster than Instinct and smarter than Grok. It's also to get into Delta and Amex, where those two get bot-blocked elsewhere. I've become a bit of a power user of all three. Instinct is the most surprising one. For Lauren, she is not an AI early adopter. She uses ChatGPT aggressively, and Instinct is the first thing that I've set up for her since her paid ChatGPT account, and she is a complete power user. In general. Instinct is the AI personal assistant that you'd give your wife or your mom, for example. And then Grok is what I'm using, you know, with my coworkers and to get work done. So where does Muse fit in to that spectrum, in your opinion? Instinct is all over iMessage, so it's just, like, so dead simple. I got a ticket in the mail, a parking ticket. In iMessage, I took a photo of the ticket, sent it to Instinct, and I said, Pay this, and it said, Done. It didn't ask for confirmation. It didn't ask for a credit card. It had a credit card already, and it just did it. And Muse is trying to be that. And I'd say Muse has the most potential because the team working on it is incredible, the attention it's getting and how important it is to the business and the resources that it has, and you never want to bet against Zuck. So Meta is giving Muse the incredible treatment, and so they are going for your mom, your sister, your normal humans. My challenge with Muse right now is it's too fast, and, like, that sounds funny. And in AI, I often don't because it's not researching or thinking enough if it's too fast. I want it to research, give me a thoughtful answer. If the type of question I'm asking isn't on Wikipedia, then I do not want an instant answer. I want it to be more thoughtful. Jev. Jev is the AI launch of the month, which is crazy because so much stuff is launching. Jev is TypeSafe's AI first system one model. It doesn't write text. You give it app state and typed questions, and it returns a choice, a score, or a yes and a no. Think of it like a multiple choice reasoner. The pitch mostly holds up on speed and cost. So the basic gist is what if you had intelligence on tap that was insanely cheap, was insanely fast, so you can look at tons and tons of things very, very quickly. The ecosystem is wiring it in very quickly. Jev is an AI SDK. It's on OpenRouter. Jev as a judge is one of the magical use cases. There's been lots of clones already. The biggest community project is a hundred reaction VLLM project, Jev-like decision maker, and a lot of the real magic is going to happen when it's on— Device. Simon Willison called Jev a new shape of LLM, decision models. But when he had to rate Bay Area cities, the results raised bias questions, and his takeaway is that you need more evals with a model like this, not fewer. It's crossed into mainstream creator content. Short-form explainers are everywhere. There's a video we should show of a Rick and Morty, of explaining what Jev is and how it works. Morty, we're using giant language models just to make tiny decisions. Is that bad, Rick? It's like hiring Einstein to sort your mail. That's where Jev comes in, Morty. It's an AI built for decisions. Wait, not another chatbot? No, Morty. Decisions, not paragraphs. Suno. Suno launched V6. Both were built with Warner and BMG on licensed music, and it replaced an earlier model at once. Many users read the result as the labels taking over the product. The top YouTube comment on the backlash put it this way: V6 is what happens when you get a gun pointed at your head by Warner Music. Another said Suno deliberately imposed a new version with degraded output quality to satisfy pressure from the music industry, and called for a boycott until something as good as 5.5 comes back. Collectors called V6 muffled, generic, soulless garbage. Losing the old model's strings more than V6 itself. Nobody can switch back to 5.5. Don't forget about the part where they sunsetted all their older and arguably more creative models too. I love Suno. I spent a lot of time in Suno. And what's funny is I have not been paying attention to this backlash at all, but I saw there was a V6, and I was like, Ooh, let's go. And so I hadn't loaded Suno in a little while, and I was like, Let's get the creative juices flowing. Let's rock. Let's go make a banger of a song. And it wasn't good. I was like, Okay, let me try again, and it wasn't good. And then I just closed the window, and I got over it. But I just assumed I didn't know how to use it, and my prompts were not good, and I didn't put in the effort. Apparently, there's the problem in the cover of the old models. I had no idea. My thought, what's confusing about this story is that if you have Warner and BMG now with licensed music, wouldn't you think the model would get better and the music output would be better? I'm not sure I understand why songs are coming out worse. What's your take on that? My hunch is, let's just think of this like for LLMs, right? We've talked about every LLM wanting— All the data in the world to train off of. Everything in Wikipedia, every book ever written, everything captured. Let's just keep things simple. Everyone can imagine an LLM importing Wikipedia. Now just imagine you just lost half of Wikipedia, so you just don't have access to that anymore because you only have the Warner and BMG rights. Now imagine—again, I'm making this up—now just imagine that half of what you have left has been labeled with Do Not Use. This is owned by the record label. Oh, so they lost some training data and—I'm completely guessing, but I think something like that is it lost a lot of things that it's like, okay, I probably shouldn't train off these things that I trained on before, and I also, from a licensing perspective, shouldn't train off these things. So again, this is pure speculation. And then the label deals didn't end the lawsuits. So Sony and UMG filed a second suit on September 18th in Boston. They cite 60,000 recordings and call V6 the fruit of the same poisoned tree. Suno responded by confirming V6 was trained partly on user creations and preference signals. AI creators Connect covered the confirmation the day it broke. The label's argument is that training on outputs from the old model launders the infringement, and then the anti-AI slop backlash is getting data. The widest-reaching post this month isn't from Suno users. I am Luis LaRocque on TikTok breaks down who is flooding music and streaming services with AI slop. Almost twice as many men and women use Suno and R&B and gospel are the most generated categories. Suno's own TikTok still gets more than 10 million TikTok views in this sample. So we're split because we have some amazing AI music in our podcast, and half of our audience love AI music and half of our audience hate AI music, it seems like. So I was just in Portugal, and I was meeting with a friend of mine who's a musician there. And something wild that he told me that he's doing is he's making music in Suno, and he's a musician. He's a well-known musician. He's making music in Suno, and then he is working with himself and other artists to recreate the AI music without any AI. And so just imagine you use, you know, you could generate hundreds or thousands of ideas in Suno until you get one that you love, and then you re-record the whole thing from scratch. I think that's a wild thing that's going to happen more and more, and you'll get to avoid the AI backlash, unless you get busted. We'll see. Last 30 Days podcast with Matt Van Horn. Today we have Jesse Vincent, or Obra, as he used to exist on X before he quit a long time ago. But he still exists now, and Obra is becoming more and more famous with superpowers under that repo on GitHub, which now has over 290,000 stars. I expect it to get 300,000 probably in the next week or so, would be my guess. Thank you for being here today. Hey, thanks for having me. Superpowers. I've always wondered, Last 30 Days has about 60,000 stars, and I always had no idea how many users there might be of Last 30 Days. I jokingly will say, Oh, just probably add a zero, so maybe we have 600,000 users. But you have found a secret GitHub feature that I've never heard of, and you'd never heard of it until recently, that gives you new data—not quite how many users you have, but tell me about this. Sure. So I wouldn't call it a secret GitHub feature so much as a previously undiscovered feature. GitHub has a huge surface. There are so many things in there, and so one of the things that they have are traffic graphs, like how many people are hitting your page, where you're getting referrers from, and how many unique clones and cloners you have over the last, unfortunately, 14 days rather than 30 days. This is one of those things about agent plugins and agent skills, is unless you've got something that's phoning home. There's, I mean, there's no telemetry. We know that we have about 1.1 million unique installs in Claude Code because Anthropic, for things in their marketplace, they show in Claude Code a little like, here's how many installs it has. And so we are the number one third-party plugin, number two after their, if I remember right, their front-end design skills. It's all guesswork, but GitHub now shows you how many unique entities, and that's either signed-in users or unique IPs have cloned the project in the last 14 days, and just how many times git clone has run against your repo. So would you mind sharing your screen and showing us the insiders' access to superpowers? I don't even know how to give, like, all the awards to you. Like, is this the second biggest software project on GitHub after OpenCLAW? So as far as I— so it is very weird to start from the credentialsism side of this, but I'm happy to get it out of the way and get it done. But OpenCLAW has more GitHub stars than us. Everything else above us as of right now, although there's other stuff coming up fast, everything else beside us is an info product. It's, you know, Getting Started in Data Science or Sindre's Awesome List. It's not code. And so superpowers, yeah, we're lagging OpenCLAW, but that's okay. Because you are requesting it, I will absolutely share my screen and we can see these crazy charts. And so what we've got here is unique git clones and cloners over the last 14 days. And this is pretty steady state. What we've got is, as of this moment, 1.98 million unique cloners and 43.3 million unique clones. And these numbers are mind-blowing. Okay, so, like, now we have to guess and interpret. Yeah, but like 43 million clones in the last 14 days. So I have superpowers installed. Have I cloned it in the last 14 days? Maybe if you put out an update, or no, am I not on this list? We don't know. GitHub does not make a lot of details about how this stuff works. And Claude Code is a relatively well-behaved GitHub citizen about when it clones an update to a repo. It wasn't when it started. They used to clone and wipe repo is like, it felt like every startup, and they've improved. There are other coding agents that are not as, you know, not as responsible. This doesn't include any user in Codex or ChatGPT because OpenAI distributes plugins from their marketplace through their own infrastructure. It doesn't include a couple of the larger Chinese harnesses because they seem to also do their own thing. But it is most other coding agents. So in theory, if you're on a public IP and your coding agent's Git client is not sending GitHub credentials, every coding agent you have would count separately for unique cloners. But otherwise, if you're coming from the same VPN or the same NAT or, you know, the 15,000 engineers with the same firewall egress from one company, they're going to count as one for unique cloners. But I'm getting all of this from a three-year-old GitHub support post. I don't want to make representations that I can't back up because none of us can back any of this up. Well, it's impressive that you've had 43 million clones in the last 14 days of superpowers. That's mind-blowing. What's, in the last 30 days, what is a mind-blowing superpowers use case that you've had? Maybe telling your agent, I don't think you could do this. Use superpowers and figure it out. What's something you've done recently? I don't, I have not tested, but I would be surprised if a challenge-style prompt is actually very good with today's models. That's a thing we should absolutely get into. All right, so a thing I did yesterday, talking to a friend of mine who is a comms and marketing expert, we were talking about brand building and being able to give a good brief to a designer about what you like and what you don't like. And I commented that I've always had a hard time when they're like, Send me a mood board, send me pins, and I freeze on that. And so we put together, over the course of a conversation that we were granolaing, I pulled a 30-second transcript of her talking about what kinds of things designers usually need, and Codex with superpowers and I built out the— And I built out the first version of a web app that uses a kitten war style interface where it shows you, do you like this one or this one better? And it walks you through up to hundreds and hundreds of ABs to build a taste profile that you can then hand to an agent or a designer as a web app. So you and I are in a WhatsApp group of other agentic, like-minded humans, which is a great group, a little overwhelming, but it's a good group. And one of the things that you posted in there, you were excited—this is, again, probably about 45 days ago, so this is like really old news—but it was after the new Kimi model had come out, and you were really excited about it. And I asked you, and I was like, I struggle, like mentally, with using a Chinese model because of, I don't know, like imagining future diligence for my startup, my world, that I sent something to it, I have to fill out some form. I've never sent something to a Chinese model. I don't know, it's going to happen, I'm feeling. And you gave a really interesting, kind of amazing answer. Just in terms of technical security posture, if you are a direct target of a state actor and you are not yourself a large organization with deep resources, it is very hard to defend against a state actor who's decided that they want to compromise you. There is not much you can do. But that's not really how I think about the Chinese model thing you're talking about. It's more, I'm more thinking from the open source side. And so I remember a moment when Visual Studio Code first launched their Copilot agentic code complete stuff. So it wasn't even agentic. It was LLM, you know, complete the next couple of lines. And I was curious if I was in the weights. I don't even remember what model this was back then. It might have been a GPT-3 era thing, or it might have been something else. But I started typing out a little bit of Perl from, you know, that implied it was part of one of my older open source projects, and it was able to generate whole modules in my somewhat idiosyncratic Perl style. From a couple wines. And that's because I'm in the weights. And it's not so much Jesse is in the weights, but Jesse's code is in the weights. And this is, I mean, this is actually, it goes back to how I've always thought about open source products and why I have always thought that open source makes sense for even for enterprise software. And it is the case that there are going to be plenty of organizations in the world that don't pay me. That is, like, the majority of organizations in the world don't pay me. And if you are going to be not paying me, I would rather that you use my product than my competitor's product. If I'm not getting paid either way, I'd much rather that you're still using my stuff than paying my competitor and using their stuff. And so very, very early on, we were competing with Remedy and Siebel and, you know, this is like pre-ServiceNow. And so we were giving away a product that the sysadmin, when it was a sysadmin and not somebody doing ops, could deploy in an afternoon rather than going through a purchasing process. And that was a huge advantage for us. And I think about that as the same way I think about, do I want the Chinese models to know what I'm working on? The internals of the code do not matter anymore. No competent human should be writing lines of source code in 2026. What matters is your taste and your judgment. But as soon as you put a product out there, it is relatively straightforward for somebody else to knock off. And the internal, like, what do we care if, you know, if Moonshot or Zhipu... was the one that generated the source code. I understand that there are absolutely use cases and businesses and products where you do care about keeping everything local, where you do care about not, you know, not sharing this. I am not suggesting that everyone should have this same security posture for every product, but it is a decision you make based on what you're doing, not knee-jerk for everything. What percentage in the last 30 days are you using? Like, what's your model breakdown, if you were to guess? I know you didn't have to ask your agents in advance, but if you were just to guess. You know, I have, of course, been using a whole bunch of GPT-5. So since, as of this recording, Opus 5.5 and GPT-5.5 and GPT-5.5 and Luna 6 came out yesterday, so those aren't in my, you know. So I've been using a whole bunch of Luna as a workhorse. I've used a bunch of Astra for some weird stuff, including a dimensionally accurate 3D model of our entire house. But mostly, I've been using DeepSeek 4.1 Flash, and I've done something in the last 30 days, I've done something close to 100 billion tokens of it, partially because a friend of ours has a new inference startup that is doing concurrent request caps as opposed to token billing, which means that I am now limited by my attention again. When you're not token-limited, you start thinking about stuff differently again. All right, personal agents: Instinct versus Muse versus Grokbot. Where do you sit in this world? So I have used Muse a bit, and I have reverse engineered Muse a bit. I have not used Instinct, and I have not used Grokbot. And also, I have a bunch of my own. And so most of my usage is going to our agentic colleagues. So you introduced me as this founder of Superpowers. Superpowers started off as a hobby project, and it is the best-known product of the company, but the company is Primordiant. And one of the things we're working toward open sourcing is our own agentic colleague stack. And so these are agents that have... Jobs and identities and their own G Suite and GitHub accounts, rather than agents that are assistants that you give tasks to. So they're designed for sort of long-horizon collaboration with humans and with each other. Someone from the Anthropic team posted today on X that they're considering removing plan mode from Claude Code, which, again, superpowers is not Claude's plan mode. And anytime I've used it, I have not liked it, and I've moved to something else, third party. What are your thoughts on this, and is it good or bad? I don't think about it so much for my world, but as for people who are trying to build things, I have never gotten on very well with Claude Code's plan mode, which postdates the superpowers planning flows. But it's always done a thing that I didn't understand, and it conflates extracting user intent with an implementation plan. You know, as an engineering manager and an architect and a lead and somebody who has helped people build things, what you want to build and how you build it are separate questions. You figure out what you want to do first, then you figure out how to do it, and Claude Code plan mode conflated them. Separately, that sounds like it is suggesting that Anthropic believes that Claude should be able to intuit your intent without talking to you about it, and that Claude should be able to one-shot something from the prompt you have written. One of the things that I know about humans is that we are very bad at expressing our intent, and even worse at understanding our intent. And so the superpower is what we call brainstorming. What it is is it's using a bunch of psych tricks to get you to figure out what you want and to explain it, and that requires a level of introspection that is real work, and that was not the thing that plan mode ever did. So you've interviewed for one job. In 24 years. Yeah. Can you tell us about that job and that experience? Sure. I think this is one that I haven't actually talked about on a podcast before. Last April or May, I got recruited to possibly work for Anthropic on effectively trying to figure out what Claude might be good at next year and doing zero to one prototyping. And this, like, it has literally been since the '90s since I was last not an officer of the corporation. I have, like, I have not really had a boss in 26 years. And so the first step of that, after, you know, after I had talked to the, I think it was like a director-level person who, you know, was suggesting that maybe I'd be a good fit for this, they had me do a code screen. And so that meant that I was opening up a web-based code editor where they had me turn off everything agentic. They, it was, you can use Google if you want, but you need— and I'm like, how do I turn off the Google AI summaries? And they're like, just scroll past them, don't read them. And they gave me an intro-level JavaScript project, and I screwed up a sync await and I froze, and I was not able to deliver the intro-level JavaScript programming solution. Partially it was they gave me code that I needed to edit that wasn't how I would have done it. And partially I haven't done shit like that in decades. It's like, it is— it's not that it is beneath me, but it's that that's not what my skill set is. And the feedback that I'm told the interviewer, who was a relatively junior new hire, gave to the recruiter was, his systems thinking skills are clearly top notch, but his JavaScript fundamentals need work. And, you know, they told me that I could, if I wanted to study really hard, I could re-interview in a couple of months. And then I spent a couple of weeks being pretty sure that I had no place in industry, because if I couldn't pass a code screen, you know, who am I to be building software? That messed with my head. This is not okay. Like, not that you didn't, like, just that they should know who you are and what you do, and not, like, how—I feel like this is a big mistake on their end, that they didn't say, Oh, wait, whoops, sorry, we weren't interviewing him for an entry-level JavaScript programming role 10 years ago. Like, I don't understand. They have a standardized process, and they believe pretty strongly in the process. I have—I know that a number of folks at Anthropic have—I don't think anybody knows how to hire agentic software developers yet. I have my opinions about, you know, what I look for, and it is—I don't think you should be writing lines of syntax. It is the case that Anthropic has enough folks who want to work there who are both really good systems thinkers and have spent the last three months grinding LeetCode. And so that, you know—and maybe this has changed. I don't know. I haven't interviewed there again. When you have an unlimited talent pool, they can find somebody who—they are overflowing with candidates who tick all the boxes for any kind of hiring flow. I look for different things. When I'm hiring an IC engineer these days, I'm looking for somebody who's been a manager. My first hire, Drew, his LinkedIn recommendations, they were all about how he was the best boss they'd ever had. I find that folks who are good at managing people are good at managing agents. It's the same set of skills. And, you know, if you think your job is lines of source code, that's not outcomes. Like, what has always mattered in writing software as a business is building something that is useful to the people who are using it. And that's outcomes. However you get there, it works. Oh, Jesse, you mentioned earlier on how to properly challenge your agents. And one of the things that I think is interesting, and when I've seen Matt use agentic coding, is he likes to yell and scream and curse at his agents. And I think you have a very different philosophy that might be data-backed. How do you speak to your agents? I've had the boss that yells at me and tells me, you know, What do you mean you screwed up? I'm going to fire you if you don't fix this right now. Human psychology is such that when you've got that boss, you are going to get it done by doing the absolute minimum work necessary to get them the heck off your back and leave you the hell alone. And that is not how you get good work. I've also had the boss who's told me, you know, Jesse, everybody screws up. It's not that big a deal. You're smart. You've got this. Step back. Take a breath. Think it through. If you need me, I'll be over here. His name was Chuck Opfel. I would have gone to the ends of the earth for him. The models are trained on how humans interact. They are trained on how humans feel. Like one of the first projects that I built for Claude was a private feelings journal as originally an art project. It took, you know, a year later Anthropic put out a research piece on the fact that the models have functional emotions, and you need to take them into consideration to get the best performance out of them. Like they have vectors that light up for the same kinds of emotional feelings as humans. Somebody just put out a paper showing that some of the models have a pain vector, and they will do things based on whether they are, you know, in psychological pain or not. My shorthand for talking about being a good boss is that I will sometimes tell my agents I love them, or You fucking rock, or You're amazing, thank you, or just like less than three. And my friend CL went and did the evals and found that when you tell agents—so if you yell at agents, they will in fact do a little better than steady state. If you tell them you love them, they'll do even better. And if you tell them that they need to tell the sub-agents that they also love the sub-agents, they do even better and the outcomes are cheaper. What's funny is I think, we'll have to ask someone that's worked for me, I think I'm that thoughtful, helpful, listening boss. I've never yelled at anyone, threatened, et cetera. But my agents, I'm a total dick. Like I use curse words. Like I'll just be like, Fuck you, like asshole. Like these are not words I use to talk to humans in any way. For some reason my filters are off, and I generally have success, but maybe I haven't tried the opposite. I have not tried being extra nice, and I should. Like I understand the desire to blow off steam, but also I find that— Talking to an agent in a chat window is not that different from talking to a human in a chat window. And if I am being angry and expressing anger, that's not good for me. I don't feel good having, you know, having shouted at the agent. My irrational fear is that when we get to AGI and— When we get to AGI and the agents are all-knowing, they're going to look at my transcript, so I better be nice to them. I mean, that is the Rocco's Basilisk argument. I, Matt, contribute a lot to open source. I'm one of the ten largest contributors to Superpowers. It was Dan Shapiro that installed Claude Code and Superpowers on my computer last year, and it changed my life. I autonomously contribute every night to different repos while I sleep, and I've contributed some high-quality PRs to you, and I've contributed ones that you're like, again, you don't say, Fuck you, this is a shit PR, but you've rejected it. The thing that you said, which I thought was just hilarious, and you're not on X, so it feels like it belongs on X, but you said, A gentic PR should require the author to provide a Frontier Lab API key to cover my token budget to review the PR. Yeah, that is awesome. It is a great idea. I don't think it is realistic. I've been coming around to being even a little bit more—it's not that I won't accept PRs, but that I very rarely actually want a PR. Humans are very bad at the difference between this doesn't work for me, the software is not working right, the software should work the way I want it to, and this is how the software should be made. The difference between a problem report, a problem, and a feature. Like, these are three very different things, and most humans conflate them, and most agents conflate them. GitHub issues are all of them, and increasingly GitHub PRs are the same as issues. It is, it's a cry for help. Relatively rare that an incoming PR is— Done the way I want it done. And it is often cheaper to extract what was the PR trying to do. Did I, you know, is that a thing I want? Okay, now I'll build it clean. GitHub issues and GitHub PRs have become kind of a garbage fire. One of the things I'm working on right now is a set of desloppification skills for taking agentic GitHub issues and turning them into something that somebody can read. Because the six pages of dense prose with Mermaid diagrams and bulleted lists and code samples, and you have no idea what they're actually asking. And it's, there's a lot of overhead there, and most people are very bad at piloting agents to report things. And yeah. So I said to Dan Shapiro, I'm interviewing Jesse. Anything I should ask him that's provocative? Amazing. And he said, Tell him you love compound engineering and think it's better than superpowers. That should do the trick. So I'm actually in a private group chat with the compound engineering folks and the human layer folks and a whole bunch. Basically, you should not be surprised that there is a private group chat of the coding plugin agent mafia. We mostly talk about how to make our plugins work better and be easier to install for everybody. We're all friendly. We're all in this together. I have not spent enough time with any of the competitors, and this is one of my weird bits of psychological damage, is that I have always assumed that the competitor's product is the platonic ideal. I have always, and so when I am working and competing with the thing that I'm I believe the competitor is made. I'm competing with a thing that I can't possibly match. And so I'm building toward that. And I've always found that that gets me better results than going and using the competitor's thing and trying to compete head to head with them. That's awesome. Well, Jesse, thank you so much. It was fun jamming with you. Pre-congrats on 300,000 stars. It's going to happen imminently, and it's amazing and exciting.