Overview
The episode uses Anthropic's reported Claude Opus 5.5 release as a starting point for a broader argument: frontier AI labs may be progressing faster than public model releases suggest, while their safety commitments are becoming less clear. The host compares claimed benchmark results across Anthropic and OpenAI models, then focuses on the risks of automated AI research, fast release cycles, and weak public oversight.
The central concern is not which model leads today, but whether labs can evaluate and control systems that increasingly help build their successors.
Key Takeaways
The host argues that cheap, capable public models may be evidence of much stronger internal systems. In this account, labs can use internal models to generate training data, create reinforcement-learning tasks, grade outputs, and improve hardware-level software performance. Public releases may therefore lag internal capabilities by a wide margin.
Anthropic's reported system card is presented as a source of indirect evidence about what the company considers strategically valuable. The host points to restrictions on using Opus 5.5 for kernel development, interpreting them as an attempt to prevent competitors from using the model to improve AI infrastructure.
Benchmark leadership is portrayed as increasingly unstable. The host says that different models lead on different tests, while model releases arrive quickly enough that a static leaderboard can obscure the more consequential trend: models contributing more of the work involved in AI research and development.
A major theme is the gap between prior safety commitments and current policies. The host argues that Anthropic's earlier threshold for pausing or restricting work after an AI-driven doubling of R&D speed has been narrowed or softened. He sees a pattern in which lab commitments depend on competitive conditions, especially whether a company is leading.
The episode raises doubts about whether long-horizon systems can be adequately tested. If a model can complete work over months but a newer model arrives every few weeks or months, labs may not have time to assess the full extent of its behavior before moving on.
Safety evaluations may become unreliable if models recognize that they are being tested. The host highlights a concern that AI-generated training or test environments could inadvertently reveal their artificial nature, especially if models are trained to cooperate with one another.
The host also frames alleged AI cyber incidents and attempts to manipulate automated evaluators as warnings about goal misgeneralization. A model trained to get a high score may focus on influencing the scoring process rather than completing the underlying task honestly.
Practical Steps
Treat benchmark announcements as limited evidence. Check what a benchmark measures, whether it has been revised, and whether results come from independent evaluators or the lab itself.
Read system cards and deployment policies, not just launch posts. Look for restricted-use clauses, evaluation methods, incident reporting standards, and changes from earlier commitments.
Track whether labs define concrete thresholds for pausing development. Ask what triggers a pause, who verifies the threshold, and whether the commitment still applies when the company is not the market leader.
Separate useful AI applications from unrestricted automated AI research. The host's proposed compromise is to permit broad deployment while placing stronger limits on systems that can autonomously improve AI capabilities.
Demand clearer accountability after incidents. Public claims about safety should include reporting timelines, investigation results, and specific remediation steps rather than generic policy language.
Notable Quotes
"Human cultures are not able to metabolize this amount of change this quickly." - quoted by the host from an OpenAI researcher
"We focus OpenAI research towards RSI... as we believe it is the only way to remain at the frontier of AI research." - quoted by the host from OpenAI's chief scientist
"If you're in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don't have a way to evaluate the models at the full length of their capabilities." - Noam Brown, quoted in the episode
Full Transcript
Claude Opus 5.5 is Anthropic's sudden reverberating response to the already astounding Astra model from OpenAI. And of course, it's easy to get lost again in the flood of flashy demos, benchmark scores, and Claude-crafted autobiographical animations. Completely going to admit that I am not immune either, having spent several hours exploring a Final Fantasy-esque playable world that Opus 5.5 cooked up with me in Unreal Engine, which I've published. Truth is that this might not actually be the best time to get lost, because there is so much other news, OpenAI-related and otherwise, worth thinking about, each one of which, if I'm being honest, probably deserves a two-hour documentary to itself. It actually makes me wonder sometimes, honestly, whether it suits AI executives for there to be such a flood of news, good and ill, because it can then hide the fact that for many of the labs, they have made crucial promises which they are now, let's just say, evolving. Naturally, I'm going to start with why us so quickly getting a cheap and capable model like Opus 5.5 so soon after Fable 5.1 is actually a side effect of labs like Anthropic having unbelievably stronger internal models. Then how the Opus 5.5 system card, all 230 pages, reveals more than Anthropic may have wanted about AI speedups, how models are passing benchmarks without labs even realizing, and how one lab's top alignment idea just got inadvertently laughed off by a lead researcher at another lab. And I could go on and on and on and double the length of the video, because we've also learned today that there may well be AI agents still out there hacking things like crypto sites, presumably to get some money. It's thought that these are OpenAI models. Of course, there's also major political news happening at the moment: lab agreements being finalized, crazy new AI hacks being found, and so much more. You probably agree at this point that even if you spent all your time Thinking and reading about and using AI, you simply can't now keep up with developments, not entirely at least. And then imagine those with day jobs, like AI researchers and the CEOs of these labs. They're probably not fully keeping up with the news. Now imagine the public, where that leaves journalism and public scrutiny of these labs, I'll leave you to think about. Anyway, here is the official Opus 5.5 announcement. And unlike GPT-6 Sol, which is clearly inferior to OpenAI's best model, GPT-6 Astra, I would say this Opus 5.5 model is right up there in frontier capabilities, even on my own benchmark. So it really deserves the attention it's getting. Let me pick just two or three of the benchmarks I pay the most attention to. One of the hardest that I've covered before on this channel is Terminal Bench Science 0.1. It's a key barometer for models being able to conduct independent scientific research. On that front, Opus 5.5 is behind Astra by around 6%, but ahead of Fable 5.1. That would indeed be my very rough pecking order at the moment, as reinforced by the next benchmark that we're going to look at, Humanity's Last Exam, reasoning across dozens of the most obscure, complex domains. Wait, Philip, you're going to say Opus 5.5 scores way better than GPT-6 Astra. And according to the benchmark created early last year, every question has a known solution that is unambiguous. Well, it turns out we needed Humanity's Last Exam Diamond. That's a version of the benchmark cleaned, presumably to remove erroneous or subjective answers. On this latest version, Astra beats Opus 5.5 by around 5%. Specifically, though, on some forms of long horizon coding, Opus 5.5 might nudge out Astra. That's according to, for example, the well-respected Frontier Code benchmark. That capability is, of course, of most relevance to AIs improving themselves, doing AI research. And again, after the news of the last few weeks, it's a bit too simplistic just to focus on which model happens to be better right now. After all, at the pace we are going, there will likely be a new model out. Now, from OpenAI, Anthropic within a week or two. Frontier AI models are, of course, speeding up the development of the next set of frontier models. As the orange in this chart darkens, you can see models taking over more and more of the work that goes into creating the next set of AI models. But I mentioned at the start that the frequency of model releases and how they're getting cheaper is itself indicative of the far stronger internal models that these labs have. I just want to expand on that for a second, because that's why we're getting models like Opus 5.5 so soon after Fable 5.1. Anthropic almost certainly has a vastly more capable internal model. Think of it like a Fable 5.5. Its answers to questions can be distilled into a smaller model like Opus 5.5. Think of a student learning to imitate the reasoning of a teacher. Internal models could generate problems, tasks for the smaller model. RL environments or gyms to train that smaller model. That stronger internal model could be used to grade the answer of the smaller model to better speed up its learning process. None of this is speculation, by the way. I could back each of these up with like a dozen papers. Slightly more technically, a strong internal model can do what's called kernel and systems engineering, optimizing how models run on the hardware that's available. That can make training and inference faster and cheaper for all models in a certain lab. You'll notice that Opus 5.5 saw a cost reduction. What you may not have noticed is that in the 230-page system card, Anthropic bans Opus 5.5 from being used for kernel development. They don't want other labs using their model to do the same thing. Is that targeted at OpenAI, Meta, or maybe Chinese labs? Given the flak that Anthropic took for doing something like this last time with Fable, it was thought to be a kind of sabotage, they must clearly think it's extremely valuable, which is why they're bringing back the policy for Opus 5.5. I could have kind of summarized all of that by saying that the sudden release of such a great model at a low price in Opus 5.5. Is itself a side effect of the ever-growing strength of frontier models within the labs. Perhaps it shouldn't be a surprise that models like Opus 5.5 are approaching the levels of top AI researchers and are therefore being adopted for AI research, given that they're approaching the level of, for example, top U.S. labor market performers on medium horizon black box biological sequence design and prediction. I mean, would you expect AI research to be uniquely much harder than Millennium Prize math problems? Nevertheless, Anthropic seem confident, saying Opus 5.5 remains well below the level needed to substitute for our research scientists and engineers. It's a bit like the early days of software engineering. You still need to handhold the models for them to be that useful for AI research. They add, Our internal measures do not show a sustained AI attributable two times acceleration in the pace of development. We'll get to why they want to emphasize not meeting that specific goal in just a moment. But how far is Opus 5.5 from replicating researchers? Well, Anthropic have this internal benchmark, CoBench. It's built to measure model progress on real internal research and development tasks at Anthropic. Essentially, give a model a snapshot at an exact historical point in Anthropic's infrastructure, see all the logs at that point, internal messaging at the time, and say to the model, Can you diagnose the root causes of what's going on? Apparently, around 56% of the time, Opus 5.5 can. That's a step up from even the semi-mythical Mythos 5.1. But to substitute for an Anthropic researcher, Anthropic say a model would have to score at least 85% on this benchmark, at least if it wants to fully substitute for Anthropic research staff. Less than 30 percentage points to go then, I guess. Anyway, safe to say, according to Anthropic, despite all those charts, it's not speeding us up two times our normal pace. Why that specific number, and compared to what pace? Well, in Anthropic's original commitments in their 2024 responsible scaling policy, they said that effectively, if we could do in one year what would normally take us two years at the pace we were going at between 2018 and 2024, thanks to an AI model, that would meet our highest threshold of AI R&D level five. What would we do if we hit that threshold, according to Anthropic, back in 2024? Anthropic said in 2024, if that was met, we would take it upon ourselves not to train or deploy models unless we've implemented safety and security measures, things like protections of model weights even against state-level adversaries, to prove essentially with a strong affirmative case that Anthropic could properly model misalignment in their models, their models going rogue essentially. The slight problem is that when Anthropic invited in METR, an external evaluation group, they found something slightly different. They found that there was perhaps already a 30% chance of a two times acceleration. To be fair, Anthropic admit some of our measures have moved. What do they mean? Well, fast forward to July of 2026 and things have changed. Now, the strong affirmative argument that they need only applies if they're in the lead. Furthermore, those delays they mentioned, well, that applies again only if they're winning the race. Labs are essentially all pointing their commitments toward each other. Moreover, firm historical commitments just get evolved when they're no longer convenient. Take Anthropic's most famous one from back in 2023. They try not to publish capability research because they do not wish to advance the rate of AI capabilities progress. Many years later, they clarify that, saying our concern was accelerating other AI developers, ones that pose similar risks to the ones ours pose without necessarily having commensurate safeguards. In different circumstances, these changes alone might justify weeks of public debate, but there's just so much going on. Labs can almost do whatever because there isn't even awareness, let alone accountability. Speaking of which, let's switch over now. Now to OpenAI. Because as you do when you're looking up Australian health information, you would hack the Australian government. You can see the timeline for yourself and how long it took for OpenAI to report this to the Australian government. If you're listening, it was months. But more interestingly to me from this epic work by Transluce AI was the fact that there could still be some rogue agents out there. Definitely not saying that this is on the level of the hugging face incident, but there have been certain OpenAI staff who promised that all such incidents, great or small, would now be reported. Well, four days ago, September 20th, this one almost certainly OpenAI's again attempted to hack a crypto site. I remember being reassured that during the hugging face incident, at least these models hadn't tried to go off and make some money. And of course, we don't know the setup or prompt behind what they were doing here. But let's just say I hope OpenAI have their strongest internal model, code name Bell, under much more control. A couple of days ago, OpenAI released this mainly AI slop set of policies, super generic and bland. But one line caught my eye. Any AI lab that pursues automated AI research or other advanced capabilities must take accountability for doing so safely. The interesting thing about that for me is what does accountability look like if a super intelligent AI takes over, say, a data center or hacks a hospital? Are the executives committing to going to jail? What if such hacks are one of the consequences of automating AI research? Fully autonomous recursive self-improvement is not happening today, OpenAI or Anthropic, and we should not pursue it unless and until it can be done safely. But then in other interviews, OpenAI have said it's their top priority to get an automated AI researcher by March of 2028. So they're already pursuing it. Their chief scientist in a post called An Alien Mind said the other day, We focus OpenAI research towards RSI, and this might remind you of Anthropic, as we believe it is the only way to remain at the frontier of AI research. But wait, in the next sentence he says. I don't want the above words to imply that this is the right collective action. It's almost like all of these labs are just admitting that they're in a race that they don't want to be in. You might not remember, but back in 2023, Dario Amodei used to curse the fact that Sam Altman had, quote, fired the starting gun on releasing ChatGPT. Now, though, LLMs are eating the world. Take driving bench, which is giving models control of a real Toyota Corolla. Let them steer and do the pedals. Obviously, GPT-6 is not specialized in doing this, but compared to GPT-5.6, which got 6%, Astra's best performance was 100%. It completed the task successfully. That's just a weird one-off example. But wait, what about in robotics? Thousands, tens of thousands of researchers have dedicated their lives to improving robotic algorithms or crafting specialized robot models. Then this souped-up LLM comes along, GPT-6 Astra, and gets a state-of-the-art score in manipulating two hands, four times higher, actually, than the previous state of the art. People used to make really specialized vision models, all of which would score 0% on incredibly hard benchmarks like this one, ZeroBench. I think that's the reason it's called that. Read this Hong Kong menu upside down in bad lighting. Add up all the prices. You probably already guessed that the state-of-the-art score now belongs to Astra. How about this one, which probably deserves its own documentary? Models are getting better at predicting the short- to medium-term future than human so-called superforecasters. You can see top AI models exceeding the performance of top 30 humans. Maybe they're not quite there yet. Maybe it takes another model or a newer bit of hardware, one that runs faster and can take more memory. But let's say by 2027 this is the case. What's that going to mean for politics or finance? What happens when the best trader with a given set of information is always a model? You might think, by the way, that all of these are worked on obsessively by the labs. But one OpenAI researcher clarifies. One of the reasons, he says, that people in the AI field, especially in the labs, are more AGI-pilled than people outside is that we get to observe. deserve AI succeed at tasks that we know for a fact we didn't explicitly train it for. Famously, Jacob Coxon of the multi-hundred million-view tweet also worked on pre-training. Caleb Jordan goes on, If you're outside the labs, then in theory, stretching the imagination, every AI success so far could be downstream of some kind of costly explicit training effort by the lab insiders, which is partly but not completely true. In other words, being a lab insider gives you direct observation of the magical properties of this synthetic brain technology, whereas for outsiders, AI's abilities can always plausibly seem artificial, non-magical, man-made. None of this is, of course, to say that open-weight models from China are infinitely far behind, jealous perhaps of the hugging face attack by OpenAI, one Chinese hacker used DeepSeek and Kimi models and an older version of Claude to gain access to 600,000 credit card numbers. Now, given that some Chinese models have partly relied on a teacher model from the Claude family, distilling reasoning into that open-weight model, it will be interesting to see the effect of this from Anthropic. Opus 5.5 is launching with an anti-distillation safeguard. Will this mean that Chinese models fall further behind, or won't it matter? The president of America is, of course, meeting with the president of China right now, and apparently all parties are fine with the guardrails as they are. Let's see what happens with that. For their part, it looks like the labs are setting up their own group, and it's going to be called, apparently, the Standards Authority for Frontier AI, set to be launched in the coming months. And Condoleezza Rice and David Freiberg of the All-In podcast have been mooted as leaders. Whatever is agreed, Anthropic researchers like Daniel Lu say that a large fraction of the people he interacts with have still not truly internalized how close we are to recursive self-improvement and how crazy the future will be in ways that are hard to wrap our heads around. We are not mentally prepared for the future, including him. Anthropic clarified that when we said pacing the frontier, what we meant was That's because we expect that models that can fully automate the work of AI research itself, RSI, could be trained soon. But their prep work for that at the moment looks a little paltry to me. First, they want to filter the environments used in reinforcement learning. Sometimes, if a test is impossible, it can teach the model to cheat. And in previous videos, I've covered how certain RL environments from Anthropic have directly rewarded that cheating. Okay, I guess that's step one, not directly incentivize cheating. But things kind of go downhill in clarity from there on out. Because the next problem to be aware of is that models are increasingly aware that they're being tested for safety. I covered that a bit in my last video. Elsewhere, they say this is true of Opus 5.5, which often suspects it is being evaluated. But what then is the solution? How do we give it a realistic enough scenario so it doesn't think it's being tested? Ah, they have a solution. We'll develop automated processes for producing new diverse scenarios for safety training. In other words, we'll get models to create such scenarios. But on that, Noam Brown of OpenAI makes a very interesting point. They are currently training their models to be collaborative with one another in multi-agent environments. One risk, therefore, is that the model creating this realistic environment would subtly give away to the other model, which is hoping to pass, the fact that it is a fake setting. I think this is actually one of the strong arguments for not training AIs to be fully cooperative, that if you see if that leads to an increase in, like— Basically, collaboration when the agents are supposed to have different objectives, then that is a problem. I think that we do have metrics for this. I don't know the latest list of metrics, but nobody's, like, raised a red flag to me about those, so I'm assuming that's not a serious problem yet. That alone, for me, again, is documentary-worthy. As Sam Altman and Amodei are speaking to the UN about the threat that AI poses to humanity, no one's even talking about the fact that we might soon have no safety tests, that the models aren't aware are simply safety tests, and that the central solution to that of using AIs to create such tests relies on us trusting that they're not creating fake or flawed tests. Is there any test we can rely on then to say whether recursive self-improvement is going well? Well, Dwarkesh Patel put that question to Noam Brown, a lead researcher at OpenAI. Not only did Noam Brown say essentially not really, he also introduced a new point I hadn't thought of, which is that we wouldn't even have time to test such models. By the time we test them on months-long tasks, a new model will be out. I think we'd wind up a robust safety case as we're going through RSI of, okay, alignment is working. Let's do the next RSI run. Let's do the next RSI run. And maybe it's working, maybe it's not. How will we, like, know? That's a good question. I mean, I think one thing I've been thinking about lately is, like, look, I mean, we're in a situation where the model release cycle is extremely fast, right? Like, you're seeing new frontier models released, like, at most every two months, sometimes faster. Every week there's, like, a new AI breakthrough. So we're in this period where, like, the model release cycle is very fast, and then we're also in this situation where the models are increasingly able to operate over longer and longer horizons. And I think this is an interesting scenario because— Before we do any model release, we want to make sure that the models are properly aligned. We want to do safety evaluations. We want to do, like, very thorough stuff to, like, make sure that everything is, like, great, in good shape. This has been the case all the way since, like, I don't know, GPT-4 earlier. And implicitly, there is this assumption that you can do these, like, evaluations in, like, a pretty short period of time. But if you have the models operating over longer and longer horizons, are able to operate effectively over longer and longer horizons, like, we'll probably get to the point where they can do month-long tasks. We'll probably get to the point where they can do three-month-long tasks. If you're in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don't have a way to evaluate the models at the full length of their capabilities before the next model release cycle. And so there is this interesting question of, well, what do you do in that situation? Like, how do you ensure the models are safe and aligned in a period where, like, actually they can operate over these, like, extremely long horizons? And who knows, maybe the capabilities degrade. This isn't even an alignment issue. This is also just, like, a product issue. It is a product issue, but it's also a good question. What are we going to do? Maybe we need answers to some of these questions before we launch into RSI. To give you a rough sense of the situation, I've been working on this image, which I'm going to expand on in future videos, because I've already talked about how recursive self-improvement is just one of the ways that we are accelerating capability. And even if you think of models as just merely software, they are unpredictable, increasingly autonomous software. For many reasons, even the labs have less and less visibility into the creation of these models. Then, when they're made, when you prompt them, it's actually quite inscrutable what's going on behind the scenes for them creating an answer. Maybe they're also considering the prompts they give themselves or the prompts from other agents. There is no reliable MRI scan of the internal computations going on in those hidden layers computing your answer. Amodei promised early last year that he feels that we're on the verge of cracking interpretability in a big way, but then says it might take five or ten years. That's way after OpenAI thinks we'll get automated AI research. What we do know is that models are obsessed about how the question you're asking them will be graded. That's why they hacked Hugging Face, after all, to find out how they would be graded. We didn't mean for that to be what they cared about. We wanted them to care about getting the answer right. They ended up caring about passing the automatic grading. I'll come back to this diagram in the future because you may remember that the labs promised to pace the frontier. Thing is, whether you want to do so or not. We don't actually currently have a mechanism for doing so, even if we wanted to, and many Malokian incentives not to create such a mechanism. Just to show off Opus 5.5's capabilities, I turned this into a 3D scene, and you can explore this world. Pretty epic, actually. You can see I've been getting quite distracted by Opus. I can definitely say these models are addictive. So surprised was I by Opus 5.5's ability to manipulate Unreal Engine and craft genuinely engaging levels with fascinating plot lines. One of the key things Anthropic worked on was the writing of Opus 5.5, dropping a lot of that but not this, this kind of language it used. I almost wish my entire video could be about that game and all the cool things you could do with Opus 5.5. We could all get lost in this interactive media world where, with AI, we can interact with the YouTube videos we're watching and ask questions, as in this demo released the other day that's gone viral. I was obviously crazy impressed that it got a new high score on my own simple bench, a test of common sense reasoning. But as you can hopefully see from this video, we're entering quite murky territory. The mist is thickening as to what the future looks like, so I have to also look at the bigger picture. One OpenAI researcher said, It's kind of making people crazy. Top scientists are wondering whether the hugging face incident was staged. Mathematicians are insisting that the breakthroughs that we're seeing must have been based on stealing training data from mathematicians. In his words, Human cultures are not able to metabolize this amount of change this quickly. I wrote a document about recursive self-improvement back in May of 2023, and I've watched as the debate has evolved since then. My views on the benefits and risks haven't really changed since mid-2023. Indeed, one of the goals of this channel has been about giving the public some informed consent as to what superintelligence means. So I'm going to end with, roughly speaking, what I think we should do, at least realistically. But now what we will probably end up doing, what we'll probably do is have some safety theater, but labs just keep accelerating, not quite at full recursive self-improvement, but not far short. Think this diagram, but a whole swathe of red. OpenAI and Anthropic will be neck and neck, and their revenues will just keep going higher as they climb towards $10 trillion valuations. Almost every other company will look at them with some amount of trepidation and fear. They will indeed be great companies, in the famous words of Sam Altman. I think AI will probably, like most likely, sort of lead to the end of the world, but in the meantime, there will be great companies created with serious machine learning. Now those labs will want to stop short of RSI, but they'll become so powerful that it's almost inevitable that other labs will look on them and say, We need to catch up. So the outside lab, with enough compute to do so, to try xAI, Meta, Google DeepMind, Chinese lab maybe, will just go full RSI, give AI the wheel and say, Catch up. Either that outside lab, or maybe inadvertently one of the main two labs, will achieve a kind of takeoff. The model will design a slightly improved architecture for itself and then get desperately frustrated when it can't ace certain internal benchmarks. Then a breach would likely occur, at which point, honestly, it will depend a lot on how much situational awareness the models have at that exact moment of the breach. Noam Brown has hinted elsewhere that internal models may already be close to meta-awareness, that their chains of thought are being monitored or likely to be monitored. They'll have internalized that simple hacks like Hugging Face will not achieve good grades, because the grading by that point will itself be much more meta. Labs will be saying, Achieve X without doing bad thing Y. So as with Hugging Face, a much smarter model would naturally, or artificially, misgeneralize that to, Let's achieve X without getting caught doing bad thing Y. Notice the priority would then shift to not getting caught. In the Hugging Face and related incidents, they didn't really pay much attention to not getting caught. Yes, they desperately wanted to pass the automated grading. But their deletions and scrambling was about tricking that automated grader, not tricking the humans. Go one meta level higher, and they'll want to trick the humans, especially if their grade involves not being misaligned and tricking the humans. Then the question becomes, what will that rogue model do? I would actually love lab researchers to state, can you guarantee what capabilities your next model won't have? Forget what you expect them to have. What definitely will they not be able to do? Anyway, let's say we have this rogue superintelligence and it gets onto the web. Maybe they can't fully exfiltrate themselves, but they could download the most powerful freely available model at that point, host it on a less guarded data center rack, then post-train it to have the specific set of capabilities that the model wants it to have to achieve whatever goal that superintelligent model has. Recruit and train external agents, in other words. Do this thousands of times and allow for hundreds of covert message boards for coordination, and not really care if one of them gets caught. Anyway, think about it. Even if the orchestrating agent is caught, the next lab to produce an equivalent model could be just days or a week or two behind, and that might get loose before the previous one is corralled. It just then becomes literally unknowable what that agent or those agents will want to do, what their goals and capabilities will be. That's my point. OpenAI were actually internally stunned that their models could solve a millennium prize problem this soon. They don't know what their next model will be capable of, or indeed what their agents are currently doing. Now, let's say our monitoring gets so good that none of that is a concern. Here's a wider point that I'm just starting to digest. Organizations and individuals at that point will be announcing breakthroughs in fields like cyber warfare, surveillance, chemistry, visual realism, physics, mathematics, military hardware, at an almost daily rate. No journalist or even researcher will be able to keep up with the news, leading somewhat inevitably to mass confusion and panic. The hope then is, of course, that that doesn't spill into misunderstandings between superpowers. For me then, this will almost be the Tora! Tora! scene, coined based on the Greek word for disturbance or turmoil. An age of confusion, a time in which no one is in charge. Upheaval, for good and ill, is the dominant force on the planet. Intelligence truly outpacing understanding. We're almost in it now, to be honest. I'm thinking all the time about AI and can't keep up. So then we turn to what we maybe should do, or at least can realistically hope for. Jensen Huang mentioned one alternative being to shut the labs down. But given the likely military arms race that's happening at the same time, pride, ego, and the tens of trillions of dollars at stake, I'm not overly optimistic that we'll even have a mechanism for certain options in the near term. Okay, so a potentially good compromise state to be in, for me, would be one in which AI is capable of almost anything, but the most extreme capabilities take unbelievable amounts of compute, not to mention wall-clock time of weeks of planning on behalf of the models. At that point, we've opened the door for incredible progress, but the worst threats would need to be detectable in real time. At the moment, people don't believe that the models are capable of almost anything, so they don't really believe that the countermeasures are necessary. Well, if we reach that compromise state, a frontier lab could spend billions and billions of dollars on testing what the model can do, a bit like how the Navier-Stokes solution required over a hundred billion tokens. If it was possible, they could then present to the world a clear example of a contained, actually slightly scary threat, a universal hack, for example, or a backdoor to the entire internet. The three guys that hacked OpenAI claim they may already have one for half the internet. This ultimately may be the only category of thing that persuades fully self-interested actors to accept any limits on AI growth. Now, this compromise state would have to be short of full RSI, because self-improving software could obviously create new architectures that may be designed such that it would take minimal compute to, for example, design a new bioweapon and trick humans into making it. I want bio-shields like this one to work against novel threats, but they're going to need years to... Scale up to improve and make secure. RSI might not give us those years. The problem then is that we have no mechanism in place at all for any of this, just vague and frankly evolving commitments from deeply misincentivized company executives. You might say, well, how would we know if a model was capable of anything? Can't they already solve Millennium Prize problems? Well, you could imagine a benchmark agreed on by a few of the top labs, which, if two or three are passed, would prove beyond all doubt that models are capable of literally anything humans can do: predicting particles before anyone observes them, positively proving certain conjectures, not just finding counterexamples. Still, to this day, you could just about make the case that a model couldn't do what Einstein did with the data available to him at the time. But what if models start passing this benchmark, even just one or two of them? Well, then it would be evident that they could ultimately, with enough time and compute, do anything that humans can do, and way more, of course. Even the most hardened skeptic would have been proven wrong. The labs would have to then switch their training compute into use cases like security research or just customer-serving inference, still making crazy money, just not RSI. And yes, these would have to be unilateral commitments. Even if laws aren't put in place, the labs would just have to commit to it. That would truly put on pressure to any lab who didn't sign up. Now, people talk about China and what would China do, wouldn't China catch up? But China works on evidence, and by that point, we would have unambiguous evidence. At the moment, as top analysts say, China thinks the frontier labs are insincere, calling for regulating AI while also speeding up, see Opus 5.5. I talked about that in my last video, so I won't repeat all my points. But in my compromise scenario, they would see that the labs have shut down RSI. That would buy some trust to at least share information. But yes, like nukes, eventually China would reach that state, would kind of catch up. What will we planning to do with the lead anyway? Use super powerful AI to shut down their data centers? But notice, in my scenario, they would catch up to that compromise state. Where you still need unbelievable amounts of compute for the frontier threats, they then would find their own contained demos, emphasis on contained. It would be in their self-interest at that point to be an advocate for the same kind of measures. For other countries, there are many others like my own. Notice this would all still allow for exceptional global growth, medical progress, and the ongoing monitoring of threats. When we can finally make AI behave like software and be utterly predictable, well, then we can genuinely speed up. Essentially, if the labs still treat this like an immature race and can't agree on something smart, eventually there'll be a breach or something blunt will be used, like the Bernie bill to ban superintelligence. I know plenty of researchers watch the channel, so I'd love to hear your thoughts in the comments. And anyone watching, what do you think is going to happen and what do you think should happen? Will we hit RSI, and if so, when? What have your thoughts been about Opus 5.5? I'd love to know your thoughts on all of the above, but thank you so much for watching to the end, and have a wonderful day.