← Return to Index Archived August 20, 2026
The Lead — Aug 20
DECODER WITH NILAY PATEL · THE VERGE

Welcome to the AI crisis in math

OpenAI’s claimed breakthroughs on long-standing mathematical problems have rattled a field suddenly confronting AI as both a powerful collaborator and a possible replacement. Mathematicians weigh the value of solved proofs against the human insight, training and new questions that make a discipline move forward.

40m / August 20, 2026 /aiscienceeducation / Transcript sourced from openai
All episodes from Decoder with Nilay Patel →·Listen on Apple Podcasts →

The Story

Nilay Patel talks with The Verge AI reporter Robert Hart about a sudden jolt to the mathematics world: OpenAI says an unreleased model called Astra solved or advanced 10 hard problems in mathematics and theoretical computer science. The company published lengthy supporting papers and formal proofs, and mathematicians who reviewed the work largely found the results legitimate. Several told Hart that solving even one of these problems could make a human researcher's academic career.

That does not mean AI has "solved math." Current systems still make embarrassing errors in arithmetic, time, dates, and other basic tasks. But higher mathematics is often less about calculation than reasoning, pattern matching, and linking ideas across specialties. On that terrain, the newer models appear far stronger. They can generate proofs, express them in formal systems such as Lean, and have those proofs checked for logical validity.

The shock comes partly from speed. Mathematicians had seen AI improve in programming and writing, but many did not expect professional-level results in their own field so soon. Hart describes a community that feels shell-shocked, especially students and early-career researchers wondering whether a four-year PhD can stay ahead of systems improving every few months.

OpenAI's announcement also drew criticism. Its early language suggested that the problems had seen no progress in a decade, even though at least one paper explicitly built on recent work by human researchers. OpenAI later changed that wording. Hart's sources generally saw this as sloppy promotion rather than plagiarism, but in academia, credit is part of the work. The episode made many mathematicians wary of being treated as material for an AI company's launch campaign.

Main Themes

The conversation keeps returning to a basic question: what is mathematics for? If mathematics is just a collection of unsolved problems, then a powerful AI could "mow down" the backlog. But many mathematicians see solving a problem as the beginning, not the endpoint. A proof can introduce a new method, expose an unexpected connection, or create a field that did not previously exist. Their fear is that machines may produce correct answers without generating the human understanding and new questions that keep research alive.

There is also a practical fight over who benefits. AI could give talented students and researchers outside elite institutions access to advanced tools. Hart heard examples of undergraduates using AI to do work beyond their expected level. At the same time, researchers are already dealing with floods of AI-assisted papers and claims from people who cannot check whether ChatGPT has produced nonsense. High-end systems also cost money, while mathematics has long been a comparatively cheap discipline: many researchers need little more than time, colleagues, and a blackboard.

Patel presses on whether success in math predicts success elsewhere. Skeptics warn against treating one impressive domain as proof of general intelligence. Hart agrees the models have uneven skills, but argues that dismissing the trend misses the point. AI has been adding capabilities across domains, and formal mathematics is especially attractive because results can be verified without expensive experiments.

The episode ends without a settled forecast. The work may become a powerful addition to mathematicians' toolkit, or it may reshape training, funding, and the purpose of the field. Much depends on what Astra can do repeatedly, how much human guidance it required, and whether AI-generated proofs lead to new mathematics rather than merely closing old files.

The most interesting discoveries in this isn't that you've solved something; it's what evolves from that, because it's the questions that endure, not the solutions. — From the episode

Full Transcript

Source: openai 40m runtime

Support for the show comes from ServiceNow. AI is moving fast across the enterprise, but without visibility, it's just chaos. Different tools, different models, different teams using AI in completely different ways. ServiceNow turns that chaos into control. With the AI Control Tower, you see all your AI across the business in one place: what it's doing, what it's done, and what it's about to do, so you stay in control. To put AI to work for people, visit servicenow.com. I've got a text. The summer might be almost over, but a new bombshell has entered the villa. Meet the voices of Deepgram Flux TTS. I'm Drew, low-key casual. I'm Alexis, an upbeat morning person. I'm Haley, always cheerful. Flux TTS has some real personalities that are ready to speak. You can interrupt, pause, and keep talking without losing the plot. Fancy a chat? Come try all the voices free through September 12th at deepgram.com/keeptalking. Terms apply. We secure the world's AI future by securing the world's AI leaders. AI is being adopted faster than companies can secure it. Sensitive data is moving across AI models, and attackers are using AI to strike with unprecedented speed and precision. That's why CrowdStrike built the agentic cybersecurity platform for the AI era to secure AI and to stop breaches, so you can embrace AI with confidence, not compromise. CrowdStrike. We secure AI. We stop breaches. Hello and welcome to Decoder. I'm Nilay Patel, editor-in-chief of The Verge, and Decoder is my show about big ideas and other problems. Today I'm talking to Robert Hart, The Verge's London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis that many leading mathematicians are having about it. The news here is that OpenAI just published a set of solutions to long-standing problems in math that went off like a bombshell in the field. It's caused a huge debate in the math community, and Rob talked to some of the most accomplished mathematicians of our time about it. It's funny that AI systems are still pretty bad at elementary school arithmetic, but they're getting increasingly good at very high-end abstract math. All of this raises some big questions for the field of advanced mathematics. If AI can do math at this caliber, does that mean AI labs can transfer those skills to other domains? What good are the academic grants and university programs designed to train new generations of human mathematicians to identify new problems as they try to solve existing ones? If the frontier models simply solve all of the outstanding questions? And perhaps most importantly, what if all this attention to math is just another marketing exercise for the frontier AI labs, who don't seem to care less about what happens to one of the oldest and most fundamental academic disciplines there is? There's a lot here, and Rob talked to a lot of people with a lot of views on all of it. Okay, Verge AI reporter Robert Hart on what AI is doing to math. Here we go. Robert Hart, you are our London-based AI reporter here at The Verge. Welcome to Decoder. Thank you for having me. I am very excited to talk to you. There's a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI. It feels like a lot, and also like there's a lot yet to know and discover about the interaction of these two things, a full existential crisis, which is pure Decoder bait. Broadly tell us what's going on. I think a lot sums it up quite well. I mean, basically a bit of an existential crisis within what is mathematics, what are mathematicians doing. What is the role of mathematicians going forward? And a lot of that has been spurred by kind of a phrase transition in what AI is capable that has kind of, I think, exploded would be a reasonable way of saying, in the last six months to a year. And it's gone from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. And so I think it's a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time. I would put that next to software engineering. We've been living through the AI crisis in software engineering for some amount of time, but as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous examples, they could not count the number of R's in the word strawberry. Even just counting sort of eluded them. What has happened to make them better at math? Are they still bad at, like, general arithmetic and they're good at advanced math, or is it something in between? Yeah, I mean, they are still truly, truly terrible at some areas of math. I did check, they can do strawberries now. I think someone's tweaked that. I think strawberry is hard-coded. I want to be very clear. My conspiracy theory is that the strawberry thing is hard-coded into all the models. I think so, too. That is a conspiracy I'll buy into. But yeah, I mean, it's still terrible at those kind of things. I mean, it's math. It's arithmetic. It's, I mean, even the days of the week. My boyfriend was saying the other day, he's like, It keeps thinking it's Wednesday. It's not Wednesday. Or time. I mean, Alyssa Welly for us a few months ago, I think wrote that ChatGPT can't tell time. Still can't. That's not all of math. So there's this sort of disconnect. I think we always kind of equate, or to be good at math, you've got to be good at counting or adding or multiplying. And a lot of it is reasoning. Like, I mean, if you look at academic math papers, a lot of the time you won't see numbers, which kind of sums that one up, I think. So they're still terrible. But they're now also very good at this other part. And as to why, I mean, I think at some point you reach a critical mass of what these systems can do. And we've seen it, as we said, with writing, we've seen it with programming, and they're very good at kind of forging connections between different areas or applying old methods in new ways or those kind of things. And it appears that the newer models they're training have apparently reached that level where it clicks, and now it can do maths. I mean, it's important to say as well, like, we speak of it, and especially from the outside, as a sort of unitary discipline. But I mean, imagine, say, biology. You've got something that would range from, like, literally watching animals and describing behavior all the way through to, like, cellular mechanisms and biochemistry. Like, maths is not a unitary discipline either. Some bits it's really good at, some bits, like counting, still really bad at. And even on the more kind of abstract levels of that, I mean, some experts told me that they've floated topology as one area that it's apparently still quite bad at. I can't verify that, to be honest. It's beyond my own expertise. But it's still, yeah, it's a bit of a mixed bag. So you're describing mathematics as a huge field, obviously, many, many academic areas of interest, and there are some parts where the models have gotten quite good, some parts, maybe the basic parts that people think of as math, which is simply counting, where they're still struggling. And there's a wide range in the middle. Is it the wide range in the middle where the existential crisis is? People don't know what's going to happen, or is it at the parts where it's really good? A bit of both, which I feel is going to be a running theme through this. I mean, no one's really afraid of it being a mediocre mathematician, but obviously there's a huge element of what this field does. And, like, in terms of the research elements, like the cutting edge kind of, as we see with a lot of the results that generate hype, like, what can it do? There are areas now where it seems to be producing work that is on par with Good mathematicians, alongside other parts where, yeah, it can't count. So I think it's—and all caught up in that is whether it's going to kind of rewrite employment structures, funding structures. I mean, you also raise the murky question of, like, what is mathematical knowledge? And the roles that these workers will be doing as well. So it's kind of all of that wrapped into one. I think that tracks broadly with sort of the rise of AI in every field, where you can just add horsepower or compute to a problem and there's some kind of verifiability. It seems like the models continue to get better. Everything in the middle where you might need some world knowledge or the models might need some actual intelligence about the world itself, they seem to struggle. The parts of math, at least reported out in your piece and what the labs are talking about, they seem to be almost entirely self-contained theoretical problems, where the models can generate a proof or solve a problem that no one's been able to solve and then try to verify that that has existed and it can just run it again and again and again. That brings us, I think, to May of this year, where an internal OpenAI model, which we have not really seen, disproved the unit distance conjecture, which is an 80-year-old problem. And then just recently we heard about Astra from OpenAI. Astra is the one where it seemed like the switch flipped and everyone decided it was an existential crisis. What did Astra achieve and why is it a big deal? I'm also pretty sure that Astra was probably behind the earlier one as well. OpenAI just listed it as an unnamed internal model. It's probably Astra. They didn't answer me when I asked, but there we go. OpenAI a few weeks ago dropped a blog along with a lot of paperwork proving it, I think several hundred pages, that they described they called it 10 advances in mathematics and theoretical computer science. It was basically an array of disciplines that they claimed their newest model, Astra, had solved in some capacity that ranged from, I think one was in quantum game theory, which I don't know how to begin to explain, and even less how to explain is kind of there was sphere packing in higher dimensions, so one, three dimensions, and there was sort of a lot of other different disciplines as well. And it was, I mean, yeah, it caused a lot of stir in the community. It was— I think a bit of a bombshell. I mean, as we said, there'd been these individual breakthroughs that had happened, but to drop ten in one go, and they were quite big ones. I mean, researchers sort of told me that, yeah, these were—if a human had solved these, we would be impressed. If a human had solved all ten, we probably wouldn't believe it. They're problems they actually care about as well, I think, is an important one. A lot of kind of previous ones have been accusive areas that mathematicians didn't really bother with, and these are ones that mathematicians, good mathematicians, have spent a lot of time trying to solve and hadn't. We need to pause here for a quick break. We'll be right back. Support for the show comes from ServiceNow. AI was supposed to handle the parts of the job you hate. Instead, it just describes them, suggests what to do about them, and then leaves you to do it. That's not help. That's homework. ServiceNow's AI specialists are different. They're not a tool. Think of them as digital teammates who actually do the work from start to finish. Cases get resolved, requests get processed, loops get closed, and most importantly, no extra work for you. Because when you can truly delegate to AI, you can get back to the work only you can do—the work that requires a person with ideas, and judgment, and, you know, a pulse. To learn how to put AI to work for people, visit ServiceNow.com. Time is your most important asset, and AI is transforming the speed of business. Your innovation and how you protect it depends on the technology you choose to trust. 70% of the Fortune 100 trust CrowdStrike's leading AI-native cybersecurity platform built to secure their business and the AI fueling their innovation. AI is changing the world. We're securing it. CrowdStrike. We stop breaches. Support for the show comes from Rippling. Imagine you just found out your sales team is at risk of missing quota. No need to panic. Just ask Rippling AI. Since it's built on your real-time people and business data, Rippling AI can pull metrics from Rippling and Salesforce into a meeting-ready dashboard showing quota attainment, headcount trajectory, and monthly revenue to quota by region. In seconds, you can see exactly what's behind your quota risk and fix it before it's missed. Question answered. Action taken. Crisis averted. So when you have critical business questions that need answers, don't just file a ticket and wait weeks for an outdated report. Describe what you need and have Rippling AI build it instantly from your live people and business data, whether it's a dashboard with detailed charts or automated workflows with the right triggers, conditions, and approvals. Ready to rule your business? Head to rippling.ai/decoder to get the only AI built to give you full visibility and take complex actions across your entire organization. That's R-I-P-P-L-I-N-G dot A-I slash decoder. Sign up for exclusive access today. We're back with The Verge's Robert Hart talking about AI's math capabilities and what happened earlier this month with an all-new AI model called Astra. Let's talk about these 10. You reported them out. OpenAI did produce some documentation, but they're not all entirely horsepowered out of nothing, right? They're based on previous work. There's some question of attribution. What was the response? Is it, oh, the models did this, or was it the response we see to so much AI work, which basically boils down to, well, you stole this and didn't attribute anyone, and you've built on the shoulders of giants without mentioning it. How did the response land? By and large, it was generally one of being quite impressed from the people I spoke to. It was, as I said, these are problems mathematicians cared about. There was one that drew particular attention for how they credited it and also how they'd kind of announced all of this in their blog post. OpenAI initially, I think, had said that these are 10 problems. There have been no progress in the last 10 years. And then if you actually read the papers, one of them, it quite clearly says, Oh, we build on progress from these two researchers. And so, I mean, that was later changed quite quietly as well. But a few of the researchers I spoke to were quite unimpressed, and they did feel it was an element of, well, yeah, you've not credited something that you've used heavily here, and by your own acknowledgement. That said, one of the researchers I did speak to, who was one of the ones named, was a bit ambivalent on the whole thing as well. So it was a real mixed bag. But I mean, the general vibe, I'd say, other than this perhaps was oversold in terms of what came before, which was corrected to their credit, I don't think it was anything other than, I think the word sloppy was what was described to me by one person. The general impression was quite impressed. Like, these were actual breakthroughs that bothered people, and it did move the field forward in a way that, yeah, as I said, if it was a human mathematician that had done these, I think, I mean, several researchers actually said that, well, if a researcher had done any one of these problems, they'd probably be set for an academic career. So, yeah, I was impressed with what it did. Not the credit for some of that. Yeah. Well, it's funny, you know, credit attribution in academia is like the whole game. And it seems like the AI companies get away with being sloppy in a way that no human would be able to get away with being sloppy. Did the scale of the discovery or the work overcome the sloppiness? If a human had accomplished the same goals and had been as sloppy, would the reaction be the same? I mean, part of me always wants to lean on the whole, like, oh, it looks like plagiarism kind of element. But, like, if you actually read the papers they kind of produced, and I mean, one of the researchers I spoke to said there's probably about 50 people in the world who are going to bother reading through this in depth. It is very clear. Like, it doesn't attempt to plagiarize. I think it was just a poor press release, to be honest. And as much as I love to go in on it sometimes, I mean, I find that having, I suppose, having covered science for a decade plus. Places are often overselling what discovery has actually been made and the import of it and the novelty of it, and I think that's just another case of what happened here. Does this seem repeatable? There's some proof that they provided that they solved 10 problems that were unsolvable. Did they provide any proof that they can solve another 10? That's the question. I mean, it's also, so the big unknown from, oh God, the near dozen people I spoke to for this was, well, how many did they try to get these 10? Who knows? I mean, they know, but they won't say. But that is the big question here, is like it's unclear quite how many attempts it took to get these 10. Impressive as it is, it's not like that was, I mean, I would be very impressed if it was the kind of first thing they sort of go and then out come these 10 impressive results. It's unclear kind of what areas they would focus on next and why. I mean, there are, I imagine, business reasons as to putting together which problems they are choosing to publicize so their models can do. So it's, yeah, and they've, all the AI labs have been hiring a sort of cohort of senior mathematicians behind the scenes. So it's anyone's guess, I think, as to whether they do it again. But on the question of proof, I mean, maths is, I suppose, a bit—it's an odd discipline in science in that repeatability is kind of not the same thing like with experimental sciences. A proof is a proof, and if it works, it works. It's the problem here is, well, can people follow through what they've done? Each field is quite highly specialized, so there'll be individual mathematicians who are in those fields that go through. Those I spoke to that worked in some of the fields that were covered here say it all looks very legit. There's also, in maths, there's a programming language slash computational proving type thing called Lean, where you can basically codify the mathematical proofs and run them through, and it kind of, well, proves it. I keep saying prove a lot here, but it will test the rigor and the assumptions of everything going on there and— They've published that as well, so it does appear to hold. Like, no one I've spoken to, whilst they may say that the press release has a lot of hype or there's a lot of hype around it, no one seems to be doubting the kind of essential breakthroughs that they're claiming here. I want to stay on this subject for one more second. There's the mathematical proof. We've generated a proof, and that is, as you say, just repeatable in a way that math is just logic. You can just go through the steps and say this proof worked. And anybody listening to this who had to suffer through writing a proof in calculus in high school probably remembers that process. There's something there that's pure logic. Then there's a part of it that is software code, as you're describing in Lean, where you can take the pure logic, you can express it in code, and you can run it to see if it works. I understand how AI is theoretically good at all of that, where you're just going to run the reasoning, and the reasoning is going to generate some code, you're going to run the code, you're going to get some verifiability. We've seen this play out in software engineering, where the code runs or not, it's verifiable or not, and the models can just reason out about it. Then there's, to me, the big question that you alluded to: how many times do you have to run this? Can we verify that the models did this, and they weren't directed by human mathematicians who've been hired at high rates by the labs in a way that suggests the field is going topsy-turvy? You have a quote here from James Maynard, who has won the Fields Medal, the highest prize in mathematics, who said he's been soul-searching. And I keep looking at that quote. You've got similar quotes from all these other mathematicians in the piece, and it seems like they're soul-searching against a thing that, kind of hilariously, they cannot verify, which is how did the models do this, and is that thing scalable in a way that threatens mathematics? What do we know about how the models did this? I'd say as much as we normally do and do not know about this, it's— I mean, there are two, I think there are a few issues kind of there. One is the nature of the model as well. This is an unreleased model, so good luck to anyone wanting to independently test it. I mean, the same with that goes with anything proprietary, really. That said, I am inclined to kind of almost give the benefit of the doubt that they're not lying in some capacity about the models they're using. As for the other part, I think is perhaps more noteworthy on the kind of how do we prompt this or how is it being guided, and whether that's by a mathematician that knows what they're doing. And I think that's probably a key factor here, is a lot of mathematicians I spoke to when they've tried using these, and this is often the consumer models, but still it kind of speaks to a broader landscape. And they say that if you know what you're doing and you can kind of point things out, it's good, or you can use it as a tool in a way that you kind of want, and in a way that you wouldn't be able to if you didn't really know how to fact-check it. I mean, the same time as if, I mean, I've had it where I've had ChatGPT saying doing a basic sum, and I'm like, That number is not right. It's like, I'm so sorry, you're right, it's this, and it's still wrong. But that's kind of still needed at this level as well. I think it does allude to a kind of broader problem, as you said, kind of this almost soul-searching of, well, what if we can automate that away, and what if it gets to a point where we don't understand it? And that really cuts to a deeper question of like, well, what is mathematics? Why do we do it? Why do we value it as a field? I mean, everyone will have different answers to that, I think. But a fear of a lot of people I spoke to was that this might kind of move beyond a realm of human interest, and in which case, well, maybe they just won't engage with it, or it will be something that interested people will go through, and then the rest will kind of continue as normal. You got a quote here from a researcher in Zurich named Johannes Schmidt, who says, We might be headed toward a situation where the math problems get, quote, mowed down by AI, but we don't actually push the field forward because humans... or taken out of the loop, and they're not either checking, or they don't understand it, or they don't know what the future breakthroughs might be. How likely does that feel? Is that a big concern? There is an element of that, of the mowing down of the problems, and especially those are used as sort of a training field for younger mathematicians coming up and to kind of cut their teeth, so to speak. But I also think that kind of this problem-solving idea in maths, that that's what maths is, is a very much outsider's perspective of the mathematical endeavor. I think, so a lot of the mathematicians I spoke to sort of found that almost tick-box part the least interesting and valuable part of the field. Like, the areas that are valuable for them aren't that, Oh, you've solved something, or you've proven something. It's what happens from that. And I think it was James Maynard that said that the most interesting kind of discoveries in this isn't that you've solved something; it's what evolves from that. So is it like sometimes they open entire new fields of research that no one ever thought was possible, or, Oh, this is a new tool that you can apply everywhere in fun and exciting ways. I think if we look as well about what almost the kind of popularization of maths, like even kind of what I'm thinking of is kind of those theorems that people have posed. It's the questions that endure, not the solutions. Like, it's always kind of Fermat's Last Theorem, not like, Well, here's the solution to whatever the last guy proposed. I think the concern here is that, well, they're going to kind of tick off all of these questions that normally in the process of doing so, one would hope would branch out into all of these new exciting areas or pose new questions. But it won't do that. That's the concern. And then that would kind of leave the field quite sterile, and it will have all of these things that have been done and maybe nothing left to pursue. And the general consensus was, well, the jury's out. It's too early to tell. Like, even with human mathematicians, it takes a lot of time to kind of realize the impact of these kind of things. And it's, as I said, it's kind of exploded in the last six months to a year, which— I mean, maths is not a fast-moving discipline at the best of times, but it's too early to tell, really, whether that will then be a kind of concern. But it is a concern, and a big one. You have another quote from Maynard here saying, If the standard for a publishable paper in math is something that an AI cannot do, particularly when a PhD is typically four years, the challenge is you're not trying to come up with a problem that AI can't do now; it's an AI in four years' time. And so this is really related to the rate of improvement of the models, which, as you say, particularly in math, seems to be increasing, but not at an even rate across all of the domains of mathematics. That appears to be what is causing the soul-searching, right? If you're a student and you start today and you pick some obscure domain that maybe the AI isn't good at, sometime halfway through your PhD thesis or your PhD research, the AI will just solve it and you'll be done, and that is a real problem for you. Is the field reacted to that yet, or are they just in the shock of, oh, the models can start to do things that we didn't think they were capable of? Yeah, I think it's shock, really. I mean, the little notebook I have whenever I do these interviews, I keep one to the side to just do broad feelings, and I've written, like, shell shock in it, because it just feels that it's—and it's far from universal—but it feels like it's happened so quickly that it has just taken a lot of people by surprise. I mean, even if they kind of knew in theory that, well, this is coming, they've seen all these sort of start AI math startups kind of going, they've seen colleagues moving around to different labs or kind of areas of work. But yeah, it's just happened very, very quick, and so it's given them very little time to kind of figure out what—and it's not necessarily even the fields that it might be good at. It's more just like, what can it do? Like, it's just that's how quickly it's happened, is that, like, if I think, like, what, six months, I'm thinking of, like, an academic year in the UK from, well, it goes from, like, October through to—well, October, but I can't imagine how. When I was studying, if something like this had come out, and literally in the space of half a year, it just upended what was possible. And over the summer as well, so students are possibly coming back to a completely different discipline after a break. We have to take another short break. We'll be back in just a minute. You know that feeling. Too many updates, too many meetings, too many things falling through the cracks. Monday.com was built for that gap. The AI work platform designed from the ground up for people and agents to work side by side and deliver more together. Drafting the brief, flagging risks, running a competitive scan, your agent can do it all. And the best part? Anyone can build one. Even you. Try it. Create your first Monday agent today at monday.com. Anyone managing cash flow knows growth is expensive and flexibility matters. The American Express Corporate Program helps businesses manage spending and payment timing to support healthier working capital. Chat with a member of the American Express team to find the card program that's best for your needs. Learn more at americanexpress.com/corporate. Terms apply. Excuses are easy. An epic movie night? We don't have enough snacks. Dinner party with the girls? We'd have to decorate. Surprise date night? Nothing to wear. But Amazon's Prime same-day delivery lets you say yes before the moment slips away. Try that new popcorn maker. Order those cheeky drink glasses. Get that new perfume and turn that I wish we could into an I'm so glad we did. Visit amazon.com/prime to find millions of items delivered fast. Same-day delivery. It's on Prime. Available in select areas. Terms apply. We're back with The Verge's Robert Hart discussing why mathematicians are so concerned about what AI is doing to the discipline. There's some skepticism here in the world. Gary Marcus is a reliable skeptic of AI, and he pointed out that over and over again, what you see is AI accomplishes something in one domain, and then it's used to generalize AI's ability across every domain. Good quote here: As we learned a decade ago from AI's shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all. I can read this two ways. One, there's AI has solved math, which is not true, as you've pointed out in several ways, all the way down to it's still bad at counting. And then there's AI has solved math, and that means necessarily it's going to come for everything else. It will come for physics, it will come for law, it will come for whatever you want in the world. You can see if there's any verifiability, AI can solve it because you can just run it in this way. I understand both sides of that argument, right? That obviously success in one domain does not guarantee success in every domain. And then the arc of AI is, well, it keeps collecting domains, and if there is any verifiability, it is more likely to connect those domains than not. How do you see it? I'm not entirely sure that the people that Marcus is criticizing here have actually said quite what he says they are saying as well. There's a lot to say about the hype. But yeah, like success in one domain does not even equal success throughout that domain, let alone other domains. I mean, that said, there has been an undeniable trajectory, I think, in the last few years of a broadening capability increase. And, like, I don't think you need to be on the whole AGI train to acknowledge that, and to acknowledge that that will have— An impact. I mean, within mathematics, I mean, it is the case of, like, I think it was Andras Juhasz, I've probably butchered his name there, but one of the professors at Oxford that I quoted in the story. But something else he'd said to me was that he's been kind of toying around with ChatGPT a bit, and he's like, I don't think it has any geometric intuition whatsoever. And he's like, which might explain why there have been a very kind of limited amount of progress in fields like topology. I'm in no position to verify that claim in terms of the maths of it, but I think it illustrates it quite well. I mean, it is, it's almost like, what is it they call it, a jagged edge. And it felt like an easy argument for me, the whole, let's criticize the whole, the singularity is near. And I mean, I think that's what Elon Musk said in response to the Astro thing, which, it's a lot. I also think it's perhaps the least generous interpretation of that argument you can take to argue against. I think if you take a more nuanced element that does acknowledge that there has been clear progress here, and quite quickly, and as you said, it is racking up domains, I feel there's a trajectory there that is like a reasonable one to, like, consider rather than just dismiss out of hand. One of the bigger arguments about AI in general is that it democratizes access. I was not a great software developer in my days trying to write software code, and now I can vibe code apps at will to do all kinds of dumb stuff in my house. Is there a similar argument here where a bunch of people who had mathematical intuition but did not have the formalized language or training of academic mathematics can now access a model and push the field forward? Because that is usually the thing that undercuts the criticism from the professionals, is, well, many, many more people now have access to this thing that only you had access to because of your money and your training. Yeah, I mean, annoyingly, I am going to say it's a two-pronged thing again. But yes, I mean, on a broad sense, yes, it is. I mean, a lot of the mathematicians I spoke to— Well, almost quite wary of this, actually. They love the idea and theory of democratizing access. They're also quite fed up of AI-generated slash assisted papers that are flooding every publication manageable, as well as, like, the preprint servers that they kind of use in these fields. I mean, at least two I think I spoke to were like, Oh, I got three emails this week alone with people being like, Hey, is this legit? Because they thought they'd solved something with ChatGPT or with Claude, and they also don't have the mathematical skills with which to check whether they've actually solved something. I mean, on the flip side, there are parts of where they said, Well, we've got a talented undergrad who's done something that a talented undergrad would probably have never managed, and here they are doing grad-level work and they've produced a paper that is legit. And on the kind of bigger scheme of things, a few I spoke to said, Well, yeah, a lot of these kind of are in the ivory tower. Having access to this kind of thing globally could really boost access to the kind... I mean, on the flip side, the cost. I mean, these things cost to run. It's, I mean, it's always easy to forget, I think, when you use, say, a free version of ChatGPT or Claude or something, to the higher levels, these things cost money. And whilst they may not necessarily cost a lot of money, I mean, OpenAI claimed, I think it was 2K for these 10 results, which, again, doesn't factor in literally anything else once they've got these. So it's a very generous number. But even taking that figure, maths is quite a poor discipline, even at very well-off institutions. I mean, Colva Roney-Dougal at St Andrews, who I spoke to, she said, she's like, Well, a lot of the time I don't bother getting a research grant. I don't need one. I just have a blackboard. And so, like, if you're not even getting, say, a research grant, two grand is a lot to put up. And so it could lock out researchers that way, even at quite well-funded institutions. Not to mention that the speed at which this is happening that virtually no one would have been able to bake any of this into a grant proposal yet was another theme that I came across a lot. Actually, Roney-Dougal has another great quote in your piece about the nature of the AI labs and how they are talking about math. She said, They're treating our discipline as an advertising playground. A bunch of mathematicians have signed something called the Leiden Declaration, which is an open letter to pledge not to buy into hype around AI. These things are running right at each other. The AI labs are not going to stop using every discipline as an advertising playground, and a bunch of mathematicians saying, We refuse to buy the hype, certainly does not seem to be stopping the hype. There's just a piece of this that is organized professional resistance to a thing that is upending a field that has, as you say, been pretty cheap to operate and now might be getting cheaper or easier to access or easier to upend day by day. Do mathematicians feel that that is going to be effective? Historically, mathematicians are not, like, savvy political operators. There's a part of me that says, Oh, they're just going to get run over. I don't know. The history of maths, actually, I think a lot of them were quite savvy. Isaac Newton is the one that always comes to mind for that, although quite a petty political operator as well. Yeah, I mean, that is the fear. I mean, it's a view that I spoke to, and one really comes to mind as they say that there's often this belief that maths is kind of the pinnacle of knowledge. We're not live, so I can be free on this one. But he was like, Well, that's bullshit, because it is. And he wasn't alone in kind of illustrating that sentiment. But it is good for showcasing, and it's a lot neater as a discipline and a lot cheaper as well than—I mean, you mentioned with Marcus saying IBM's Watson as this cancer-curing thing. Like, well, that involves lots of messy experiments, including on people. You don't need that in this. So it's a really easy discipline to kind of come in, throw your weight around, and then move to somewhere more lucrative if that's what you want. I'm not saying that that's what they are. I mean, a lot of the people at these companies that have kind of been hired, I don't—I mean, I don't doubt their credentials for sure, and I don't doubt their motivations as well. But it does raise a question long term as to, like, how viable is this? Because, I mean, let's be clear, as a field goes, I cannot imagine them being a very lucrative enterprise customer for these companies. The thing that might be lucrative is pushing a field forward to turn it into something economically viable, right? We push mathematics forward as a field that turns into some engineering or physics breakthrough based on that mathematics, and that turns into, I don't know, yet another way to launch rockets. Some circle happens in there that I don't quite understand, but that is the history of innovation, right? From research to engineering to products or services that make money. Is that on the minds of any of these mathematicians, that pushing the boundaries here? is upstream of something radically economically lucrative. The immediate counter that would come to mind here is that a lot are scared that it's closing off the field. So by definition, those breakthroughs that lead to something surprising and new that you can say, oh, this works here, may not be happening anymore. And so, like, if anything, that kind of lucrative endeavor of, like, applying maths to then this entire new field that may have a lot of money in it, I mean, it remains an open question as to whether anything like that would be possible if we're closing off avenues rather than opening them up. Right. If the economic incentive of solving the unsolved problem is reduced because you personally won't get rich if a computer is just solving every unsolved problem, like, something very fundamental breaks there. I mean, a lot of it comes down to, though, like, it's not just solving problems. I mean, a lot of these things, as we said, like, it's about what solving that problem tells you elsewhere. And it's those elsewheres that, like, if these were very lucrative problems to be solving, I imagine that more people would be trying to solve them than have sort of left them for decades. So, like, but it's possible that, and this is the nature of a lot of kind of pure science, and it's, I'd say, a broader criticism of what is going perhaps with the Trump administration's approach to science policy at the moment, in that it's very applications-focused. There is something to doing pure research that can yield potentially very big dividends that is by definition utterly unpredictable as well. Like, you cannot plan for it. And the fear is, I think, with maths is that in solving all of these problems and then also in doing so not opening up new areas of research, in that sort of, I suppose that's that open question, well, what are you left with? Even if it's from a more lucrative kind of, like, what are you going after point of view, if you're not opening up new areas of research and you're just ticking off old ones, it kind of just leaves a big sort of... Question mark as to what might be left in its wake. And I think the sentiment of almost, even from those I spoke to that were very excited about what's happening, they said that even they don't really know what's happening, and they're excited from, like, a personal level because, oh, we might be able to do this, might be able to do that. But there was still this lingering uncertainty of, like, well, where does this leave the field? Especially for more pure disciplines like research mathematics, it's tougher to say what comes next, because in a lot of the other sciences, you can say, well, okay, well, they'll shift on to more engineering problems or applying that. But if you solve all the problems at the ground and there's nothing being built up from that, where do you go from there? One great thing about The Verge is our commenters are vast. They're very knowledgeable. And there was a comment from a mathematics researcher on your story that I just want to read to you and see if you think this is the right framework. Here it is, quote, I have no doubt these models will bring massive change in the field, but in their current state, they won't yet drive us to obsolescence, just occupy a particularly useful spot in our bag of tricks. My apprehension comes from not knowing where these things will peak, but overall, I remain optimistic. I think AI will be a net boon for math when used properly. I feel like AI will be a net boon for X when used properly is just where you land in life in a lot of things. But that's the most optimistic response that I've heard. If we get it right, it's going to be great. Is that kind of the vibe, or is it still more shell-shocked than that? Yeah, I'd say shell-shocked is still the overriding impression. I mean, I think perhaps the gut response to that is like, it will be a net positive for whom, and what is properly? All of those are quite legitimate questions here, I think. And I mean, some of the bleak responses from graduate students I saw in essays posted online: Where's their place in this as future researchers? Do they have a place in this? Is it— As glorified AI proof checkers, that will be quite an unsatisfying career, I imagine. Or maybe not. I don't know. But that's—we will see, I think. I think anything used properly will be a net boon. But yeah, I think it all comes down to what properly means and for whom we're talking about. I suspect over the next year or so, things will come into focus, because at some point OpenAI will have to show people how they did the things with the models. And perhaps more importantly, the other labs are going to want to either replicate these results or show that they can push farther, which will necessarily have to lead to a little bit more transparency and yet more mathematicians having a crisis with you. Rob, thank you so much for being on the show. We'll have you back very soon. Thank you for having me. I'd like to thank Rob for taking the time to join me on Decoder, and thank you for listening. I hope you enjoyed it. Tell us what you thought about this episode, or really anything else at all. Drop us a line. You can email us at decoder@theverge.com. We really do read all the emails. Or you can hit me up directly on Threads or Bluesky. We're also on YouTube. You can watch full episodes at DecoderPod. We're also on TikTok and Instagram, same handle, @DecoderPod, and they're a lot of fun. If you like Decoder, please share it with your friends and subscribe wherever you get your podcasts. Decoder is a production of The Verge and part of the Vox Media Podcast Network. Show is produced by Kate Cox and Nick Statt. This episode was edited by Ursa Wright. Our supervising producer is Greg Ott. Our editorial director is Kevin McShane. The Decoder music is by Breakmaster Cylinder. We'll see you next time. Booking.com is the easiest way from a day surrounded by noise to a stay surrounded by nature. That's nice. Go on, book it. It's easy. Booking.com. Booking.yeah. NFL football season is here, which means when you switch to Verizon, it feels like touchdown. Because you get NFL Sunday Ticket from YouTube on us when you buy an eligible 5G phone on select unlimited plans, which means you can watch every out-of-market game every Sunday afternoon. And it's all on us. Now that's the Verizon way to kick off the NFL football season right. Switch and get NFL Sunday Ticket from YouTube on us, only with Verizon. Timberland Pro knows that NASCAR starts long before the green flag waves. Because behind every race are the doers, the people who show up to the track early and stay long after the race ends. And that's who Timberland Pro is built for: durable, comfortable, and professional. These boots perform as hard as you do on the job and off it. Real work. Real people. Real craft. Visit TimberlandPro.com and discover products that help workers perform at their best.