Overview
This episode explores two intertwined questions: how the competition between OpenAI and Anthropic is evolving as their products increasingly converge, and what AI agents will realistically look like in enterprise and consumer settings. Aaron Levie argues that the market is moving beyond simple chatbots toward agents that can execute tasks across software, files, and workflows—but he also stresses that adoption will be slower and messier than Silicon Valley often assumes.
Key Takeaways
One of the central insights is that OpenAI and Anthropic are no longer meaningfully separated by “consumer” versus “enterprise.” While Anthropic built strong momentum in coding and enterprise use cases, Levie notes that ChatGPT has also spread deeply into companies, even outside API usage. In practice, both companies are now competing across coding, enterprise knowledge work, and general-purpose assistant products.
Levie sees the next big shift as agents that use coding ability not just to write software, but to operate computers and complete knowledge-work tasks. In his framing, the breakthrough is not “AI that can code,” but “AI that can apply coding and tool use to any domain”—law, marketing, research, finance, and more. That expands the market dramatically, from developers to essentially every knowledge worker.
At the same time, he makes a counterintuitive point: coding may be misleading people about how fast agents will improve. AI coding advances quickly because code can be automatically evaluated—did it run, did it pass tests, is it clean? Many real-world tasks, like editing a video or making judgment calls, do not have such clear feedback loops. That means automation outside software development will likely take longer than many expect.
A major practical obstacle is not just model intelligence, but enterprise data chaos. Agents need context, and most large companies store documents, contracts, research, and assets across dozens of disconnected systems. Levie compares this to hiring a brilliant new employee who has zero tribal knowledge: intelligence alone is not enough if they cannot identify the real source of truth. In that sense, many “AI problems” are actually data organization problems.
He also highlights trust, compliance, and liability as major barriers. For agents to be truly useful, users must let them access files, inboxes, and business systems. But once agents take real actions—especially in regulated industries like healthcare, finance, and law—questions of responsibility become unavoidable. Existing legal frameworks assume a human is ultimately executing the transaction.
Finally, on who captures value, Levie declines to pick a single winner. His view is that the foundation model labs will win regardless because they supply the intelligence layer, but there will still be room for vertical, domain-specific products that package that intelligence into trusted workflows for legal, healthcare, finance, and other industries.
Practical Steps
- Treat AI adoption as a data-readiness project, not just a model selection project. Consolidate key documents and define clear sources of truth before expecting agents to work well.
- Use AI for structured support before full autonomy. Ask for pros/cons tables, drafts, and research summaries rather than immediate final decisions.
- Be careful with prompting on sensitive topics. Rephrase questions in neutral and opposing ways to test whether the system is simply mirroring your framing.
- Match task type to tolerance for speed versus accuracy. Fast answers may be useful for low-stakes queries, but higher-value work should allow more time and compute.
- In enterprises, begin with narrow, auditable workflows rather than broad “do everything” deployments—especially in regulated environments.
Notable Quotes
“A new employee has a PhD, but they joined your company one minute ago.” — Aaron Levie
“An AI problem is really a data problem.” — Aaron Levie
“We sort of extrapolate most things from how good AI coding is.” — Aaron Levie
Full Transcript
How is the battle between OpenAI and Anthropic shaping up now that they're both basically building the same product? And what is the future of AI agents? Let's talk about it with Box CEO, Aaron Levy, right after this. This episode is brought to you by TruDiagnostic. I've been trying to get more intentional about my health lately. Not just how I feel day to day, but what's actually going on under the hood. That's why I checked out TruDiagnostic. They offer at-home tests that measure your biological age, not just how old you are, but how your body is aging on a cellular level. Their TrueAge test looks at things like your pace of aging, organ system health, and even risk factors tied to lifestyle, giving you real data to act on. What I like is that it's not guesswork. You can track changes over time and see how things like sleep, diet, or exercise are actually impacting your body. And taking the test at home was so easy. If you're serious about optimizing your health and longevity, this is a really powerful tool. Right now, Big Technology Podcast listeners can get 20% off at trudiagnostic.com. Use code BIGTECH at checkout. That's trudiagnostic.com and use BIGTECH for 20% off today. Choose TrueAge, TrueHealth, or the combo kit as a one-time purchase or a subscription. Welcome to Big Technology Podcast, a show for cool-headed and nuanced conversation of the tech world and beyond. We have a great show for you today. We're going to unpack the battle between OpenAI and Anthropic now that their product roadmaps have pretty much converged. And we'll also talk about the future and the present of AI agents and where that technology is heading. And joining us is Aaron Levy of Box, CEO of Box. Aaron, thank you. Welcome. Yeah, good to be here. I certainly like the framing on the battle. You know, I think it's – to some extent, it was sort of an inevitable outcome because if you think about it, like if you have this AI model that is super intelligent packed into a model, it eventually has to converge on, you know, all of the same use cases will be represented by that. And so then I think the labs eventually need to compete head-to-head, you know, for all those use cases. Yeah, I'm glad to get this discussion going even before the first question comes out. Yeah, sorry. Okay, okay. I was like, I'll frame – your intro was basically a question, so why not? That's right. But it is really what's happening. So just to frame it, we saw Anthropic take the lead in enterprise. And OpenAI seemed satisfied. For coding, yes. For coding, but also they were selling into enterprises through the API. Yeah. And that was where my belief initially about Anthropic came that as Anthropic goes, so goes AI because if this technology is useful to businesses, that means that the cap on the amount of money that it can make is going to be higher. So Anthropic made this big bet on enterprise and on coding and crushed it. And OpenAI made this big bet on consumer. ChatGPT, by the way, is probably at a billion users right now, even if it's not announced. Yep. And they did very well there. But then something interesting happened where the coding models in December became good enough to code for kind of long time horizons without interruption. And they became useful to even the non-technical folks. And then we saw this emergence of both these companies wanting to build this super-app style thing that basically, that's sort of what the question is. Is it going to be an assistant for you? Is it going to be something that does your work? They both want it to do kind of everything for you. Where do you see that going? And how do you see the battle shaping up? Yeah, so let me just inject two couple quick thoughts in your initial framing. And then I'll answer the question more directly. I think probably to represent both sides of Anthropic and OpenAI on this, I think probably the story might be even more kind of complicated than even that initial framing. Because I actually think ChatGPT leaked into the enterprise and has had actually a lot of enterprise traction of enterprise deployments, which is separate from the API business. And so if you go to a lot of enterprises, they actually will have ChatGPT as their corporate standard for kind of their corporate LLM for employees to use. So, you know, it's hard to kind of decide what data do you end up looking at. But I would generally argue that both have done actually extremely well in the enterprise. And ChatGPT, obviously, even more focused on the consumer historically. And now, obviously, you have this increased battle for enterprise dominance, both with coding, the APIs, and the end-user kind of corporate knowledge work use case. This kind of co-work use case as well. The co-work use case being that kind of third one. And the big breakthrough that has happened recently, you know, literally just recently in the past few months, is this idea of what if you could give, what if an agent was really, really good at coding, but the use case wasn't to build software. The use case was to use its coding skills and general kind of tool calling skills and the ability to run scripts. What if the agent was really good at all of those capabilities, but was applied to the rest of knowledge work? And what kinds of use cases would that open up? And, you know, kind of the mental model is, like, what if everybody was, like, truly an expert at using their computer, and they could write code for any task they wanted to do? But that same, you know, person that was the expert at using their computer and, you know, writing code was a lawyer and they were a marketer and they were in life sciences and they did research. That's basically the power of agents today more and more in terms of where we're going. And so the idea and co-work kind of, you know, best manifested this early on. I think we'll certainly, you know, you know, see based on the rumors, OpenAI have a presence in this space and other players is, you know, what if you had an agent that was your general purpose knowledge worker agent, but again, it could use every tool on your computer. It can write code on the fly for a new problem that it hasn't seen before. It can use things called skills to be able to leverage existing kind of ongoing scripts and code that it needs to be able to use. What kind of now superpower would that be, you know, to be able to have as, you know, this kind of this workhorse that you have next to you? That's kind of the next frontier of AI agents. And so I think we're clearly moving from a world where you will use AI as this thing you chat back and forth with. And that was kind of the first manifestation of the chatbot to now a paradigm where the agent is given a task. It has a set of resources it has access to. It has access to maybe your data, your software, tools on your computer, tools in the cloud. And it can go off and work for minutes or hours or maybe even days and go and generate, you know, some, you know, effective work output that you can then go and use, review, and then incorporate into your broader work. So this is kind of the big prize because it goes from the TAM, the total addressable market being, you know, all of engineers, to now the total addressable market is every knowledge worker. And that's probably about a 30 to 50x larger market in terms of, you know, humans on the planet and their use cases. So you see this as business first. This is, this is going to be primarily business. I think. But it's interesting because Greg Brockman when I had him on described it as like a laptop where you could use your laptop for your personal stuff. You could use your laptop for your enterprise work. Yeah. And I, I fully agree with that framing. And I actually think that will suck it into the enterprise. I think, I think what we're going to see is that the, the value and the ROI on those tokens, you know, the tokens are not going to be cheap anytime soon. And so the ROI on those tokens will just be much higher in the enterprise because it'll be, you know, generating something that is sort of, you know, impacts the GDP in some way. And so I think that we will probably prioritize a lot of these systems toward those types of activities. But, but I totally agree with his framing that, that you'll just use it in a general purpose way. And, and probably the more that you're the kind of person that already likes to automate your life and, you know, do a bunch of automation things in your personal life, you'll use this also in a personal capacity. But I think most of the, the true economic value of it will come from the enterprise. Is this stuff going to work? I mean, there's two things to it, right? There's the capabilities side. And then there's also the interest in using it. So again, just going back to one of these examples that I spoke about with Greg last week, basically what, what Codex, OpenAI's, you know, new coding app that can do your work for you tool. I still don't really know how to refer to it. But what it can do is just for one example, it can, if you need to edit a video, it can go into Premiere and like put chapters in your video. But I also think like, do we really need like software to do that? Or aren't people just going to be, aren't people just going to prefer to do it the old way? And how deep can it get? Like, can it, do you think this will actually So the question is, like, you know, how deep can that go into work? How long running can that work, those agents be across work before you have to sort of review the output that the agent is doing? How well do these models work on much more subjective tasks, like editing a video is like, you know, going to be actually, in many cases, a harder task than coding. Because, again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding, and because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like, the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from, like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the training process, rather, you can instantly evaluate, did the code run? How clean was the code? We have a bunch of areas of work that don't have, they don't have that ability to instantly sort of verify. So the reward function is a lot trickier for the agent. And then thus in the real-life workflow, it's kind of hard to then go and automate that task. So I think this is actually going to take a lot longer to play out than maybe what we and some think in Silicon Valley. Because what was happened in Silicon Valley is we sort of look at all of the power of AI coding. And because that's like the most economically useful task within Silicon Valley, we sort of extrapolate most things from like how good AI coding is. And because again, the code right now is like, it has this great property of in the eval process, in the I think there's plenty of issues with the idea of how much of our life do we put into these systems? How much do we rely on them for every little thing? Andre Carpathy had this funny tweet where he sort of said, you know, I had an AI go and review something and I asked for it to critique me, but then I had it do the exactly the opposite and it sort of found, it created just as good of a justification on the exact opposite of what it had said on the other side. And we see this a lot, which is, you know, I'll mostly represent myself. I don't know if my wife wants to be pulled into this, but you know, I slash we use ChatGPT for parenting a lot. And it's funny because like, you just know how you could prompt it and get a completely 180 different answer on the facts of the situation. And so you actually have to like, you really have to understand how these systems work so you can ensure you're not just getting, again, what is the sort of, you know, mean response based on your prompt, you really need to pull out of it. What is it? What really, you know, should you do in this particular situation? So you have to like do like, you have to, you know, you know, sometimes word things in like, in a negative fashion versus a positive fashion. You don't, you don't wanna like bias the agent as you're writing the question. You have to do a bunch of this kind of stuff. And that, that'll be, I just think that'll be like a thing we generally learn over time in society, just as we eventually learned how to use search engines and other tools. Right. And I think when you try to get a response on a big life question from these things, something that's important to keep in mind is its goal is to get you to write another prompt. Yes. That reward function is definitely tricky. In general, what you really want is the as much as possible. You want the agents to do things like generate me a table of the pros and cons of this thing and make sure that you make arguments for both sides. And then you want to be really in the position of interpreting that and making a decision based on what you think is relevant in your situation. I, I do things, I have to do these things sometimes, like even for like medical questions where I know that I've in my prompt, I've, I've sort of, I've, I've over kind of biased the direction that I know the agent is going to go in or the, the, the chat will go in. So then I, I do a different prompt, which is just like, under what circumstance would you, you know, imagine this type of, of, you know, kind of medical issue would show up. And then I, and then I kind of see, okay, is, are those things showing up here? Versus if you just give it your symptoms and then you'd be like, and do you think it's this? And it'll be like, yes, it's definitely that. Like, do a happy bowler. Exactly. The big question though for this stuff to work is, and I think you talked a little bit about how useful you want it to be in your life, you have to trust it and you also have to give up a lot of control. Like, to make these agents work really well, like, think about any example we just, we just went through. You have to be like, here's my computer. I have my files. Take actions on my behalf. And honestly, they work better when you take the guardrails off and trust them to do things for you. Do you think we're like, again, for this product vision to work, that has to happen. Do you think we're in a place where it's feasible for people to give up that type of control to these bots? Well, so this is, this is where the diffusion, this general category is where the diffusion will be longer than, than where people in Silicon Valley think. So if you're in Silicon Valley and, you know, every tweet that you and I read, you know, that goes viral in the Valley is, is often, it's coming from like a 10 person startup. They have, they have basically like, they, they started from a completely clean slate of, of the way that they work, the environment, the tools they use, the data that they have. And they can just, they can build their organization around, around getting output from agents. And you go to the rest of the world, take a company that has, you know, 10,000 employees, been around for decades, their data is in, again, 20, 30, 50, a hundred different systems. If you go and ask that company, where are your latest, you know, contracts for this client? It could be in five different places. If you go and say, where's the latest marketing campaign assets? It could be in 10 different places. If you say, where's the research for the new, for that new breakthrough that you're working on? It could be in, you know, five different repositories. So the challenge is if you're in, if you now want to go deploy an AI agent in that environment, you can almost think about it like, like a new employee joining that company. And that new employee is like insanely smart. Like they have a PhD, but they just joined your company one minute ago. You've given them access to your tools. And you say in 30 seconds from now, I need you to go and find me the research for this new product we're building. The problem is that person is going to go and they're going to go look through all your systems, but they're not going to know like, well, which is the one that, that really is the authoritative copy of that research plan or that marketing asset or that contract. They're not going to know where that is because that came through kind of tribal knowledge. It came through, you know, you knowing over like, you know, 10 different meetings that you pulled the wrong thing, or you had to ask your colleague, where is that right source of truth for something? So that new employee has, doesn't have any of that context. It doesn't know any of the, any of that tribal knowledge or the work patterns that, that have existed at the company. The agent is in that exact same situation, but they're even worse off because they, they are, they really don't know when they don't know something. And so what happens is the agent gets access to those 10 systems and it, it says, Hey, you, you say, Hey, when's the, you know, uh, when's the launch of that new product? The first document or set of documents it finds that, that seemingly talk about that thing, it's just going to pull from those. It's not going to know that actually maybe there's two other systems I should go and check and then compare the answers to the first ones that I found. It's just going to go and deliver that answer to you. And so the challenge though then is that you're at the mercy as an enterprise. Uh, you're at the mercy of, of how well is your information organized? How well did you document, you know, your, your underlying processes? How easy is it for an employee or an agent to get access to the true source of truth to any project or, or thing going on in your business? The harder it is for a person to be able to go in and find the right thing, it's gonna be 10 times harder for the agent. And so the real world, not the 10 person startups that get to, you know, get started without any of that, uh, that history in the real world, most enterprises are dealing with all of those challenges. And so they, they go in and they try and deploy an agent and the agent has to, first of all, connect to all of those systems. Then it has to try and figure out again, where is the, where is the right information that needs the right answer? Then you're reliant on that system having been kept up to date with exactly the right information, the right data, that right, you know, the right copy of the, the, the document. Um, and that's the big challenge. And so we are gonna be in for again, years and years of enterprises realizing that an AI problem is really a data problem. And to get the AI the right data, they need to make sure they have infrastructure, software tools, systems that all are in service of giving the agent context. Um, and some companies are, are ahead of the curve on that, but a lot of companies are still kind of reckoning with, I have a lot of infrastructure that's legacy. Agents don't work well with that set of legacy tools. And so I can't, you know, easily get agents to access that data. We see this every, you know, every day in our business because we're helping customers sort of move to a modern way of managing their information. But where we come from in our, in our industry of, of, you know, with enterprises managing enterprise content, companies have 20 or 30 different systems where their enterprise documents are. And that just simply won't work with agents. So that's, that's probably your biggest challenge is the agents need context. The context is everywhere. How do you ensure that the agents have exactly the right context they need to do their work? That will be the big challenge for knowledge work automation. And, but there, you know, beyond getting them access to that context, it's do you trust them with that context? Like I need an agent in the worst way. I mean, I think OpenClaw would be great for me if it could go through my inbox, if it could read all my emails, draft the responses it thinks that I need to send that Another kind of security adjacent issue, which is really just kind of regulatory and compliance oriented, which is, you know, who's liable when the medical practice has an agent that does, you know, prescriptions and the wrong prescription is filed? Like, that's a really—That's going to be a new novel problem that we face in the world. And right now, that liability, you know, that labs are not going to, you know, take on the liability for every single use case that you do. They're going to have very narrow liability that they have around copyright and IP protection and stuff like that, but they're not going to, you know, they're not going to be able to, you know, handle every medical claim that is as a result of misuse of AI. And so then, does it go to the company? Does it eventually go to the doctor or the user of the tool? So we have, like, massive, you know, a hundred plus years of legal frameworks that that sort of, you know, that just always assume that a user, a human, is on the other end of every transaction and representing, you know, some part of that transaction to a client or a patient or a citizen. And so when agents are doing that, this opens up a whole new field of of questions. And so in finance, in health care, in legal, we have just incredible amounts of updated laws that will have to get written and case law that will be generated over the coming years. So that in its own way is a point of friction for rollout in enterprises. We just have to figure out a lot of these types of things. OK, a few more questions about this. Yeah. Are you sure this is the right bet for the labs? I mean, maybe this will go a certain way and then they might be like, well, actually, the chatbot was the best application of our technology. I don't know that there is as much of a tradeoff between those two as they could basically do both. And if it I think the right manifestation actually is just is a let's just say ChatGPT or Claude, you should go to either of those applications and you should give it a task. And if that task is like, what was the sports score from the game last night? Just answer it. And if the other task is like, you know, I want to get a dashboard from my Salesforce data connected to my box documents. And then I want you to, you know, generate JIRA or linear tickets based on some, you know, workflow that happened there. It should be able to execute that. And so that's just all one system of there's a fast search. There's a capability where the agent has access to tools. There's a mode where the agent sets a plan and then can, you know, talk to your software. Like, I think that's just one continuum, one very long continuum of ways that we will use agents in the future. So I don't consider it a sort of a bet or or something in that kind of classic sense. This is like inevitably guaranteed where, you know, any kind of agentic system is going. But it doesn't trade off from any of the simple, fast chatbot stuff as well that you will just continue to use in your in your daily life. Yeah, it could be a thing also where you're asking it, let's say it realizes you're asking it for a certain team's sports score. It can say, well, let me send you like an email as soon as it's done or build you a widget on your phone or even an app tracking that and some news stories you always ask me about. Once it has that ability to code, that sort of merge between your interests and building things for you, it can it can end up producing stuff. A hundred percent. Actually, I would say one of the biggest in my personal kind of use cases for AI, one of my biggest challenges has been the chatbot modality was would just happily give up on tasks too easily. So you would say, like, you know, give me the top 100 companies that do X and it would return like here are 25 that I found. I don't know where to go and find the next, you know, 75, but if you'd like, you could do it. You know, you could ask me this and it'd be like, well, that wasn't my question. I wanted the top 100. And now you go to, you know, a great example is perplexity computer. This is working great on this dimension. You say, hey, perplexity computer, give me the top 100 companies that do XYZ and it will just it will. It's just a workhorse. It does not give up until the task is complete. And so to your point, that when I do that query, that's hard. It should just prompt me and say, do you want to be notified when this is done? And I know it's going to take 15 minutes. That's fine. This is sort of an asynchronous task, but it's way better to, you know, get the right answer than in the kind of very fast chatbot mode. You're just not going to get the answer ever. Yeah, the lazy chatbot stuff to me is really funny. Like I've had it like edit transcripts before and I'm like going through the transcript. I'm like, you dropped an entire thing. Yeah. Or you decided or yeah, you decided to shrink it in half, but also summarize parts of it after I said do it verbatim. And it's like, sorry, I wasn't supposed to do that. But yes, I mean these things. There is one thing in AI that is is just like, like there's just no free lunch, which is, which is that you can have something fast, like insanely fast, but like moderately accurate or pretty accurate and insanely slow. And like you just get to choose. And like, do you want the thing to So, so, you know, we have a bunch of use cases within Box where we, we built a new agent that works across your entire box account. This is Box agent. This is the Box agent just came out last week. And the Box agent is basically this evolution to more of a full agent that, that has all of your Box account that it has access to as a search tool. It has a document reader tool. It can generate content. It can create folders, you know, all, all of these sort of, you know, kind of core capabilities within Box. And so the Box agent, you know, is you know, is just like a, a user of Box in terms of what it has access to. But you have this really interesting trade off that you have to give the agent. And we try and do this centrally when we're designing the agent, but we actually had to expose this choice to customers. We have a pro agent and a regular agent. And the, and the decision point is, you know, we can have the agent, very simple one, you ask the agent as we were testing this and kind of just cranking on this for over months. You ask the agent, what are the top, what are the top sort of box offices in, you know, around, around the world? And basically, or maybe something even more precise. What are the, the box offices? What are the addresses of box offices in the following locations? And we'll we'll do this trick where we, where we give it a few fake addresses, fake locations and, you know, a bunch that are real. And you have this dilemma, which is the agent has to go in and run this query. The user wants this really fast, right? And so what, what you should do is just the agent should just go and search for, for all these offices and find the locations. But what happens when it doesn't find two or three of, of the addresses? You basically have this, this, you know, choice point for the agent that the agent has to go through, which is, do you stop at one search? Do you do three searches? Do you do five searches? Do you do 10 searches? How does the agent know what it doesn't know? How does an agent know when, when the task is truly complete? And the way that we'd sort of test this is like, again, we give it fake, fake locations. And so you basically have to figure out like, when does the agent decide to give up on it? It couldn't find those locations or not. And the challenge is, is that that is a that is like a task where you just have to, you have to decide how, how much compute do you want in this process and that will generally correlate with how long the task goes for. So I can get you that answer back in five seconds, but it'll be wrong half the time. Or I can get you the answer back in 15 seconds and it'll be right 95% of the time. So how does the user sort of, you know, understand and interpret those trade offs? This is one of the big challenges in AI. Okay, we need to take a break, but when we come back, I definitely want to speak with you about who's going to get the value from this new set of use cases, whether it's going to be the big labs or those building upon the technology. And I also started this podcast saying we're going to talk about how OpenAI and Anthropic stack up in the competition. And I've yet to get you to weigh in on who's going to win this. So let's do that right after this. If you think about it, most work isn't actually hard. It's just repetitive. Status updates, routing tasks, answering the same internal questions over and over again. These are the things that quietly eat up your team's hours every week. That's where Notion's new custom agents come in. Notion is an AI-powered connected workspace for teams. Notion brings all your notes, docs and projects into one space that just works. It's City and County of Denver saw a 50% reduction in false claims against them and a 94% reduction in safety events overall. This is the kind of visibility that every operation manager needs. Don't wait for the next accident to take action. Head to samsara.com slash big tech to request a free demo and see how Samsara brings visibility and safety to your operations. That's samsara.com slash big tech. Samsara, operate smarter. And we're back here on Big Technology Podcast with Box CEO Aaron Levy. Aaron, before the break, I mentioned that I was curious to hear your perspective on who's going to get the most value from this technology. Is it going to be the labs or is it going to be the people, the companies building on top of their technology? And it does really seem like there is some competition there. I mean, they want a lot of this agentic stuff to happen within their super apps. So how is that battle going to shake out? It's very different than like, I have a chatbot and I'm applying that chatbot technology inside like a legal app. Yeah, yeah. So I think, first of all, I would say, unfortunately, I'm going to give you kind of some lame answers here because I think the jury's out. I don't think we know, you know, ultimately what happens because you can kind of argue your way into a couple of different outcomes. One is that you could argue pretty easily that eventually domain-specific agents end up being the best way for these agents to manifest in an enterprise because the domain-specific agent deeply understands the context of that industry. It can wire up to data systems, proprietary or public data that is just purpose-built for that particular industry. They can do the change management of the workflows of that industry because they will just have people that are just like dedicated in their focus in a particular industry use case. And they're just, again, like you have a full complete solution just applied to your vertical. Conversely, you know, the kind of bitter lesson people would just argue that actually everything I just described is like two or three model generations away from getting eaten away. And to the bitter lesson side of this, I think that the part that I would just argue is like there's always domain-specific context. If for no reason other than just the model can't know what all the different projects are that somebody is working on and the data that they have access to, the model has to tap into that. And so then the only question is like how much is the value created by the products that allow the model to tap into that information? Or is it actually easier and easier to do in a kind of purely horizontal way over time or with some of the skills that you just pull into the agent? And I think like the classic debate that you'll see on kind of social media around this is, you know, Harvey or Legora versus the kind of more horizontal cloud co-work style agent. I just think it's a really great debate. And I don't know that I just don't know that you can totally simulate out what's supposed to happen here. Because even in, you know, kind of traditional SaaS software, we saw $30, $40, $50 billion vertical software companies emerge in categories where there was already plenty of horizontal products that could have solved those problems. But just that relentless level of deep vertical focus led to customers being much more willing to trust the vertical player because they just know that every morning that company wakes up thinking about their workflows. And so I think that it's just, it's too early to see how this is going to play out. The good news is there are going to be value in both sides because even the vertical domain specific players will be riding on top of the intelligence from the horizontal labs. And so in both, in all the scenarios, the labs win, you know, a very big prize. Like that, that's the thing. So the labs are fine either way because they're going to have, they will be the intelligence layer of any of these outcomes. Then the only question is how much value is created on top of the labs for the applied layer. And, and we just, it's just very early to see how that plays out. Right now, I think it's going to cut differently by industry. I think there's some industries where the customer has such either regulated or just like high value work that they need to do that they just want an off the shelf solution that just thinks about that work day in and day out. And then there'll be a lot of things that are just like, okay, you know, writing an email, you know responding to my calendar request, putting that in email and then adding that to a Salesforce record. That's very general purpose. Like that, that's going to be something much more, you know, suitable for like a pure horizontal agent. But like I have to go super deep in some legal workflow or I have to go super deep in an M&A transaction. These things are pretty tailored use cases that I would, I would, you know, probably more often than not bet on the applied kind of layer. Okay. And so just for clarity, the bitter lesson folks are the ones that say, you add more compute, the models will get better and they'll basically like, they will be able to handle any use case that, you know, someone who's building on top of the model could with, you know, specificity. So, yeah. And the way to think about it is just like, imagine you have that much, let's say this is like your bar chart. And three years ago, if you were a rapper on an AI model and you actually were like, like successfully delivering a high value outcome and you, you know, the bar chart was this, the top of the bar is the kind of, you know, full solution. The rapper companies would have needed to, you know, do like 80% because, because, you know, the models were pretty weak. Now the models have gotten good and it kind of moves up, up the sort of rapper upward. Then you can just vibe code a wrapper studio. Now, now, now here's, here's the thing though, that it's important though. It's important to not think about this as a static, you know, sort of dimension. What's happening is as the models get better and better, one would think, well, the rapper should shrink until the point where the rapper is just like that big, right? But what's happening is, is that actually as these capabilities get better and better from the models, the use cases start to expand that the customer wants to go do. And so then there's basically another set of things at the rapper layer that is, that is sort of needed to get built out. And we'll just have to again see how, how rich and deep is that ecosystem. But I think there's going to just be, I think there'll be hundreds of successful, thousands of successful products at that layer simply because again, enterprises, they, they just want to, they want to wake up. They want to get their job done. They want to have some alpha relative to competitors and they don't want to be thinking all day long about how do I go implement a new technology solution? So the company that can show up at their, at their offices and basically say, I have, I have the purpose-built solution just for your use case that they're going to have a leg up. assuming that there's no other trade-off in like, it's worse intelligence or it's vastly more expensive or it's, it's you know, it's so minutely, you know, useful that it's just not worth adopting another vendor for, but there's a lot of reasons why you still buy, you know, vertical or domain specific technology. So there are, speaking of like making things bigger and then getting better, there are some new models that are on the way. So we hear OpenAI has this spud model that I spoke with Brockman about. Anthropic apparently has a bigger model coming out as well that just finished training. Brockman actually said something interesting that spud was built on two years worth of research. And, you know, we've talked a little bit about these models getting better with more compute. Well, actually the compute buildout started like crazy maybe two years ago. So we're going to start to see what's what the product of building on these bigger data centers actually is. Turn it to you. What have you heard about these new models? What are they going to do? I think we're probably, you know, reading the same, same conversations. I'm listening to the same clips that of your interviews. And, and I, I do, I appreciate that, that this round of model improvements seem to be more public than, than other ones. I, I would say the, you know, it's, it's always hard. There's always these like viral leaked images now online and like you can't tell which ones are, are actually real or not. I think there's a lot of, a lot of generated content out there, but you know, for all intents and purposes, it's pretty clear that we have two gigantic, you know, capability models coming out in the, you know, weeks and months ahead. And I think, I think certainly probably the biggest takeaway is just like, we are nowhere close to hitting a wall. I remember it was probably only about a year ago where there was a lot of, a lot of talk on like, Oh, have we hit a wall and these things are only kind of eking out, you know, tiny little improvements in, in capability. That's just obviously not the case anymore. We saw that through the winter. I think we're about to see that in the, in the next, you know, two major model drops. I think that's incredibly exciting. And, and, you know, on every dimension that I think is going to matter agentic coding, agentic tool use, domain specific kind of applied areas of knowledge work, life sciences, legal financial services, consulting, et cetera. I would expect that I love that. No, no, no, this is great. Journalists' rules is let the subject talk and sit back. Yeah, but media training says don't answer any further and just let the interview ask more questions. Listeners and viewers, Aaron and I will sit here for the remainder of this podcast. This is the ultimate end state of two sides of training. So I think I'm not going to answer it in the way that you'd obviously like. What I would say is that you have two just incredibly competitive, insanely talented, well-funded, very motivated companies in both of those companies. And I think I've probably used this kind of analogy in your podcast before. I can't shake it from my head, so I do mean this fully. It's sort of like trying to predict anything about the cloud wars in like 2008. It's just like we are still so early in the total sort of evolution of the market. And, you know, I ran this stat recently, actually. I think my numbers are like mostly correct. You know, they came from AI, so, you know, bear with me. I did some extra Googling to check on them. But in 2010, the cloud revenue of AWS, 2010 is like kind of like yesterday. Like, I remember 2010 pretty perfectly, right? Like, it wasn't that far away, which is scary. So, 2010, AWS was about $500 million in revenue. Azure launched that year or had just launched. GCP was called Google App Engine. That's how early this was. Their logo was like a jet engine, like a little cartoon jet engine. So, needless to say, like, not a serious contender in the cloud infrastructure wars. So, that was $500 million was like the dominant player. The past year, you know, I think the total spend on cloud infrastructure is a couple hundred billion dollars, you know, range. So, just think about that scale in 15 years to go from $500 million to a couple hundred billion dollars. And so, if we were doing a podcast in 2010 and we're like, how is this going to all play out? And it actually, the answer just should have been, it doesn't matter. Like, literally, like everybody ended up with a $50 to $100 billion revenue business at the end of all of that 15-year period because, because of how valuable cloud infrastructure was. So, I think of intelligence more as like a multiple on that. And so, it's kind of like the daily skirmishes that we have to kind of pay attention to and get excited about. It like probably just doesn't amount to as much as just you fast forward five or 10 years and all of these products are five to 10 to 20 to 50 times larger. So, that would be my answer. Maybe, though. I mean, it does matter, I think, and to a degree because if you're able to command this lead, you can maybe get more funding, more infrastructure. And that all compounds on each other. But I agree with your central point, though, is that where it's early and like even if, let's say, Anthropic, just to use one company as an example, has a lead now, it doesn't mean they'll be holding it forever. Well, and even in the cloud, like cloud was the kind of the original CapEx-dependent, you know, sort of, you know, CapEx-heavy form of software. And you would have thought, like, will there be this major compounding thing, like whoever can build the most data centers gets the most workloads and then they'll build more data centers and then they'll get more workloads. And yet, 15 years later, from that point in time, we now have four in the U.S., including Oracle, four at-scale gigantic cloud providers. We now have Neo cloud providers. We have international cloud providers. You know, China has its own ecosystem, as an example. So you basically have, you know, at a minimum, 10 very, very good businesses that are in cloud infrastructure from what you would have thought, you know, should have already had this sort of like escape velocity kind of return. So I think AI has a lot of similar properties, which is unless there's some so kind of closed proprietary research event and breakthrough that happens that just simply nobody else knows about, and we have no evidence that we've ever had one of those in AI, like, you know, these things just eventually sort of emerge across the ecosystem. Unless that happens, I think, you know, any one lab probably has a six-month to one-year lead on like a, on the breakthrough AI model. There's lots of network effects, like, like the more people that build on your APIs, then your tools, you know, work with those APIs. So we're not only in an intelligence-only competitive battle, so there's lots of reasons that you're going to see network effects in ChatGPT, in Codex, in Claude Code, and so on. But these markets are just so big that, again, I'm just not worried about kind of who wins in this simply because all of these companies will be much bigger in the future. Aaron Levy, always great to speak with you. You're always welcome on the show. Thanks for coming on. All right, everybody. Thank you so much for watching and listening. We'll be back on Friday with Ronjan Roy of Margins to break down the week's news. And we'll see you next time on Big Technology Podcast.