Back to all episodes

Opus 5.5 launches: is Anthropic back?

September 29, 2026

Anthropic might be back. Opus 5.5 performs at Fable 5.1 level for 40% less, and Sonnet 5.5 dropped the same day we went live, beating Opus on Terminal-Bench 4. OpenAI answered with GPT-6 Sol and Luna and a rumored long-running agent called Aeon. Meta launched Muse at Connect, then hired MongoDB's CEO to lead a new enterprise platform. Then there's safety. SNL took a shot at Dario, a report says the labs are looking into tens of thousands of incidents, and NVIDIA launched an agent safety platform with 100+ partners. We also explain why Mastra doesn't believe in model routers, and close with quick hits: Anthropic may kill plan mode, one company rebuilt Devin in two days, CI costs are exploding, and Google is sending TPUs to space.

Episode Transcript

00:00

Shane: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents, not dozens, in which their frontier models took steps that outside evaluators would consider problematic. So then of course someone updated the Felony Bench, and Anthropic and OpenAI are now benchmaxing. How do Anthropic and OpenAI really not have any control or any like visibility into what they're doing? Maybe they should be using some basic security practices because it sounds like a lot of the stuff's pretty trivial. Welcome everyone to Agents Hour. If you're just tuning in, it's Monday, September 28th. We're doing the news. There's a lot of things to talk about. We got some new model drops. We have people building model routers. We have TPUs in space. I mean all kinds of things. So we're gonna get into it. We got the drama. We got the news. Let's go. Let's go. I thought this was funny because it's kind of true. And of course this is poking fun at Dario a little bit. In some ways it's also a compliment. 'Cause it's saying like Opus 4.6, then people said Opus 4.7 was maybe not quite as good. It was a little bit of a backtracking, Opus 4.8 felt worse. People hated Opus 5, but now Opus 5.5.

01:19

Abhi: It's a Chad, dude, Giga Chad. I would agree, dude. Opus 5.5 is a Giga Chad.

01:24

Shane: And I'm sure a lot of you are like me. I know a lot of people on our team are like me as well. We kind of bounce back and forth between Codex or Claude Code or Claude Desktop. I use Mastra Code to write code, but I use the Codex Desktop app and the Claude Desktop app just to like see how they do things, see how well it works. I have run some like long-running projects in there and Yeah, this got me to open up the Claude Desktop app for the first time in probably a month, and I was running some long running tasks with it. And Opus 5.5 is legit.

01:51

Abhi: Legit. Feels good. Feels very good. Honestly, it feels better than Astra and for like way less the price. I haven't got rate limited. Usually I get rate limited in a day. I think I've run seven days of Opus 5.5 work not getting destroyed by the rate limits. It's amazing. So.

02:09

Shane: Anthropic is back, maybe yes or no?

02:12

Abhi: What do you think? You know, I was worried about, you know, all of us dying in the next decade, but you know, maybe I should be worried more now because 5.5's so good.

02:20

Shane: So this was came out on September 22nd, introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus. From what I've heard internally and what I've seen online with some of the, you know, people always show the demos, right? You can never really trust 'em. But I think the vibes of what we've seen internally match you know the hype that we've seen externally as well.

02:48

Abhi: Agreed. It passes the vibe check.

02:50

Shane: A few other things from Claude. Cloud sessions are now officially available. And Sonnet 5.5 dropped. This is from September 28. So they launched Opus 5.5, you know, last week Sonnet 5.5 came today. I haven't actually tried had a chance to try it. This came out just a couple hours ago. But from what I've seen online, people are saying it's giving the same vibes. Like it feels like you're talking to the same model and the way it's responding. It gives you kind of or pre-Sonnet vibes when it was you know people felt like it actually was good to converse with. I did see a benchmark. I think this was specifically on Opus, but Opus doesn't use em dashes anymore, so maybe they're trying to improve the writing 'Cause I know it felt like when you're talking to Opus, especially after 4.7, it just felt it was a task to talk to Opus. It talked in a certain way and it was really annoying and it could not write at all, it didn't seem like for me. So maybe they're actually trying to improve its ability to like summarize, generate ideas, not be so verbose. I haven't confirmed all this, but in the little testing, at least with Opus, it feels oh, quite a bit better.

03:49

Abhi: You know, it stopped using these dumb catchphrases too, like load-bearing and like belt and suspenders and it's way more chill. I think it called me dude today. So that I mean I appreciate that. So.

04:01

Shane: Probably 'cause you talk into it that way, right? Trying to match your vibe. Which it should do, right? You know, if you're talking to it, you know, "Thank you, sir, for doing this." Well then maybe it should talk to you professionally. But if you're saying just do the thing, and then it can respond back. It should meet you where you're at.

04:16

Abhi: Yeah, meet me where I'm at, dude.

04:18

Shane: It's funny because even online you see all these like AI bots on Twitter and I by the time I get to like the second sentence, I know what I almost know what model they use to write that response and it's so much Claude slop that I'm hoping that, you know, as they're improving the models, they're going back so at least the bots will sound a little bit better. And if you're using a bot to post your all your Twitter posts, just don't do it. It's annoying. Alex Albert posted, Sonnet 5.5 has the same feel I liked about Opus 5.5. It writes clearly, it's very fast, and it's a major capabilities jump over Sonnet 5. Really great model to iterate with. This also came out today. It says Opus 5.5 cheats less than prior Claude models in Drone-Bench. It's also number one getting a better score than both Astra and Fable. Really impressive. There's a lot of benchmarks, but I do think that Opus 5.5 seems to do well on them.

05:07

Abhi: Oh, one thing about Sonnet that we just saw before the show, we were looking at Terminal-Bench 4. And Sonnet is actually scoring higher than Opus 5.5 on agentic coding. I believe the number is 70.6% compared to Opus 5.5 at 66.4%. I mean I don't know. Maybe the maybe something's wrong with Terminal-Bench, but usually it's not. We'll see. We'll dogfood it ourselves. But just to compare, Sonnet 5 scored 10.6%. So this is definitely a step up forward.

05:39

Shane: Here's my theory. You know like sometimes you meet that really smart person that's too smart for their own good. And...

05:44

Abhi: They get bad grades.

05:45

Shane: Yeah, they get bad grades or they're or they're like they don't have the street smarts to do like real world things. You know, so I wonder if like Terminal-Bench is judging like real world things and Opus is just sometimes, you know, sometimes too smart for its own good. Obviously it's a great model. Like I think it's a very good model. But you know, sometimes you need someone a little dumber. Just come in and get the job done.

06:02

Abhi: Yeah, get the job done.

06:03

Shane: Yeah. So maybe Sonnet's like the grinder. You bring in, get the job done. And maybe sometimes that means do a better job than like the really smart nerd sitting in the corner who can answer every question, but maybe can't get the real work done.

06:15

Abhi: I'm gonna have to do some experiments with Sonnet after this.

06:19

Shane: Alright, so OpenAI did try to respond. Yeah, on the same day, I think, they welcomed GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. So they have these kind of like cute model cards where they have Astra, Sol, and Luna. I guess Terra has just gone, you know, is gone now. I don't know. Maybe Terra will come back, but they want to delimit it at three. I thought that was kind of interesting. With Anthropic, you know, you have Haiku, Sonnet, Opus, and Fable. That's kind of like four feels like a lot. To choose from. Three's a little bit better. Do you want like small, medium, or large? I guess that's the you know don't need maybe you need the XL, but I think it's better just to have three. My hot take is like Astra is the new Sol, Sol is the new Terra, and then Luna is still Luna, but Yeah, we'll see. I think overall though, Astra's still good. It seemed like the tide was shifting towards OpenAI for quite a while. But now maybe with this latest round of launches, Anthropic is kind of taking the lead again.

07:12

Abhi: Yeah, I've been having so much trouble with GPT-6 Sol because I have a long horizon task that I expect to essentially, you know, check in with me when you're done. But what's happening a lot in my sessions is it'll complete one piece and then wait for me to tell it to continue or give it, you know, ask permission. Should I continue down the path? And I kept having to literally my chat, if you just scroll, I hit up it's a hundred continue. Continue, man. Just keep going. And that is so frustrating, which did not happen in 5.6. It didn't happen in Astra. So either maybe it's a skill issue. But it's annoying.

07:50

Shane: This came out on September 26th. OpenAI is set to reveal a long-term agent at DevDay. So maybe this is what you need, the long-term long horizon agent. Codenamed Aeon, built for long-running tasks similar to Grok Bot or Manus, works for hours, days, or even weeks. It's likely built on Astra. So it sounds like it's their kind of Grok Bot or Muse or whatever that everyone's trying to build. Let's talk about Meta, because Meta's making a move. There was the Meta Connect conference last week. This was September 23rd. First they announced Muse, personal agent. It has voice in real-time video. I'm assuming this is kind of like to compete with Grok Bot and all the trends of around having a your own personal assistant. So you can get things done when you're talking. You can also bring Muse to glasses so you can talk hands-free throughout the day. The Meta glasses can now serve as FDA-cleared hearing aids. The glasses are very light and they had a video with some different celebrities wearing the glasses and Yeah, even Palmer. Yeah, that was that was the big thing, right? Because they had huge. Yeah, because he sold to Meta, right, his kind of VR. Was it Oculus, right? Is that? And then yeah, kinda had some falling out.

08:58

Abhi: Got fired for being a Republican.

09:00

Shane: Yeah. And then now it came back and said, you know, like these glasses are actually pretty good. So for them, I don't know how much they had to pay him to do that, but but either way, like people seem to like it. So they have the Meta glasses. They also have the new VR glasses. So that's where they brought Palmer Luckey in and kind of shared it. And then they also have this little like Muse charm, which is literally Tamagotchi. Tamagotchis are back.

09:23

Abhi: They're back, dude.

09:24

Shane: The real life Tamagotchi. So that's all the things that happened at Meta Connect. They released a lot.

09:29

Abhi: Some people are excited about it. Muse assistant, you know, Meta's really trying to rewrite the story about themselves as like we're AI native. Obviously Horizon is dead, like you know, they're the AI company. They're also care about your safety and your privacy, which I don't necessarily think is true, but they are trying to build sandbox primitive so you can get your own Linux machines. You'll have your own API key vaults and you can pay for that, which will cost a lot more to run Muse inside of that. And then you'll have zero day retention, so you're not getting trained on that. I'm assuming in this market for Muse, most people don't know what a Linux machine is, right? So like why would they buy one? So they'll probably just be the consumer and get trained on and all that type of stuff. But for those who are gonna use it for like maybe heavier workloads, they'll probably maybe pay for the Linux box or whatever. But yeah. It seems like Meta's really switched the narrative in the last couple of weeks.

10:29

Shane: Yeah, I mean with the Muse launch and the model being pretty good and now all the consumer apps around it. I do wonder too, Meta's so well positioned to tackle and to compete on the consumer side. And I still, as much as, you know, we on this show, we kind of throw some shade at Google, I still have friends who use Gemini, right? Because they are more on the consumer side. They don't care about the benchmarks. They just care, can it do my simple tasks I throw at it? And for the most things it probably does a reasonably good job. If you look at the benchmarks, Muse is probably already surpassed Gemini in a lot of things, or at least you know close in others. And they have such a consumer market that it feels like they could easily capture a good chunk of that, you know, consumer AI. So if I was OpenAI, I'd be a little worried about Meta. If I'm Anthropic, I'm maybe slightly less worried today, but Google, I would also be worried as well, right? Because they have obviously huge distribution just like Google does.

11:21

Abhi: Maybe Grok should be a little worried too. Because like the everything app type of agent where it's gonna purchase things for you, it'll find you products, you know, help you in your real life. That's what these other auxiliary models and companies are doing. So now like there's another person who's trying to do that with the Rolodex of a bazillion users that use it. Also my parents still use Facebook, like so they'll probably get mused up.

11:45

Shane: Yeah. I mean if they integrate it with Facebook and Instagram and things, I could see people really adopting it. There's more from Meta though. And so this is where if you're Anthropic, maybe you worry a little bit more. Mark Zuckerberg posted this is today, so

11:59

September 28th: We believe superintelligence will create significant new opportunities for all people in businesses. Meta already serves billions of people at scale and helps hundreds of millions of businesses reach customers. Today we are starting the next major pillar of our business. Meta enterprise platform.

12:13

Abhi: I like how he said superintelligence. He's trying to use the word. Yeah.

12:17

Shane: Meta enterprise platform to help businesses use AI to grow and transform in new ways as well. So Meta's going full enterprise. They obviously have consumer. Now they want to be in the enterprise game. So it'll use our strengths that few other companies have, advanced models, leading agents, large scale infrastructure, and years of working closely with many businesses. They'll focus on bringing our full technology stack, including the Muse Agent, Meta Business Agent, Muse API, Muse Code, and more to businesses and developers to help them grow. To lead this effort, and this is the part that I thought was pretty interesting. I'm excited that Chirantan Desai, CJ, will join Meta as chief enterprise platform officer, reporting directly to me. So CJ is the CEO of MongoDB. Was the CEO. Was. Now, you know, that they pulled him. I mean Meta's been good at this, right? Pulling big names from places. And he was obviously at Cloudflare before. But yeah, that's pretty interesting because MongoDB is a very much an enterprise sales organization, right? Like that's where they make their money. And I think Meta's wondering, you know, and thinking, how do we make money out at enterprise. You know, CJ can probably help us quite a bit with that.

13:22

Abhi: Changing the game. I guess Anthropic should be scared. Maybe there'll be more FDEs coming into these places.

13:28

Shane: If you're an FDE, you're gonna have lots of options. Sheel Mohnot says Mark really buried the lede in his third tweet. He poached a CEO of a $30 billion public company to lead it. So pretty big get if you think about it, because MongoDB is no small company.

13:41

Abhi: I wonder what that comp package looks like. It's gotta be good. More than what I got. Like, share, and subscribe.

13:47

Shane: And follow us on X.

13:48

Abhi: And tell your friends.

13:50

Shane: And their friends.

13:50

Abhi: I mean, we're not begging. Well, maybe a little bit. Subscribe to Agents Hour every Monday, noon Pacific.

13:57

Shane: Alright, so everyone reckons with safety. Let's talk about safety. This is so funny. Alright, I mean we gotta watch it, right?

14:03

Abhi: We gotta watch it. We have to.

14:05

Jan the Producer: Hey, editor's note, I can't play this for copyright reasons. It's an SNL skit about Dario, and I will link it in the description so you can watch it yourself.

14:13

SNL: I want to assure you that if we can pressure lawmakers to create guardrails, we will be able to stop me.

14:19

Shane: So there's that, and the funny thing about it is if you are paying attention, that's kind of how you felt, and I think most consumers, most people out there that aren't in the space, this is probably the first time they've been exposed to the fact that this is, you know, at some level kind of wild that you control the models and you're saying it could potentially kill us all. Couldn't you just stop it from doing that thing that is worrisome? So I think it is kind of nice for SNL to show the more like common sense realistic side.

14:53

Abhi: Yeah. ...are like viewing it pretty gravely. Maybe this doesn't really help though, but I think it's quite funny.

15:06

Shane: But then we have this

15:07

Scoop: OpenAI, Anthropic, and security researchers are investigating tens of thousands of incidents, not dozens, in which their frontier models took steps that outside evaluators would consider problematic, which I don't know what problematic means, but obviously a much larger amount of potential incidents. So then of course someone updated the Felony Bench. And Anthropic and OpenAI are now benchmaxing. And I obviously I don't believe that they're all felonies. I think this is just a joke, of course. But it is this idea of like how do Anthropic and OpenAI really not have any control or any like visibility into what they're doing. Maybe they should be like using some basic security practices because it sounds like a lot of the stuff's pretty trivial things. It's like these are some basic security expertise on the team would have like solved not all these things maybe, but a large percentage of these things. And so at some level maybe you're just acting irresponsibly.

15:59

Abhi: Maybe. I mean it's like Jensen, both Jensen and Satya said it all in, that the labs are transforming into engineering organizations, which makes me think that or makes it imply that they are not currently operating like one. So.

16:13

Shane: Yeah.

16:13

Abhi: That's what happens when you send an auditor or evaluator. You just get a bunch of red lines.

16:17

Shane: And then Jensen, you know, right on cue came out today. This kind of announcement. It says today, with over 100 industry partners, we introduce the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. And then they have a whole bunch of logos that kind of like on the journey with them I guess. Alright, so let's talk about the model router era. Jev came out last week. Everyone's using Jev for everything to try new things. That's cool. You should. You should see what it can do. OpenRouter released Jev Router, a cache-aware model router powered by Jev and Typesafe AI. Picks the best model and reasoning effort for each request, balancing quality, speed and cost. Theo then came back. Tyler on our team said, Theo spent a bunch of money to do what I could have already told you he would have got for a result, which I thought was funny. But he said I stayed up till 2 AM and spent a thousand dollars benchmarking Jev Router, so you don't have to. Performance on DeepSWE was roughly the same as GPT-6 Astra on low. It costs slightly more, and it took almost five times longer to run. What are your thoughts on model routers? There's a whole bunch of these. Some have released model routers and then rolled it back.

17:55

Abhi: Yeah.

17:55

Shane: This is before Jev, and so maybe people, you know, now are thinking Jev is a faster way to do a model router.

18:00

Abhi: So many products came out with this like special model routing. And then they have benchmarks that say that they're as good as Astra or Opus. You know, Devin had one. A bunch of people have these. And I think that's cool. I think it's cool, but I don't know if it actually works in practice. Our kind of as Mastra's standpoint against it is we don't think it works. Because the prompt cache is too delicate and you might as well just continue using the model you're already using based on your cache so from that perspective we don't I don't know, we don't really do it. We did release a model router processor in Mastra, but that was more of an experiment because previously the barrier to entry of doing model routing was you have to figure you have to make a decision on what's the task trying to do and what models should it pick. So you're actually spending money on reasoning. We thought with Jev we could see if it actually makes things better. And you know, Jev can classify what the thing is, but even in my own experimentation it didn't actually do the job properly. Maybe I need n closer to infinity to make something more substantial, but still it's just cool as a pattern if you want to like play around with it. But just never really thought it actually made sense with Jev even included.

19:16

Shane: Yeah, I mean here's my take. I think that there are probably some use cases where you could save a lot of money. I think 99% of the use cases, 99% of us aren't gonna have the sophistication to know if it's actually gonna work. Yeah. And so if you don't have the sophistication to know, you probably shouldn't be using it. Because you can trust some provider that says it works. But are you actually gonna benchmark it like you know to Theo's credit at least he did? The answer is no, you're probably not. You don't have the time, right? So if you had all the evals, if you can benchmark even like a model change that like using this model and then switching to upgrading to the new model gives me better performance. Well then you could benchmark a model router and you could decide for yourself if it actually works. And if you're doing that, by all means like go crazy, make a model router. You might be able to save some money, but I think for most people, it adds a ton of extra complexity, right? Like now you're like have this routing step in there, it's deciding between models based on usually the the way someone builds a model router is its classification of how complex is this. A real human that's a good engineer, sometimes they get it wrong. Like quite a bit, they get it wrong, right? You get in there and you're like, oh shoot, I didn't realize that this area was actually touching this area of the code. This two hour task is gonna take two days now, right? Or sometimes you get in there like, dang, this is gonna take a week, and you get in there and it takes a day. Humans have struggled to classify engineering tasks, and of course sometimes you're right, and the majority of the time maybe it can make a decent effort, but so often you're wrong. And so I think that's where adds a ton of extra complexity and unless you can actually measure it, I don't think you should be doing it.

20:54

Abhi: Yeah, and like, you know, Jev Router from OpenRouter, most model routers are naive and you know they just use the last human input to decide what the complexity of the task is. What you should be doing is looking at the conversation. Then, you know, if you're gonna change a model, then you know, you need to have the ability to understand is it worth destroying the cache or not. That's already a freaking decision. How do you make that decision?

21:18

Shane: And Jev can't really reason, right? It's a system one model, which means it's meant for like quick decisions, not deep thinking about a problem. So in some ways I'd argue maybe Jev isn't the best model router anyways. Because you actually want to reason around about it and decide this is a kind of a big decision. Do I send it to a cheaper model or do I keep the cache warm and keep up with this more expensive model? Let's go to the quick hits. We got a bunch of them. We'll run through these kind of rapid fire. Thariq from Anthropic says, we're thinking of killing plan mode and using a Shift+Tab hotkey to adjust effort levels. So this is a hot take. This got some people wondering, should we be using plan mode? So what do you think?

21:55

Abhi: Dude, I don't even use plan mode anymore in Mastra Code. I just plan with build mode and I just make moves.

22:01

Shane: So the one area where I still use it in Mastra code is when I want to use a different provider to plan versus build. Yeah. Which I still will occasionally do. Sometimes I use a better model to plan and then a cheaper model to build, especially if I know it's gonna knock it out of the park. But I'm doing that less and less these days. So usually my solution is just to say with two or three words that I tell it I want to make a plan and it'll just make the plan first, right? So you don't need to make the plan because you're the few times you want it to, you want to like stop and think about it, you just tell the model and it will go through the process with you. You end up with a plan and you make it anyways. There are some things in Mastra Code if you... You know, we have like a plan goal mode, right? Which is you can make a plan and it'll run. So if I'm doing a really long running task, I'll still maybe break it up. But I do agree in general that it's not needed as much as it used to be. Peter Steinberger says our next version of OpenClaw uses a decision model to automatically decide between steer or queue, which is kind of a cool idea. Right now if you use most agents, when you send a message it either like interrupts and steers it or queues it to the end. And it's kind of cool if to think about like you just send the message and it will know if it should do it now or it should do it later. This was a pretty popular post that got some views from Ian Sefferman. He says, I know Cognition doesn't care about little ol me and this is just my story, so take it with a grain of salt, but here's why I don't think all these crazy scaling revenue numbers by coding agents are going to stick. And then Ian goes on to just tell the story around as Devin bill was getting out of control with his company. It was hundreds of thousands of dollars, maybe. I'm trying to see the amount. Yeah, their usage was 100K plus run rate, so annual bill to Devin. And they wanted to just pay quarterly rather than monthly or something. And there was some hangups and so he just got frustrated and so he went and basically tried to build his own version of Devin and he said he kinda benchmarked it and it was basically as good. And he did it in two days. And so even if it wasn't as good, if it saves you a hundred thousand dollars and he did it in two days, I think that is again, who knows what benchmarks, who knows if it was just vibe tests, but this was kind of one of the reasons behind us building Mastra Factory, right? We use Devin, it was pretty good for some things, but we started building and we were using Mastra Code and we just said what if we put Mastra code and had it just be able to do all these tasks that we were sending Devin. So I think this has been kind of our experience too.

24:10

Abhi: Yeah, I've seen a lot of projects too, like OpenDevin, open this. So it's definitely hit a nerve. Devin is not cheap at all. Yeah. You pay per review, you pay for this, pay for that.

24:21

Shane: Yeah, I mean I do think it's reasonably good. But do you want to rent your intelligence or do you want to kind of have more ownership of it? Yeah. I think that's the decision that a lot of people are facing. Cursor introduced rollouts. So rollouts write a monitoring plan, then watch changes as they deploy. Deployments are verified, so regressions are caught before users see them. This is kind of cool, it's kind of tied. Yep. It reminds me of like OpenAI's when we saw their internal factory, right? They had this kind of concept of an agent that watches the rollout. So Cursor released it as its own kind of standalone thing. That's cool. I think this is early. We'll see more of things like this. Not only did the agent write the code, but the agent will monitor its release as well.

24:58

Abhi: I think the more distributed systems get built with AI, this is gonna become super super important.

25:06

Shane: skills.sh reached a million agent skills and nearly 280 million installs. So a lot of people using it. There's some jokes in there about like prompt engineering people to do install and do things that they shouldn't, but overall I think it's crazy how many people have been using this. That's awesome.

25:21

Abhi: It's like half of them are prompt injection skills. So be careful. Yeah, be careful what you install. Not all million are valuable. There's some porn ones in there for sure. Yeah, yeah.

25:32

Shane: Yeah, not all million. Yeah, there's like 0.1 percent. But still, congrats to Vercel on that. That's a lot of people using it. DigitalOcean is in the game, I guess. They've released managed agents in public preview. So you can run, you know, your own agent in a runtime environment that pauses when idle. It's this really more and more moves to just having cloud-run agents. So you have cloud-run agents that write your code for you. So even DigitalOcean is trying to get into the managed agents game. If you are watching the whole show, if you're watching live, you know we talked to Daniel from StarSling earlier. If you're not, you can catch that episode. It'll be released as a standalone episode. But this topic came up. CI has become the top bottleneck of every engineering team I talk to, including Lindy. Our CI spend has become stratospheric. The idea is, you know, we've experienced this too. You're writing more code, more PRs, means you're running more CI checks. And agents are really good at writing tests, right? Or sometimes they're not even good, but they'll just write a lot of tests, way more tests than humans would write. So your CI, so you not only do you have more runs, but your runs last longer. You can even write, you know, the end-to-end tests that you always dreamed of and never took the time to write. It'll write those for you. So people's CI processes are getting fatter and they're running more often. So it's becoming incredibly expensive. There are tools like StarSling that we mentioned. But I don't know, what do you think about this?

26:51

Abhi: It's gonna continue being a problem, you know. That's why I'm looking forward to the next gen of CI tools. Maybe CI goes away like we were talking with Daniel. And it's more agent CI and agent verification. Should be a lot cheaper hopefully.

27:05

Shane: Chris Tate made a CLI version of and if you read this, this is just kinda cool. A lot of people struggling with losing all their disk space because of agents. So Chris Tate made a cool little tool, a CLI tool for seeing where all your disk space is going. So if you have tons of work trees and your local coding agents are doing some random things and just eating a ton of disk space. You can reclaim some of that back. This is the last one. Google's launching TPUs in space next week on SpaceX Falcon 9 to test AI data centers in orbit. This was from September 24th. So supposedly it's happening this week. I don't know when, but this is crazy if true. Elon's been talking about compute in space, you know, GPUs in space. And now apparently Google is doing it. They have some TPUs that they think can run in orbit.

27:51

Abhi: And using the Falcon 9. Yeah.

27:53

Shane: The rich get richer. Yeah, it's wild. I think people thought this was like years out right? That people would actually do this. I think I saw Elon Musk have a tweet that said, you know, his prediction is the amount of processing on Earth will be just a rounding error compared to what's going to be in space. That's a pretty bold prediction. I think we're early days. But it will be interesting to see if this actually works. I mean at some point I also think of like if you are in the US, you know there's a lot of people that just are really anti-data centers. I'm just gonna say uneducated people, very uneducated people with all the reasons why we should not have data centers. And I am fairly educated. I haven't read everything, but I think you can kind of tell when something doesn't smell right. And most of these people just have no idea. And so my question or my thought is like they probably got what they wanted. They didn't get any economic activity in their state or cities from these data centers. And so maybe that's good for them, but also the the companies are gonna find ways to go about it anyway. So if they can't do it in the United States they're apparently gonna do it in space. So, yeah.

28:52

Abhi: I mean what's weird is the CEOs of these companies are saying that the cities that they're putting the data centers in have these huge economic growth, new jobs, people get richer. You get a nice mall and all that stuff. And maybe it's for the people who don't want to live in the data center city that it doesn't really improve your life.

29:11

Shane: Yeah, but why do you fight it? Like, let's not pretend there are no drawbacks. There absolutely are, right? Like you maybe have some noise. You do use more electricity, but if they bring their own electricity, right? You use some more water, but much less than a golf course, and I don't see people halting golf courses in my area. You know, it's a twenty-seventh of a golf course or something ridiculous like that. So I don't know, I feel like a lot of people are uneducated and that's just caused people to have to get creative and so now they've pushed this agenda of getting TPUs and GPUs and space up, which maybe's cool. Like maybe that's a good thing. Maybe that's where it should have been all along. Yeah. But I do kind of feel bad, like in my community, I actually think it would have helped a lot of small towns that are not doing well. They used to have factories, and maybe it doesn't bring on a huge boom of workers like those old factories did. But you do need plumbers and electricians and maintenance and all these things, right? Like those are real jobs. And now I feel like those jobs are the you don't have access to 'em because you're uneducated and you didn't bring in data centers when in my opinion, I wish at least in my state, like where I live most of the time, which is South Dakota, when I'm not in San Francisco. I wish we would have brought more data centers in because I think it would have helped the smaller communities. So that's my take at least. And yeah.

30:21

Abhi: Those people have families who need teachers and grocery stores. And a whole economy, right?

30:26

Shane: And if you've ever been to smaller towns, you know that they just don't have a lot of industry, right? They all get pushed to the big cities. And so if you go to these small towns, they're in terrible shape. And if you just had a data center which could still be within like an hour of a bigger city, it could have maybe made a difference into some of these smaller communities and give them a bit of a revitalization. Would it have worked 100% of the time with every community? I don't know.

30:47

Abhi: Yeah.

30:47

Shane: But if I was the decision maker, I would have said it's at least worth a shot. I don't know, just me.

30:52

Abhi: I like how Elon's letting Google go first and so he can like see what happens and then he can revitalize his plan for when he goes puts his data centers in space.

31:01

Shane: And that's it for the show today. Follow us on X @mastra. You can follow the show on YouTube and check out all the other videos at mastra-ai. I'm Shane. You can follow me @smthomas3. You can follow Abhi @abhiaiyer.

31:14

Abhi: If you're in SF next week, come check out Tech Week. It'll be really fun. Come to Mastra events and say what's up.

31:22

Shane: Alright everyone, have a great day. This is Agents Hour. We'll see you next week. Peace.