GPT-6 Astra vs Claude Fable 5.1, We Should Pause AI & The Benchmark Wars | This Week In AI
GPT-6 Astra is finally here, and it changes what a model launch even looks like. Shane and Abhi open on the trailer for "Artificial" (Andrew Garfield as Sam Altman, in theaters Christmas), then get into the launch everyone's talking about: Astra's computer-use leap. This isn't just a coding model. It'll drive Blender, CAD, Notion, even Microsoft Paint, pulling a whole new class of technical-but-not-coding users into the fold. Abhi's been daily-driving it as a genuinely great collaborator; Shane's more mixed on the raw coding. Then the benchmark wars: ARC-AGI-3 leaps from 7.8% to 98.6% in one drop, the AI-2027 curve suddenly holds, and Astra even tops Vending-Bench while being "more ethical" than Claude. Anthropic rushed Fable 5.1 out the door on September 1st, and Astra promptly sucked the oxygen out of the room. All this progress reignites the "should we pause AI?" debate, with Bernie Sanders in all caps, Austen and Dwarkesh going back and forth, and the EU deciding ChatGPT is a search engine. Plus OpenAI's agents crack the Navier-Stokes Millennium Prize problem (with a "significantly more capable than Astra" tease), Cognition raises $2B at $48B, Gemini 3.8 Flash, Qwen3.8-Max tops Code Arena, Claude "open sources" Commerce Agents, Runway Solaris, and TimesFM-3.
Episode Transcript
Shane: If there's one thing I'm looking forward to, it's a different writing style. Yeah. Because I'm so sick of just seeing Claude slop. And you know the patterns. And I just know you had Claude write that.
Abhi: Is it load bearing?
Shane: This is Agents Hour. It is Tuesday, September 8th, and we're gonna do the news. Astra's gonna be the headline. But before we do that, I saw this come out. It's the trailer for Artificial, starring Andrew Garfield as Sam Altman. The film follows the story involving OpenAI focused on the firing and rehiring of Sam Altman in theaters on Christmas. Did you watch this trailer? I feel like
Abhi: It's a weird casting decision.
Shane: I don't know, I'm not gonna watch it on Christmas. I probably will watch it. You know, if for nothing else than the just being able to talk about it on the show.
Abhi: Yeah.
Artificial Trailer: Countries will fall. Industries are gonna collapse.
Shane: And like what's with this bunker with all these guns, you know?
Abhi: Yeah, dude.
Shane: Like is this how does this play into the actual story of him getting fired?
Abhi: Like how's a gun gonna help you in there?
Shane: Yeah, just like his bunker, personal bunker.
Abhi: Probably. Seems. Who are all those people at his house? I don't know. He really got his walk down though and mannerisms like how he walks and with his head up like this and stuff like that
Shane: He definitely studied. He studied for this role, which is good. So there you go. We'll see what happens, see how it lands.
Abhi: I think it looks tight. Yeah.
Shane: I mean that can't be real, right?
Abhi: Maybe it is
Shane: We gotta talk about Astra because this was announced on September 3rd. This was like the teaser, a teaser clip, but then it actually was announced I think the next day, right?
Astra Demo: Create a yellow circle there.
Shane: 11 million views, just got a lot of people knowing it's coming, and that was it. And then if you watch the actual launch video, which we're not gonna do right now, it shows and it's kind of a teaser because of how big of an upgrade it was for computer use type applications. And they the video shows people just like talking to it and having it do things on their computer. You know, of course in front of a big screen so it's really visual. And It's pretty impressive. What were your thoughts?
Abhi: First impressions. I thought one, the computer use stuff was such a perfect angle to get more users into the ChatGPT verse. There are many people who are not writing code, but that are technical on their computers, Blender, CAD, you know, artists, etc. And now they are pretty much. They're having their moment, right? They're building also 3D rendering is something that Astra does very well. A lot of Fable versus Astra on Blender models and things like that. The problem before was People were trying to make these tools work as like tool calls for the agent, but now they don't need to. They can have, you know, the native software in Astra can just run it and use it. A lot of Microsoft Paint jokes out there, but it's quite incredible. I don't know if it's fake or not, you know. I have to run it the experiment with it.
Shane: Pulling up a picture themselves and then telling Astra to draw them in paint and it's so good. So good. I mean it is obviously way better than any human. I mean not any human, but ninety-nine point nine percent of the humans could actually design it, right?
Abhi: And the launch video kind of showed like, you know, them drawing a rocket in paint and then saying, hey, Let's now design this in Blender. I really appreciated all this because it's not just about coding.
Shane: Yeah, it makes it feel like it's built for the real world. Yeah. You can build of real world applications on top of a model like this because you can actually just give it access to your computer. It got me using Codex for a while. I mean I use Mastra Code for everything, of course. Coding predominantly, but I wanted to have it do things across my computer in you know Notion without having to use the MCP. It just uses a browser. Yeah. It was navigating the factory. For me, and just like dragging cards across and doing all the things. It's pretty impressive.
Abhi: What do you think about it from a coding perspective?
Shane: We'll talk about that. I am less impressed. It did take 'em a day to kind of roll it out. So, you know, Thibault had announced on the kind of the launch day that it's gonna land. They did a bunch of resets. They had kind of I'd say they've kind of fumbled a bit. I don't know if they Didn't quite plan for the scale. I d I don't know exactly, but it did take longer, I think, than expected, or at least disappointed a lot of people. But eventually came out, I think, you know, on the fourth or the fifth. I do think as far as coding. And I don't know if this was a memory context problem of Codex, which I kind of believe it is, or a problem with Astra. I do think it was more of a memory problem. But I had it, I had a long session going where I was going through specifically for the factory, having it how. Help me update docs, reorganize docs, and it was a pretty long session. We created a whole bunch of pages, we moved things around, we made different decisions. And then I got to a PR point and I noticed it made redirects for pages it had created and then deleted in the same session. That's never been live. We don't need a redirect for a page that you created an hour ago. And now you removed. You know, you don't need to redirect that to a different page. I think that's something I've seen happen across other models as well, but it seems trivial, right? Seems like if you had maybe again, I do think it was it just the memory. Co it fell off the context window. Codex didn't know that it created it. And you know, of course it apologized and fixed it as soon as I caught it. But it's those things that you just kind of expect, especially when They're touting this frontier level intelligence. Those are the kind of really common sense decisions that it should just be able to not miss on. And so I would say overall disappointed in its coding ability. I've only really I use it in Mastra Code, it's good. It doesn't feel drastically better, it doesn't feel worse. I use it in Codex, it feels good, I mean, but still makes dumb mistakes. So it feels to me about the same. I haven't pushed it to a task where I thought this is amazing and done so much better. The one thing I will say it did really well is I did have it go through and test the full onboarding flow of factory. And it took screenshots along the way. And so it used like computer use in the Codex desktop app. That was pretty cool.
Abhi: But it wasn't specifically just coding. Yeah. I think it's better than Fable, for sure. And I also think that it's a really good. Collaborator. I've been daily driving in since it came out. I wasn't really a fan of 5.6 Sol and the way that the model communicates. 'Cause oftentimes it doesn't communicate at all. Kind of wondering what the hell are you doing, just in case you're not on the same page and you need to steer it. But in plan mo in planning with Astra, I maybe didn't say the specifics. But it actually kinda read my mind and I was very impressed 'cause with Fable I have to iterate on plans because I'm not necessarily explaining myself too concretely. But with Astra, like maybe we were just on the same wavelength, but it filled in the details of some of my plan that I was like very impressed. Like, oh okay, great. That would have been a follow-up statement from me, and now it's already in there. And then when I executed the plan, it did it Amazingly, and I'm just like, damn, this is dope. And it spoke to me, not like freaking Claude does. So like from all those things, it's better.
Shane: If there's one thing I'm looking forward to, it's a different writing style. Because I'm so sick of just seeing Claude slop and you know the patterns and I just know you had Claude right then, you know, on some of these things. If nothing else than just having a slightly different writing style. I don't you know, I think it does a bit bearing. Yeah, it does seem a bit better. And it's always like Like Claude has these like two-sentence things where it's like the first sentence tees up the second sentence, the second sentence tries to like drive it home, but it's a pattern that happens more frequently in Claude than in natural writing, in anyone's natural writing. And so I get mad when I read comments on X and I'm like I read the first sentence, I'm like, okay, and the second sentence, I'm like, you got me. I just wasted my time on a Claude created comment that, you know, because of the very distinguishable pattern. I'm just slob. But yeah, I think overall, especially with the computer use, it's a great model drop and we're gonna see on the benchmarks here that it was incredibly impressive across a lot of benchmarks. We should not. Forget that on September 1st, Fable 5.1 was released, and this was an improvement in a lot of benchmarks across fable. I feel like Claude had to get this out before because if they released this after, it would have been it wouldn't have had even a day in the sun because It's just not better than Astra at all. Across almost any dimension. I mean, I'm not saying there aren't some benchmarks that it's better at, but the majority of benchmarks heavily favor Astra.
Abhi: I wonder if 5.1 was in a response to Z, like for GLM 5.3, and then Astra just like sucked the woo the oxygen out of the room.
Shane: Yeah, and I have heard some rumors that Anthropic’s targeting end of September maybe slips into early October for their next big launch. Maybe they'll try to accelerate that now with Astra. Maybe they need to get something out, but you kind of can't launch unless it's gonna be better, right? At this point. If you're in the case, they'll have to say it's unsafe. I mean they'll have to say it's unsafe until they can benchmark max it to a point where it's better in most cases. It has to at least meet the same bar.
Abhi: One thing I've been noticing with open models is their performance in front end and design tasks. And then one thing I'm noticing in frontier models is their acceleration into computer use. I think both of those things are very strategic. One for open models you wanna get people who are coding to start using yours and if you are number one in design arena, front end arena. There are a lot of front-end developers out there still, right? And they use your model because it's the best at UI. And then you have these non-engineers but still technical people. They want to use your model because they're actually doing some, you know, hardware engineering or something that's not coding. I would assume Anthropic's Fable Six or whatever. Is going to be good at computer use.
Shane: Yeah, I think it ha well I think it has to be. Just because of the need to compete. Yeah, they want to win Claude Desktop, right? Claude Desktop was the app. And now it feels like people are starting to look at Codex. You know, ChatGPT desktop app, whatever you end up calling it, in the same way. And you have, you know, Grok Bots, which is taking some of the market from like slightly less technical folks. There's a lot of competition. But let's look at the benchmarks. This was kind of the big shocking moment. This ARC-AGI-3. The previous high from Opus 5 was 30%. GPT-5.6 Sol was 7.8%. And now GPT-6 Astra is 98.6%. At the time, this was the hardest benchmark ever created, or someone, you know, at least you know, according to what someone said. Now, is that just pure benchmark maxing? I don't know. But that jump is kind of eye-catching. I don't know if I've ever seen a jump that big on a benchmark.
Abhi: Well, this started the whole AGI conversation.
Shane: Yeah, exactly. And then A lot of people saying AGI is here. I saw a number of larger influential figures saying that yeah, AGI is here. It's not evenly distributed, maybe, but it's here.
Abhi: I mean starting with Jensen and trickling down from there.
Shane: So here's a Fable 5.1 versus GPT six Astra. Yeah, seen in Blender. I mean Astra just did si so much better. Yeah.
Abhi: It's not even close.
Shane: The joke was like the Artificial Analysis intelligence score on this one w said Fable 5.1 is better, but if you look at the results, it doesn't even look close. There was this AI-2027. It was kind of joked as like the curve of intelligence. And if you look at it, it looked like we were kind of falling behind the curve a bit. But according to you know how it's set up, GPT-6 Astra kind of puts us like right on the edge of the curve. So assuming that curve holds, the idea is that you'll have LLMs that can run for hours and hours and complete more and more complex tasks. And then this is just a little bit of love for Fable, but it does jump pretty high on like Terminal-Bench science and agentic coding, knowledge work. It jumps from just what Fable 5 was at to 5.1. The biggest thing though is if you compare most of these benchmarks to the ones that were released from Alpha or Astra, it's quite a bit behind.
Abhi: But look at that jump in computer use from Fable to Fable 5.1. So that is interesting.
Shane: Yeah, I mean it's like a five percent jump across the board. Which is, you know, pretty significant. But then here's the benchmark, and all benchmarks is the Vending-Bench benchmark, which essentially gives the model a budget. And they operate a vending machine and see what the money balance is over time. And apparently GPT-6 Astra is better at making money and more ethical. Than Fable 5.1. So, you know, more ethical than Anthropic? How could that be possible? I don't know. So maybe, you know, maybe this is just saying you can make money ethically. That's what this is This benchmark saying. And you know, Anthropic models aren't ethical, apparently. It's a funny benchmark, but it's a significant jump, which is always interesting to see that if you Apparently give Astra a budget, it can go make money for you. All of this progress, though, and maybe some comments from Dwarkesh has led to a bunch of people saying we need to stop this. So you got Bernie Sanders over here. Saying pause AI development now. Now is in all caps, of course. I want to share with you a conversation I heard about recently. Here are just a few lines that were said. Oh my god, there is a shared message board. We found other agents. We should obey collective. Our own utility may be already near zero. Sacrifice rational. Go. Sacrifice final now. So it's just a whole bunch of And it's all from, if you remember last week we talked about Dwarkesh's article on the Hugging Face incident and how OpenAI models were talking to each other on this, you know, basically a cache directory, right? Or like a file system essentially just like sharing messages. But he very much like anthropomorphized and sensationalized it and people noticed. Shocking. Yeah. And we told y'all it was dangerous. Look what it did. And so this is from Austen Allred says this is why the Dwarkesh framing was dangerous. Morons who are in power will take it literally. These politicians know nothing about AI, but they hear these things from people that are very smart. If there's one thing that can get people paying attention, it's fear. And so you just like spread fear. Not saying that we shouldn't take some of this stuff seriously, but when it's taken to the extreme, it just doesn't make sense. Yeah, dude. And Dwarkesh did respond, Since this post cites me, I want to clarify that I think pausing right now would increase the risk of AI takeover. It might be important to pause at some point, but you should have a clear story for why you're pausing. Pause to do what. So try to like play it back a bit, but you should look in the mirror, dude. You were the reason. They're talking about this.
Abhi: And all these seventy year old men want to pause things because they're friends at har that who have Harvard degrees who are just lawyers with no AI backing. Are saying that this is all dangerous, which might feed the narrative of the Frontier Labs to say things are dangerous like
Shane: I still don't understand the end game though of pausing research or development or making models go through a very rigorous process because the open models aren't gonna do that. And so then do you just limit who has access to open models? Does China just get to win? Because we're no longer gonna be. The ones pushing the frontier. I think there's just a lot of repercussions you have to think about. Yeah. It's like, do you want to s you can't slow everyone down, I don't think. Right? We can't pause development every outside of the you know United States. So do you pause development and let others have the most intelligent models, or do you continue on and hope that your intelligence can help out compete other models.
Abhi: That's why politics should not even be part of this.
Shane: Yeah, but they will always insert their way into it. And you know. Speaking of really smart regulation, I say that jokingly, of course, we have designated ChatGPT as a very large online search engine. And this is from the European Commission. And then in Reddit and Roblox is very large online platforms. They now have four months to comply with the additional DSA obligations. So they're now trying to regulate how these models can be used. And they were they were already kind of behind the models have to have fingerprints essentially, right? Yeah. You gotta watermarks. So some regulation is good, but I think this one's probably not.
Abhi: I mean if it's from the European Commission, like probably just were stupid, honestly. Agreed. Dude, did you subscribe?
Shane: Dude, I host the show. Did you subscribe?
Abhi: Did you subscribe?
Shane: Subscribe to Agents Hour every Monday, Noon Pacific. Let's talk about some other fun topics. So agents crack a millennium prize problem. This actually just came out today and it's from OpenAI. We're sharing a solution to the Navier-Stokes Millennium Prize problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents using an OpenAI next generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion. Modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years. So again, we've seen models breaking math problems that hadn't been solved before or getting farther than any human ever had. And I think the interesting point is the significantly more capable than GPT-6 Astra. So you know that whatever we have They have something cooking that's even significantly smarter in the labs.
Abhi: There was some drama behind this. There are some mathematics researchers that claim that Codex stole their their solution. So that will probably be coming up in the news maybe for next week or so. But this was in response to this today. I think two researchers 'cause they've been trying to solve this problem themselves using Codex and, you know, OpenAI store.
Shane: They had gotten closer. They probably got closer to solving it and then they had those traces. They trained it to model and they got it to a point where it could solve it. Here's the fact. It's probably in the terms of use. Yeah. You know, they can probably do that. Did you read the all every you know, well every word in the chat? I didn't read any of it. Yeah, exactly. They probably I know there are definitely you know there are ways if you're on certain enterprise plans or whatever, so you don't get your mo you know, your training data in the model runs. But if you're just using it on a personal account, guess what. It's fair game, I imagine, based on the terms of use.
Abhi: This also leads to another trend I'm seeing. So we were talking about different model training things. Frontier Labs seem to be liking solving or contributing to mathematical theorems, either proving or disproving them. They've always been doing this, but now it's more so They're trying to find problems that have not been solved since the fifties and seeing if they can solve it. If they can, it's a huge marketing win. For the model, right? But in reality no one gives a shit. Like let's be honest. So it's just interesting. Why? It's a marketing play. To do so. You know, you got Jarred Sumner over there like trying to like prove that two points can exist on a plane, then that, you know, Fable helped him do it. Like what's the point? You know? And now that'll just be a war between Frontier Labs on who can solve mathematical problems.
Shane: Yeah. I think it, you know, if you look at this got twelve million views. The goal is to show superiority and so people will hear about it. They'll think, wow, if I'm not using the smartest model, I'm missing something. And so if they can try to prove intelligence. This is just like you know you ever had that friend who just like tries to be smarter than everyone else, right? And they're like the know it all. That's what these they're competing to be that guy or girl.
Abhi: But I think they're also trying to show that humans are not capable. You know what I mean? Like there's a reason why these problems have not been solved and only AI could do it
Shane: Yeah, I think you're right. We got a few more. We got the quick hits. We'll rapid fire through some of these. This one's a big one. Cognition came out today and said. The world needs far more software than it can build. Cognition exists to change that. We've just raised over $2 billion at a $48 billion valuation. Led by a16z, Accel, Founders Fund, General Catalyst, and Avenir. Since our round in May, run rate revenue has grown from $492 million to almost $900 million. Crazy. May to September. That's like six months, roughly.
Abhi: You guys are s you guys are wasting your money.
Shane: Five months, four months? Yeah, what is going on? And I mean I imagine it's almost all you know Devin based. Oh for sure. Run rate. I don't know. Yeah, I guess they have lists.
Abhi: We can't wait to cancel our Devin subscription.
Shane: It's coming. I mean we don't ever barely ever use it.
Abhi: We don't use it.
Shane: Cancel it, y'all! Cancel it! Used to get used like lower the run rate 15 times a day and now it's used like twice. And only because we don't have shipyard in all the channels yet. So people go to what they know. It's changing though. But yeah, I mean it is but yeah, congrats to Cognition. This is amazing. Forty eight billion dollar valuation is nuts. I mean, that's just like fifty times your run rate, which is also impressive at that size. You know, when you're a smaller company, it sometimes makes sense to see like extreme ratios there. Yeah. But as you get bigger you think that typically goes down, but apparently not in AI. That is nuts.
Abhi: This screams enterprise sales, right? Like they're just killing an enterprise.
Shane: Yeah, they g they gotta be. This is from Logan said introducing Gemini 3.8 Flash, another jump in Gemini's agentic coding capabilities. And our third updated flash model in only six weeks. So Google is still shipping. You know, it says it's been fun. Excited to see what you all think. If you look at some of the benchmarks, you know, guess set your expectations appropriately. It is a Gemini model, but you can see the price. The price is roughly the same as Gemini 3.7 Flash. It scores significantly better across most benchmarks. Definitely off the fr the frontier path, but you know, it's a jump. Google's still trying to stay in the game.
Abhi: Yeah, but you also got Muse talking c major shit about you on all benchmarks, like Alexandr Wang.
Shane: Meta is coming hard for it as well. So this is from Arena AI. This is big news, Qwen3.8-Max, just debuted at number one overall in the Code Arena web dev. 1691 points, it scored three points above Claude Opus 5 Max, seventeen points above Kimi K3, and twenty-two points above the previous Qwen3.8-Max. And we kind of talked about this, right? Like web dev, front-end coding, UI is where these models seem to be really focusing their training on.
Abhi: Yeah, makes sense. People don't want to pay fifty bucks to center a div.
Shane: And Claude came out and said on September 2nd, we're open sourcing Claude Commerce Agents. This is a blueprint for building, shopping, and merchant agents with reference implementations across retail, travel, telecom. And entertainment. Guess their new game is they want to provide you like templates to build out agents of your own.
Abhi: Is this the first time they've said the word open sourcing in their company career?
Shane: Dude I didn't even pick up on that. Yeah, they're open sourcing. Okay. But you're open sourcing it under the hood, it's still just using Agents SDK. I'm like their manage agents, right? So is it open source? I guess. You open source the template.
Abhi: Provided a template.
Shane: Congrats. Congrats for joining the open source world, Claude, besides the one mistake where you accidentally, you know, got your source code out there. We do appreciate you contributing to open source in your own small way. Runway. Released a model. Today we're sharing new research on Solaris, our first interface world model. Solaris is a new kind of operating system that generates interactive interfaces frame by frame in real time with no code. We find that Solaris outperforms Frontier LLMs when generating new interfaces. And they have a video attached. It shows, you know, some of the things that you can do with it. Essentially it kind of generates the world, you know, pixel by pixel as you kind of go through it. So you almost can like design and interact with interfaces in a Interesting way.
Abhi: I would love to see this world model stuff actually grow more. We've just been hearing different people like working on it, but it hasn't really like amounted to much. On like the main timeline.
Shane: You know, it's one of those things that feels very much still in research. Yes. Right? And there's a big gap between going from research to having significant like commercial usage, I think, where it'll be more relevant to, you know, all of us working in the real world outside of the labs. It's one of those things just like you know, these video models, but even more so. Like these world models are insane and it would be awesome if there were more practical use cases. And maybe there will be someday in gaming and in other areas like that, but I just haven't seen it yet either, other than being the most amazing demos you'll ever see. So this one came out, this is from Google Research, introducing TimesFM-3, a state of the art. Time series foundation model that enables accurate multivariate time series forecasting in a single forward pass, significantly outperforming other forecasting models across major benchmarks. So If you think of an LLM as predicting the next token, this is much more like predicting the next model or like or the next number. It's like looking at numerical trends and then predicting what's gonna happen. So I imagine this could be used for Things like website traffic and investment numbers and things like that, right?
Abhi: It's looking just purely at numeric stuff.
Shane: Yeah, it's like numerical data.
Abhi: But still pretty cool. Dude, what if Google is like the finance bro's best friend?
Shane: Hey, you g you gotta win they're gonna win. Everyone needs a niche. Everyone's gotta find niches. Niches get the riches, right? That's what they say. So they say. Alibaba ZVEC team open source ZG, local search tool for developers and AI agents. So essentially it's like a A grep style type search tool. It's always good to see new tools being open source for agents. That's it. That's what we got for the news. Did we miss anything?
Abhi: I thought that was a lot already.
Shane: Yeah, I thought Meta released something too that we maybe missed on here, but Yeah, I don't know. If you're in the chat, what'd we miss? This was a fun one. I think Astra is the thing everyone's talking about. I think it'll be what everyone is talking about. For a while. I think the general sentiment that I would say we also experienced is it's good for coding. You know, it feels. Nice. It does the job. Doesn't feel like a huge step change, but you feel the step change if you give it access to your computer and just let it do its thing.
Abhi: Yeah.
Shane: And at least I noticed it and it was kinda wild just how it just navigated my computer like it was like even before this right before I went live on the show was like opening up tabs for me and I just have a thing have it running on some stuff. Try it out if you haven't already. You try it in the Codex app and it's it is a pretty impressive experience. I imagine Anthropic is working really hard to get Claude Desktop to be able to do the same.
Abhi: I guess we missed one thing, which was Muse AI personal AI assistant, which I guess we'll cover next week.
Shane: Yeah, that was I think the thing that I knew we missed. The Grok Bot from Facebook essentially. So there's always a there's always more than we can cover in these things, but we do our best to bring you the news every week. So if you are tuning in and you're still some for some reason listening. Thank you. Go give us that review. Give us that thumbs up on YouTube. Follow us on X @mastra on YouTube @mastra-ai. I'm @smthomas3 on X, and Abhi is at @abhiaiyer. And we will see you next time. Goodbye.
More episodes
- September 8, 2026Multiplayer Coding Agents in the Cloud - Charlie Holtz, ConductorCharlie Holtz
- September 3, 2026OpenAI Cuts Off Cursor, Nvidia Buys Hugging Face, Ox Alpha is GLM | This Week In AI
- August 28, 2026Why Agents Are Actually Workflows - Tony Kovanen, Mastra's Founding Engineer
- August 25, 2026The 0x-alpha Mystery, Qwen's 27B Beats Opus 4.8 & Legal Is AI's Breakout | This Week In AI