Back to all episodes

The 0x-alpha Mystery, Qwen's 27B Beats Opus 4.8 & Legal Is AI's Breakout | This Week In AI

August 25, 2026

A shorter but juicy one. Shane and Abhi open on the debate everyone's having — do you want to use your agent to reach into someone else's tool, or do you want that tool to build the best agent and meet you where you already are? Then the week's biggest mystery: a stealth model called 0x-alpha that appeared on OpenRouter beating Fable and Sol on coding, with a wild one-hundred-trillion-tokens-a-day capacity that had everyone convinced it was a secret Google model — until someone made the server name itself, and it turned out to be Z.ai. From there: Qwen breaks the size curve with a 27B model that outranks Opus 4.8 on the agentic index and runs on your own hardware, open-weight tokens cross 62% on Vercel's AI gateway, and Salesforce ships Slack Code so humans and agents can work in the same channel. Plus the case that legal is AI's breakout use case — Harvey trained its own model, Tenet, on a Kimi K3 base, and legal is the fastest-growing Codex adopter, up 108x since February. In the quick hits: agents now burn 5x the tokens humans do, and Tobi Lütke vibe-codes a Git-at-scale binary over a weekend.

Episode Transcript

00:00

Abhi: This like stealth release stuff bewilders the community, takes them by storm. Will the frontier ever give it out in stealth?

00:06

Shane: I feel like they don't have to. Their marketing is it got banned by the government. That's their marketing stunt. Welcome to Agents Hour. Today we have A pretty short episode considering some of the past weeks we've had, but there are some big juicy topics that we do want to dive into. Let's start with this post this is from someone named Ali and she says I really don't want to use your agent I want to use my agent to use your thing and this got a ton of Engagement, a lot of people, you know, were commenting, resonating with it. Some disagreed, a lot of people agreed. But the question is in this. New age, do you just want to use your agent and connect to someone el else's tool, maybe through MCP or whatever? Or do you want that tool to just build the agent themselves and you use that.

01:06

Abhi: I agree with this statement for things where like the boundary of the product is also on my end, right? Like if it's a coding thing or infrastructure or whatever, but here is where I disagree. I would rather have an agent in Google Slides than have my agent talk to Google Slides. You know what I mean? Where these products are encapsulated. Like Descript. Maybe I wanna be have my own agent, but no, I'd rather just be in Descript just making moves in this thing. But where I think is comes into play if I were to make it more general. If I also have industry experience in the product, I want my agent in there. If I have no business doing that in the first place, like making music or anything, I want to use their agent. That's where I kind of draw the line.

01:50

Shane: Yeah, so you're saying because they would conceivably be better at building an agent in the domain that you're not experienced in. And you wouldn't be able to guide the agent to the result. It's probably better if they just build it, they control the context, rather than you try to control the context. Correct. I think both. Use cases are valid and I will ultimately think it depends, very similar to you. For instance, you know, for example is PostHog. Like I don't use PostHog every day. I use it quite a bit, but not every day. And I've noticed that the agent in PostHog’s actually pretty good. It does what I need it to, and now I don't have to connect the PostHog MCP that probably has fifty tools to my agent. I do worry that over time your agent might get worse and you're not validating it or you don't know it because you're just slamming a bunch of MCPs in there and it which Adds to the context and every what the average person doesn't understand, if I have a hundred MCPs connected to my agent, that's not very good, right? Like that's going to impact the results. It's gonna slow down the agent, there's more tokens, all that stuff you know, factors in and the average person doesn't really understand that. So I think there is this risk of, at least today, and maybe this changes over time, but if you just connect your agent to every tool possible. You now your agent has to try to do every single thing. And I think that you might get worse results. And you might not know why your agent seemingly got dumber over time. You might think it's the model. You might think it's you know the way you're using it, but maybe it's just your tool context is getting polluted because you're connecting so many damn MCP servers to it.

03:16

Abhi: Yeah.

03:17

Shane: So that's why I think the specialized agents are useful because you don't have to worry about that. You only go in there some of the time. I don't use PostHog every day. If I did, I would much rather connect it to my agent and just use the MCP. Another thing is like Notion. Notion's agent. I think it kind of sucks. And I use Notion almost every day. So yeah, I'll just use the MCP and it works. Right. I understand what I want it to do. I can guide it. But I'm taking on that, you know, all those extra tools in my agent, right? If especially if I just like connect it to the my agent and I don't have different, you know, settings or profiles or whatever or projects that have like di different configurations. So I think in general. I agree with Ali for a lot of the tools, but not for everything. Like they're specialized tools, just let them build the agent. They're the experts. Make the agent really damn good. Yeah, maybe someday my agent can talk to your agent and not have to care about the context, just like passing messages across. But for now, I'll use the agent when I need to.

04:12

Abhi: If your agent is built in Mastra, I'll use it

04:15

Shane: There you go. There you have it. Alright, we gotta talk about I don't know if it's 0x-alpha. I've been saying 0x-alpha, but we gotta talk about 0x-alpha. Because it's a bit of a mystery and it kind of just hit the timeline later last week around what is this model that just seemingly popped up within OpenRouter. So this is from MTS Live. It says situation explained. A stealth model on OpenRouter is beating Fable and Sol on coding and nobody knows who made it. 0x-alpha has a 1 million token context window, text image, and video input free for a week week with capacity for a hundred trillion tokens a day. And I think that's what blew people's minds because they didn't understand where could a hundred trillion tokens a day come from. That compute doesn't just exist lying around. There's only a few companies that in theory could have that capacity.

05:01

Abhi: I love the conspiracy theories that have come from this.

05:03

Shane: Yeah, there have been a lot, and we'll talk about some of them. Brandon Carl says whoever is offering 0x-alpha had the ability to offer the entire token capacity of Google immediately, 100 trillion tokens a day. Which led to a lot of speculation that Was this a Google model? The Gemini team was posting some kind of obscure things, making people think maybe this is a Gemini model. Could this be the new Google? How good is this model? People were starting to run benchmarks on it. This post says WTF. Ben ran this mystery model through 10 DeepSWE tasks and has scored over 80%, versus 65% for Fable and 52% for GPT-5.6 Sol. This is insane. Probably a Chinese company. Either a new GLM or Kimi model, I reckon. I did see some posts that said they ran more of the tasks. It didn't score this well on everything. I think this was a bit cherry picked. I think overall it seems like th what I've seen it's actually performs. Not as good as Fable in 5-6, so maybe it is Gemini. I don't know. But it's not. At least not according to this post which says breaking the stealth model 0x-alpha is Z.ai so it's a glm everyone else is guessing it from tokenizer vibes and emoji rates i made the server say its own name out loud. Sent it one malformed request and it threw a Java stack trace, naming its own internal. What do you think?

06:19

Abhi: Is it GLM? It'd be good for the narrative that it is, but I'm really hoping it is Google. I think it might be Z A I and this is all like some speculation, right?

06:29

Shane: But it seems that it's a smaller model, but it's doing it does really well on at least the benchmarks that people have run. Not fable level though. But their theory is that this is almost like a GLM flash type model or a smaller model that maybe is on par or better than the last like GLM 5.2, but is a smaller model and so that is how. Z.ai can serve a hundred trillion tokens a day is that it's just significantly more efficient because it's so much smaller.

06:58

Abhi: There are other indicators that would lead for this to be Z. So like Open Code offered this through Open Code Go for free or whatever. And they were tracking the token usage. People are using it. If it was Google, anybody would have to sign a bunch of NDAs. We have done this. I don't know if I'm not allowed to say that I've done it, but okay. But you have to sign NDAs to use future models. So I don't think it's Google for that reason alone. They're not gonna give this away in stealth that other people can use without tons of bureaucracy. So it's gotta be an open model. That's what my guess is. And you know, if this evidence is damning enough, then it probably is Z.

07:37

Shane: And I think we will eventually find out. I think the hype has died down from the model a bit because I think people realize it wasn't as good as what they originally thought. So I don't think it's frontier, it's not a new frontier model. But if it is, you know, a GLM model and it's much smaller, you know, if you think back to, you know, the Qwen model that came out, which we will talk about and how that's, you know, like a small, basically on device, you know, you can run it on your own computer model and it benchmarks. Very well. Well then maybe GLM has done the same type of thing. Where it's a smaller model. We don't know the size yet, but maybe that can compete close to the frontier without being frontier sized.

08:14

Abhi: This like stealth release stuff, it like bewilders the community, takes them by storm. So I'm curious how many stealth models we can go through before it like loses its luster. Or is it always gonna be a really hype way to debut a model. To give it out in stealth first.

08:29

Shane: Yeah, I mean I think there's just a marketing aspect to it for sure.

08:32

Abhi: Will the frontier ever give it out in stealth? I wonder.

08:36

Shane: I feel like they don't have to. Their marketing is it got banned by the government. That's their marketing stunt. Qwen breaks the size curve. We talked about this a little bit last week, but if you look at the artificial analysis agentic index, Qwen3.8-27B is actually ahead of models like Opus 4.8, ahead of 5.6 Luna. So it is yeah, ahead of GLM 5.2. It's ahead of these much larger models, right? And this is only a twenty-seven billion parameter model.

09:09

Abhi: Yeah.

09:09

Shane: If you told someone eighteen months ago, right? Not even you told someone six months ago that they could have, you know, an Opus level model on their device, like on their computer running. You know, obviously still needs to be decent sized, right? You still gotta have some hardware to run that. But I think people would freak out. Running an opus level model on consumer hardware is pretty insane. It makes me wonder like how many tasks do we even need to pay for tokens anymore? Will teams just have like their own reasonably good model on their own machine? That they run most tasks to and then the hard tasks get sent to the frontiers? I don't know. Dude, did you subscribe? Dude, I host the show. Did you subscribe? Did you subscribe? Subscribe to Agents Hour every Monday, noon Pacific. Let's talk about open weight models a little bit more. So this is from Guillermo from Vercel. Says today is a record day for open weight share of tokens on Vercel AI Gateway and OpenRouter’s posted. Similar things so this is not just Vercel this is across a lot of the popular AI LLM routers and so August 22nd 62% of tokens were open weight models and in June 24th. So about two months ago it was only 28. 4%. That's a pretty wild increase. I am curious on is it because there's two ways you can kind of read this. Maybe, you know, a lot of people moved from frontier models to open models. Or maybe there's just a lot more tokens being sent through and a higher proportion of those tokens are now becoming open models. So they don't really share like total token growth count or anything. So you can't really determine our Closed models in you know, are they total are the did the total count go down? I doubt it, but it probably just didn't grow as fast as the open models. And I think if you're using Frontier tokens, y a lot of people, the majority of people are going right through the providers, right?

10:57

Abhi: Not through a router. I mean I would suspect this too. Like it's gonna be more open 'cause you know, you are not paying max coding plans on these gateways. And there are a lot more open models to choose for price.

11:12

Shane: We gotta talk about Slack code. This is from Mark Benioff on August 19th. It says, don't code alone. Slack code is live. Humans and agents. Same channel, same work. Launching today with agents from Anthropic, GitHub, Cognition, Vercel. This is real multiplayer coding. See it at Dreamforce. What do you think of this?

11:33

Abhi: I mean if their website was up that day, I would have been more happy about it, but it was down for like the whole day.

11:38

Shane: Yeah. I mean they said see it live and then people were saying it wasn't quite live. But if you watch the video, there's basically a code panel that pops up in a conversation, right? You're basically prompting in Slack and you're just having it write code. You know, we have mentioned that coding, you know, from your mobile device is the dream and Slack has a pretty good mobile app, so now you can in theory code from anywhere. Within Slack. I'm not convinced that Slack's gonna win this on its own. I feel like if you ever use Devin, Devin Slack agent is pretty good. I know Linear's investing heavily in their agent as well within Slack. I think obviously Slack is ki trying to compete with all the apps that integrate in with it and trying to just do it themselves. They wanna maybe Slack wants to own a piece of the coding pie in some ways. But we'll see how many people actually use the Slack version or there's always those things where there is a version in the app itself, but people use third parties because third party ones just end up being better. And so I think we'll probably see some of that as well.

12:35

Abhi: Could you imagine a world where Slack bans third-party agent apps just so you can use Slack code?

12:42

Shane: I feel like they can't. Right. It's like Salesforce is kind of built on this ecosystem of you can build sales I mean they have proprietary language, of course, but it's all built on like this marketplace and all these different apps. It's like so complicated. They have all these different companies that just provide Salesforce type services. I feel like they wouldn't want to go back and change that specifically for Slack. I think they'd want Slack to be open, Slack to be the communication tool, but maybe not. Maybe they're looking at all the you know the revenue from all these coding agent companies and saying maybe we can grab a piece of that.

13:12

Abhi: Is this G A yet? Can we use this? I don't see it in our slide. Yeah, well we didn't go to Dreamforce this year. Oh that's why.

13:29

Shane: Recently and especially the last week around just legal as a breakout use case for AI. So Harvey has introduced Tenet which is the first model post-trained for legal. So Harvey's actually not just using frontier models, they're actually training their own or post-training their own. So Tenet is a Kimi K3 base that we post-trained with Fireworks on a corpus of publicly available legal data, synthetic data, and human expert data simulating long horizon legal work. This is pretty wild. That means. There are actually companies that started as just you could call them like a model wrapper, right?

14:05

Abhi: Yeah. They were just an application on top of a model.

14:07

Shane: And now I would say cursor was the same way, right? Like cursor didn't start as building its own LLMs. It just was kind of an agent that used whatever model you gave it. I think Harvey was kind of similar, I just it picked a model under the hood or whatever. But now I think they're seeing one. They don't want to compete with Claude. They know if they send all their things, all their data through Claude or OpenAI, that they might eventually be competing with them. So they kinda wanna own their own destiny. And now with some of these open weight models, it's becoming easier and more possible for teams to do that

14:38

Abhi: I was on a panel with a principal engineer from Harvey and we were talking about the eval loop and Harvey has so much data now. On different cases and different users. They have so much. They have huge eval suites. They have this whole like discipline in making sure their evals are good because they're what we're what we call a real world agent. They impact the real world through I mean obviously law I guess you know does impact people. And I think anyone who is a real world agent that has a lot of data, eval suites, and actually takes observability seriously. Will be doing the same thing.

15:16

Shane: Yep. And I think there's different methods for doing it, right? Are you just doing like traditional like fine-tuning, post-training, reinforcement learning? I mean there are different practices for how you could pull this off, but I do think that it is getting easier and especially if you have a real world use case where you're collecting data first, eventually you can use that data and train potentially a smaller model or a cheaper model to do it as well, or maybe, you know, again, of close to frontier model that can then yeah perform better and you own the outputs, right? You own that

15:47

Abhi: No. Yeah. Kimi K3 base is not a bad model to start from to then train with your company, let's say your company swag in there. And then you can now make more money, right? You obviously you put some money into training, but now your inference costs are way lower than they what they used to be if you're using a fable or something. I would say that's kind of a threat to frontier usage in companies with data. Obviously you have to have a practice in your company to harness this data and do things with it. You probably have to have an AI squad like Harvey does. There are a lot of costs that go into this. But it's not about what happens now, it's about the horizon of once you have this, your margin will get better over time.

16:29

Shane: And if you think about models getting smaller, right? We talked about. Qwen 27B. The smaller the model, the easier it is to train it, right? So the price of training your own model is gonna continue to be compressed downwards. So I think over time you're gonna see even more of this. And Harvey's has been, you know, kind of on the leading edge of a lot of a lot of things, right? As being a prime example of a successful. Agent out in the wild in a specific niche or specific vertical. But continuing to talk a little bit about law and the legal. Use case. This is a post from a16z.AI power users are showing up outside of tech, the fastest growing Codex adopters since February. Legal is 108x. So basically they're just figuring out enterprise job title when they sign up for Codex. And in the legal use case, it's 108x since February. Sales is 41x. Recruiting is forty one X, marketing twenty-six X, healthcare twenty-four X, and then coding, which you know is what we're all probably using it for, is still impressive. It's five X, but that's tiny compared to 108x.

17:31

Abhi: Yeah. Dude, I want to see all of these even go even higher. Like healthcare needs to be in the hundred X. Actually don't really care about sales, to be honest. But law, healthcare, real world shit, they need to be in the thousand X. That's where we should put our energy into.

17:46

Shane: Yeah. I mean I think the reason coding is lower, the reason sales is lower is because those are the com those have been. Kind of really common use cases.

17:53

Abhi: Yeah.

17:54

Shane: Sales agents have been around for a while now. A lot of people they've already been using them, so they're no the growth rate isn't as high. But it still is wild to think that even in coding. If it's been it was five X. It means that most people aren't using these tools on a daily basis. There's still a lot of room for it to grow. We have not hit the pe by any means and I think it's going to continue to sp expand outside of just engineers and developers into other places. So let's cover some additional quick hits. This is another post from a16z. Again, some play on the show today. But it says humans are the minority user of AI. Agents burn nearly five times the tokens that people do. Up 14x since February. So just like, you know, we've said in this show, you know, you're not writing docs for humans and then agents are consuming docs more frequently than humans. Agents are consuming websites. I think Cloudflare announced, right, more than humans. Agents are now. Using tokens more than humans.

18:53

Abhi: Wasn't that always the case though? Like don't we use AI through our agents?

18:56

Shane: Yeah, but I mean I mean I think a lot of cases it was like single-turn call and response for a lot for the longest time with ChatGPT. Like and then you had some tools, right? But it still was just like user message. Now you get a response. And I think what this is probably charting is that an agent is deciding to like use tokens on its own, right? It's like calling to the LLMs on its own. So again, don't know ex you know, you can read the post and see exactly how it's measured, but I think it's because the agents are running for longer, they're doing more autonomously, they're calling sub-agents, right, that are doing work on the humans' behalf. Humans are becoming further disconnected from the initial query to the result that you get, right? Where in the past it was you didn't want the agent to run more than twenty seconds because you knew if it did, it probably was going to go way off track. Where now people are trusting it for longer horizon tasks. This post came out on August 18th. This is from Cursor we talked about last week. Cursor Origin, but they wrote this really detailed blog post called Git at Any Scale. And it's really about why it's so hard to scale. Git, the way it was built, and then their solution for how they have basically said they solved Git's scaling problem. They can infinitely scale it with this approach.

20:10

Abhi: Did you read through this? Yeah, it's very interesting.

20:13

Shane: It definitely is a very detailed post and it's very well done. So if you if you're curious on like scaling Git and it's not just about it's just scaling in general, I think. It's a really in instructive post.

20:25

Abhi: Yeah, the main thing after reading it, the main thing, main takeaway is it's kind of the what we said on the show where it's not entirely GitHub's fault that it's had these issues. Mainly because Their volume is super high and their architecture started in two thousand eight. And so to evolve over the years to then suddenly change your architecture to supply this mass scale that they've never seen before, which is an exponential growth. It just doesn't make sense for them to take all the blame. Obviously the market is huge factors. But it also makes this case where maybe you do need to have Something like Origin or something built with current technology that's meant for this scale. Because it breaks down like every architecture piece of GitHub. Or Git in general and then what you need to do to scale that when it gets loaded. And so yeah, I've I had more empathy for GitHub after reading it.

21:21

Shane: Yeah, and I think if you read it you'll say that Cursor definitely doesn't blame GitHub, but I think this is cursors like write a really detailed post a w around why. Kit is hard and show that you have a very clear understanding and you can handle the scale. And it's honestly a marketing tactic as well, right? It's like showing your expertise. Showing why cursor origin has to exist, showing how you solve the problem so people can't just say, oh, like yeah, it's easy now because you don't have the traffic, but when you get to GitHub scale. But they're trying to prove ahead of time. No, this actually is the architecture that will scale. And you know, unfortunately for Git, they've they kind of have this legacy architecture that's been set up for years now. And so Cursor wants to show itself as really have deeply thought about this problem and solved it in a way that can inspire confidence for people who want to make the shift or who are considering moving off of GitHub.

22:12

Abhi: It's highly not to say that Cursor Origin has proved that they can scale yet, but they think with the architecture that they've done, they will.

22:19

Shane: Yeah, I mean i the theory is there. Now it's does it work in practice? That's always the question.

22:29

Abhi: We'll see.

22:30

Shane: And then this is Tobi. Lutke from Shopify said based on that last post, Git at Any Scale has been one of the most interesting blog posts I've read in a while. It came right when I was frustrated with Shopify's internal Git system as an exercise I've implemented over the weekend as open source. It's a single Rust binary that you can point at At any S3 type object store, it uses WAL and CAS primitives and requires no other data store. It also implements. Bundle URI so large git repos like the Shopify monorepo are very fast download as a chain of static bundles. So there you go. They basically, you know, open source their own kind of version of it. And again, it's spin-o it's vibe coded over a weekend.

23:08

Abhi: I don't know if I'd trust it, but put it in production right now.

23:11

Shane: Yeah. Yeah, I don't think Shopify's running in production on it. But it would be cool if someone did, you know. Take something like this, open source it, so a cursor origin of sorts could exist. And that's the show. You can follow us on X @mastra. You can subscribe on YouTube @mastra-ai. You can follow me on X at @smthomas3. Follow Abhi at @abhiaiyer. We do the show every Monday, usually every Monday at least, right around noon Pacific time. We do the news, we bring on guests, talk some smack around some AI concepts or SF isms occasionally. You know, we have some fun. If you are tuning in, in live. Thank you. We appreciate all the comments, all the love and sometimes the hate that you give us. Any parting words before we close out, Abhi? Don't be a meat proxy. Don't be a meat proxy. And with that, let's get out of here. Peace.