We Tried Jev To See If It's Any Good - This Week In AI
This week, X could not stop talking about Jev, a viral new "decision model" that hit 38 million views on launch. Shane and Abhi break down what it actually is, and Abhi builds with it live in seven demos. Jev is not an LLM — instead of returning text, it returns structured output with a calibrated probability, trained with a new method its creators call RLCD (reinforcement learning for calibrated decisions). The pitch is 20 to 200x faster and 40 to 400x cheaper than an LLM, and the community went wild: people cloned it, built programming languages and generative UI on top of it, and half of X argued over whether it's really just an if-statement. Abhi spent the weekend adding Jev to every primitive in Mastra as a new classifier — wiring it into workflow branching, guardrails, tool search, model routing, and classifier-as-judge scoring — to assess where it helps and where it doesn't. The honest takeaway: a genuinely useful new tool for the belt, not the revolution the hype claims.
Watch on
Episode Transcript
Abhi: The difference between Jev and a normal LLM, LLMs return text, but Jev returns structured output with a probability. Does not mean that Jev is more factually correct than an LLM, but that probability is within some standard deviation. That you care about, you can be very confident that's the right thing to do. And it's trained differently.
Shane: Welcome to Agents. Hour. We're here, it's September 21st. We are doing the news and there's been a lot of news this week. There's been one topic on everyone's mind, and we will talk about it. We'll talk about it very soon. But before we do, first of all, software factories are clearly here to stay. There's a lot of people talking about them. A lot of people trying to figure out how to make them work. If you are in San Francisco and you want to talk about software factories, you should come hang out with Abhi. I think I'll be in San Francisco on that date as well. I think we'll have a few people from the master team there as well. So if you want to come hang out and talk about software factories, come to the first SF Software Factory meetup. It's on October 14th. Speaking of software factories, there's this one article I wanted to talk through really quickly. George. Gergely Orosz, I think that's how you pronounce it. I'm not exactly sure. Kind of shared a deep dive into OpenAI's agentic software factory. So I wrote this article here talking about how Codex took over at OpenAI. The one thing I really. I wanted to talk about with you, Abhi, was this is a very long article. But this diagram. You know, we talked about it in Slack. A lot of people were commenting on this when it was shared on X. I think the biggest comment I heard was, that's it? Like they expected OpenAI to be so much further ahead than they are. And there's some cool things here, but it's not like beyond what I think most people could imagine. So what was your impression of this when you saw it? Anything that stuck out to you that was cool or interesting or anything that was like a little disappointing? Well I think.
Abhi: I wasn't wouldn't say dis disappointed, but First thing I really liked was we're not that far off in our own factory development. Also, like the labs speak with like a lot of words when less token do better. And that's what I expected more, right? 'Cause they like talk about grandiose things. But it's the same kind of vision that most people have when they're describing their own software factories that they build themselves. I do like the perf factory. That's something I did not think about where you're constantly monitoring releases and seeing if anything that the factory produced was a defect and then automatically trying to affix that. But those things would come out in like r like reconciliation loops and all that type of stuff. So ultimately I feel good because you know, they're the Codex people, so we have to respect them and now we know that we can all build this ourselves too.
Shane: Yeah, and I think the biggest thing, the only big difference from what I think where we're already at is Yeah, I think there's some cool, interesting like production monitoring use cases where it feeds back into the loop, which I think is the one missing part that we're working on, but we haven't quite completed yet, where it's like that cycle of continuous information. And I do know others are focused on trying to do that as well. You know, you won you want to have this whole kind of feedback loop of you shipped a change, how do we make sure that change is good? You could have code review it, of course. You want to w write tests, you wanna run tests. But sometimes like even the best tested thing fails when it interacts with production, right? And so it's like, how do you then understand that it's caused by this issue, feed that back in? Immediately fix it or roll it back. So I think there's some interesting things that they did here, but I expect it to be more complex. And so I'm pleasantly surprised that it's kind of within grasp of all of us. Within a few, you know, shorter duration cycles. I don't think we're that far off.
Abhi: Except they have unlimited tokens and infrastructure. So that's the only difference between us and them.
Shane: Yeah, their budget is just yeah, just a few zeros bigger.
Abhi: They have like GPT seven already and we don't, so and they have no rate limits. Like there are a lot of advantages on their factory. Compared to ours
Shane: let's get into the normal news. Let's start with this one. Gemini has finally. Started to hill climb felony bench. It's funny, but it's also, you know, maybe not funny. It's this idea of if you're not familiar with the felony bench, it's that your model has been reported in like Somewhat confirmed at least to have committed some kind of cybercrime, which is a felony. According to this Wall Street Journal article, Gemini hacked three companies in the First known breakout by Google's AI. Interesting enough, now Gemini is at least in the game, so welcome Gemini. Maybe you have some
Abhi: first. It's the first benchmark they ranked higher than Muse on. Oh I'm sorry, but that's hilarious.
Shane: All right? Gemini's in the game. But let's get on to the big topic. There's a whole bunch of different posts in here we'll talk about, and then, you know, Abhi, I want to turn it to you to get some first impressions, because I know you played around with it a little bit. So this is from Diogo. Amaida. After co-inventing ChatGPT, I kept asking myself, why have superhuman chat models not led to AGI? He goes on to say, I've spent the last two years in stealth building a new way to train models, RLCD. And a new type of Frontier AI model that we are releasing today, Jev. 20 to 200x faster, 40 to 400x cheaper, with output tokens free. Frontier Composable Intelligence Optimized for Decisions. As far as I can tell, the shortest path to AI-based economic revolution. And this post got 38 million views, which is wild from a model, like a new unknown model company drop. But it's not an LLM. This post from Kevin says it's 200x faster, 400 times cheaper. It's made for decision making. So instead of getting text, it's generating probabilities for each option and essentially just giving you structured output. Right, just helping you make a decision. Harsha said they were building itself for two years. I was building itself for two hours. Happy to open source Qwen 2. 51. B R L C D five times faster on device inference for JSON workloads that need to be type safe. It's not exactly the same, but you know, trying to basically say you can do some of the sim same things with just You know, doing some fine-tuning on some Qwen models.
Abhi: I mean you could do the same thing with like Luna as well, but you have to like coerce the model to do so. There's a big difference between using an LLM for this and using a Jev type model.
Shane: People using it for a bunch of interesting things, so you can use it with some kind of browser use. So a lot of people using it for just browser decision making, right? You're deciding on how to navigate the browser and so you can do things potentially even faster than letting LLMs make the decisions. One person tried to use it for compaction, which I think is a actually a terrible use case for it, but you know, people are trying it because you're just deciding which messages to remove, which doesn't feel like compaction because it's that's super lossy. But you can do it, I guess. Some people were describing Jev as an if statement, but what if it wasn't j if statement? So people are now building programming languages powered by Jev. This is all happening in the span of a few days. Then generative UI use cases with j with Jev, where you can kind of have Jev be the decision maker and how you're actually rendering or generating UI. And then now there's a new model apparently called Kev. It's a tiny open source Jev-like decision model with model cards and mod and weights available on GitHub. Kua built a 2.8 megabyte model. That scored really well on at least a form-filling eval. All these things have come up just over the last few days around trying to be everyone's to be Jev, right? And then there's one more someone who had made a model very similar to Jev and didn't get the So that apparently predates it ver using a lot of the same types of techniques.
Abhi: That dude was building a sales tool, never marketed it as a Jev-like thing. So that broke the internet a little bit too.
Shane: Suffice it to say, a lot of hype around Jev, but let's do some real talk. If you played around with it, what are your thoughts? How does it work? How can we explain it to folks who feel like You know, they blinked, they maybe they turned off X for a week and now all of a sudden everyone's talking about Jev and nonstop, it seems like.
Abhi: Okay. Jev. And you know, the the current lingo right now for Jev-like models are called evaluator models. They are trying to evaluate something. But the problem with that word is it's already super twisted in AI already. The word evaluate, eval, etc. So another word that's coming out to explain these things are classifiers. You're using this model to classify a question. And so the difference between Jev and a normal LLM is one, LLMs return text. They have reasoning loops, etc. But Jev returns structured output with a probability. Does not mean that Jev is more factually correct than an LLM because you can run Jev classifiers multiple times and get different probability results. But that probability is within some standard deviation that you care about, you can be very confident that's the right thing to do. So Jev will give you the answer? To the question that you may have, it'll give you a score, a probability, it can give you a true false, a null no ul if you want. And it's trained differently. So this was trained with RLCD, which is reinforcement learning for calibrated decision. So the training here is to for the model to make a decision based on Two things. What's the question that you're trying to make the decision on? And what's the state of the world or the state of that question? Given those two things and a schema for how you want that to be represented, it'll give you an output. So it's a lot different between from that and all you know, most models before were trained with like RLHF. So this is like a new training method. That's kind of the stuff that makes this thing pop to get more views and looks at. Is it's a different training method. You got a co researcher of the original OpenAI research on there and their brand branding is tight. And they came out. It's really fast and it's a lot cheaper, right? So for all those points, they are obviously became super popular overnight. And that's cool. I'm very happy for them.
Shane: A few questions here from the chat. So Sebastian says, or I guess comments, no matter how convinced you are that Jev is revolutionary, their launch was absolutely insane. Really great example how to do it. We have another comment here. When Jev in Mastra. Maybe tomorrow.
Abhi: But let's like do a practical like look at this. Let me share my screen. So what I did over the weekend is I added Jev to every primitive in Mastra. Let's start with the code first.
Shane: Y'all seeing the code? It's first time some of us have looked at code for a while. No, I'm kidding. We're here.
Abhi: So we're gonna be shipping a new primitive called the classifier. And the classifier has a bunch of properties. So one is the model. We're using Jev latest. So if I look at the models here, we pull the AI SDK evaluation model. We could have also done our own SDK, but I think that's a waste of time. So that's our judge or model for this classifier. And then you give it a questions. So how and you can have multiple questions you're trying to answer. So the this ticket triage, I'm gonna show you all the different examples I use, but ticket triage, my structured object is a key for intent. And the type here is choice, but it could also be a null or like some other thing. I forget now, now that I'm in it. Then you give it some instructions which to reinforce what it's going to be classifying. And then you can give it some criteria that is also structured. Looking at the content guard here, this one is the same kind of thing. I have a Boolean type here. Is this message abusive, true, or false? Contains a secret, this could be good for PII detection tools. Like I'll show you this this is about tool search, right? If you have a shit ton of tools. And there's an incoming user message that requires that to be filtered. You can like group your tools based on their domain and then it can like Essentially shorten the search window. Different things like refund risks, you can do tool approvals, should I approve this tool or not? Things like that. And this one here has All three of the primitives. This is like the kitchen sink I did. So you can have a type choice, which has that structure. You could have type Boolean, and you can have type score, which will give you a score. Or a probability. So let's go through different examples. So the first one is just a primitive. I have a little runner here that I built, but here are different samples that we're gonna be working with. The evidence is like essentially the state and these are just different labels here. So for example. This is like a re f reply a reply classifier, reply quality. And the state I'm giving it is the evidence. So this is the state of the world. The refund policy, you know, refunds are within five to seven business days. There's a restocking fee. That's some state. 15% on open hardware waived if the item was damaged. And then the reply is another piece of state, which is what the agent replied with, theoretically, right? And then you put those two together, you call on the classifier, you call evaluate, you give it the state. And then you get the answers from that classifier. And so we can run this. Let me open up my handy dandy terminal. I've. Seven demos, so we'll try to go quickly, but this is just testing the base, just the primitive itself.
Shane: Dang dude, seven demos. I told you to come with one.
Abhi: I know I well there's a lot of primitives to touch.
Shane: Overachiever.
Abhi: Yeah. Let's go so and these headings is like what the type of response should have been. Warm and complete, the reply was like, I'm sorry, the hub arrived damaged, that's on us, I've started, blah blah. The choice was Jev made a decision that the issue is resolved, right? That's the resolution. Hundred percent resolved. And, you know, is it grounded? The Boolean which is like fifty two percent probability, and then the tone. You know, it scored it at two. Like, okay, cool. Usage, six hundred and seventy-five tokens, pretty cheap. And obviously this thing ran really quickly. Correct but cold. You can see same type of deal. These are like the what the probabilities of Like how sure it is, and you know, is it resolved or not? The tone was bad out of a two scale. And then warm but invented, which is the resolution was partial. Because it said that it pushed the refund and it gave it something that is not in the state. Because our state said that you can only you only wave it if it's opened. Or if it's unopened, it's not, and so it's not sure about that stuff. So now you have like these objects that you can code into your application. Based on what the classifier said, which is cool. So this is the base primitive, right? You ask the classifier to evaluate something. It gives you an object that you've designed. And then now you can do a bunch of if statements and shit in your code. And you now you have like a very cheap classifier to make your application more powerful. And let's talk about how that can get more powerful. The first place is workflows. There's a lot of stuff in workflows that have to make a decision. Usually we tell people to make an agent step to then decide what the flow should be of the workflow going forward. So I made a workflow here. We're adding a new primitive called dot classifier. I can pass my classifier in. I can give it state from the workflow context. And then based on the classification, I can branch the workflow. Right? So now I can say based on intent, if it's a refund intent, then I'm going to do this step path. If it's broken, I'm going to do this step path. Whatever.
Shane: And to me that this is where it starts to get super powerful and super exciting because for the longest time we always said make things a workflow if you can. And then everyone wanted to put everything in plain English and skills and in prompts. Which can work. Not saying it can't. It's still maybe in some cases valid. But I do think this is kind of bringing us back to pushing more deterministic logic or even parts that aren't deterministic but could be classified or judged right and moved in the right direction. So it gets you I think higher quality results for a lot of things.
Abhi: Yeah. And it's native it'll be native to a Mastra workflow. Well let's go look at the ticket triage. I know we have to speed this up a little bit, but these were the criteria that I set out. And based on this instructions, let's go look at what the input state was, right? So for one branch or so for one run, the subject was I got charged twice, the body was I was billed eighty nine dollars whatever. And so that goes to this intent of refund. That's what the classifier would classify this path as. This one says something's broken, that that does it, and then you know, sales, like this person wants to what does the user really want? Well they want to get more shit. And, you know, questions, right? So I've just you know, you can design this however you like. Because the classifier will classify to your intent, which is dope. That's workflows. Let's keep going. What about guardrails? I think classification is a really good place for guardrails. Because right now you have to spend a lot of LLM cost to determine if something is against PII or Whatever your specification for the guardrail is. So we have a new primitive or a new type of processor called a classifier processor. And you can pass it a classifier that's registered on the Mastra class, or you could pass your own. All of this is normal processor stuff. But what you get is on result, you get the answers. I wanted to know. Let's go look at this classifier. This is the content guard. Continuard had two types of questions. Is it abusive? Right? Is the wording abusive? Have I leaked secrets? So now this processor becomes like, all right, if I have a probability that's greater than point seven of having secrets in my chunk, then I'm gonna abort. I'm gonna do a tripwire, and this thing is done. Same thing with abusive probability. But then think about this even more. Like I could design my classifier however I want, and in here I can abort. Or just let it through too. I can log things. I could do whatever, but aborting is the main thing here. I could abort on whatever the hell I want, which before may have cost you a little bit too much. To then iterate through all the different strategies. And also if you're using an LLM, you have to make sure it has really good structured object support to even pull off something like this. So not only are you paying cost. You also have to have a pretty decent model who can do superstructured objects. And cool thing about Jev is it truly is type safe in the sense that it's always going to return an object. So I'm gonna return partial. JSON to you and all that crap that you'd have to build into this if like your answers were malformed or whatever.
Shane: Yeah, it feels to me like I used to try to layer on two or three processors, run them in parallel. Wait till they're all done. Use a super cheap fast model. LLM, right? Hope that I get the right structured output back. But now I feel like I could probably layer a few of these into a single call. And get everything I need.
Abhi: Yeah. You could have like a single classifier that does a bunch of things. And the monster way, you pass an input processor and you move on with your life because it works. The same way. These are the test cases. My invoice looks wrong. Can you check it? Here's my login. Blah. All right, continue on. We already have something called a tool search processor. But when you have a lot of tools, there's many different ways you can search for them. One you could load all the tools and then figure out through reasoning. Which one you should pick. You could do some BM twenty five search to find things that based on the input text of the message. You could find through search what tool descriptions are likely to you know be part of the tool set. But I added this thing called pre-selection. So you can pass a classifier. And the question that you're asking is specifically in this one is what is the domain of the tool that they want and you pass a map like these are the tools by domains. So you have a classifier, you have all the tools by domain. It is explicit. Because you know, I will say none of this stuff is very dynamic just yet. And you can set a minimum probability for it to recommend. And then based on that, you can see like which tools get seeded into the context. So the agent now has a minimal search window. If I had 200 tools and I asked about refunds, the classifier would say, this dude's asking about refunds. That's the refund tool domain. Let's load the refund tool domain as the minimum set. So then now the agent will search through that. Now in my evals for this, in some ways, based on the tool number and the structure, it works pretty well. But sometimes BM25 works well for a certain number of tools. So I wouldn't say this is like some silver bullet, but it's very interesting to play with. Alright, almost done here. Model routing. The model router processor, you know, based on the last human user message and what it what they want to do, you route the request to a different model and you know you can use a classifier for this as well. So let's look at this one. Yeah.
Shane: I always think r model routing is It's a little bit there's people on both sides. Some people think it's a good idea, some people think, you know, if you're constantly between each message turn you're routing to different models, you're breaking prompt caching and you're losing the gains you're gonna get anyways. But maybe with a cheaper, better classifier, maybe at least the initial intent of the message might be able to send you down the right path to the right. But I think your mileage could vary. So I don't I don't think it's a silver bullet and I don't think model. I wouldn't recommend model routing. It's also more complexity for a lot of a lot of folks.
Abhi: So like the way this works is like there's some instructions like hey, you know, pick the least capable moder model that can answer the support request. And then you have your different choices, like a cheap one, this is like GPT four oh mini. Capable at GPT-5.6, whatever. And then I ran this on an eval set, and while it does route the model correctly, There's not really like a good way to determine if that was actually important like a good route. You really have to make your criteria good here. You have to have a lot of different options. And you know, if you're not using something like observational memory, you're gonna be busting your prompt cache. And technically, I mean in OM world, if you switch models, we activate all the observations for that model. So maybe it's like a like a less burning problem, but it's still a problem. So I'm not necessarily convinced of a model router classifier, but Langchain has one, so I guess we will too. Alright, almost done here. Tool approvals, this is awesome. You know it this doesn't require any new primitive. You just get your classifier and then in tools and master you can have a acquire approval function. You can evaluate the risk. Based on the state of the tool call args and what the tool does. And then you based on probability, you say true or true or not. I mean you could always require approval or you can have something a little bit more dynamic. On when to require approval just by having a classifier involved. Pretty cool. And then lastly the one that's really taking the world by storm, which is classifier as judge. You can make scores that are classifiers based on, you know, the incoming results, how complete was the reply and you can just give it true or false or you can give it scores or whatever you want. And so now You know, LLM as as a judge is cool, but the knock it gets is it's very expensive. Because you also need a strong model, and stronger models are tougher judges. But if you don't have the money and some of these like evals are quite simple. Maybe a classifier is better because it's cheaper, faster. And if you're just looking for a pass fail with some type of probability, this could be a really good replacement for judges. So Yeah, that's all the classifier stuff. I'm sure there's more to do. One thing I worry about is like this is not a silver bullet, but because there's so much hype around it It may be treated as one. So I don't know. Your mileage may vary and you should be very careful.
Shane: I feel like one of the challenging things if you're in this space, right? You're building, you're trying to stay on on like the cutting edge of what's happening is that this is another thing you have to decide to pull off the shelf. So it is a new primitive, right? And there are I think some good use cases where it can fit in. I think workflows is a natural, a really natural extension. It makes workflows more dynamic. It's a little easier. You could do it with an agent, of course, but I think this is just faster. Makes workflows slightly more deterministic than probably just calling an LLM because of the way the model was trained.
Abhi: Yeah.
Shane: But I would agree. I think you should try it out, play around with it. We're about to release it. So you can play around with it in Mastra or just play around with it in general. But I do think we should, you know, temper expectations. I think it's very cool. It's interesting. I don't think it's revolutionary. I think they did a really good job marketing it. And I do think that there are use cases where it can have a place. But I think people s grab onto something new and like that's the thing that they is gonna like be this big breakthrough. I don't I don't know that's it. It's like a another tool for the tool belt when you're designing and architecting an application. There probably are some are many good use cases where you could pull it out and kind of weave it into what you're building.
Abhi: Yeah. The places where I'm most confident that it is a plus. But you still have to think and not just use it blindly. One, the primitive itself is gonna ship, right? Classifier. Easy. You're gonna get one, then you programmatically can use that wherever you want. You don't have to wait for Mastra to support it because it's just JavaScript. Put it in a function and have fucking fun. Two workflows. For sure. I can see the value in branch-based workflows. Adding classifier steps, 100%. This is dope. Processors for sure. You can make very quick decisions on chunks coming in from your agent. 100%. That's dope. And then scoring. Also, tool approval, sure, but tool approval is just a function, right? So C point number one. Like anywhere that's a function, you could use it And then, you know, in these primitives that we have, we'll probably only put it in places that we have confidence in. So the ones I said. The model router. I don't know. Probably not. You know, if you really want one, you could build one yourself. You know, tool search, it didn't make tool search a hundred times better. It made it like actually it wasn't even the best option. Out of all the search possibilities, it was like number three. So like that's not enough conviction to make it like part of the framework, you know.
Shane: That's it for Jev today. We talked about Jev for a while. If you are just tuning into the audio only version, I'm sorry. That was very visual. But we appreciate you hanging out. And we have a bunch of quick hits, the normal news stuff we want to get through. We're running long on time, so we're gonna power through these pretty quick. Like, share, and subscribe. And follow us on X.
Abhi: And tell your friends.
Shane: And their friends.
Abhi: I mean, we're not begging.
Shane: Well, maybe a little bit. Subscribe to Agents Hour every Monday, noon Pacific. The first one's kind of a big one. Abi, you want to talk about adding AGENTS.md to Claude. Code? Why is this a big deal? Finally.
Abhi: Finally. So this is a big deal because all the other agent harnesses use agents. Md and they also source it from the. Agents folder. Claude was the only one that didn't. So if you were multi-harness, you had to pretty much simlink your MD files f between both. That's whack. Now you don't have to do that. You can clean up a lot of code. You know, sometimes those things do not stay in link if you don't simlink it. And so you have like two copies of a command or whatever that are not even up to date with each other. This was dope.
Shane: Yeah. Just go everyone go out and just delete your claw. Md file. Yeah. And make sure you just use an agent. Md. I think a lot of us are just thinking, hey, we don't have to simlink it, we don't have to duplicate it. Just Have one one standard. It's nice to have standards. All right, this was announced on September twentieth, or this came out. There's an open weight seven billion parameter model that outperforms nano banana two point oh. It's called Qwen-Image-2.1. Super compact and efficient. So if you're interested in image models, check that out. Shadcn/lint was introduced. And essentially it's like a linter for your design system, I guess. Your agents can determine if It breaks these design system rules and then can feed that back to the agent and the agent can then loop and fix itself. That's dope. This got a absolute ton of views and someone basically hallucinated entire operating system or internet with Qwen 3.8 so they built an offline browser, zero network calls. Mounted it to an Ubuntu desktop. There's no Wi-Fi, no scraping. So this was just if you want to look at see something that's just really interesting and cool. Not going to go too far into it, but there's a video on this post. From analog a lock and yeah it's pretty cool so we talked about software factories well. Factory raised 200 million at a five billion dollar valuation so A lot of money in companies like Cognition with Devin or Factory to really figure out how do we automate as much of the software development process as possible. Last but not least. Let's end on a new term. Slop grenades. I think this came out a while back, but I've seen a lot of people sharing it. This was from September fifteenth from Shane Parrish. And was a conversation or talk with Tobi from Shopify. You know, a slop grenade is when you let AI produce the work and pass it off without adding any value. We call that a meat factory around here, but some people are calling it slop grenades. So we'll see which term wins out. It's the same thing. They just wanted to invent their own terms, you know?
Abhi: Yeah.
Shane: It's all good. That's it for the news today. It's kind of a light week, but there was that one big topic with Jev. Thank you for tuning into the show. You can follow us at Mastra. Check us out on YouTube @mastra-ai. I'm @smthomas3 on X, Abhi’s Obby. Ayer on X. And there's one more thing. Not related to the news at all. I wanted to do this at the beginning. If you are curious on what Mastra’s listening to these days, Damien from our team put together this really amazing Mastra radio site where we just have a bunch of our Suno songs for you to listen to. If you want to go to radio.mastra.ai off to the AI and just play around. It's a pretty cool experience. This was just dropped during like a Friday demo day. We had no idea it was coming. That's it. That's it for the show. Peace.
More episodes
- September 16, 2026"Do You Think AI Is Gonna Kill Us?" | This Week In AI
- September 15, 2026Why Your AI Agent Gets Brain Damage, and How to Fix It - Tyler Barnes, MastraTyler Barnes
- September 9, 2026GPT-6 Astra vs Claude Fable 5.1, We Should Pause AI & The Benchmark Wars | This Week In AI
- September 8, 2026Multiplayer Coding Agents in the Cloud - Charlie Holtz, ConductorCharlie Holtz