Is Google finally back? (This Week In AI)
Anthropic's IPO prospectus leaked, showing a $2 trillion valuation target on $4.6 billion in revenue, with a quarter of that coming from two customers. Google is back in the frontier race with Gemini 4 Argon, which tops the Vals Index and can output up to a million tokens at once. Decision models are suddenly everywhere: OpenAI's Decisions API, OpenRouter's new category, Cloudflare's Clef, GLiDE, Respan's Span-01 and Kev. The White House got the big labs to sign an accord on superintelligence, Factory fired an advisor who then became Cognition's CRO, and Vinod Khosla picked a side in public. Shane and Abhi also talk about why maintainers like Sindre Sorhus and Hono are closing external PRs, and go through quick hits including Turso joining Supabase, OpenDots, Conductor Mobile, Karpathy on explainer videos, Claude Code mods, and a video model that passes the Turing test.
Watch on
Episode Transcript
Shane: I saw some kind of prediction that now at the end of the month Google will have the best model. It's like flipped. It went from like 1% to like 60%.
Abhi: Yeah. Well I mean we've been talking so much shit about Gemini. I want them to have a win.
Shane: Hey everybody, welcome to Agents Hour. I'm Shane here with Abhi. It's October 5th. We're doing the news. We do this every week. We have a bunch to cover today. There's a lot of little topics, a few big ones, so we're gonna fly through it pretty quickly. The first is Anthropic released, or at least it had leaked, that they had filed for their IPO, I think, a while back at this point, but now the prospectus of it kind of leaked out over this last week, and they talked about their revenue, they talked about their operating losses, and I think some people are giving them a pretty hard time because they're targeting a very high valuation. And their revenue isn't even close to what some other tech companies have. I think they're like targeting Meta level valuations with a tenth of Meta-level revenue or something like that, right? But yeah, what do you think?
Abhi: I think, you know, every, the multiples are crazy now. But I think a lot of people are giving them flak too because of where the source of revenue is coming from, it's not really distributed. Right? You have outliers that are making up a ton of your revenue. And if you ever lose those to Muse or OpenAI, then you'd be screwed, right? So it all depends on if they can continue the trajectory of getting whales to keep the revenue high. Think it's gonna be hard. But you know what? The layman person in the market will probably buy on IPO day for a premium, like crazy premium.
Shane: Yeah, I can't imagine they get to $2 trillion. That seems wild. I do know the growth rate though is tremendous, right? That's what people buy on. It is a tremendous growth rate never before seen. And so if it continues it becomes, you know, potentially, you know, you could argue, potentially becomes the biggest company in the world, right? Or very close. Is that gonna happen? I think it's still a long shot, but that's what people potentially buy on. And we will see. You know, I don't think there's any more news on when the date would be, whether it gets held back. I've heard rumblings that they might push it back a little bit. I think all these new model launches actually hurt them. This doomerism, I think probably doesn't help them, but it's coming from Dario, so maybe he thinks it does. I don't know. It's like a bit of a game. I don't know what game they're playing, but I do think that they have a lot to show to hit that valuation. It's gonna be very hard. It's an uphill battle for them to actually reach that valuation. But if people buy, you know, that's all you need. Well.
Abhi: Not financial advice. Don't buy at IPOs. Just, that's my rule of thumb personally, because the market will correct. You know, you're operating in a private market where your valuation, it's like Whose Line Is It Anyway, all the numbers don't matter and stuff. But then once you get into the actual public market, you'll get screwed on that IPO price. So, you know, if you're really an investor, just wait, personally.
Shane: Not financial advice.
Abhi: Yeah. The only thing I didn't get burned on is that Cerebras IPO, so what's up? But other than that, you know...
Shane: You didn't follow your own advice is what you're saying.
Abhi: Yeah. I was screwed for a while though. So then, you know, finally it came back.
Shane: Alright, let's talk about decision models. This is a new genre. We've been talking about it on and off, basically on, for the last couple weeks. This was at OpenAI DevDay, which we're going to talk about OpenAI DevDay a little bit on this show. We did a whole episode on it last week. So if you want to know what we think about what they launched at OpenAI DevDay, go back, watch that episode. It's on YouTube, Spotify, all the places. But OpenAI released the Decisions API for lightning-fast, constrained decision-making powered by Luna. Supports visual inputs and tuned to be able to make decisions in less than a few hundreds of milliseconds end to end. To which the founder of Jev said, begun, the clone war has. Because if you've noticed, there's a ton of these now. Jev basically created a new category. On the one hand that's good. On the other hand, you know we'll see if Jev makes it out. They have their own category that they defined.
Abhi: And you know, it's interesting that it's not only Luna as like, you know, a decision model, but you know, a lot of people are training a Qwen model. So they're trying to get this similar concept going.
Shane: OpenRouter now has an entire decisions model category. So they're building out these options. So if you wanted what they're calling decisions, we've called them classifier models, people are calling them decision models. I do know the creator of Jev says using "decisions" as the only classifier for it is a little short-sighted, I think. But we'll see what they have cooking. They might have more things coming, or maybe they don't. I don't know.
Abhi: But then how do people understand what a System 1 is, right?
Shane: Yeah, exactly.
Abhi: They need some word for it.
Shane: Yeah, they're calling it a System 1 model. If you are just hearing System 1 model for the first time, all it means is you don't actually think, you just act on like simple decisions. So that's why it's good for simple things. If you need it to like reason about and make a decision, it's actually not good at that. So you gotta break down a complex problem into simple decisions. And then these models are good at making very quick, fast decisions without having to do a lot of deep thinking. Cloudflare released Clef. They say today we're releasing two fast and accurate decision models that top the benchmarks for quality and latency. Use them hosted on Workers AI or grab their weights from Hugging Face. So this was on October 1st. So Cloudflare is getting in the decision model game. SGLang also has decision models, which they, you know, you mentioned like the Qwen models. That's what they have available. So you can, again, answer quick yes-no questions, get probability. It's just more proof that everyone's building decision models into their products, gateways, frameworks. There is something called GLiDE, the first decision model that thinks, you know, which is interesting. System 1.5. Also known as just a normal model, but I guess it is fast, right? Smaller. It beats Jev by 6.9 points. So it uses adaptive thinking. So it calculates, you know, a fast probability distribution. If it's uncertain, then it does some reasoning. The founder of Jev says he doesn't like benchmarks, but people are going to benchmark and figure out where does Jev fall short, where do these other models potentially beat it, and then that's how people are going to make decisions on what models to use.
Abhi: Before we move on to the next thing, we have to mention the homies at Respan have a decision model, Span-01. Yep. And then Devin made Kev, which was trained on Qwen as well. So the space is heating up for sure.
Shane: The one thing I will say about this is I think it's a good thing for the industry because for so long there was this push to get everything into your skill file, your prompt, and just let one big model do it all. And I think what we've seen and what we've probably known, but we've kinda had to push against the narrative a bit, or people have pushed against the narrative. Is that if you can break down a complex problem into more deterministic chunks, you're gonna have better results long term. If you're using Mastra, it's workflows, right? If you can break something down into a workflow rather than just like a prompt, that's probably gonna give you better results. Now sometimes a skill file is fine. Just put it in a skill, let it do its thing. It's gonna be accurate 98% of the time and it's good enough. But if you need it to be accurate 99.9% of the time or whatever, a more, you know, directional workflow with maybe certain decision points is actually the better way to architect something like that. If you can define the process and you need to make smart decisions throughout the process, but you can break down those things to where, you know, a decision model or a System 1 model could decide between, you're probably gonna get better, more reliable results at least. Alright, we're gonna breeze through OpenAI DevDay because again, we did a whole livestream on it last week. Go check it out. It's been launched. It's on YouTube, all the places. OpenAI introduced Dots. There was a pretty big demo failure because of some Wi-Fi issues. So yeah, kudos to the individual that did the demo. She did a great job of working through it, but Dots are, you know, just like Grok Bot, like Muse, it's very much a trend as well. They released GPT-6.1 Sol and they introduced a new plan, a more expensive plan. And then they also made changes to the $200 plan, which they basically cut your limits a bit. They also kind of improved the ChatGPT extensions. Right. So you can build and kind of extend ChatGPT almost like you're extending VS Code, or you know, you basically build a plugin, you can extend the UI, you can make changes to the app and you can use Codex or ChatGPT right on top of it. So that was also a cool launch. Anything else? Any other things we should highlight besides telling people just go watch the full episode?
Abhi: Yeah, go watch the full episode. I did play with Dot. Yeah, it's pretty cool. I don't think I'll ever use it after playing with it though. That's just me personally though. I'm just gonna stick to my one ChatGPT thread that I've had for two years. That's my dot right there.
Shane: OpenAI did release a new Dots demo after the first one on stage didn't go well. So if you want to see what it is, you can go check it out. This was from October 1st.
Abhi: This got so much controversy though, which is not fair, but I think the industry as a whole is sick of seeing booking flights as demos. And then there's a lot of people like arguing like what airport you should be going to, whether, if you're trying to go to Silver Lake, should you go to Burbank or LAX? Like that's like some minutiae BS. I think the bigger point that was made was Why are we always doing simple things for these demos, like booking a flight? And then in the chat, or in the threads, it's like, yo, I want, I wanna see Dot do my taxes. Like I wanna see Dot do some real stuff.
Shane: Yeah, a lot of people also hated on it because it took quite a bit of explaining for how to actually do it the first time. And I would just, you know, encourage you to say, like, yes, you probably have to explain yourself at least once, you know. If me and Abhi are working on something together and Abhi works a certain way, he has to explain it to me once, but then hopefully he never has to tell me again. Or he doesn't have to tell me every time. So you know, I would give any demo like that. Imagine you could explain it once and then it knew your preferences. It knew, you know, I like window seats, right? It knew I want to sit by the window. So I think it'll be figured out. I don't know if it's going to be OpenAI or Muse or Grok Bot that does it best. I do agree though. Needs to be better demos.
Abhi: Dude, also the funniest roast on this post, which is so funny. The camera was like shaky like this, like this all the time. And then someone said like, oh yeah, it's like watching The Blair Witch Project. The internet never ceases to amaze me.
Shane: Why would they have shaky cam on when like the first demo didn't go well? You think the second one they'd want to make it like as professional as possible, not like there's an indie, you know, camera crew holding the camera. All right. Sam Altman came out on October 2nd and said there's some speculation about our partnership with Cerebras. Cerebras is a close partner and we have a deep engagement pushing on the frontiers of speed. That was the entire tweet. That was it. So then...
Abhi: Not financial advice, but...
Shane: Yeah. But that led me to think, okay, what's going on here? So I did a little digging and apparently it's because GPT-6.1 Sol is, you know, using NVIDIA GPUs rather than Cerebras for inference or something like that. And so it caused a lot of speculation of okay, why are they not using Cerebras for this? Are they going back on that? Are they not caring about that partnership anymore? But it sounds, you know, according to Sam, he's trying to assuage the fears that they're not gonna use Cerebras. Maybe they just haven't yet. Maybe they will later. I don't know. But I thought that was interesting.
Abhi: I mean the stock, the stock went up since this tweet. It was trending down. The tweet comes up and now the stock's trending back up. In a month overview. Not financial advice though.
Shane: All right. Probably the biggest news from last week, and we kind of hit it here in the middle, is Google dropped a new model, Gemini 4 Argon. And this is from September 30th. It says introducing Gemini 4 Argon, our new frontier model. It's built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense, rolling out today to a set of trusted testers through our Fairwind Program. So again, not rolled out to everyone yet. But if you look at the benchmarks, they're very good. You can see Vals Index, AutomationBench, DeepSWE, Terminal-Bench 4, and it doesn't win on everything, but for the longest time, Gemini were just releasing new Flash models because they couldn't really compete on the frontier. Or they at least didn't want to talk about their benchmarks. So their benchmarks always seemed to perform pretty poorly. I think this was the first Gemini model in a long time that caused people to stop and say, damn, they might actually be back in this race. Now the caveat here is I haven't tried the model. Not yet. But we have not tried the model, and I'm sure most of you listening to us have also not tried the model. So are the benchmarks true? Does it actually live up to the hype? That's to be determined. But if it does, then I do think Google has kind of reinserted itself back into the frontier conversation at least.
Abhi: Yeah. I'm curious how they'll position it marketing wise and, you know, will they talk shit about Muse? And like pay back some of the shit talk from Muse. If everything is good, Google is in a really interesting position to capitalize on the model itself.
Shane: Yeah, and we already talked about this in the past, when we were kind of hating on the fact that they didn't perform well in the benchmarks, but that they have such huge consumer adoption, right? There's no company better positioned. Facebook is pretty well positioned with Muse, but as far as like getting consumer market adoption, yeah, they are way better positioned than X, I think, because everyone's using, you know, Gmail and Google, and you have access to all these things. If you could build especially for the workplace, like a personal assistant for an organization, and you tie it into G Suite and all the tools that workplaces use, you could probably win that space. You could win the consumer space. I think that if they can be competitive on the model layer, they can win a lot of other places.
Abhi: Yeah. I mean we've been talking so much shit about Gemini. I want them to have a win.
Shane: Yeah. I think I saw some kind of prediction that now at the end of the month they think that Google will have the best model. It's like flipped. It went from like 1% to like 60%. Which is, that's an insane jump. Vals AI released that Gemini was on top of the Vals Index, so you can see the accuracy rating compared to, beating Sonnet 5.5, Opus 5.5. Yeah, GPT-6 Astra, GPT-6.1 Sol. So it's beating those on the Vals Index. And then someone came out and said, wait, did I read that right? Gemini 4 Argon has a 1 million token output limit, so not a 1 million token context window, but an output limit, so a lot of frontier models cap it at, you know, like in the hundreds, two to 200K range for like output tokens, right? So it's essentially like five to seven times larger output. You can basically write entire books almost in one shot, which is insane. And then I thought this response was very funny. They named it that because after four prompts all your tokens Argon. Because you can just output a ton of tokens. You're gonna spend a lot of money potentially. So it'll be interesting to get access to it, to play around with it, and to see how good it does on various tasks. Of course everyone's gonna test it with coding, I think, right? That's the first task we probably all do if you're watching this show. But I would love to test on some of the other tasks as well, just you know, writing, things like that.
Abhi: I'm stoked, you know. It'd be nice if Google's back in the game. I hope it's not a flop. Like, share, and subscribe.
Shane: And follow us on X.
Abhi: And tell your friends.
Shane: And their friends.
Abhi: I mean, we're not begging.
Shane: Well, maybe a little bit. Subscribe to Agents Hour every Monday, noon Pacific. So let's talk about the White House AI luncheon. I don't know if it's a summit, a luncheon, I don't know what they're calling it, but this is a post from David Sacks. But to set the stage, President Trump got together a bunch of leaders in AI for this kind of lunch dinner thing and they came together and came up with kind of some resolutions that they're all gonna like hold themselves to. They signed it. So we'll read through what it exactly entails here in a bit. But I think there's some critiques of it that it was just like all the net worth billionaires in one place, like agreeing, you know, kind of backslapping and agreeing that they're gonna do this thing, whether they do it or not. There's really no like guarantee that they can do it, but I think it's a step in the right direction. What were your thoughts on it? And then I'll read through what they actually signed here in a second.
Abhi: I mean I think it was cool, given the doom week we had and people saying different things and the president also, like, disagreeing. So it was nice to see them all together. The amount of memes that were created from the pictures is comical. It's like pure comedy. So I think that blessed us with that too. So two good things happened.
Shane: And like if you saw the video at the press conference with Dario in the background and he's just like moving around in the video and you're just like, what are you doing, dude? Someone comparing Dario to Mr. Bean in the background. I thought that was pretty funny. You know, like Dario's just in the background of this. And I do think one thing I will give the president credit on is I don't think, you know, you could tell like Dario would not traditionally like President Trump, just based on politics, I'm imagining. But being able to get them in the same room, President Trump said a lot of nice things about Dario. I think Dario kind of said some reasonable things back. I think they played nice, and now it's hopefully getting some of the resolution that I think Dario wanted in the first place, which was at least getting all the labs to agree to this audit. So I feel like Dario hopefully can hang his hat on that he made some progress in that, which sounds like it was always very important to Dario. But I feel like everyone else was kind of in agreement. And then, you know, Sam Altman wasn't there, I think, because of OpenAI DevDay. So I think Greg Brockman was there. But I feel like it was this whole thing was kind of targeted towards like Dario being a bit more of the doomer narrative. It was a little bit, but then the other labs are a bit less doomerism coming out of them. So I feel like it was very much like, how do we get everyone to agree that we're gonna take ownership, but we're gonna also play nice and we're gonna have kind of independent audits. So if we read what it says, it says White House Accord on Superintelligence, Joint Commitment on Frontier Responsibilities. In order to build a positive future for the American people and the world, we believe every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public. Starts with every company that is training and deploying frontier models having robust internal processes and controls to ensure that their technology behaves as intended. And that any issues are promptly identified and resolved. Therefore, in addition to any other precautions, we believe each company should implement the following four layers of controls and audits. So these are the four things that people are agreeing on: robust internal controls to monitor the capabilities and alignment of its models during training and deployment around areas like cybersecurity, biosecurity, and chemical threats. And to ensure that its models do not hack or access technical systems in unintended ways. Two, empower an internal team to ensure all the controls, monitoring, and detection are operating as intended. And that any issues are remediated. Three, partner with an independent external auditor or evaluator to carry out independent assessments of whether the controls, monitoring, and detection are operating as intended. So that was a big one, is you have an external party that has access to validate and audit. And then this is the fourth one that I think potentially gives it a little bit of teeth,
where you can hold people accountable: designate an independent committee of the Board of Directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators, as well as to ensure any issues identified are remediated. So essentially getting it to like board level visibility that if there's issues, the board has to at least sign off that they're gonna release this model knowing that there's these potential issues, right? So I think it puts more responsibility on the board, which means there's, you know, potential for lawsuits and all those things, right? There's a lot of like responsibility that comes with that.
Abhi: Especially for them to disagree on release.
Shane: Yeah, exactly. So I do think the goal is like hopefully it's enough to slow the teams down from shipping something they shouldn't have, but not so heavy-handed that there's this huge process that needs to be in place where we have to slow down so much that we can't get new models out the door.
Abhi: Yeah. Dude, startup idea: model third-party evaluator.
Shane: Yeah, I mean, there's gonna be a lot of those, right? Make some money, y'all. Alright, more drama. Factory vs. Cognition. For those of you, you've probably heard of Devin, you've heard of Factory Droid. This is from Matan from Factory. It says, we're terminating Chris Degnan for unethical conduct involving Cognition. The last three months have seen incredible progress in AI capabilities. San Francisco has flourished as new companies that solve new more ambitious problems are finding great success. Generally it's a wonderful time to be building. So basically it goes on. I'm not gonna read the whole post, but a person that has been like an advisor to Factory was also starting to talk to Cognition and apparently got hired by Cognition without Factory knowing and they're obviously very direct competitors.
Abhi: I think it was after he got removed. Then he joined.
Shane: Yeah, but then there's speculation that he wasn't actually removed, he just resigned, and now Matan's saying he was removed. But this also came out half an hour or an hour after: I'm super excited to join Cognition as its chief revenue officer. Going from the first sales rep at Snowflake to CRO at a $100 billion public company was the thrill of a lifetime, basically wants to do it again, excited to join Cognition, loves its growth, all those things. The issue is if he was an advisor to Factory, you know, at like board level advisory, should he have done this? Was this, like, clearly... Factory thinks he was doing something wrong. He thinks he, you know, had disclaimed everything. He had not shared any secrets with Cognition. But of course Factory is accusing him of sharing trade secrets of what they're doing, what they're building. And then this is where it gets really wild because then you have an investor in both companies. So this is Vinod Khosla from Khosla Ventures, who has investments in both Factory and Cognition, picking a favorite here. This is like picking a favorite between your two kids publicly. This is targeted at Factory, was a response to their message. You are a struggling second tier competitor that is more unethical and lying just because you have no decency or sense of proper behavior and shows your desperation. Straight out lying about if Chris being fired I thought would be below even you. Shots fired. This was shots fired. And then of course then everyone saying you should never raise money from Khosla because they'll literally tear into you in public. Just a lot of drama on the timeline. And then Shaun Maguire says, what Cognition doesn't understand is that we're sitting on a nuclear weapon to end them. Hopefully we don't need to use it. Okay.
Abhi: Causes so much more drama.
Shane: So much drama on the timeline if you are paying attention to what's going on in the AI world. As far as like coding agents.
Abhi: Yeah. You know, it's true though, like people in the valley, they move differently based on who they've invested in. We've met people who've invested in our competitors. And treat us different ways because, you know, their allegiance lines. It's kinda whack, right? Like as a human, that's super whack 'cause maybe you are indifferent. You wanna be friends with homies and like in every company. But typically that doesn't necessarily happen. You know, just your investment is your alignment. I don't act that way, so y'all can be friends with me. But yeah, it's just whack. We've seen it happen. This is another example of that happening.
Shane: Because it sometimes feels like that, right? Allegiances. Yeah. Yeah, there's allegiances and there's you know, falling outs and it's all happening in public and Yeah. X is free, you know, or you can pay, but X is free for the most part, right? So you view all this, grab your popcorn and read the thread like we're seeing in the chat.
Abhi: Yeah, like when you're at a smaller stage like seed round or A, you don't really see it much because at that stage investors are making multiple bets and you know unless they're like taking big swings at certain companies. But once you get into serious rounds like these companies in particular have raised hundreds of millions of dollars now. And so things get a little Game of Thronesy.
Shane: Yeah, they're throwing in big numbers and they want to protect their investments. All right, let's talk about open source. Mastra's open source. We love open source, but there's this trend of open source closing the door. So this is from Sindre Sorhus. This guy's a legend, by the way. Due to AI, I have disabled external pull requests on all my repos. Open source as we have known it was fun while it lasted. I will still maintain projects and handle issues. So what are some of the open source...
Abhi: P-map, my favorite freaking thing. Execa, my other favorite thing. And the list goes on. Sindre has done so much for TypeScript and open source in general. But those are like my two favorites. But he's done a bunch of stuff.
Shane: Yusuke Wada says, sad news, we disabled PRs from external contributors on Hono. Hono has not stood here without PRs. I will never forget the PR usualoma created for RegExpRouter, a damn fast HTTP router we have never seen, but PRs don't work in this era. Contribute in other ways. Thanks. I guess we can say we concur, right? We concur. In some ways. I don't believe open source is dead. But I do believe and we've said this, the contribution model to open source is changing. I do believe that quality issues with, you know ideally reproductions and enough details without trying to enforce a specific decision is the right way. And then that way the team that is, you know, behind the open source project can use their models with their context to actually fix them in their way. So I think issues are the new PRs. But I do believe that people should get contribution credit for issues.
Abhi: Yeah, and both of these authors have in their posts, pretty much saying the problem is your agent opens up a slop PR and then the maintainer is using an agent to review it and that's going back and forth models of unknown quality on the other end, and you know the question you ask yourself is, I could just do all this stuff myself. We've been talking about this for weeks though. And hoping that the community has discipline to listen to our wishes. For the most part they have.
Shane: We do let them open PRs if they have issues, right? Like there are some requirements. You have to have an issue. It has to be triaged as like a valid issue. But I would much prefer you just open the issue and make it detailed. Follow up if we have requests or questions about it, and then we will prioritize it and get to it. It's probably gonna get done faster than if you give us, especially if it's a big PR, because we gotta review it. It takes more back and forth time. You gotta get through our CI, right? There's a ton of CI.
Abhi: Gonna wait to approve you to run it. It's like no.
Shane: Yeah. You're better off just opening an issue and making it detailed. The more detailed the better, with, you know, reproduction, and that's gonna get it fixed much faster.
Abhi: Like, we probably should close PRs. I just feel like that is an extreme thing. Ideally, and it's been working so far, the community knows that, hey, we'd rather have issues. And we do say, when you do open a PR for an approved issue, oftentimes I close it and say, hey, Factory's about to handle this anyway, so don't worry about it. Then people got mad at us, you know, oh, you know, you're not letting us contribute, but the issue was the contribution all along. So.
Shane: Yeah. I think that's, as long as we can change the model where we value issues as much as we used to value PRs, especially good issues. I think people should still feel good that you're contributing to open source, but it's changed a little bit.
Abhi: Dude, you know what I'm seeing? Last point on this. I've seen this a lot lately where we get slop reviews. So instead of contributing an issue or a PR people are sending their agents to review code, which is even the worst thing that you could do. You know what I mean? So don't do it.
Shane: Let's go through the quick hits. We're gonna go through these rapid fire. Turso is joining Supabase. So Supabase raised a pretty good sized round. They acquired Turso. If you've used Turso, you now are gonna be using Supabase, I guess, because they're gonna keep it going, but within Supabase, it sounds like.
Abhi: Many people have used Turso because of Mastra.
Shane: Yeah, Mastra uses Turso. The homies. Yeah, they are the homies, so congrats. Atai from CopilotKit says, introducing OpenDots. Self-hostable, always on AI coworkers that works with any agent harness. Computer use, browser, terminal. Bring agents to Slack, Teams, Spaces and Pages for projects, voice calls, web and mobile. So if you want to try to build your own dot rather than using OpenAI's, you want to use OpenDots. That's cool. Charlie, friend of the show, been on a few times, introducing Conductor Mobile. Run a team of cloud agents from your iPhone, live now in the App Store. So you can code from the beach. Cua says we're reimagining what it means for your agents to work with all your computers. So they're excited to share Cua Spaces, built on Cua Driver and our virtualization stack, rolling out to macOS today. It's free and open, free and source-available, so you can control your desktops from anywhere. What do you think of computer use in general?
Abhi: I think it's the next frontier, let's say.
Shane: You think so? I mean, I do think that computer use is really cool when I can watch the agent control my computer and so I would know how to do it or I can like interact with it. I feel like that's one way to even like learn how to use an application. It's like just to watch the agent do it.
Abhi: Yeah. But you know, all the new models are saying specifically that their computer use is better, right? So this is all like a natural progression of the frontier. Cua, friend of ours, YC. They're doing it open source. So.
Shane: All right. Introducing e2e, the agentic testing framework for any app, npx e2e init. So it's a way to better write end-to-end tests for your application. Let's talk about this Andrej Karpathy post. I'm gonna share the whole full thing. We'll go through it really quick, but it is pretty interesting. Karpathy has kind of been silent for a while.
Abhi: He got neutered.
Shane: Yeah. But he's back. He at least was back to post this. I think there are a few good tricks in here though. So we'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips, and tricks. So it says, ask your LLM to explain something in ASD-STE100. Okay, if you want it to have better writing. Use diagrams and images instead of writing. Ask your LLM to create a diagram. You know, I do that quite a bit. I think you know a lot of us probably do, but that's a good tip. Ask for output in HTML rather than just text. This is the one that resonates with me because this is what I want. Explainer videos. The output format I am most bullish on is fully custom bespoke explainer videos generated on any arbitrary topic. Here's why I think this is interesting because if my agent would know my preferences, it could make a video based on my learning style. So I go back, and someone's probably built this. I haven't seen any good examples. But sometimes I want like I'm looking at this PR. It's a lot of code. It's 400 changed files. I have this wall of text to read, explaining what it is, because my agent condensed it. But if you could just explain it to me on a two-minute YouTube video in terms that I understand because you know the parts of the code base I've touched, that would be pretty cool. Could explain it to me and show how it interacts with the things I already know about. So I think that there's room for personalized explainer videos as these models get better, which is, you know, I'm using code as an example, but I think there's a lot of examples of like I want to learn this topic, I have this specific camera, I want to learn how to do it. Build me an explainer video and show me how I can actually like use this camera to do this one task. And it can customize that video specifically for you. I don't know, things like that. Alright, let's continue on. You can now mod Claude Code, change how it behaves, customize the UI, swap in your own features, write one with a few lines of TypeScript, or even have Claude build it for you. Mods ship inside plugins. So you install them with /plugin in the CLI or desktop app. Did you see this thing called Griffin?
Abhi: Yeah, dude.
Shane: Alright, so introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Which is wild. It's not so much better than what was there before. Like I feel like I can tell, but only because I know now. If I was in the moment, maybe I wouldn't have been able to tell. The latency is pretty good. It looks pretty real. It's pretty hard to pick out. If you weren't paying attention and you just had a conversation, you would probably think it was real unless someone told you ahead of time.
Abhi: Yeah, watch the video. Because in our YC batch we had a company called Pickle that was trying to do this in Zoom meetings and I thought it was pretty impressive back then.
Shane: Yeah, the thing about Pickle though is it was kind of like it was still your voice, right? You would talk, but it would just be your avatar matched to your voice. This is like 100% synthetic, which is a different level of game.
Abhi: Yeah, dude.
That's the next level: not attending standup in the morning and staying in bed.
Shane: Yeah. ElevenLabs on September 28th said, introducing Eleven v4 and Eleven v4 Turbo, our fastest and most emotive voice models yet. You can watch the video. They're very good. They've always had really good voice models. I think they're getting even better. I saw some comments after both this and Griffin and ElevenLabs, like, now's the time where you talk to your parents and your grandparents around, like, having a passphrase or knowing that it might not be you. I do worry about that. Like, you know, our voices are out there on the internet. If you've ever recorded anything live, it's out there. Like it doesn't take a lot to like clone someone's voice, clone someone's appearance. And you can definitely scam people. So have a passphrase with your family. Yeah. Recommendation. This post came out from Kilian. He says, can LMs discover and fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do. SWE-sweep benchmarks this on 100 repos, 22 languages, 4,000 real bugs. But anyways, that's a way to just like scan your code base for bugs. What do you think about this as a topic? Like using some kind of external tool to scan your code base, try to find bugs proactively before users even do.
Abhi: I think it's amazing and we are starting to use one from our homie Dan Robinson. Detail. If you guys want to try Detail, go to detail.dev. It can run a scan on your whole codebase and then start opening issues with the proposed fixes. So I think I'm very bullish on this because we ship a lot. We also need to fix a lot of bugs. So it'd be nice to have like early detection on a lot of things.
Shane: Yeah, it's like proactive monitoring. It's like pre-monitoring before users see it. And then you can evaluate which ones are gonna be the biggest impact, which ones users are most likely to hit, and which ones are just edge cases that no one will probably ever see. I think as a category it makes sense, especially if you think of a world of basically unlimited compute, which we're not there, right? Like we're definitely far from it. But if you had that, well then why wouldn't you want a bunch of agents to try out or look through your code? Try your application out in all different kinds of ways and find all the bugs before the user did. That's what I would want. And I think like tools like this are gonna be steps to get there. Tools like, you know, SWE-sweep, Detail and others will help. Yeah. And that's the show today. This is Agents Hour. You can follow us on X @mastra. Follow us on YouTube, mastra-ai. I am @smthomas3 on X. Abhi's @abhiaiyer. Any parting wisdom before we close the show out this week?
Abhi: I think I'm gonna start using the word superintelligence.
Shane: No more AI.
Abhi: Just throwing it out there. I'm gonna give it a try.
Shane: Should we rebrand the show Superintelligence Hour? Probably.
Abhi: You know, mastra.si, you know, maybe. Well I guess in California you're not allowed to say superintelligence, but who cares? Let's give it a try this week.
Shane: Throw superintelligence around at some of the meetups and yeah, see how it lands.
Abhi: I might do that tonight and be like, hey, you guys are all interested in superintelligence, right?
Shane: Alright, everyone. Thank you for tuning in to Agents Hour. We'll be back again next week. We should be doing it on Monday. It is a holiday in the US, but I think we'll still be here, so hopefully you will be as well. We'll see you next time. Peace.