AI broke CI. Now what?
AI agents write code faster than CI can check it. GitHub reports a 4x jump in Actions minutes since 2023, and Anthropic says CI became its biggest bottleneck after job volume grew 25x in six months. Daniel Worku, co-founder and CTO of StarSling, joins Shane and Abhi to talk about what CI looks like when agents write most of the code. Faster hardware helps, but StarSling also runs a background agent, built on Mastra, that A/B tests changes to your CI and opens a PR when it finds a speedup.
Guests in this episode

Daniel Worku
StarSlingWatch on
Episode Transcript
Daniel Worku: The amount of code being generated has grown exponentially. GitHub themselves, I think, shared that they've seen a 4x growth in GitHub Actions minutes. Anthropic recently published an article about how CI became their biggest bottleneck as a company. They had 25x CI job volume growth in six months.
Shane: Welcome to the show, Daniel. Before you jump in, I do want to just set the stage for people of why we kind of pulled this together last minute. We weren't planning on having a guest today, but a bunch of things came across my feed and I'm sure your feed as well since you're way in the space. The first one I think was this that I saw. My GitHub Actions spend is going out of control. And then from there, Flo, friend from Lindy, said, CI has become the top bottleneck of every engineering team I talk to, including Lindy. Our CI spend has become stratospheric. And then, you know, I saw another post which is kind of tied to that, which I responded to, which is saying the cost of coding has plummeted, but long builds have become a bottleneck for many companies. There's a bunch of startups trying to solve it. I, of course, chimed in that we use StarSling to help solve it over here, and we've done a case study. But I thought what better way to talk about this, because it is very relevant and very timely, of you know, CI in the age of AI, who better to bring on than you? So introducing, this is Daniel from StarSling. Tell us a little bit about StarSling, about yourself, and yeah, love to talk a bit more about CI.
Daniel Worku: Well first, thanks for having me on, Shane and Abhi. We're in the code verification bottleneck moment where all the factories are churning out code and now everyone's like, oh no. Either I'm waiting on CI or I'm spending a bunch on CI. So I really appreciate that you two brought me on. I'm a co-founder here at StarSling. Previously I worked in Big Tech and a bunch of startups, most recently at Netflix. I led the team that built the most used internal tool for developers, called Netflix Console. But at StarSling, we do two things. We give you faster, self-improving CI, and we just launched last week our new product, which is cheaper code review. So we have a Review Runner that makes it much cheaper to run code reviews.
Abhi: We use StarSling at Mastra. We've been talking about the CI problem. And we had this problem even before we had our Factory running. You know, we were just churning out so much code that, you know, our CI times just went through the roof. You know, thankfully as an open source project you do get the subsidy of GitHub runners. But you know, many people are not in that world. They're paying for Actions minutes, etc., and so, you know, when Yonas and you pitched us for using StarSling, I was wondering like, you know, what was the biggest edge that you guys could provide? And I was just wondering, could you explain to the viewers what's the problem with CI today, and you know, how are you guys tackling the problem?
Daniel Worku: Yeah, that's a great question, right? Because there's a lot of solutions in the CI space out there. It is an old problem. We've been running CI as long as people have been writing code and now all of our agents write the code. So most solutions to give you faster or cheaper CI at their core are really just about faster hardware. I shouldn't assume what's going on at GitHub for the reason why, but they use very, very old machines when they run your CI jobs. Whether you're private or open source, you're getting like three or four generation old CPUs, slower disks, slower network. And so most solutions in the CI space said, okay, we'll just use top-of-the-line machines. We can charge you less per minute than you would pay GitHub for your actions. You'll save time and you'll save money in two ways, right? You'll execute your CI faster, so you pay for fewer minutes, and then the pricing per minute is lower. So that's a traditional solution. It's just hardware. And so StarSling does have that, right? We run fast CI on ephemeral sandboxes. But really the differentiator and our edge is we make your CI a self-improving system. So after each run, we track the telemetry on the load that's running for our customers' CI jobs. And then we have a background agent that runs in a loop. It's a performance optimization agent. It's actually built on the Mastra framework. And it creatively comes up with different proposals for how it can make your CI run faster. And then it A/B tests those using the exact same machines, the same environment that would appear in GitHub Actions. And so what this looks like to you as a customer is you know, every few weeks you'll get a PR from StarSling that's like, hey, you can make this test like 5x faster if you merge this in, right? And here's the results. I've run it, you know, ten times on the A and the B leg. And what really changes it for people is you get a CI system that adapts and improves the code, not just hardware.
Abhi: Yeah, well I was just looking, like we have 20 PRs closed by StarSling. And I think it's like hours of time gotten back just from it running and helping us figure that out. But you started this journey with a different open source project, right? To really benchmark everything. Tell us that story.
Daniel Worku: Yes. So we also open sourced something called High Performance Sandbox Benchmarks. And this was actually our internal benchmarks that we used to determine which compute providers to use. This was kind of our shortcut so that we could focus on, remember I said, we want to focus on the self-improving AI and giving you PRs. We don't want to do the hardware, right? And the way that people usually solve the hardware is you'll go out to a colo and you'll do capex, right? You'll like buy some machines and like wire them up. Or you go to a cloud, right? You go to AWS and you get the fastest machines. So we wanted to shortcut that whole process and we said, well, our agents already want to use sandboxes when we run this verification loop. What if we ran your CI jobs in a sandbox themselves? And so that's what our high performance sandbox benchmarks project was really about is we went out, we looked at all the different sandboxes out there. And we ran real CI jobs and real workload that represents what happens during code verification to identify what blend of sandboxes to use. Unbeknownst to our customers, you know, the sandboxes are still working on their reliability and resilience. So we actually instantly fail over between the different sandbox providers when they have an outage or an infrastructure issue, but we're just focused on how can we get the fastest speed for our customers without having to go out and buy those machines ourselves.
Abhi: Do you think there's a reason why CI is getting more complex these days? Is it because agents are just adding a bunch of tests or maybe people are YOLO writing their workflow YAMLs?
Daniel Worku: I think all of the above. The amount of code being generated has grown exponentially, right? So GitHub themselves, I think, shared that they've seen a 4x growth in GitHub Actions minutes from 2023 until, I think it was like April when they published their report. And if you're on the cutting edge, right, so Anthropic recently published, I think one or two weeks ago, they published an article about how CI became their biggest bottleneck as a company. Right. This is a firm literally on the frontier. They had 25x CI job volume growth in six months. Their engineers were pushing on average like an order of magnitude, like 8x more code. So that's one. The amount of PRs, the amount of code coming out is more than ever. I agree. I think another piece of this is agents love to write unit tests. And this may sound heretical as the co-founder of a CI company, but I actually think most unit tests are not useful. Right. So I think there's another aspect of this, which is we just gotta steer the agents to produce high-value code verification, right? Like a single end-to-end test where you actually run the full feature, I think is much more valuable at catching a real production regression than you know, like 200 unit test cases that can easily be vibed and changed. They're basically like a tax every time you write code. And then I think probably like another aspect of this is people are shifting to stacked PRs, which I think is very healthy if you want to review a lot of code really quickly. The downside is most CI systems are not designed to do incremental caching. And so if you have a stack of, let's say, 10 PRs, your CI job is just gonna do 10x of the compute. And I think that's probably a third factor that's really growing this bottleneck for you.
Shane: I would say that seems like one of the things we really disliked about stacked PRs. I mean, there's some good things, but a lot of bad. And one of them was every time a change came where you had to like restack, it'd run CI so much more often. And our CI, you know, as much as it's significantly faster than it used to be, it still is one of the bigger bottlenecks. Yeah.
Abhi: Yeah.
Shane: I mean it used to be really, really slow. Now it's just, you know, it's faster, but still kinda slow, you know? Yeah. So it's...
Daniel Worku: Probably we also had the problem of queuing, right? If I remember right. So Mastra, you had free minutes from GitHub, but they capped you at 50 concurrent runners. And so you had these queue times of like, you know, 30 minutes or an hour to run per iteration on a PR.
Shane: Yeah, and we, you know, we have, you know, with us having our own Factory, which does a lot of you know, fixes a lot of issues, we get a lot of PRs from that. Obviously the team's shipping a bunch of PRs all the time. We have a monorepo, so that causes conflicts as the PRs are all going to the same repo. We've done a lot of the tricks, right? All the things you need to do to help speed it up. And obviously StarSling's helped us out a lot as well. I mean I think at one point it was like, I don't know, it wasn't quite an hour-long CI time, but it felt like an hour at some times. And now it's, you know, significantly faster.
Abhi: I think it was like 30 minutes or so. It was. But then you find a small change and you just wasted a whole run. You know, you're finding changes, your agent keeps updating it, responds to CodeRabbit or whatever. You're paying the tax again and again and again. And you know, for some projects it's really hard to find the dependency graph of what you need to change, what you actually should run in CI. It almost becomes like a full-time job just to figure out how to run the CI efficiently, you know, and that's not something most people want to do. They just want to run some shit and get their PR out there. Exactly. Which is why we don't want to solve that problem. And I'm glad StarSling is, for sure.
Daniel Worku: Yeah, exactly. Right. It's just time consuming, right? And no one wants to do it. I mean we do have a bunch of open source skills, so you can optimize your CI. It's definitely a tractable thing that folks can do. But like you said, Abhi, you want to spend your time shipping. You don't want to spend your time like, oh, let me check in on how the A/B test went for my optimization to shave off, you know, like a half a percent of a second or something.
Abhi: Yeah, that's why you just auto-accept the StarSling PR and then it'll be marginally better, 'cause they'd already proved it out for you.
Shane: As we are just shipping so much more code, anytime we can decrease CI by 30 seconds, a minute, that's huge. I don't even know what our CI times are now. I know... Yeah, they're much better than they used to be, but they still feel long, right? Anytime you have to wait.
Abhi: Which is tremendously better.
Daniel Worku: I think your slowest job is like eight minutes, and then because it's a multi-stage pipeline, that's how you get to the full 16:10. So we're working on crushing that down. So, crazy idea here, right? But I think in the future, let's say that agents are the only ones writing and reviewing PRs. Because if you look at where we've gone in the last two years, it seems like we're heading to a world where engineering is gonna be much more of like a spec and strategy level activity rather than the nuts and bolts of a PR. So, a crazy idea is, I think in the future there's probably not even going to be CI as we know it. It will probably be some sort of agent-driven thing where you just ask, you know, like StarSling or another provider, is my code good? Like, is it verified? And then you don't have to run all this compute and run all these jobs because you have this like super smart, you know, like IQ 150 agent that is only responsible for telling you like, yeah, the tests would have passed if we ran all of them.
Shane: Or it's smart enough to know that anything it wasn't sure on, it could run those specific tests, but it doesn't need to run everything, right? It's like a hyperintelligent engineer where if you think about a codebase where, you know, back when we used to look at code, you really knew. You would know. If, especially if it's a smaller codebase, you'd probably know if it's gonna pass or not, right? You just knew if you were in it. As the codebase gets larger, of course that breaks down, and that's why you need all the CI, because you don't know every aspect of different parts of code. But if you can imagine an agent that could know all the aspects, well then it would be smart enough to know which tests need to run. What's the risk of merging this without the full test suite? How do we get this to just be merged quickly, right? How do we speed up the merge so we don't have to wait and rerun the tests 15 times as other PRs come in?
Abhi: Exactly. And you know, the thing that we're paying, we're paying the cost right now as running a Factory, which is we're running CI on the PR, then we are also running our own review agents on the PR, and that infrastructure is not shared, but reviewers need to verify things and run code to check and things like that. And so now you have these two disjointed systems, one that's maybe already cached and efficient. One is a sandbox that just got spun up for the review agent. Is that kind of your goal for the StarSling Review? Did you want to share the infrastructure there?
Daniel Worku: Yes. So we just launched StarSling Review Runners. And so yes, sharing the infrastructure is part of it, so that you don't have this separate system where, like you said, you're gonna check out the codebase, you might need to run tests during the review, or you wanna make a code change when you're proposing. Most reviewers now have proposed edits and you want to verify those are good. So yes, sharing the infrastructure between both is part of it. And that's what's powerful about starting at the infrastructure layer, not
just building an agent: we can expand into all these runner use cases. But the big thing for us was we had a lot of customers come back to us and say, Hey, I love what you're doing for CI, but my biggest bottleneck now is code review. I'm spending a ton on all of these hosted code review solutions. You know, these are folks that work at companies that are already starting to shift their coding from Anthropic and OpenAI onto open weight models. And they're like, I want to do the same thing for my code reviews. Right now I'm using all these hosted platforms. How can I do that? And so that's the second thing that we're really going after with code reviews: we think that you should be able to bring your own model. If you wanna, you know, pass in your keys, we're just a runner. We can give you the harness to help you do the code review, but give you the freedom to define the skills, define, you know, is there a pattern that you want to match to only run this subagent on, you know, this package for these reviews, stuff like that. And so that customization and giving people ownership over their code reviews, that's the other aspect that we're really going for, right? Because the hosted code review runners, they're a black box. You don't know what model they're using. Most of them don't expose the memory. So they're just learning and getting better at your codebase and locking you into their ecosystem, right?
Abhi: Also, I mean some of them are quite slow, right, to perform, because they have to boot up the world. Some of them have to run and validate the stuff that you're doing already in... You know, and if you add on another review agent that needs a sandbox, you're just doubling the amount of compute and infra you're spinning up to do the same thing. Yeah, so I mean I'm really excited about it. We'd love to try it. I think we can even hook this into our Factory. So our review sessions that we're running in our Factory actually start in a StarSling Review Runner, you know, like those are the cool things that you can do once you have access to the primitives, right? Exactly. Yeah.
Shane: All right, Daniel, well we appreciate you coming on the show and talking a little bit about speeding up CI and what CI looks like in you know kind of the age of AI. If someone is interested in learning more about StarSling, seeing what you're doing, where should they go? What should they do?
Daniel Worku: starsling.dev, you can get started for free. Dude, did you subscribe?
Shane: Dude, I host the show. Did you subscribe?
Daniel Worku: Did you subscribe?
Shane: Subscribe to Agents Hour every Monday, noon Pacific.