Back to all episodes

Should You Train Your Own Model?

September 28, 2026

Should you train your own model? Maybe — but probably not yet. If you haven't nailed your prompts, workflows, and evals, training won't fix your agent, and it only starts to pay off once you're running millions of tokens a day. In this episode, Osmosis co-founder and CTO Andy Lyu breaks down when it's worth it and how to do it. Andy covers the two main techniques for training your own models: supervised fine-tuning (SFT), which teaches a smaller, cheaper model to copy a stronger one using thousands of good examples, and reinforcement learning (RL), which rewards an open-weight model like Qwen, GLM, or Kimi when it gets a good result so it figures out its own path there. The hard part of RL is the reward — score it badly and you get reward hacking, where the model learns that doing nothing is the safest way to avoid a penalty. They also dig into RL environments and Harbor, unit tests vs. LLM judges as verifiers, how much data you actually need, why good evals are worth having even if you never train, and where to start learning on your own.

Guests in this episode

Professor Andy

Professor Andy

Osmosis

Episode Transcript

00:00

Professor Andy: Training is not something we would recommend until you've exhausted every other option to improve your agents. That includes things like prompt optimization, having better workflows, having evals. If you don't have any of these set up, quite frankly, you are not ready for reinforcement learning. We are not even ready for SFT. You are still in that builder stage.

00:21

Shane: We're gonna be doing something that we like to call builders learn. ML because you know if you know anything about Abhi and I, you know that we're not. I wouldn't say formally trained in machine learning. But we've picked up quite a bit over the last few years and some of the reasons is we've been we ha make made friends like Professor Andy who has helped us along.

00:40

Abhi: Yeah, so super stoked. It's been a while since Professor Andy is Graced us with his knowledge.

00:45

Shane: Yeah. And so for those some of you who've been around some of the OGs, you might have seen him. I don't know, it's probably been almost a year since. Yeah. He was on the show. But for those that haven't, you're in for a treat.

00:56

Professor Andy: Hey guys.

00:57

Abhi: Professor!

00:58

Professor Andy: It's been a while.

00:59

Abhi: For the new viewers, why don't you introduce yourself? Sure.

01:02

Professor Andy: Yeah, so I'm Andy. I'm the co-founder and CTO of a company called Osmosis. The high level what we do is we help companies train their own AI models. You know, many companies come to us and say that Hey, we're spending a ton of money on OpenAI and Anthropic every month, but our agent use case is just a routing use case or a customer support use case. So what we do is we help them take their data, take their infrastructure, their harness and then create a model that is more predictable, malleable, and as well as more performant than if they were to use off-the-shelf models.

01:31

Abhi: That's very cool. Let's start from the very basics again. When people say train your model, what does that even mean?

01:38

Professor Andy: So right now there are two common ways to train your model. The first one might be more easy to understand. It's a technique called supervised fine-tuning. The idea is You want to gather a bunch of traces or gather a bunch of agents doing work that you know is good and then take those trace and teach another model which may be weaker but it's smaller and cheaper to do the same work. So that is supervised fine-tuning. You're essentially asking a model to mimic a more powerful teacher. The other thing that we like to talk about in the industry is called reinforcement learning. This has the potential to really excel at you know what the frontier is currently i capable of so i think in a news story later you're you guys are gonna talk about Ox alpha and related models. These models are mainly trained with reinforcement learning and in fact most of Modern day LLMs like GPT-5.6, Claude. Opus 5, these are also improved with reinforcement learning. The TL;DR is inspired from nature. When when you touch, you know, a really hot stove or when you do something that's, you know, not the best of judgment, your body will tell you, hey, don't do that, right? Like if you touch a hot stove, your hand will pull back. That's instinctive. Because there's a penalty from the environment around you. And similarly, when you eat food, you know, when you go to bed, these have a reward, right? Like you feel better after doing so. If you're really hungry and you have, you know, a burger right in front of you, you're gonna feel good eating that, right? So to the machine, this is kind of the same thing. We want to reward the model for doing the correct action in an environment. And similarly, if they do something that we don't want them to do, we penalize them. We scale them up to hundreds of thousands of these similar interactions. And over time the model learns a good idea or a good strategy of how to solve really complicated questions like coding or office work.

03:21

Shane: Has the process, I guess, changed? Over the last year on how you do this reinforcement learning. I know typically people think, you know, I'm just taking a step back from the outside looking in, this sounds really. Challenging. When would I actually want to train my model? How would I make the decision of when to do supervised fine-tuning or reinforcement learning? How much data do I need? When does it start to make sense? Because Yeah, obviously there's a cost barrier to, you know, you might be sending a ton of tokens to a larger model, but there's also a cost and a time to actually do this training process.

03:54

Professor Andy: Yeah, so for training for agent builders, it is not something we would recommend until you have sufficient scale. And what I mean by that is You've exhausted every other option to improve your agent until you reasonably can't anymore. That includes things like prompt optimization. Having better workflows, you know, having evals, like being able to evaluate your agent to see how well it's doing in production. If you don't have any of these set up, quite frankly, you're not ready for reinforcement learning. You're not even ready for SFT or supervised fine-tuning either. You are still in that builder stage, right? To get your get your harness off the ground and get your you know agent product up and going. In terms of data scale, let's say you reached millions, hundreds of millions of tokens per day. That's the skill that you would need. Then you it may can you can probably think about starting to collect samples from what your agent's doing in production and using those traces for supervised fine-tuning. You typically would need anywhere from 10 to 20,000 examples, like good examples, to really have an impact there. So Supervised fine-tuning is usually quite a bit more data hungry compared to other algorithms. But it is easier to set up once you have it going. For reinforcement learning you would need a different unit of data. You would need a thing called an RL environment. The environment is essentially your agent in a sandbox and basically in some compute it can execute itself in and then some task to ask the model to do. So if you're creating an agent that answers deep research questions, well ideally you should have a bunch of environments with you know interesting deep research questions as well as the correct answer alongside so you know how to reward or penalize the model based on how it did. That's kind of the high-level overview but the general idea is that if you just started building your agent don't even worry about it don't even think about it it's only when you have a ton of users using it on a daily basis. And they realize that, hey, it's not no longer sustainable for me to pay OpenAI, you know, six figures every month. That's maybe where you want to think about, you know, training your own models.

05:48

Abhi: Let's say you have the trace volume and you have the need. Should you always just do SFT? First, is there like a rules to the trade here?

05:56

Professor Andy: Well the real answer is that it's it's a very research-based process. Most folks try SFT first because that is the easiest thing to try. There's a lot of platforms already that allows you to upload a trace and then just perform SFT, right? But typically. SFT doesn't work that well because you are limited by you know, your data. Let's say you only have traces that tackle one particular problem, but your agent product needs to solve a whole suite of problems. Well then the model you train would be really good at that one problem, but it may actually devolve or degrade on the other problems you care about. So what kind of data you collect and how you define it is also very important. I was recently in China talking to a lot of the Chinese researchers from MiniMax, you know, Z.ai and Alibaba. And they were telling me that the researchers spend fifty percent of the time training and then fifty percent of the time actually cleaning that data, like fine-tuning that data. So that data there is really, really important. If you want to make a model that actually works.

06:51

Abhi: Does there need to be a distribution of the examples within the trace? Like let's say let's wear make it very simple. Like let's say I have fifty examples. Forty of them are one class of problem and the other ten are the other class? Does that mean the trained model is biased towards the example data that it has? Yeah, totally.

07:10

Professor Andy: Your model will be heavily biased and probably decent at doing the task at forty examples. But it will definitely underperform or even regress on the ones you have ten examples for. And it will definitely regress on the use cases that you have zero examples for.

07:23

Abhi: So what's like the best distribution? Should they all be even in your set?

07:27

Professor Andy: Yeah, so there's two ways of thinking about it. One is the safe bet, which is keeping your distribution even. The other other way to thinking about it is if you have the ability to actually monitor your use case, like what people are using your models for you can model that distribution. So let's say in production, 70% of your users are using your agent for one thing, 30% of users are using it for another thing. That could be your data distribution. That's actually okay. But ideally, if you don't have that information, try to keep it balanced if you can.

07:54

Shane: So when you think about reinforcement learning and then kind of how it's working under the hood. I know you have these RL environments, you have this kind of idea of some kind of reward function, or but I is it the idea that you don't have to have every trace of every possible scenario coded in because if you get the right environment and the right reward, the agent can kind of learn it s on its own or figure itself out. Is that roughly how it works? Can you go into more detail on that?

08:19

Professor Andy: Exactly. Reinforcement learning is a much more flexible algorithm and technique. Because you're giving the model freedom to do whatever it wants because you don't care about what the model thinks or how it got there. You just care that the result is good. And supervised fine-tuning. You're actually forcing the model to think in a particular way. So if you're training an open source model like Qwen, but you're giving it traces from Anthropic, Anthropic models think and talks in a different way than Qwen does. So what you're actually teaching the model is how Claude thinks plus how Claude acts. But in reinforcement learning, because you're allowing Qwen to kind of you know thrive in its natural habitat it is only learning a new skill without having to learn a new dialect of English or how to talk.

09:01

Abhi: Yeah, so you're exactly right in that case. And typically when people do model training and they want to get frontier performance, are they their base model is an open source model? Is it is that a requirement?

09:12

Professor Andy: Yes, today it is a requirement. It used to be a base model, but we have we've actu actually had pretty good success training frontier models with current open source models. If you look at the recently released GLM 5.3 or the Kimi. K3 models, they're actually very, very close to the frontier. We've done training runs for customers on those model scales as well. And we've been able to show that these models can outperform frontier models. In addition, for very specialized use cases, smaller models like Qwen 3.8 27B. Which is only a fraction of what Opus or Sonnet is, but with training, it can also outperform them by quite a bit. So yeah, training frontier models will require open source. But that doesn't mean that it's impossible. It's actually quite the opposite.

09:52

Abhi: Very cool. So how's Osmosis doing these days?

09:56

Professor Andy: Yeah, we've been r really, really busy. I think from last time I was on a show. The problem persisted, but it's larger in t in the sense that compute is harder and harder to get. We recently onboarded the data center in Cape Town, South Africa. Just because of how limited the resources are. But yeah, it's been good. I think we are busy in the sense that a lot of people are asking us to do training runs of You know, trillion parameter model scale, which is very resource intensive, but at the same time is a good problem to have because you know you literally cannot meet demand. Right. What's better than that?

10:26

Abhi: Last time we talked, you were trying to do like a self-starting product where you know y you would not need like too many FDEs or anything. Was that still the plan? Has that has that kind of Formed now.

10:38

Professor Andy: I think we realize that with the current compute situation, it's unlikely that we can support the self-serve kind of situation. We are actually working on something where if the user comes in with a Modal API key or with a base 10 API key and they have the requirements, we can actually directly plug into their compute. So this is something that we are working on for the self-serve motion. In terms of B2B kind of sales process, we have simplified it down where we have pre-built templates and pre-built executions. That you know an AI agent now can spawn in you know trillion parameters training runs and actually get a good result. So that's been very exciting so far. Can

11:14

Abhi: we Talk more about oral environments. Sure. You know, when you said it's just a sandbox, right? But that word sandbox means a lot of things these days. Right. Yeah. And it really depends on what the agent's task is. So what are some examples of RL environments that are maybe more practical?

11:30

Professor Andy: Yeah, let's get more concrete here. So One of the most popular frameworks that most Frontier AI labs use today and most data environment companies use today is called Harbor. Harbor framework you If you guys have heard of terms like Terminal-Bench, you probably wouldn't know this. There is the same organization behind Terminal-Bench that created this RL environment framework. What that means is it runs in, like I said before, a sandbox, and you're defined in a sandbox more concretely. You can think of it as a virtual machine with some CPU, some memory, and some disk. And you can pre-initialize it with whatever software you want. So you can load in your agent harness, you can load in some data, basically whatever you need. And that's your RL environment. The environment itself is some compute plus some software plus your agent harness. And then Harbor helps you orchestrate all of that. Now the task itself is going to be what you want the model to do. So in a concrete example of a RL environment could be, let's say you load up the environment with a code base, you know, some open source code base, like a Linux kernel or something. And then you load it with all the you know. C compiler tools or all the tools you need to do the kernel development. And then you will ask the model, hey, I want to implement this new feature in this Linux kernel help me do it and validate your findings. Harbor will then launch that run into the virtual machine and let the agent do whatever it wants. And after it finishes you have an artifact solder, right? Because the agent's gonna make some code changes and then you come out with a new repo. Then there's a stage called the verifier, which is what you mentioned before at the reward function. The verifier can consist of two types of verification. One is if you pre-write unit test, right? So If you guys are familiar with test-driven development, this is basically what that is. I write a bunch of tests, my developer, i. E. The model, goes out and do it and then my test to see if it passes or not. And based on how many tests it passes, I give it a reward. This is a verifiable case. Another case where it 's less verifiable is OpenAI has a very popular dataset called GDPVal. GDPVal is I believe it's a hundred tasks of economically viable or valuable task. Ranging from you know tax calculation to finance accounting to understanding schematics. These are very like difficult problems that don't have a ground truth. There is a general right idea, but it's really hard to say that if you got it right or not. So then we would use something called a rubric verifier. A rubric verifier is really you want to ask another LLM with a correct question to grade the LLM that did the task. So if you're calculating taxes for a company, you would have a rubric to say like, hey. Did the model account for these deductibles or did the model come out with this tax amount? And because the judge model has the correct answer, it can very easily verify that result and then give you an answer. That's the whole anatomy of our environment and some RL tasks.

14:21

Abhi: Totally makes sense that it's a virtual machine, but if you're RLing browser agents, is the environment a headless browser that It's gonna be doing the tasks on?

14:31

Professor Andy: It doesn't have to be. Right now, sandboxes, because it's just compute, you can actually install a Linux operating system with a desktop manager. With an actual chrome instance if you want to connect it just via like the simple like chrome cdp extension you can in that case it will be like a headless way. But now there's frameworks that allow you to use computer use on top of these browsers. So you can either stream a video directly to your agent or take screenshots to your agent or let your agent decide when it wants to see the screen. This is all very flexible here. It's really up to you and how you want to do it.

15:06

Abhi: And in these cases, like the app itself needs to be like a sandboxed kind of app, right? Which gets challenging, let's say.

15:14

Professor Andy: This will get challenging if you have many applications, especially because sandboxes are typically not very resource rich, like They're typically in a range of two. CPUs, four gigs of RAM, you know, like your low-end device kind of thing. If you're using Chrome, like the full Chrome, it could be very memory hungry, right? So you will need your sandbox to be much much larger. Yeah, and that becomes challenging at scale because if you're training them, you can imagine during training we spin up thousands of sandboxes concurrently. Multiply two. CPUs by a thousand, that's already two thousand CPUs, right? If you need more, especially on the memory side, with how expensive RAM are these days. Yeah. It's gonna get expensive for sure.

15:51

Abhi: And then, you know, there may not even be these mock environments for what you're training against, like Let's say you wanted an agent to use Amazon. Com.

15:59

Professor Andy: Yeah, I think to use the simple browser use cases, there there are frameworks built out to help you. That's why I recommend Harbor again. There's a whole suite of RL environment companies that build these environments for Frontier Labs. All of them kind of congregated on this Harbor format as a standard. So a lot of people I have already built out. Very helpful frameworks, very helpful bootstrapped environments with things like, you know, Chrome browser pre-installed or some other software we need pre-installed, which makes it easier for you. Dang.

16:29

Shane: So I do have questions on so it sounds like you If you do plan on doing reinforcement learning, you need the environment, right? You're gonna need a lot of compute, obviously, but you also need to design this thing called a reward function. Right. And I feel like there's probably some nuance and trickiness to getting that right. If the reward function is good, I imagine the results are gonna be better.

16:49

Professor Andy: Is that right? Yes, yeah. So there's one really common problem in reward design called reward hacking. This is a behavior we've seen in real-life production use cases. Let me give you an example. So I mentioned before if the model is something bad, you want to penalize it. So what some folks we've seen have done is let's say the base score is zero. The model does nothing, it gets a zero score. If the model does everything perfectly, it gets a score of one. But if the model hits a lot of penalties, like it keeps messing up, it keeps doing the wrong thing, you can potentially get a negative score, right? So it can you can potentially go on to negative one? That's the range of scores you have. Now, this sounds reasonable in theory, but what actually happens in production or during training is the model realize, hey, as long as I don't make a move, I won't get penalized. Right? So I'm happy with a zero because if I do anything, I have a potential chance of risking for a penalty, which gives me a negative score. So this is the example of reward hacking. Our intention is for the model to learn itself to, you know, the 1.0, the perfect score. What the model realizes halfway through if it's just lazy and does nothing at all, it at least gets a zero. It won't get a negative one, right? So that's an example of reward hacking and which is poor reward design. Right. And this happens really more than you think in even in very mature organizations.

18:02

Shane: Can you describe what a re typical reward function is? Would be maybe for any kind of example y you've seen or looked at. So this is just like a simple function that says, you know, one or zero or negative one. Or how how complex do these reward functions get?

18:17

Professor Andy: Yeah. The output of the reward function you can think of it as a scalar between zero and one. There's actually no hard requirement for the number, but typically we just do zero and one. Because in training, that's the range we normalize it. So you even if you give it a hundred or even a thousand at train time we will normalize that down relative to Other scores into that zero and one number. So you might as well just put give me that number in the first place. The reward function typically consists of a verifiable reward and a rubric-based reward. A verifiable reward is like I mentioned before. Very simple unit test, right? That you can just run once on the CPU, give it the basically the trace of the agent and then the output of the agent, and then this code will run and tell me how many units has the agent has passed. That's very simple, very cheap. The other one is more up to design and is more dynamic. It's the rubric design. The rubric design is essentially you having the correct answer in a bullet point list and you're asking another AI model to go and look at the agent that you just evaluated. And then tell you how well that agent did during its execution. Like for financial tasks, it could be, did the model, you know, account for the correct pitfalls or did the model account for the correct loopholes in your contract? If the model found it, your LLM judge will know because you already told the judge what the correct answer should have been. So it's usually a combination between these two.

19:34

Abhi: So I'm starting to kinda get this at least to my understanding. The big blocker for SFT is the amount of volume of traces. And so if you don't have that, you shouldn't be doing it. And for RL you should think about your agent more sophistic in a sophisticated way and really think about how to judge it. And once and that's not something you just do overnight, right? You have to really put effort into that as well. So there's no quick win in training.

20:01

Professor Andy: That's why I recommend people to really exhaust the prompt optimization layer and, you know, the agent design layer. The iteration speed for these techniques are relatively quick, maybe a few minutes to a few hours at most. But then if you're doing training, this is a week-long, month-long effort. But then if you change your harness midway, you may have to restart your training progress. So that's just not something that's worth it.

20:23

Shane: It does sound like if you have a really good eval suite it's at least helpful if you do ever kind of graduate out to when needing to do this training process, right? So it's like you have a well defined idea of what does a good result look like, what does good agent responses look like. And then I'm assuming that helps feed into the the judge function of the reward function or like designing the reward function to optimize for the right behavior.

20:49

Professor Andy: Yes, having good evaluations for your agent is incredibly helpful if you were to train your model later on. And I will argue even if you don't plan on training, having good evals is still very good for you to understand. If you want to use a new off-the-shelf model. What we realize is that even as newer and newer models come out, and you know the benchmarks are all going higher and higher. They may actually not be a good fit for your own use case. We've seen really large companies, you know, just a few months ago still want to stay on GPT. Because GPT five and later iterations have actually regressed on their own internal evaluations. So evaluating models will be a really interesting problem right now and going into the future. And having good evals is not only good for training but it's also good for you right now to understand where you're at.

21:33

Abhi: Do most of your customers that come in have this good data sets or are are they struggling to do this as well?

21:41

Professor Andy: I think everybody is struggling to do this, apart from the data companies themselves, but even the data companies that we work with, work with us to, you know, fine-tune their reward pipeline. This is a very difficult thing and I think this is one of the most AI proof tech job honestly like how to design good reward functions. We've tried using a lot of AI agents to help us see if we can generate synthetic rubrics, but we've noticed in training. These rubrics are always almost too noisy and it's really hard to get a good result with them. So we help companies fine-tune their reward or we help them build out their evaluation suite because a lot of these agent companies come to us without having a robust enough eval. Most of the eval is the eye test at the moment.

22:22

Abhi: Yeah. I can't imagine anyone that has really good data hygiene right now that then just wakes up and wants to train models right away.

22:30

Professor Andy: Yeah, it it usually takes a little bit. I think the industry moves so fast right now that 's. You can't really afford to sit down and really eval everything. It's just go, go, go, right. Like you're trying to add new features and new things. And you still just trust the LLM to be this magic black box. And you know, it can get you pretty far. But Eventually once you scale up, and then you have to, you know, sit down and look at how your agent's actually behaving.

22:51

Shane: Andy, we appreciate you coming on the show. What's the best way for if people are interested in I guess two questions here. If they're interested in learning more and they want to just from the you know, the scientific perspective, maybe consider doing this at fr on their own in the future, where would you send them? Like what's the best resources that that you can say for someone who's just looking to get into this? From like a builder or engineering perspective.

23:15

Professor Andy: For consumer grade, you're probably not gonna train production models, but you learn. I think Hugging. Face has a really good resource on this. So you can look at the Hugging Face handbook on reinforcement learning. That should give you a pretty solid like start on what RL is and how that field has evolved. And to run. Simple RL experiments on your own, you can use frameworks like Unsloth, which makes it very easy for you to use consumer grade GPUs to run the same algorithm on you know like a small, small, dense model like Qwen 3.5 four billion parameter or twenty-seven billion parameter. But for more mature like enterprise level model training, if you want to get into like Kimi K3, GLM 5.3 or Deep. C V4. Pro. Yeah, you can feel free to contact us. We’re at Osmosis.ai. Ai. Awesome. Good seeing you, dude. Yeah, good to see you guys too. Dude, did you subscribe?

24:04

Shane: Dude, I host the show. Did you subscribe?

24:07

Abhi: Did you subscribe?

24:09

Shane: Subscribe to Agents Hour every Monday, Noon Pacific.