Back to all episodes

Why Your AI Agent Gets Brain Damage, and How to Fix It - Tyler Barnes, Mastra

September 15, 2026

Tyler Barnes is a founding engineer at Mastra and the person behind its memory systems. In this chat with Shane, he explains observational memory — Mastra's fix for the thing every coding agent does where it fills up its context, compacts, and forgets what you were doing. A background agent watches the conversation and writes short observations, which replace bulky tool calls and messages in context, so you keep what matters and drop the noise. It scored state-of-the-art on the LongMemEval benchmark. Tyler also covers recall, extractors, subconscious memory that builds a knowledge graph, and agent signals for steering an agent mid-task.

Guests in this episode

Tyler Barnes

Tyler Barnes

Mastra

Episode Transcript

00:00

Shane: You're working for a long time, a long session, and all of a sudden you hit this imaginary window that then now causes this massive compaction event and you lose context and you feel like you have to just like restart the session.

00:10

Tyler: A lot of people have hated compaction just because it's like your your agent, it's like they got brain.

00:19

Shane: I'm with Tyler, founding engineer at Mastra. And we're gonna be talking about memory, specifically observational memory. We're gonna talk about agent signals, but first, you know, a lot of people have probably seen you before, but it's been a while since you've been on the show. What were you doing before Mastra? And then what was, you know, what are some of the things you've worked on since you've joined Mastra, which you know was I mean you joined when we were in YC.

00:40

Tyler: I joined Mastra, like you said, early on. I had been working on like my own agents and stuff, and I think you and I would get on, you know, calls just to hang out sometimes and we kind of demo stuff to each other. I was getting pretty s pretty excited by the stuff you guys were working on. So

00:54

Shane: I remember you telling me it was very it was very early, but you had you'd basically wired up a camera. Where if your cat came, it would detect that this is a cat, so feed it. And if your dog came, it would detect that it wasn't a cat, don't feed it, or something like that.

01:10

Tyler: It would play like an alarm to scare the dog away. So it was always eating the cat food.

01:14

Shane: So you're using like some AI there to like detect. It was

01:17

Tyler: like GPT-4o or something, you know.

01:19

Shane: Yeah, I think it was before that actually, but yeah, yeah, but it was somewhere in that range of models. It was, yeah. Yeah. We always talked you know, just models and what we were building. But then you joined Mastra and you've obviously worked on all, you know, all different areas for sure, but you've specialized in A couple.

01:37

Tyler: Yeah, I guess memory was the first thing.

01:39

Shane: So tell me a little bit about the first version of Mastra Memory, because we've gone through some iterations now. There's still a lot of people that use the initial memory system that we built.

01:49

Tyler: Yeah, there's quite a few people. I think so it started with there's three types. Maybe we just started with like the message history, very simple, you know, the last. X number of messages like ten or twenty or however many you want. And then we added working memory, which is sort of the agent can update a tool or Use a tool to update a chunk of context so over time it can kind of keep track of something.

02:08

Shane: Did that context basically just like sit in the message history or sit in the system prompt? Or how did the context actually what was the underlying mechanism to make that happen?

02:18

Tyler: It did sit in the system prompt, which is not great for prompt caching. We've actually fixed that since with a newer version of memory. But

02:23

Shane: when we built that, no one cared about prompt caching. That's true.

02:26

Tyler: It wasn't the thing people were talking about. Yeah. I think it it existed, but not all the providers even had automatic prompt caching at that time. It was like you're a lot of the time you're just paying uncached prices all the time. So

02:37

Shane: you hadn't. Message history, working memory, what else?

02:41

Tyler: And then the other one was semantic recall, which is just rag. So every new user message and assistant message you basically do a rag query and then insert into the system prompt again with some relevant context. At the time we ran LongMemEval on it and got like a really high score and it was it was somewhat controversial because it was like a oh you you just use rag and you can score very highly on you know this benchmark. We've since gotten a much higher score, but it was quite interesting that such a simple system could work so well.

03:08

Shane: Yeah. Huge message history and only insert the parts that mattered so you'd actually have less context which at the time you know for folks that was great. You didn't send as much context. It was cheaper. You had the extra lookup, of course, so maybe you had some extra latency, but you probably made up that latency because you're sending less tokens to the LLMs. But over time we've realized that maybe there's other approaches that could even be you know even be better, especially as prompt caching became more prominent. Yeah. Walk us through like how did you go from okay, this first memory system, which was pretty good, and arguably there's still a lot of memory systems that use RAG today, but what was the next iteration?

03:48

Tyler: I guess we Sort of had like quite a big jump into the next one which is observational memory. I'd been doing a lot of experimenting with coding agents and prompt caching with coding agents is very important. They just eat tokens like Just nonstop calling tools and you know reading big files and things like that. So the memory systems that we had really didn't work very well for that use case. So through a lot of iteration, you know, just trying things out. I did a I eventually came up with observational memory, which is a prompt cacheable system. We ran LongMemEval on that as well, and we got like state-of-the-art at the time. So that was like a very big jump as well.

04:25

Shane: Yeah, and can you tell people so we talk a lot about benchmarks on the show, but can you tell people a little bit about LongMemEval and I know we want to talk more about observational memory of course, but what is LongMemEval and is it a good benchmark?

04:37

Tyler: Some people don't like it. I think it's an okay benchmark. We actually don't have a lot of great memory benchmarks, but I think it's quite good at testing the recall for you know a single turn for agentic use cases. You really want to be able to test across many turns, and you want to be able to test prompt caching, how well the memory system can guide the trajectory of an agent as it's working on something. We're probably gonna end up, you know, running out some more benchmarks in the future. LongMemEval is a good one and it does test some decently sized conversation histories. So s you know, semi real world.

05:08

Shane: Can you tell a little bit more about how does observational memory work? How does it what's going on under the under the hood? How is it prompt cacheable? How does the system actually work over long conversation periods so you know you don't lose context? Because I think before one of the frustrating things with coding agents, and they've gotten a little better, but some still suffer from this is this idea of you're working for a long time, a long session, then all of a sudden you hit this imaginary window that then now causes this massive compaction event and you lose context and you feel like you have to just like restart the session. But how does observational memory differ?

05:44

Tyler: A lot of people have hated compaction just because it's like your your agent it's like they get brain damage as soon as it happens. I think with Codex it's gotten a lot better, but it's still not quite as good as you know like a better memory system. Observational memory is sort of you know it's almost like a hybrid of compaction and a better memory system. So as your agent is working in the background there is an observer. So this is another agent which is taking in all of the turns and it is creating condensed observations of what happened.

06:10

Shane: So this is running kind of in parallel to if I'm talking to an agent, there's another agent that's just watching the conversation.

06:17

Tyler: Exactly. It's sort of buffering these observations in the background. So each chunk of observations maps to some set of messages in the conversation history. And they just continually build up until you hit a certain threshold. And then those messages get replaced with the observations. So you're very, you know, token heavy tool calls and messages. Suddenly get replaced with a very dense representation where the information is not lost. What you lose is really the a lot of the context rot. You know, the things that didn't matter contextually to the conversation.

06:49

Shane: So you have this observer, right? And then what happens if it continues to grow even past that? Does it just continue to run? How does it know that it can basically go forever, right? I think that's part of that's one of the benefits of observational memory. You know, one example is I have a an email agent that has run through at this point tens of thousands of emails, right? And I I've kind of steered it to what I want and it keeps track of the decisions I've made in some of the context, but it doesn't. I can just use that same memory, you know, I've been going on like three months in the same like I

07:21

Tyler: Just

07:21

Shane: never changed.

07:22

Tyler: Does it run on a cron or something? Or is it when emails come in?

07:24

Shane: Every day I basically will just like run it and it'll just like run a script that processes all my emails. I'll eventually I'll put it on. We we have Mastra schedules now. I'll put it on a schedule. But it this is before schedules existed. So I just had it, you know, I just have a script that I run. But it's uses the same thread. It's used the same thread for you know probably four months, five months at this point.

07:43

Tyler: Yeah, so I think that is one of the the coolest parts of observational memory is that the chat just feels g like it goes forever. This background buffering means you never need to like stop and wait for compaction to happen and then suddenly the chat is much worse. The quality just stays sort of consistent. And you never notice the memory system doing anything. It's just sort of happening in the background and then swapping out these big heavy tool calls and stuff for with observations. I guess yeah, you just asked like how does that how does that actually like keep going forever? Like eventually you know these observations are gonna get so long that they won't fit in the context window. And that's where we have a second agent, which is the reflection agent. And at a certain you know token size of observations you will go ahead and look through all of them and figure out what are like the overall like themes and the important parts of this and what are the things that didn't actually really matter that were observed and it will sort of condense that even further down. By default it actually condenses the first 50% of observations, which makes it, you know, even more sort of consistent feeling. So you get like Condensed observations, raw observations, and then the raw conversation history.

08:42

Shane: So then that's kind of how the context window stacks. You have condensed observations. Then like an another tier of the second fifty percent of observations that haven't been reflected. And then any new conversation history that comes in. Until it hits another level where observations run again?

09:02

Tyler: And then the observations are being reflected at token thresholds essentially.

09:06

Shane: And so what are the big differences? Because it's still not completely lossless, right? There still is the chance you could lose some information. But it does and if you use it, you know, and I've obviously use it extensively, it does feel better than just, you know, early compaction, right? Where it could get this huge context and then eventually just trim it way down and you felt like you lost a lot of information. So what makes it what are the things that the characteristics that make it better?

09:31

Tyler: We've sort of dogfooted it for like months basically before we released it or before we even benchmarked it or anything. We were just sort of tuning this prompt based on like vibes essentially, right? Like we were like manually evaling it through. Dog fooding it daily in Mastra Code actually, which was originally the first versions of that were created to dog food observational memory. But, you know, we just sort of worked on it, iterated on it till till we got to the point where it's like this thing works really well. Let's benchmark it. And then it scored really high. And we sort of released it at the same time that it r released the benchmark results.

10:02

Shane: Yeah, fun fact about Mastra Code, I remember like early versions of Master. Code and obviously it's changed. See it wasn't even called Mastra Code. It was something else. Reese or something that you were working on. It was just like a markdown coding agent that you're using to test this. And then those like those like early experiments turned into what became Master. Code, which was like because of all the context, encoding agents really good way to test that memory was working well.

10:24

Tyler: Yeah. Yeah, 'cause they just eat tokens, right? Yeah.

10:27

Shane: I know recently we've been improving observational memory a bit. You know, we released this, I think it was what, f? February maybe or something. February, March time frame when it first came out. What are some of the things you're excited about? What have we released since and what's coming, potentially coming next?

10:43

Tyler: A couple minutes ago you asked a question which I didn't answer, which is like, what are the downsides, I guess, of observational memory? And that sort of feeds into your question now. The new features that we've been releasing and that we're about to release sort of make up for any of those little downsides. A big one is, you know, over a very long period of time, eventually it is going to be get so compressed that you'll lose details. So one of the first things we added was recall. So that's again, that is just RAG, right? You're actually doing RAG against the observations rather than raw messages. So you save quite a bit of space by doing that. But the agent actually has a tool that it can use to search for something. You know, if it has a hint in its observations that there's some something happened in the past but it doesn't fully understand from those observations what happened, it can do a search intentionally. Which retains the prompt cache. And then it has some tools to sort of page back and f forward through like the raw messages at that point.

11:33

Shane: So it is you it is searching through raw messages but based on context that it sees in its observations? Yep. And so then it tool call comes in, it gets stacked on the end of the message list. So it preserves prompt caching, I'm assuming. Yep. And then the results get pulled in. So it's just like feeding it's almost it's like an agentic search, right? It's like feeding new results to the agent so it can kind of search its own history essentially.

11:56

Tyler: Yep. That's exactly it. Well that actually works quite well. We want to go even further than that. We want this thing to be like a perfect memory system. You know, you can just throw literally anything at it and it will always remember everything. That's the next thing that's coming. So we have two features. So one we did just release, it's a lower level feature called observational memory extractors. And what this allows you to do is provide some kind of schema, and then as the observer is running, it can extract some structured information out and then you have a callback you can do whatever you want with that. So we actually used that to fix working memory, right? Which earlier we just we said that was invalidating the prompt cache continually every time it get got updated. Now we can use the extractors. There's another feature actually we were going to talk about agent signals. It uses that as well. There's a combination of them, but the gist of it is that you can extract some data and do something with it. So you not only rely on, you know, you keep the prompt cache of the main agent, but you use that prompt cache of the observer to make a follow-up request and extract the data. So you're sort of piggybacking on the observer's cached context.

12:56

Shane: Awesome. So this would allow you to basically define certain types of information that might be important for your application. Later we're gonna talk with Jan about video production. So I'm building a video production agent. It might pull need to pull out types of settings that I use or types of clips that I want to take and it would then be able to extract information if I message something around those parameters and save that for later. So it gets pulled back into context as needed, something like that.

13:23

Tyler: Exactly, yeah. I think maybe like the simplest to understand example is that we have a thread title generation. As a observational memory extractor. So the observer is continually checking does the thread title title need to be updated and then it just updates it basically. So

13:37

Shane: yeah, so if I start talking about one thing and then I steer it in a different direction, it can update the title of the thread. So exactly. If I have thread history in a chat bot I can know that it's actually correct.

13:49

Tyler: Yep. Okay. That one it is a little more. It's sort of hard to explain to people. It's not that you know complex, but it can be a little bit hard to you know kind of mentally understand w when you would want to use it. But the reason that we're we've shipped that one is a upcoming feature which is subconscious observational memory. So this one's gonna be you know the actual memory feature that pushes OM a little bit further.

14:11

Shane: When you first told me about observational memory? And I think the last time you were on the sh show talking about observational memory, you said you kind of thought about like, you know, my brain seems to work very well, like try to make it human-like, and subconscious OM sounds even more human-like.

14:24

Tyler: Right. Yeah, I mean yeah, the r we didn't even mention that, but that was like the original inspiration, was just thinking like you know, as as a human is working writing code or doing whatever. You don't need to like choose to s to remember things or to save memories, right? It's just you're s have something in you that's observing. And I guess what you're hinting at similarly, right? There's you have a subconscious behind that even, right, which is sort of probably like keeping track of things longer term and storing things in different ways. And

14:50

Shane: I've seen other people, you know, say that agents need to dream and I think it's kind of a similar vein of something that's happening in the background at certain m times, right, where you can like pull out the right information. But tell me a little bit about subconscious OM.

15:05

Tyler: It's gonna use the extractors, so we'll you know, every you know, chunk of observations get that get made. It will have extracted some information which gets sent to some background agents. There will be some a few different types of them. Every chunk of observations there will be an agent which is taking that and converting it into a graph structure, right? So you'll have some like entities with relationships between them like facts and things like that. It'll be stored in a database.

15:31

Shane: Like a knowledge graph type thing?

15:32

Tyler: Yeah, a knowledge graph, exactly. It'll be a knowledge graph as well as

15:36

Shane: I knew we were gonna have to talk about graph.

15:39

Tyler: Graphs.

15:39

Shane: Graphs. So it always it always everyone always wants graphs and now we're gonna have them. Tell me a little bit more. How does the graph work? How how will the agent interact with the graph? How does that process work?

15:49

Tyler: So at observation time we'll have sort of a more of a stateless graph. Extraction happening, what that means is the agent that's doing the extracting isn't going to need to know all of the history going back. It'll just recognize entities and then relationships between them and just store those in the database. Then when you get to reflection time, we will have more of more stateful agent that's able to go back and look through all the previous memory. Graph structures that already exist and then it will take all of the you know sort of stateless graph structures that were created and rework those to fit into the existing structure. So sort of again like right observation and reflection, taking that metaphor and continuing to go go forward with it, basically.

16:31

Shane: And then how does the agent then use that graph to You know, potentially answer questions, right? Assuming you want to keep that data and then recall it or use it when the agent needs to answer a question that might be stored in this like knowledge graph of sorts.

16:44

Tyler: That actually goes into the other feature. Maybe a little bit. Agent signals.

16:48

Shane: That's something else that was released. What tell me what are agent signals? How do they work? And maybe then we can connect the dots on the graph side of things.

16:55

Tyler: So there's like a bunch of types of agent signals. I'm not going to go into all of them because there's some that are more advanced, but essentially it's just a way to decouple sending context into the running agent loop. From needing to own that loop. So before agent signals, you would send a prompt and you would immediately get a stream back. So like the client or consumer or whoever sending the prompt essentially was the one owning the stream.

17:15

Shane: I kick it off, I own it, right? That's. I think what most people expect when they talk to an agent is if I kick off the stream, I need to own the stream.

17:22

Tyler: Exactly. So with signals we've decoupled that. You can send any kind of context. There's multiple different kinds, like notifications, regular messages, and a few other things. You just send that into a specific thread. Right. So like a conversation. And then separately you're subscribed to the conversation and you're just listening for whatever's happening. So many, many clients could be subscribed and many clients could also be sending messages in.

17:47

Shane: Okay. I think you're gonna have to break that down because there's for me the thing that made the most sense, the thing that allowed me to kind of get it, is If I kick off a long-running task and the agent's doing some things, I can essentially add a message or send a signal, you know. And then it'll just get inserted in the next at the at the exact right moment the next time the loop happens. As soon as it's able to do that. So it doesn't actually interrupt the loop, it doesn't stop the loop. It just is sent as a signal. So you can essentially like steer an agent. If you're you know, if you think about a coding task, if you're watching the agent and you you're wondering why the hell are you going that way? Like don't do that. Go look at this part of the code because you know it. You might send and steer the agent to the right thing. Or often what I do is I s. I send a long message. What I'd like to think is a carefully curated plan and then I realized oh there was one more thing. Also don't forget to do this and then it just gets inserted as the agent's running. I think that's one example of a signal but you said like external systems as well. What would be external signals or notifications that come in and how would those work?

18:49

Tyler: Yes, maybe maybe your agent is working on like a pull request, right? And it needs to subscribe to any info from the pull request, right? Like if CI is failing or you get new. PR comments, that would be a notification signal. Your agent would be subscribed and anytime. This external notification happens, it'll just be dropped into context. Essentially in the same way as what you were just describing, but with a little bit of different formatting so the agent understands this is like a system sort of piece of context.

19:16

Shane: So it's like an external webhook of sorts?

19:18

Tyler: Yeah, basically.

19:19

Shane: But it gets inserted into context. What does that context look like? You know, so how does the agent know that this is I received a notification from GitHub or whatever.

19:29

Tyler: So we wrap each signal in some XML, right? Because agents are trained with XML, like user tags, assistant tags. With this one, we'll wrap it in like a notification tag. Maybe it'll say like type GitHub or something like that, right? And then it'll have some info about, you know, oh maybe CodeRabbit left some comments or whatever.

19:48

Shane: And then it just gets inserted as so this could happen. The agent has stopped or this could actually happen while the agent is still working on something else.

19:54

Tyler: It's both, right? So then that's actually a good point. That's a w a really nice part about it is if your agent goes idle, it can just wake it up again, right? So We can immediately begin addressing those pull request comments.

20:05

Shane: So the agent is basically sitting there, CI fails, pull request comment comes in, whatever, GitHub sends a webhook to your agent, wakes it up, your agent keeps running. And then while it's running, I could then say, oh, actually ignore that comment. I don't care about it. And I could steer it as it's as it's running as well, right?

20:21

Tyler: Exactly. Yep. Yeah, it's very flexible in that way. I guess where this ties back to observational memory is we have a type of signal called a state signal, which is just like a chunk of state that is in context and anytime it drops out of context it gets added back in. Typically people would be having a dynamic system prompt for that, right? Like we were talking about that earlier. But that invalidates the prompt cache. So with this new state signal, since we can just inject context at any point, like you were just saying, we'll be able to surface pieces of relevant memory from the graph. Maybe here's like the top you know five nodes in the graph that are updated or like were most recently updated or

20:58

Shane: that's how it connects to the graph. So you have this the signals get sent in from your Potential like nodes from the graph that might be useful for whatever the agent's trying to do. And then it's at the bottom of the context. And then when observations happen, I'm assuming that information, if it's needed, just gets moved back into observations.

21:14

Tyler: Yes, exactly. And if the observations remove the state signal from context, right, like we're compressing the messages, it'll again just get injected.

21:23

Shane: I feel like I'm learning more stuff today. I knew a lot of it, but I didn't even know all this. Okay, so we're in London today. We're doing this l live in London. Some of you watching this live. This is the third conference we're doing. What is one of the your m most fond memories of Mastra? You've been here a long time. Whether it's one of these events, whether it's, you know, even something we launched.

21:42

Tyler: I think it was probably my first week, which would be we had like a meetup in SF. You know when Mastra was in YC. I think we all just worked in I think it was the dungeon is what it was called, right? The dungeon

21:53

Shane: which was a two bedroom apartment near in the dog patch of San Francisco.

21:58

Tyler: Yeah.

21:58

Shane: And we had probably what Ten people working in that office.

22:04

Tyler: Yeah, that would be shit done.

22:06

Shane: It was not it was not comfortable. It was definitely packed. But we got a lot done.

22:09

Tyler: And it was a lot of fun still, you know.

22:11

Shane: I don't know if it was that week, but a lot of the team was at hotels, but I was sleeping on like a fold-out cot. I think Tony was sleeping on a mattress on the floor. I mean, it was just it was a fun week. Yep. All right. Well thanks, Tyler. Thanks for coming on. Thanks for talking about observational memory, subconscious observational memory, and telling us a little bit more about agent signals. We appreciate you spreading some knowledge with the rest of us.

22:33

Tyler: Well yeah, thanks for having me. It's a lot of fun.

22:36

Shane: Every week in AI, something insane happens. And there's so much drama. Every Monday we break it down live. We do the news, we bring on guests building in the space, and we go deep into the stuff that actually matters. Agents Hour. Every Monday, noon Pacific. Follow, don't miss it. Peace.