Skip to main content

Observational Memory

Added in: @mastra/memory@1.1.0

Observational Memory (OM) is Mastra's memory system for long-context agentic memory. An Observer watches conversations and creates observations, which a Reflector restructures by combining related items and condensing overarching patterns. Together, they maintain an observation log that replaces raw message history as it grows.

Usage
Direct link to Usage

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5-mini',
memory: new Memory({
options: {
observationalMemory: true,
},
}),
})

Configuration
Direct link to Configuration

The observationalMemory option accepts true, a configuration object, or false. Setting true enables OM with automatic model selection. When passing a config object, set model at the top level or on observation.model and/or reflection.model. When all model fields are omitted, each role defaults to 'auto'.

Observer input is multimodal-aware. OM keeps text placeholders like [Image #1: screenshot.png] in the transcript it builds for the Observer, and also sends the underlying image parts when possible. This applies to both single-thread observation and batched multi-thread observation. Non-image files appear as placeholders only.

OM performs thresholding with fast local token estimation. Text uses tokenx, and image-like inputs use provider-aware heuristics plus deterministic fallbacks when metadata is incomplete.

enabled?:

boolean
= true
Enable or disable Observational Memory. When omitted from a config object, defaults to true. Only enabled: false explicitly disables it.

model?:

'auto' | string | LanguageModel | DynamicModel | ModelByInputTokens | ModelWithRetries[]
= 'auto'
Model for both the Observer and Reflector agents. Cannot be used together with observation.model or reflection.model. When all model fields are omitted, each role defaults to "auto". Automatic selection prefers google/gemini-2.5-flash when GOOGLE_GENERATIVE_AI_API_KEY is configured, then a low-cost model for the active provider, then the active main model. Use "default" for the legacy fixed default.

autoModels?:

Record<string, string>
Overrides the model automatic selection picks for a provider, keyed by provider ID. For example, { openai: "__GATEWAY_OPENAI_MODEL_NANO__" }. Providers not listed use the built-in pick.

resolveModel?:

(modelId: string, context: { requestContext?: RequestContext }) => MastraModelConfig | Promise<MastraModelConfig>
Creates the model automatic selection picked, so it can use your credentials and routing. Receives only concrete model IDs, never "auto". When omitted, the picked ID goes through Mastra model routing.

scope?:

'resource' | 'thread'
= 'thread'
Memory scope for observations. 'thread' keeps observations per-thread. 'resource' is deprecated and will be removed in a future release. It shares observations across all threads for a resource and works much worse than thread scope for prompt caching and agent understanding. For cross-thread continuity, use thread scope with retrieval or resource-scoped working memory. See resource scope (deprecated).

activateAfterIdle?:

number | string | false | "auto" | Record<string, number | string | false | "auto">
Time before buffered observations are forced to activate after inactivity, even before observation.messageTokens is reached. Accepts a numeric millisecond value such as 300_000, duration strings like "5m" or "1hr", "auto" for a provider-aware prompt cache TTL, false to disable inherited observation idle activation, or an object of per-provider TTLs such as { default: "auto", anthropic: "1h" }. Object keys match the actor model provider before the first ., case-insensitively, so anthropic covers the direct Anthropic provider but not Claude through another provider such as amazon-bedrock; default covers every other provider. Reflections do not inherit this setting. Use reflection.activateAfterIdle to opt reflections into idle activation.

activateOnProviderChange?:

boolean
= false
Force buffered observations to activate when the actor provider or model changes. Reflections do not inherit this setting. Use reflection.activateOnProviderChange to opt reflections into provider-change activation.

shareTokenBudget?:

boolean
= false
Share the token budget between messages and observations. When enabled, the total budget is observation.messageTokens + reflection.observationTokens. Messages can use more space when observations are small, and vice versa. This maximizes context usage through flexible allocation. shareTokenBudget is not yet compatible with async buffering. You must set observation: { bufferTokens: false } when using this option (this is a temporary limitation).

temporalMarkers?:

boolean
= false
Insert temporal-gap reminder markers before new user messages when the previous message in the thread is at least 10 minutes older. The marker is persisted in memory, emitted as an inline reminder event so clients can render it specially, and shown to the observer so it can anchor observations to when events occurred.

retrieval?:

boolean | { vector?: boolean; scope?: 'thread' | 'resource'; instructions?: string }
= false
Let the agent look up the raw message history behind its observations. Observation groups keep durable pointers to the original messages, and a recall tool is registered so the agent can browse them. true enables cross-thread browsing by default. { vector: true } also enables semantic search using Memory's vector store and embedder. { scope: 'thread' } restricts the recall tool to the current thread only. Default scope is 'resource'. { instructions: '...' } appends application-specific recall guidance after Mastra's built-in retrieval instructions.

hooks?:

ObserveHooks
Lifecycle hooks fired for every observation/reflection cycle — the manual observe()/reflect() APIs, turn-driven synchronous observation, and fire-and-forget async buffering. Callbacks receive threadId/resourceId/trigger call context ('manual' | 'turn-sync' | 'async-buffer'), and the end hooks (onObservationEnd/onReflectionEnd) additionally receive the OM model call's token usage and providerMetadata (where providers such as the AI Gateway report per-call cost), so apps can account for OM model spend without wrapping the observer/reflector models in middleware. Config-level lifecycle hooks are non-blocking by default: their errors are logged without failing the cycle. With hookExecution: 'await', config-level lifecycle errors can reject observe() or reflect(). Per-call observe({ hooks }) lifecycle hooks are also awaited and can reject observe(). A failing awaited start hook skips the model call; an end-hook failure can reject the call after its work has completed. Async-buffered cycles always use non-blocking config-level lifecycle hooks, regardless of hookExecution. Their cycle failures are reported through onObservationEnd.error or onReflectionEnd.error, rather than thrown to the caller. ObserveHooks also accepts transform hooks that intercept and can replace cycle data: beforeObservation({ messages, ...context }) runs on the messages about to be observed (return { messages } to filter/redact; an empty array skips the Observer call), afterObservation({ observations, ...context }) runs on the Observer's output before it is persisted, beforeReflection({ observations, ...context }) runs on the text sent to the Reflector, and afterReflection({ observations, ...context }) runs on the Reflector's output before it is persisted. Transform hooks return void to pass data through unchanged and are always awaited on every path. A thrown error fails the cycle before committing its transformed observation or reflection text, but doesn't roll back extractor callbacks or other side effects that already ran. The after hooks replace only text, not the separate structured extractor results stored in thread metadata. Reflection extraction and its callbacks run before afterReflection; this hook neither recomputes extracted values nor reruns callbacks. After hooks aren't a redaction boundary for all cycle data.

observation?:

ObservationalMemoryObservationConfig
Configuration for the observation step. Controls when the Observer agent runs and how it behaves.
ObservationalMemoryObservationConfig

model?:

'auto' | string | LanguageModel | DynamicModel | ModelByInputTokens | ModelWithRetries[]
Model for the Observer agent. With 'auto', each observation uses the Gemini default when GOOGLE_GENERATIVE_AI_API_KEY is set, otherwise a low-cost model for the calling agent's effective provider, then the calling model or configured model instance. Cannot be set if a top-level model is also provided. If neither this nor the top-level model is set, falls back to reflection.model.

instruction?:

string
Custom instruction appended to the Observer's system prompt. Use this to customize what the Observer focuses on, such as domain-specific preferences or priorities.

continuationHints?:

boolean | { currentTask?: boolean; suggestedResponse?: boolean }
Which continuation-hint sections the Observer emits during synchronous observation. Async buffered Observer calls do not generate continuation hints. Pass false to disable both, or an object to disable them individually. Agents that drive their own control flow generally want { suggestedResponse: false } so memory does not compete for what the agent says next. A previously stored hint stops being injected into context once both observation and reflection disable its section.

threadTitle?:

boolean
When true, the Observer suggests short thread titles and updates the thread title when the conversation topic meaningfully changes. This is opt-in and defaults to disabled.

extract?:

Extractor[]
Custom values to extract after observation. Schema-less extractors are requested inline in the Observer output. Schema-backed extractors run as a follow-up structured output call and are stored in thread OM metadata.

manageWorkingMemory?:

boolean
Let the Observer manage working memory through OM extraction. Adds WorkingMemoryExtractor, defaults workingMemory.agentManaged to false, and defaults workingMemory.useStateSignals to true. See Working memory updates.

observeAttachments?:

'auto' | boolean | string[]
Controls which image/file attachments are forwarded to the Observer model alongside their placeholder text lines. true (default) forwards all attachments. false drops all attachments while keeping placeholders visible. 'auto' uses the provider capabilities registry to decide: attachments are forwarded when the Observer model supports multimodal input, dropped otherwise, and forwarded when no capability data is available for the model. An array is a case-insensitive mimeType allowlist supporting exact matches ('application/pdf'), wildcard subtypes ('image/*'), and bare '*' for everything. Useful when the Observer model is text-only (e.g. some DeepSeek endpoints) while the main agent uses a multimodal model. Tool-result attachments are filtered using the same rule.

maxRetries?:

number
Retries after the initial Observer model call on transient provider errors. Governs OM's own retry ladder; the Observer model call itself is configured with no provider-level retries.

failurePolicy?:

'abort' | 'continue'
Terminal policy once Observer retries are exhausted. 'abort' aborts the agent turn. 'continue' emits the existing failure diagnostic, keeps the failed input pending for a later cycle, and allows the main agent turn to continue. Persistence, indexing, transform, locking, invariant, and explicit abort failures remain fatal. This setting doesn't prevent the underlying provider error or change blockAfter or attachment handling, and it has no backstop for a sustained outage: pending messages keep accruing in the main agent's context until they reach the model's context limit.

messageTokens?:

number
Token count of unobserved messages that triggers observation. When unobserved message tokens exceed this threshold, the Observer agent is called. Text is estimated locally with tokenx. Image parts are included with model-aware heuristics when possible, with deterministic fallbacks when image metadata is incomplete. Image-like file parts are counted the same way when uploads are normalized as files.

maxTokensPerBatch?:

number
Maximum tokens per batch when observing multiple threads in resource scope (deprecated). Threads are chunked into batches of this size and processed in parallel. Lower values mean more parallelism but more API calls.

modelSettings?:

ObservationalMemoryModelSettings
Model settings for the Observer agent. The temperature: 0.3 default is only applied when the resolved model is known to support temperature. The maxOutputTokens: 100_000 default is only applied with default model selection (no model set, "default", or a ModelByInputTokens selector). Custom models get no maxOutputTokens default.
ObservationalMemoryModelSettings

temperature?:

number
Temperature for generation. Lower values produce more consistent output. The 0.3 default is only applied when the resolved model is known to support temperature.

maxOutputTokens?:

number
Maximum output tokens. Set high to prevent truncation of observations. The 100000 default is only applied with default model selection; custom models get no default.

providerOptions?:

ProviderOptions
Provider-specific options passed to the Observer agent, such as Google thinking configuration.

bufferTokens?:

number | false
How often background observation buffering runs. Values between 0 and 1 are fractions of messageTokens: 0.25 buffers every 25% of the threshold (7.5k tokens with the default 30k). Values of 1 or more are absolute token counts: 5000 buffers every 5k tokens. Buffered observations are stored until the messageTokens threshold is reached, then activate instantly without a blocking LLM call. Must resolve to less than messageTokens. Set to false to disable all async buffering (both observation and reflection).

bufferOnIdle?:

boolean
Run background observation buffering when an agent turn ends and the agent becomes idle. This is separate from bufferTokens, which controls step-time async buffering. Set this to true to buffer short idle turns without waiting for the next turn or the messageTokens threshold.

bufferActivation?:

number
How much of the message window to clear when buffered observations activate. Values between 0 and 1 are the fraction of messageTokens to remove: 0.8 removes ~80% of the message history and keeps ~20% (6k tokens with the default 30k). Values of 1000 or more are the token count to keep: 4000 keeps ~4k message tokens after activation. Note the direction flips: a higher ratio removes more history, while a higher token count keeps more.

activateAfterIdle?:

number | string | false | "auto" | Record<string, number | string | false | "auto">
Time before buffered observations are forced to activate after inactivity. Accepts milliseconds, a duration string, "auto" for a provider-aware prompt cache TTL, false, or an object of per-provider TTLs. An object replaces the top-level value and is not merged with it. If unset, the top-level activateAfterIdle value is used for observations. Set false to disable the top-level idle setting for observations. Currently only applied when using the standalone ObservationalMemory class; new Memory(...) applies the top-level activateAfterIdle only.

activateOnProviderChange?:

boolean
Force buffered observations to activate when the actor provider or model changes. If unset, the top-level activateOnProviderChange value is used for observations. Currently only applied when using the standalone ObservationalMemory class; new Memory(...) applies the top-level activateOnProviderChange only.

blockAfter?:

number
Safety net for when background buffering can't keep up. Values from 1 up to (but not including) 100 are multipliers of messageTokens: 1.2 resolves to 120% of the threshold (36k tokens with the default 30k). Values of 100 or more are absolute token counts and must be greater than messageTokens. Above this point, activation uses the smallest set of buffered chunks that reaches the retention target, even when that overshoots the target by more than the usual safeguard allows. It never activates more chunks than are needed to reach the retention target, so it removes only slightly more history than a normal activation. It changes the result only when the retention floor is above roughly 20,000 tokens; with the default bufferActivation (a 6k floor) it has no observable effect. Activation usually keeps a minimum remaining context (the smaller of 1000 tokens or the retention floor), but a single buffered chunk that covers the whole pending window still activates and can leave less. Crossing blockAfter does not trigger a blocking observation. A synchronous (blocking) observation runs when the messageTokens threshold is reached and activating buffered chunks doesn't bring pending tokens back under it, for example when a large tool result arrives after buffering has stopped. All remaining buffered chunks are activated first, so the synchronous observation only covers messages no chunk has observed. Only relevant when bufferTokens is set. Defaults to 1.2 when async buffering is enabled.

previousObserverTokens?:

number | false
Optional token budget for the observer's previous-observations context. When set to a number, the observations passed to the Observer agent are tail-truncated to fit within this budget while keeping the newest observations and preserving highlighted 🔴 items when possible. When a buffered reflection is pending, the already-reflected observation lines are automatically replaced with the reflection summary before truncation. Set to 0 to omit previous observations entirely, or false to disable truncation explicitly.

reflection?:

ObservationalMemoryReflectionConfig
Configuration for the reflection step. Controls when the Reflector agent runs and how it behaves.
ObservationalMemoryReflectionConfig

model?:

'auto' | string | LanguageModel | DynamicModel | ModelByInputTokens | ModelWithRetries[]
Model for the Reflector agent. With 'auto', each reflection uses the Gemini default when GOOGLE_GENERATIVE_AI_API_KEY is set, otherwise a low-cost model for the calling agent's effective provider, then the calling model or configured model instance. Cannot be set if a top-level model is also provided. If neither this nor the top-level model is set, falls back to observation.model.

maxRetries?:

number
Retries after the initial Reflector model call on transient provider errors. Governs OM's own retry ladder; the Reflector model call itself is configured with no provider-level retries.

failurePolicy?:

'abort' | 'continue'
Terminal policy once Reflector retries are exhausted. 'abort' aborts the agent turn. 'continue' emits the failure diagnostic and allows the turn to continue, leaving already-persisted observations committed and deferring reflection to the next threshold crossing. A provider outage usually takes out both stages, so set the policy on observation too if the turn should survive one.

instruction?:

string
Custom instruction appended to the Reflector's system prompt. Use this to customize how the Reflector consolidates observations, such as prioritizing certain types of information.

continuationHints?:

boolean | { currentTask?: boolean; suggestedResponse?: boolean }
Which continuation-hint sections the Reflector emits. Pass false to disable both, or an object to disable them individually. A previously stored hint stops being injected into context once both observation and reflection disable its section.

extract?:

Extractor[]
Custom values to extract after reflection. Schema-less extractors are requested inline in the Reflector output. Schema-backed extractors run as a follow-up structured output call and are stored in thread OM metadata.

observationTokens?:

number
Token count of observations that triggers reflection. When observation tokens exceed this threshold, the Reflector agent is called to condense them.

modelSettings?:

ObservationalMemoryModelSettings
Model settings for the Reflector agent. The temperature: 0 default is only applied when the resolved model is known to support temperature. The maxOutputTokens: 100_000 default is only applied with default model selection (no model set, "default", or a ModelByInputTokens selector). Custom models get no maxOutputTokens default.
ObservationalMemoryModelSettings

temperature?:

number
Temperature for generation. Lower values produce more consistent output. The 0 default is only applied when the resolved model is known to support temperature.

maxOutputTokens?:

number
Maximum output tokens. Set high to prevent truncation of observations. The 100000 default is only applied with default model selection; custom models get no default.

providerOptions?:

ProviderOptions
Provider-specific options passed to the Reflector agent, such as Google thinking configuration.

bufferActivation?:

number
When background reflection starts, as a ratio (0-1) of observationTokens: 0.5 starts reflecting in the background once observations reach 50% of the threshold (20k tokens with the default 40k). When the full threshold is reached, the buffered reflection replaces the observations it covers, preserving any new observations appended after that range.

activateAfterIdle?:

number | string | false | "auto" | Record<string, number | string | false | "auto">
Time before buffered reflections are forced to activate after inactivity. Accepts milliseconds, a duration string, "auto" for a provider-aware prompt cache TTL, false, or an object of per-provider TTLs. Reflections do not inherit top-level activateAfterIdle; set this explicitly to opt reflections into idle activation. Currently only applied when using the standalone ObservationalMemory class; this setting has no effect through new Memory(...).

activateOnProviderChange?:

boolean
Force buffered reflections to activate when the actor provider or model changes. Reflections do not inherit top-level activateOnProviderChange; set this explicitly to opt reflections into provider-change activation. Currently only applied when using the standalone ObservationalMemory class; this setting has no effect through new Memory(...).

blockAfter?:

number
Safety net that forces a synchronous (blocking) reflection when background reflection can't keep up. Values from 1 up to (but not including) 100 are multipliers of observationTokens: 1.2 forces a blocking reflection at 120% of the threshold (48k tokens with the default 40k). Values of 100 or more are absolute token counts and must be greater than observationTokens. Between observationTokens and blockAfter, only async buffering and activation run. Only relevant when bufferActivation is set. Defaults to 1.2 when async reflection is enabled.

Token estimate metadata cache
Direct link to Token estimate metadata cache

OM persists token payload estimates so repeated counting can reuse prior token estimation work.

  • Part-level cache: part.providerMetadata.mastra.
  • String-content fallback cache: message-level metadata when no parts exist.
  • Cache entries are ignored and recomputed if cache version/tokenizer source doesn't match.
  • Per-message and per-conversation overhead are always recomputed at runtime and aren't cached.
  • data-* and reasoning parts are skipped and don't receive cache entries.

Extractor API
Direct link to Extractor API

Extractor defines a value that OM should extract during observation or reflection. Built-in OM values such as current-task, suggested-response, and thread-title use the same extractor pipeline as custom values.

src/mastra/agents/agent.ts
import { Memory, Extractor } from '@mastra/memory'
import { z } from 'zod'

const memory = new Memory({
options: {
observationalMemory: {
model: 'openai/gpt-5-mini',
observation: {
extract: [
new Extractor({
name: 'User profile',
instructions: 'Extract stable user profile facts that should be remembered.',
schema: z.object({
name: z.string().optional(),
timezone: z.string().optional(),
}),
}),
],
},
},
},
})

name:

string
Human-readable extractor name. OM slugifies this value into the extractor slug. Names must be unique after slug generation.

slug:

string
Read-only property derived from name — not a constructor option. Generated stable identifier for persisted values and XML tags. Slugs use lowercase letters, numbers, and hyphens. Built-in slugs and reserved XML tags cannot be used by custom extractors.

instructions:

string | (context) => string
Instructions for what to extract and when to update the value. Use a function to derive instructions from runtime context.

schema?:

ZodType<T> | (context) => ZodType<T> | undefined
Optional Zod schema for structured extraction. When provided, OM runs a follow-up structured output call after the main OM operation. When omitted, the extractor is an inline string extractor emitted in the Observer or Reflector response. Use a function to derive the schema from runtime context.

includePreviousExtraction?:

boolean
= true
Controls whether the previous extraction is shown to the extractor on future OM runs. Set to false for values that should only come from the current OM run.

metadataKeyPath?:

string | false
= 'extracted.<slug>'
Dot-separated OM metadata path used to persist the extracted value. Set to false to skip OM metadata persistence entirely.

onExtracted?:

(context) => T | void | Promise<T | void>
Optional hook called after a custom extractor returns a value and before metadata is persisted. Returning a value replaces the extracted value. Throwing records an extraction failure.

Extraction behavior
Direct link to Extraction behavior

  • Extracted values are stored in thread OM metadata under om.extracted.
  • Built-in extractor values are also mirrored to the compatibility metadata fields currentTask, suggestedResponse, and threadTitle.
  • thread-title updates the thread title only when observation.threadTitle is enabled.
  • observation.extract runs during observation. reflection.extract runs during reflection.
  • Schema-backed extractors add a follow-up structured output request.
  • Schema-less extractors are inline string extractors emitted directly in the Observer or Reflector output.
  • Dynamic extractor functions receive runtime context, including source, threadId, resourceId, mainAgent, memory, and requestContext when available.
  • WorkingMemoryExtractor uses the normal extractor pipeline to update working memory through the active Memory instance. It uses structured extraction when working memory has a JSON schema and skips OM metadata persistence, so the working memory payload isn't duplicated under OM extracted metadata.
  • When workingMemory.schema is set, WorkingMemoryExtractor validates each update against that schema before saving it. If an update doesn't match, it's skipped and reported as an extraction failure while the previous working memory stays in place. Because the schema isn't sent to the model as a structured output constraint, one invalid working memory update can't fail other extractors. A null in an optional field is treated as not provided.
  • observationalMemory.observation.manageWorkingMemory adds WorkingMemoryExtractor and defaults workingMemory.agentManaged to false. It defaults workingMemory.useStateSignals to true when working memory is enabled.
  • Extraction failures are reported in OM marker data and don't discard other successful extracted values.

Examples
Direct link to Examples

Working memory updates
Direct link to Working memory updates

Use observationalMemory.observation.manageWorkingMemory when OM should update working memory.

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'

const memory = new Memory({
options: {
workingMemory: {
enabled: true,
},
observationalMemory: {
enabled: true,
observation: {
manageWorkingMemory: true,
},
},
},
})

Set workingMemory.agentManaged: true if the main agent should still receive working memory tool and instruction injection.

Custom thresholds
Direct link to Custom thresholds

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5-mini',
memory: new Memory({
options: {
observationalMemory: {
model: 'google/gemini-2.5-flash',
observation: {
messageTokens: 20_000,
},
reflection: {
observationTokens: 60_000,
},
},
},
}),
})

Shared token budget
Direct link to Shared token budget

When shareTokenBudget is enabled, the total budget is observation.messageTokens + reflection.observationTokens, which is 100k in this example. Observations that use only 30k tokens leave up to 70k for messages, while short messages give observations more room before reflection starts.

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5-mini',
memory: new Memory({
options: {
observationalMemory: {
shareTokenBudget: true,
observation: {
messageTokens: 20_000,
bufferTokens: false, // required when using shareTokenBudget (temporary limitation)
},
reflection: {
observationTokens: 80_000,
},
},
},
}),
})

Custom model
Direct link to Custom model

By passing a model in the config, you can use any model from Mastra's model router.

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5.6-sol',
memory: new Memory({
options: {
observationalMemory: {
model: 'openai/gpt-5-mini',
},
},
}),
})

Different models per agent
Direct link to Different models per agent

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5.6-sol',
memory: new Memory({
options: {
observationalMemory: {
observation: {
model: 'google/gemini-2.5-flash',
},
reflection: {
model: 'openai/gpt-5-mini',
},
},
},
}),
})

Custom instructions
Direct link to Custom instructions

Customize what the Observer and Reflector focus on by providing custom instructions:

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'health-assistant',
name: 'health-assistant',
instructions: 'You are a health and wellness assistant.',
model: 'openai/gpt-5.6-sol',
memory: new Memory({
options: {
observationalMemory: {
model: 'google/gemini-2.5-flash',
observation: {
// Focus observations on health-related preferences and goals
instruction:
'Prioritize capturing user health goals, dietary restrictions, exercise preferences, and medical considerations. Avoid capturing general chit-chat.',
},
reflection: {
// Guide reflection to consolidate health patterns
instruction:
'When consolidating, group related health information together. Preserve specific metrics, dates, and medical details.',
},
},
},
}),
})

Async buffering
Direct link to Async buffering

Async buffering is enabled by default. It pre-computes observations in the background as the conversation grows: when the messageTokens threshold is reached, buffered observations activate instantly with no blocking LLM call.

The lifecycle follows buffer → activate → remove messages → repeat. Background Observer calls run at bufferTokens intervals and produce chunks of observations. At the threshold, activation moves observations into the log and removes raw messages from context. Above blockAfter, activation may overshoot the retention target instead of activating fewer chunks. A synchronous observation runs if pending tokens are still at or above the threshold after all buffered chunks have been activated.

Default settings:

  • observation.bufferTokens: 0.2: Buffer every 20% of messageTokens (e.g. every ~6k tokens with a 30k threshold)
  • observation.bufferActivation: 0.8: On activation, remove enough messages to keep only 20% of the threshold remaining
  • reflection.bufferActivation: 0.5: start background reflection at 50% of observation threshold

Async buffered Observer calls don't generate continuation hints (suggestedResponse, currentTask), and activation clears any previously stored hints.

To customize:

src/mastra/agents/agent.ts
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5-mini',
memory: new Memory({
options: {
observationalMemory: {
model: 'google/gemini-2.5-flash',
observation: {
messageTokens: 30_000,
// Buffer every 5k tokens (runs in background)
bufferTokens: 5_000,
// Activate to retain 30% of threshold
bufferActivation: 0.7,
// Above 1.5x the threshold, let activation overshoot the retention target
blockAfter: 1.5,
},
reflection: {
observationTokens: 60_000,
// Start background reflection at 50% of threshold
bufferActivation: 0.5,
// Force synchronous reflection at 1.2x threshold
blockAfter: 1.2,
},
},
},
}),
})

To disable async buffering entirely:

observationalMemory: {
model: "google/gemini-2.5-flash",
observation: {
bufferTokens: false,
},
}

Setting bufferTokens: false disables both observation and reflection async buffering. Observations and reflections will run synchronously when their thresholds are reached.

note

Async buffering isn't supported with the deprecated scope: 'resource' and is automatically disabled in resource scope.

Streaming data parts
Direct link to Streaming data parts

Observational Memory emits typed data parts during agent execution that clients can use for real-time UI feedback. These are streamed alongside the agent's response.

Read extractor results
Direct link to Read extractor results

Both completion events carry extractor output in their data payload. The extractor fields are:

interface DataOmObservationEndPart {
type: 'data-om-observation-end'
data: {
/** Whether the completed work was an observation or reflection */
operationType: 'observation' | 'reflection'
/** Values extracted during this OM operation, keyed by extractor slug */
extractedValues?: Record<string, unknown>
/** Extractor failures from this OM operation. Successful extractor values are still included */
extractionFailures?: Array<{ slug: string; error: string }>
// ...other fields documented in the tables below
}
}

Both extractor fields are optional. A completion can include values, failures, both, or neither. data-om-observation-end reports synchronous completion. data-om-buffering-end reports completed background work whose buffered content still awaits activation, although extractor metadata is already persisted. DataOmBufferingEndPart carries the same extractor fields, and both types are exported from @mastra/memory/processors. See Read extracted values from a stream for a consumer example.

data-om-status
Direct link to data-om-status

Emitted once per agent loop step, before model generation. Provides a snapshot of the current memory state, including token usage for both context windows and the state of any async buffered content.

interface DataOmStatusPart {
type: 'data-om-status'
data: {
windows: {
active: {
/** Unobserved message tokens and the threshold that triggers observation */
messages: { tokens: number; threshold: number }
/** Observation tokens and the threshold that triggers reflection */
observations: { tokens: number; threshold: number }
}
buffered: {
observations: {
/** Number of buffered chunks staged for activation */
chunks: number
/** Total message tokens across all buffered chunks */
messageTokens: number
/** Projected message tokens that would be removed if activation happened now (based on bufferActivation ratio and chunk boundaries) */
projectedMessageRemoval: number
/** Observation tokens that will be added on activation */
observationTokens: number
/** idle: no buffering in progress. running: background observer is working. complete: chunks are ready for activation. */
status: 'idle' | 'running' | 'complete'
}
reflection: {
/** Observation tokens that were fed into the reflector (pre-compression size) */
inputObservationTokens: number
/** Observation tokens the reflection will produce on activation (post-compression size) */
observationTokens: number
/** idle: no reflection buffered. running: background reflector is working. complete: reflection is ready for activation. */
status: 'idle' | 'running' | 'complete'
}
}
}
recordId: string
threadId: string
stepNumber: number
/** Increments each time the Reflector creates a new generation */
generationCount: number
}
}

buffered.reflection.inputObservationTokens is the size of the observations that were sent to the Reflector. buffered.reflection.observationTokens is the compressed result: the size of what will replace those observations when the reflection activates. A client can use these two values to show a compression ratio.

Clients can derive percentages and post-activation estimates from the raw values:

// Message window usage %
const msgPercent = status.windows.active.messages.tokens / status.windows.active.messages.threshold

// Observation window usage %
const obsPercent =
status.windows.active.observations.tokens / status.windows.active.observations.threshold

// Projected message tokens after buffered observations activate
// Uses projectedMessageRemoval which accounts for bufferActivation ratio and chunk boundaries
const postActivation =
status.windows.active.messages.tokens -
status.windows.buffered.observations.projectedMessageRemoval

// Reflection compression ratio (when buffered reflection exists)
const { inputObservationTokens, observationTokens } = status.windows.buffered.reflection
if (inputObservationTokens > 0) {
const compressionRatio = observationTokens / inputObservationTokens
}

data-om-observation-start
Direct link to data-om-observation-start

Emitted when the Observer or Reflector agent begins processing.

cycleId:

string
Unique ID for this cycle — shared between start/end/failed markers.

operationType:

'observation' | 'reflection'
Whether this is an observation or reflection operation.

startedAt:

string
ISO timestamp when processing started.

tokensToObserve:

number
Message tokens (input) being processed in this batch.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

threadIds:

string[]
All thread IDs in this batch (for resource-scoped).

config:

ObservationMarkerConfig
Snapshot of messageTokens, observationTokens, and scope at observation time.

data-om-observation-end
Direct link to data-om-observation-end

Emitted when observation or reflection completes successfully.

cycleId:

string
Matches the corresponding start marker.

operationType:

'observation' | 'reflection'
Type of operation that completed.

completedAt:

string
ISO timestamp when processing completed.

durationMs:

number
Duration in milliseconds.

tokensObserved:

number
Message tokens (input) that were processed.

observationTokens:

number
Resulting observation tokens (output) after the Observer compressed them.

observations?:

string
The generated observations text.

currentTask?:

string
Current task extracted by the Observer.

suggestedResponse?:

string
Suggested response extracted by the Observer.

extractedValues?:

Record<string, unknown>
Values extracted during this OM operation, keyed by extractor slug.

extractionFailures?:

Array<{ slug: string; error: string }>
Extractor failures from this OM operation. Successful extractor values are still included.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

data-om-observation-failed
Direct link to data-om-observation-failed

Emitted when observation or reflection fails. The system falls back to synchronous processing.

cycleId:

string
Matches the corresponding start marker.

operationType:

'observation' | 'reflection'
Type of operation that failed.

failedAt:

string
ISO timestamp when the failure occurred.

durationMs:

number
Duration until failure in milliseconds.

tokensAttempted:

number
Message tokens (input) that were attempted.

error:

string
Error message.

retrying?:

true
Set when a reflection attempt did not compress below its target and OM is retrying at a higher compression level. A new start marker follows, so this failure is not final. Absent on final failures.

observations?:

string
Any partial content available for display.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

data-om-buffering-start
Direct link to data-om-buffering-start

Emitted when async buffering begins in the background. Buffering pre-computes observations or reflections before the main threshold is reached.

cycleId:

string
Unique ID for this buffering cycle.

operationType:

'observation' | 'reflection'
Type of operation being buffered.

startedAt:

string
ISO timestamp when buffering started.

tokensToBuffer:

number
Message tokens (input) being buffered in this cycle.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

threadIds:

string[]
All thread IDs being buffered (for resource-scoped).

config:

ObservationMarkerConfig
Snapshot of config at buffering time.

data-om-buffering-end
Direct link to data-om-buffering-end

Emitted when async buffering completes. The content is stored but not yet activated in the main context.

cycleId:

string
Matches the corresponding buffering-start marker.

operationType:

'observation' | 'reflection'
Type of operation that was buffered.

completedAt:

string
ISO timestamp when buffering completed.

durationMs:

number
Duration in milliseconds.

tokensBuffered:

number
Message tokens (input) that were buffered.

bufferedTokens:

number
Observation tokens (output) after the Observer compressed them.

observations?:

string
The buffered content.

extractedValues?:

Record<string, unknown>
Values extracted during this buffered OM operation, keyed by extractor slug.

extractionFailures?:

Array<{ slug: string; error: string }>
Extractor failures from this buffered OM operation. Successful extractor values are still included.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

data-om-buffering-failed
Direct link to data-om-buffering-failed

Emitted when async buffering fails. The system falls back to synchronous processing when the threshold is reached.

cycleId:

string
Matches the corresponding buffering-start marker.

operationType:

'observation' | 'reflection'
Type of operation that failed.

failedAt:

string
ISO timestamp when the failure occurred.

durationMs:

number
Duration until failure in milliseconds.

tokensAttempted:

number
Message tokens (input) that were attempted to buffer.

error:

string
Error message.

observations?:

string
Any partial content.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

data-om-activation
Direct link to data-om-activation

Emitted when buffered observations or reflections are activated (moved into the active context window). This is an instant operation: no LLM call is involved.

cycleId:

string
Unique ID for this activation event.

operationType:

'observation' | 'reflection'
Type of content activated.

activatedAt:

string
ISO timestamp when activation occurred.

chunksActivated:

number
Number of buffered chunks activated.

tokensActivated:

number
Message tokens (input) from activated chunks. For observation activation, these are removed from the message window. For reflection activation, this is the observation tokens that were compressed.

observationTokens:

number
Resulting observation tokens after activation.

messagesActivated:

number
Number of messages that were observed via activation.

generationCount:

number
Current reflection generation count.

observations?:

string
The activated observations text.

triggeredBy?:

'threshold' | 'ttl' | 'provider_change'
Whether activation was triggered by threshold crossing, activateAfterIdle expiry, or a model/provider change.

lastActivityAt?:

number
Unix-ms timestamp of the last assistant message part used for TTL checks.

ttlExpiredMs?:

number
How long activateAfterIdle had been exceeded when activation fired.

previousModel?:

string
Previous assistant model identifier that triggered activation (e.g. openai/gpt-4o).

currentModel?:

string
Current actor model identifier that triggered activation.

recordId:

string
The OM record ID.

threadId:

string
This thread's ID.

config:

ObservationMarkerConfig
Snapshot of config at activation time.

data-om-thread-update
Direct link to data-om-thread-update

Emitted when the Observer updates the thread title. Only emitted when observation.threadTitle is enabled.

cycleId:

string
Unique ID for this observation cycle — shared with observation markers.

threadId:

string
The thread ID that was updated.

oldTitle?:

string
The previous thread title. Undefined if the thread had no title.

newTitle:

string
The new thread title.

timestamp:

string
When this update occurred.

Standalone usage
Direct link to Standalone usage

Most users should use the Memory class above. Using ObservationalMemory directly is mainly useful for benchmarking, experimentation, or when you need to control processor ordering with other processors (like guardrails).

The ObservationalMemory class is the engine; to attach it to an agent, wrap it in an ObservationalMemoryProcessor, which needs a Memory instance for loading and persisting messages. Note that stores.memory is typed as optional on storage adapters, so a non-null assertion (or a runtime check) is needed:

src/mastra/agents/agent.ts
import { ObservationalMemory, ObservationalMemoryProcessor } from '@mastra/memory/processors'
import { Memory } from '@mastra/memory'
import { Agent } from '@mastra/core/agent'
import { LibSQLStore } from '@mastra/libsql'

const storage = new LibSQLStore({
id: 'my-storage',
url: 'file:./memory.db',
})

const memory = new Memory({ storage })

const om = new ObservationalMemory({
storage: storage.stores.memory!,
memory,
model: 'google/gemini-2.5-flash',
observation: {
messageTokens: 20_000,
},
reflection: {
observationTokens: 60_000,
},
})

const omProcessor = new ObservationalMemoryProcessor(om, memory)

export const agent = new Agent({
id: 'my-agent',
name: 'my-agent',
instructions: 'You are a helpful assistant.',
model: 'openai/gpt-5-mini',
inputProcessors: [omProcessor],
outputProcessors: [omProcessor],
})

Standalone config
Direct link to Standalone config

The standalone ObservationalMemory class accepts all the same options as the observationalMemory config object above, plus the following:

storage:

MemoryStorage
Storage adapter for persisting observations. Must be a MemoryStorage instance (from MastraStorage.stores.memory).

onDebugEvent?:

(event: ObservationDebugEvent) => void
Debug callback for observation events. Called whenever observation-related events occur. Useful for debugging and understanding the observation flow.

obscureThreadIds?:

boolean
= false
When enabled, thread IDs are hashed before being included in observation context. This prevents the LLM from recognizing patterns in thread identifiers. Automatically enabled when using resource scope through the Memory class.

Recall tool
Direct link to Recall tool

When retrieval is truthy, Mastra registers a recall tool that pages through raw messages behind observation group ranges. With the default resource scope, the tool can list threads (mode: "threads") and browse another thread through threadId. It also supports cross-thread search. Set retrieval: { vector: true } for semantic search (mode: "search"), or use scope: 'thread' to restrict the tool to the current thread. The tool is automatically added to the agent.

When retrieval is enabled, OM also injects instructions telling the agent how to use recall: when to verify details, how to search and page, and when to read raw messages. They're included from the first turn, including read-only runs, and only describe the modes that are configured. Reflected groups are marked _kind: reflection_ so the agent treats them as lossy summaries. Use retrieval: { instructions: '...' } to append your own guidance after the built-in instructions.

Observation pages
Direct link to Observation pages

With vector search configured, mode: "observations" reads the original observation groups around a search hit, including groups that reflection condensed away and buffered groups that haven't been activated yet. Copy the hit's groupId exactly as shown and, in resource scope, its threadId:

{
"mode": "observations",
"threadId": "source-thread-id",
"groupId": "group-id-from-search",
"direction": "after",
"limit": 5
}

Use direction: "before" to page backward, or omit direction to include the anchor group itself. Pages return { results, count, hasMore } with continuation calls, and say when newer messages haven't been observed yet.

Search results
Direct link to Search results

Search results are listed by date and share a 2,000-token text allowance, so long groups are shortened, starting at the line that best matches the query. Results already in the agent's context come back as short references, and lower-ranked hits fill the freed space. limit counts excerpts, not references, so count can exceed limit.

Storage requirements
Direct link to Storage requirements

Group IDs in search results include the record that holds them (groupId@recordId), so paging reads a single record. Groups indexed before this change still work, but page by scanning the thread's observation history. Paging needs a storage adapter that supports observation history filters, so upgrade it together with @mastra/memory. Convex users need to redeploy their Mastra server functions.

Parameters
Direct link to Parameters

mode?:

'messages' | 'threads' | 'search' | 'observations'
= 'messages'
What to retrieve. "messages" (default) pages through message history. "threads" lists threads for the current user. "search" finds messages by semantic similarity (requires vector store and embedder). "observations" pages original observation groups around a search hit.

groupId?:

string
Required for observations mode. Copy the observation group ID exactly from a search result or continuation call. IDs shown as groupId@recordId let paging read the record directly; a plain group ID also works.

direction?:

'before' | 'after'
For observations mode: return groups strictly before or after the anchor group. Omit to include the anchor and following groups.

query?:

string
Search query for mode: "search". Finds messages semantically similar to this text across all threads for the current user.

cursor?:

string
A message ID to anchor the recall query. Extract the start or end ID from an observation group range (e.g. from _range: \startId:endId\_, use either startId or endId). If a range string is passed directly, the tool returns a hint explaining how to extract the correct ID. When both cursor and threadId are omitted for mode: "messages", the tool browses the current thread from the position set by anchor.

threadId?:

string
Browse a different thread by its ID, or pass "current" for the active thread. Use mode: "threads" first to discover thread IDs. When provided without a cursor, reading starts from the beginning of the thread.

anchor?:

'start' | 'end'
= 'start'
For mode: "messages" without a cursor, page from the start (oldest-first) or end (newest-first) of the thread.

page?:

number
= 1
Pagination offset. For messages: positive values page forward from cursor, negative values page backward. For threads: page number (0-indexed). 0 is treated as 1 for messages.

limit?:

number
= 20
Page size, from 1 to 20. Defaults to 20 for messages and 5 groups for observations. Search defaults to 10 excerpts, plus compact references for evidence already in context. Bounded backfill can return fewer excerpts.

detail?:

'low' | 'high'
= 'low'
Controls how much content is shown per message part. 'low' shows truncated text and tool names with positional indices ([p0], [p1]). 'high' shows full content including tool arguments and results, clamped to one part per call with continuation hints. Image and file parts show their filename, media type, and stored URL or provider file ID so the agent can reuse the attachment. Inline file data such as base64 or data URIs is omitted. Pass viewAttachment with cursor and partIndex to send inline attachment data to the model.

partType?:

'text' | 'tool-call' | 'tool-result' | 'reasoning' | 'image' | 'file'
Filter results to only include message parts of this type. Only applies to mode: "messages".

toolName?:

string
Filter results to only include tool-call and tool-result parts matching this tool name. Only applies to mode: "messages".

partIndex?:

number
Fetch a single message part at full detail by its positional index. Use this when a low-detail recall shows an interesting part at [p1] — call again with partIndex: 1 to see the full content without loading every part.

charOffset?:

number
= 0
Character position to continue reading a truncated single part. Only applies with cursor and partIndex. When a part is larger than the token budget, the result includes nextCharOffset — pass that exact value here in the next call to read the following chunk. The chunks concatenate to the original part text.

viewAttachment?:

boolean
= false
Show the attachment itself rather than only describing it. Requires cursor and partIndex, and applies when that part is an image or file. Inline data such as base64 or data URIs is sent as a native media part. Attachments stored as remote URLs or provider file IDs, media types outside image/* and application/pdf, and oversized payloads come back as an explanation instead.

before?:

string
Filter to threads created before this date in threads mode, or observations dated before it in search mode. Accepts ISO 8601 format (e.g. "2026-03-15", "2026-03-10T00:00:00Z").

after?:

string
Filter to threads created after this date in threads mode, or observations dated after it in search mode. Accepts ISO 8601 format (e.g. "2026-03-01", "2026-03-10T00:00:00Z").

Returns (messages mode)
Direct link to Returns (messages mode)

messages:

string
Formatted message content. Format depends on the detail level.

count:

number
Number of messages in this page.

cursor:

string
The cursor message ID used for this query.

page:

number
The page number returned.

limit:

number
The limit used for this query.

detail:

'low' | 'high'
The detail level used for this query.

hasNextPage:

boolean
Whether more messages exist after this page.

hasPrevPage:

boolean
Whether more messages exist before this page.

truncated?:

boolean
Present and true when the output was capped by the token budget. The agent can paginate or use partIndex to access remaining content. When a single part is itself too large, the partIndex result includes nextCharOffset for continuing within the part.

tokenOffset?:

number
Approximate number of tokens that were trimmed when truncated is true.

charOffset?:

number
On single-part results (partIndex), the character position this chunk starts at. 0 unless the call passed a charOffset.

nextCharOffset?:

number
On single-part results, present when the part was truncated and more content remains. Pass this value as charOffset in the next call to continue reading from where this chunk ended.

note?:

string
On truncated single-part results, the exact follow-up call for retrieving the next chunk.

Returns (threads mode)
Direct link to Returns (threads mode)

threads:

string
Formatted thread listing. Each thread shows its title, ID, and dates. The current thread is marked with ← current.

count:

number
Number of threads returned.

page:

number
The page number returned.

hasMore:

boolean
Whether more threads exist on the next page.

Returns (search mode)
Direct link to Returns (search mode)

results:

string
Search hits displayed by observation date, with excerpts or compact references for evidence already in context. Full entries include thread details, a relevance score, an observation group ID, and source message ranges.

count:

number
Number of displayed observation groups, including compact references. Can exceed limit when search fills covered hits with additional excerpts.

ModelByInputTokens
Direct link to ModelByInputTokens

ModelByInputTokens selects a model based on the input token count. It chooses the model for the smallest threshold that covers the actual input size.

Constructor
Direct link to Constructor

new ModelByInputTokens(config)

Where config is an object with upTo keys that map token thresholds (numbers) to model targets.

Example
Direct link to Example

import { ModelByInputTokens } from '@mastra/memory'

const selector = new ModelByInputTokens({
upTo: {
10_000: 'google/gemini-2.5-flash', // Fast for small inputs
40_000: 'openai/gpt-5-mini', // Stronger for medium inputs
1_000_000: 'openai/gpt-5.6-sol', // Most capable for large inputs
},
})

Behavior
Direct link to Behavior

  • Thresholds are sorted internally, so the order in the config object doesn't matter.
  • inputTokens ≤ smallest threshold → uses that threshold's model
  • inputTokens > largest threshold → resolve() throws an error. If this happens during an OM Observer or Reflector run, OM aborts via TripWire, so callers receive an empty text result or streamed tripwire instead of a normal assistant response.
  • OM computes the input token count for the Observer or Reflector call and resolves the matching model tier directly

Methods
Direct link to Methods

resolve:

(inputTokens: number) => MastraModelConfig
Returns the model for the given input token count. Throws if inputTokens exceeds the largest configured threshold. When this happens during an OM run, callers receive a TripWire/empty-text outcome instead of a normal assistant response.

getThresholds:

() => number[]
Returns the configured thresholds in ascending order. Useful for introspection.

skillResultRedactor
Direct link to skillResultRedactor

skillResultRedactor builds a beforeObservation transform hook that redacts the results of the built-in Agent Skills tools (skill, skill_search, and skill_read) before the Observer model sees them. Each result is replaced with a placeholder while the tool call is kept, so the Observer still records which skill was used and what it was called with, without the skill text.

import { Memory } from '@mastra/memory'
import { skillResultRedactor } from '@mastra/memory/hooks'

const memory = new Memory({
options: {
observationalMemory: {
model: 'google/gemini-2.5-flash',
hooks: {
beforeObservation: skillResultRedactor(),
},
},
},
})

Parameters
Direct link to Parameters

toolNames?:

readonly string[]
Tool names whose results are redacted. Defaults to the built-in skill tools: skill, skill_search, and skill_read. Use this to redact a different set of tool results.

Returns
Direct link to Returns

(input: { messages: MastraDBMessage[] } & ObserveHookContext) => { messages: MastraDBMessage[] } | undefined

The hook returns messages with matching tool results replaced by a placeholder, or undefined when no message matched (which passes the payload through unchanged). Tool calls, arguments, and every other message are left as they are.