Skip to main content

Search and indexing

Search lets agents find relevant content in indexed workspace files. When an agent needs to answer a question or find information, it can search the indexed content instead of reading every file.

Search works with both mounts and a workspace-only filesystem. With mounts, paths from every mounted filesystem are available through one composite filesystem. With filesystem, search uses that provider directly. In both cases, queries search the workspace index rather than reading live files on demand.

Use workspace search when your agent needs to:

  • Find exact terms, filenames, or error messages with BM25 keyword search
  • Find conceptually related content with vector search
  • Combine keyword and semantic results with hybrid search
  • Search a large set of files without reading each file in full
  • Index content from files, databases, or APIs

Quickstart
Direct link to Quickstart

Enable BM25 search, index content, and query it through the workspace:

src/mastra/workspaces.ts
import { Workspace, LocalFilesystem } from '@mastra/core/workspace'

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
})

await workspace.index('/docs/guide.md', 'Reset passwords from the account settings page.')

const results = await workspace.search('password reset')
console.log(results)

This configuration also gives agents tools for searching and indexing workspace content.

How it works
Direct link to How it works

Workspace search has two phases: indexing and querying.

Indexing
Direct link to Indexing

Content must be indexed before it can be searched. When you index a document:

  • The content is tokenized (split into searchable terms)
  • For BM25: term frequencies and document statistics are computed
  • For vector: the content is embedded using your embedder function and stored in the vector store

Each indexed document has:

  • id - A unique identifier (typically the file path)
  • content - The text content
  • metadata - Optional key-value data stored with the document

Querying
Direct link to Querying

When you search:

  1. The query is processed using the same tokenization/embedding as indexing
  2. Documents are scored based on relevance to the query
  3. Results are ranked by score and returned with the matching content

Workspaces support three search modes: BM25 keyword search and vector semantic search, plus hybrid search that combines both.

BM25 scores documents based on term frequency and document length. It works well for exact matches and specific terminology.

src/mastra/workspaces.ts
import { Workspace, LocalFilesystem } from '@mastra/core/workspace'

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
})

For custom BM25 parameters (k1 is term frequency saturation, b is document length normalization):

src/mastra/workspaces.ts
const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: {
k1: 1.5,
b: 0.75,
},
})

Vector search uses embeddings to find semantically similar content. It requires a vector store and embedder function.

src/mastra/workspaces.ts
import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
import { PineconeVector } from '@mastra/pinecone'
import { embed } from 'ai'
import { openai } from '@ai-sdk/openai'

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
vectorStore: new PineconeVector({
apiKey: process.env.PINECONE_API_KEY,
index: 'workspace-index',
}),
embedder: async (text: string) => {
const { embedding } = await embed({
model: openai.embedding('text-embedding-3-small'),
value: text,
})
return embedding
},
})

Batch embedding
Direct link to Batch embedding

The embedder above takes one text at a time. Indexing a workspace with hundreds of files calls the provider hundreds of times, which is slow and expensive.

When the provider supports batching (for example, OpenAI's embedMany), pass an embedder that takes an array of texts and accepts many embeddings back in one call. To opt in, set a batch: true property on the function. Mastra checks for that property at runtime and switches to the batched path.

The following example replaces the single-text embedder with a batched one. The embedder function takes an array and returns an array of embeddings in the same order, plus carries two extra properties:

  • batch: true: marks the function as batch-capable. Without this property, Mastra calls it one text at a time.
  • maxBatchSize: the largest array the provider accepts in one call. Mastra splits larger requests into chunks of this size and sends them in parallel. Set this to your provider's documented limit. For example, OpenAI accepts 2048 and Cohere accepts 96. Voyage accepts 128. Omit it to send every pending text in one request.
src/mastra/workspaces.ts
import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
import { PineconeVector } from '@mastra/pinecone'
import { embedMany } from 'ai'
import { openai } from '@ai-sdk/openai'

const model = openai.embedding('text-embedding-3-small')

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
vectorStore: new PineconeVector({
apiKey: process.env.PINECONE_API_KEY,
index: 'workspace-index',
}),
embedder: Object.assign(
async (texts: string[]) => {
const { embeddings } = await embedMany({ model, values: texts })
return embeddings
},
{ batch: true as const, maxBatchSize: 2048 },
),
})

Object.assign adds the batch and maxBatchSize properties to the embedder function. Mastra reads them as metadata and never passes them to the provider.

Single-text embedders still work. The function signature (text: string) => Promise<number[]> is unchanged, so existing code keeps running without modification.

Configure both BM25 and vector search to enable hybrid mode, which combines keyword matching with semantic understanding.

src/mastra/workspaces.ts
const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
vectorStore: pineconeVector,
embedder: embedderFn,
})

Custom index name
Direct link to Custom index name

By default, the search index name is derived from the workspace ID. To set a custom name, use searchIndexName:

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
searchIndexName: 'my_workspace_vectors',
})

The index name must be a valid SQL identifier: start with a letter or underscore, contain only letters, numbers, or shows, and be at most 63 characters long.

Indexing content
Direct link to Indexing content

Manual indexing
Direct link to Manual indexing

Use workspace.index() to add content to the search index programmatically. The file paths become document IDs. You can also pass metadata for each document.

// Basic indexing
await workspace.index('/docs/guide.md', 'Content of the guide...')

// Index with metadata for filtering or context
await workspace.index('/docs/api.md', apiDocContent, {
metadata: {
category: 'api',
version: '2.0',
},
})

Manual indexing is useful when:

  • You're indexing content that doesn't come from files (e.g., database records, API responses)
  • You want to pre-process or chunk content before indexing
  • You need to add custom metadata to documents

Auto-indexing
Direct link to Auto-indexing

Configure autoIndexPaths to automatically index files when the workspace initializes. Each entry can be a directory path (indexed recursively) or a glob pattern for selective indexing.

autoIndexPaths works with a static filesystem or with mounts. For mounts, include the mount prefix in each path, such as /docs/**/*.md. Resolver-backed filesystems aren't auto-indexed during workspace initialization because the provider is only selected for a request; index their content manually instead.

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
autoIndexPaths: ['docs', 'support/faq'],
})

await workspace.init()

When init() is called, all matching files are read and indexed for search. The file path becomes the document ID.

Glob patterns let you index specific file types:

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
autoIndexPaths: ['docs/**/*.md', 'support/**/*.txt'],
})

Searching
Direct link to Searching

Use workspace.search() to find relevant content. Results are ranked by relevance score.

const results = await workspace.search('password reset')

for (const result of results) {
console.log(`${result.id}: ${result.score}`)
console.log(result.content)
}

Search options
Direct link to Search options

You can customize the search behavior with options:

const results = await workspace.search('authentication flow', {
topK: 10,
mode: 'hybrid',
minScore: 0.5,
vectorWeight: 0.5,
})
OptionDescription
topKMaximum number of results to return. Default: 5
modeSearch mode: 'bm25', 'vector', or 'hybrid'. Defaults to the best available mode based on configuration.
minScoreFilter out results below this score threshold (0-1).
vectorWeightIn hybrid mode, how much to weight vector scores vs BM25. 0 = all BM25, 1 = all vector, 0.5 = equal.

Search results
Direct link to Search results

Each result contains:

interface SearchResult {
id: string // Document ID (typically file path)
content: string // The matching content
score: number // Relevance score (0-1)
lineRange?: {
// Lines where the match was found
start: number
end: number
}
metadata?: Record<string, unknown> // Metadata stored with the document
scoreDetails?: {
// Score breakdown (hybrid mode only)
vector?: number
bm25?: number
}
}

Understanding scores:

  • Scores range from 0 to 1, where 1 is a perfect match
  • BM25 scores are normalized based on the best match in the result set
  • Vector scores represent cosine similarity between query and document embeddings
  • In hybrid mode, scores are combined using the vectorWeight parameter

When to use each mode
Direct link to When to use each mode

ModeBest forExample queries
bm25Exact terms, technical queries, code"useState hook", "404 error", "config.yaml"
vectorConceptual queries, natural language"how to handle user authentication", "best practices for error handling"
hybridGeneral search, unknown query typesMost agent use cases

Agent tools
Direct link to Agent tools

When you configure search on a workspace, agents receive mastra_workspace_search and mastra_workspace_index tools. Use WORKSPACE_TOOLS.SEARCH constants to configure them independently. For example, remove indexing when the agent should search existing content without changing the index:

src/mastra/workspaces.ts
import { LocalFilesystem, Workspace, WORKSPACE_TOOLS } from '@mastra/core/workspace'

const workspace = new Workspace({
filesystem: new LocalFilesystem({ basePath: './workspace' }),
bm25: true,
tools: {
[WORKSPACE_TOOLS.SEARCH.INDEX]: {
enabled: false,
},
},
})

See workspace class reference for the complete tool list and WorkspaceToolsConfig for shared settings.