createScorer
Mastra provides a unified createScorer factory that allows you to define custom scorers for evaluating input/output pairs. You can use either native JavaScript functions or LLM-based prompt objects for each evaluation step. Custom scorers can be added to Agents and Workflow steps.
How to create a custom scorerDirect link to How to create a custom scorer
Use the createScorer factory to define your scorer with a name, description, and optional judge configuration. Then chain step methods to build your evaluation pipeline. You must provide at least a generateScore step.
Prompt object steps are step configurations expressed as objects with description + createPrompt (and outputSchema for preprocess/analyze). These steps invoke the judge LLM. Function steps are plain functions and never call the judge.
import { createScorer } from '@mastra/core/evals'
const scorer = createScorer({
id: 'my-custom-scorer',
name: 'My Custom Scorer', // Optional, defaults to id
description: 'Evaluates responses based on custom criteria',
type: 'agent', // Optional: for agent evaluation with automatic typing
judge: {
model: myModel,
instructions: 'You are an expert evaluator...',
},
})
.preprocess({/* step config */})
.analyze({/* step config */})
.generateScore(({ run, results }) => {
// Return a number
})
.generateReason({/* step config */})
createScorer optionsDirect link to createscorer-options
id:
name is not provided.name?:
id if not provided.description:
judge?:
model:
instructions:
jsonPromptInjection?:
inputProcessors?:
outputProcessors?:
errorProcessors?:
processAPIError and can inspect LLM API rejections and signal a retry, e.g. StreamErrorRetryProcessor. Legacy model adapters use generateLegacy() and do not run error processors.maxProcessorRetries?:
type?:
prepareRun?:
This function returns a scorer builder that you can chain step methods onto. See the MastraScorer reference for details on the .run() method and its input/output.
The judge only runs for steps defined as prompt objects (preprocess, analyze, generateScore, generateReason in prompt mode). If you use function steps only, the judge is never called and there is no LLM output to inspect. In that case, any score/reason must be produced by your functions.
When a prompt-object step runs, its structured LLM output is stored in the corresponding result field (preprocessStepResult, analyzeStepResult, or the value consumed by calculateScore in generateScore).
Retry judge requestsDirect link to Retry judge requests
Use the existing judge errorProcessors configuration to retry transient failures inside a failed judge request. This doesn't retry the scorer workflow, trace target, batch item, score write, or a completed scorer step.
@mastra/core 1.49.0 doesn't include scorer error-processor configuration. Upgrade to a version with scorer processor support, or backport that focused change, before using this configuration.
The following example uses one bounded retry budget. Set the processor maxRetries and judge.maxProcessorRetries to the same value. Keep the internal judge agent's model retries at the default of 0 so model retries don't multiply processor attempts.
import { createScorer } from '@mastra/core/evals'
import { StreamErrorRetryProcessor } from '@mastra/core/processors'
const isTransientNetworkError = (error: unknown) =>
error instanceof Error && /ECONNRESET|ETIMEDOUT|socket hang up/i.test(error.message)
const retryProcessor = new StreamErrorRetryProcessor({
maxRetries: 2,
maxRetryAfterMs: 30_000,
delayMs: ({ retryCount }) => Math.min(1_000 * 2 ** retryCount, 30_000),
matchers: [isTransientNetworkError],
retryUnknownErrors: false,
})
export const responseQuality = createScorer({
id: 'response-quality',
description: 'Scores response quality',
judge: {
model: myModel,
instructions: 'Return a score and concise reason.',
errorProcessors: [retryProcessor],
maxProcessorRetries: 2,
},
})
.generateScore({
description: 'Score the response quality.',
createPrompt: ({ run }) => `Score: ${run.output}`,
})
.generateReason({
description: 'Explain the score.',
createPrompt: () => 'Explain the score.',
})
With this configuration, a failed request has at most three provider attempts: the initial request plus two processor retries. If generateScore completes and generateReason receives a retryable failure, only generateReason retries.
StreamErrorRetryProcessor honors retryable provider metadata and narrow custom matchers. It keeps retryUnknownErrors disabled by default, so authentication, invalid-request, and context-length errors fail immediately unless you explicitly match them. It bounds Retry-After values to 30_000 milliseconds by default. Use maxRetryAfterMs to change that cap.
Avoid adding outer scorer or workflow retries. Avoid combining a nonzero model retry setting with this processor unless you intentionally accept additional attempts.
Override retries for one stepDirect link to Override retries for one step
A step's judge configuration overrides scorer-level judge fields. Processor arrays replace the scorer-level arrays. Omit maxProcessorRetries in the step configuration to inherit the scorer-level numeric cap.
Coordinated processor retries require a judge model that uses Mastra’s current generation API. Legacy model adapters call generateLegacy(), bypass error processors, and use that API's separate AI SDK maxRetries default of 2.
Type safetyDirect link to Type safety
You can specify input/output types when creating scorers for better type inference and IntelliSense support:
Agent Type ShortcutDirect link to Agent Type Shortcut
For evaluating agents, use type: 'agent' to automatically get the correct types for agent input/output:
import { createScorer } from '@mastra/core/evals'
// Agent scorer with automatic typing
const agentScorer = createScorer({
id: 'agent-response-quality',
description: 'Evaluates agent responses',
type: 'agent', // Automatically provides ScorerRunInputForAgent/ScorerRunOutputForAgent
})
.preprocess(({ run }) => {
// run.input is automatically typed as ScorerRunInputForAgent
const userMessage = run.inputData.inputMessages[0]?.content
return { userMessage }
})
.generateScore(({ run, results }) => {
// run.output is automatically typed as ScorerRunOutputForAgent
const response = run.output[0]?.content
return response.length > 10 ? 1.0 : 0.5
})
Custom Types with GenericsDirect link to Custom Types with Generics
For custom input/output types, use the generic approach:
import { createScorer } from '@mastra/core/evals'
type CustomInput = { query: string; context: string[] }
type CustomOutput = { answer: string; confidence: number }
const customScorer = createScorer<CustomInput, CustomOutput>({
id: 'custom-scorer',
description: 'Evaluates custom data',
}).generateScore(({ run }) => {
// run.input is typed as CustomInput
// run.output is typed as CustomOutput
return run.output.confidence
})
Built-in Agent TypesDirect link to Built-in Agent Types
ScorerRunInputForAgent- ContainsinputMessages,rememberedMessages,systemMessages, andtaggedSystemMessagesfor agent evaluationScorerRunOutputForAgent- Array of agent response messages
Using these types provides autocomplete, compile-time validation, and better documentation for your scoring logic.
Trace scoring with agent typesDirect link to Trace scoring with agent types
When you use type: 'agent', your scorer is compatible for both adding directly to agents and scoring traces from agent interactions. The scorer automatically transforms trace data into the proper agent input/output format:
const agentTraceScorer = createScorer({
id: 'agent-trace-length',
description: 'Evaluates agent response length',
type: 'agent',
}).generateScore(({ run }) => {
// Trace data is automatically transformed to agent format
const userMessages = run.inputData.inputMessages
const agentResponse = run.output[0]?.content
// Score based on response length
return agentResponse?.length > 50 ? 0 : 1
})
// Register with Mastra for trace scoring
const mastra = new Mastra({
scorers: {
agentTraceScorer,
},
})
Step method signaturesDirect link to Step method signatures
preprocessDirect link to preprocess
Optional preprocessing step that can extract or transform data before analysis.
Function Mode:
Function: ({ run, results }) => any
run.input:
[{ role: 'user', content: 'hello world' }]. If the scorer is used in a workflow, this will be the input of the workflow.run.output:
run.runId:
run.requestContext?:
results:
Returns: any
The method can return any value. The returned value will be available to subsequent steps as preprocessStepResult.
Prompt Object Mode:
description:
outputSchema:
createPrompt:
judge?:
analyzeDirect link to analyze
Optional analysis step that processes the input/output and any preprocessed data.
Function Mode:
Function: ({ run, results }) => any
run.input:
[{ role: 'user', content: 'hello world' }]. If the scorer is used in a workflow, this will be the input of the workflow.run.output:
run.runId:
run.requestContext?:
results.preprocessStepResult?:
Returns: any
The method can return any value. The returned value will be available to subsequent steps as analyzeStepResult.
Prompt Object Mode:
description:
outputSchema:
createPrompt:
judge?:
generateScoreDirect link to generatescore
Required step that computes the final numerical score.
Function Mode:
Function: ({ run, results }) => number
run.input:
[{ role: 'user', content: 'hello world' }]. If the scorer is used in a workflow, this will be the input of the workflow.run.output:
run.runId:
run.requestContext?:
results.preprocessStepResult?:
results.analyzeStepResult?:
Returns: number
The method must return a numerical score.
Prompt Object Mode:
description:
outputSchema:
createPrompt:
judge?:
When using prompt object mode, you must also provide a calculateScore function to convert the LLM output to a numerical score:
calculateScore:
generateReasonDirect link to generatereason
Optional step that provides an explanation for the score.
Function Mode:
Function: ({ run, results, score }) => string
run.input:
[{ role: 'user', content: 'hello world' }]. If the scorer is used in a workflow, this will be the input of the workflow.run.output:
run.runId:
run.requestContext?:
results.preprocessStepResult?:
results.analyzeStepResult?:
score:
Returns: string
The method must return a string explaining the score.
Prompt Object Mode:
description:
createPrompt:
judge?:
All step functions can be async.