MastraScorer
The MastraScorer class is the base class for all scorers in Mastra. It provides a standard .run() method for evaluating input/output pairs and supports multi-step scoring workflows with preprocess → analyze → generateScore → generateReason execution flow.
Most users should use createScorer to create scorer instances. Direct instantiation of MastraScorer isn't recommended.
How to get a MastraScorer instanceDirect link to how-to-get-a-mastrascorer-instance
Use the createScorer factory function, which returns a MastraScorer instance:
import { createScorer } from '@mastra/core/evals'
const scorer = createScorer({
name: 'My Custom Scorer',
description: 'Evaluates responses based on custom criteria',
}).generateScore(({ run, results }) => {
// scoring logic
return 0.85
})
// scorer is now a MastraScorer instance
.run() methodDirect link to run-method
The .run() method is the primary way to execute your scorer and evaluate input/output pairs. It processes the data through your defined steps (preprocess → analyze → generateScore → generateReason) and returns a detailed result object with the score, reasoning, and intermediate results.
const result = await scorer.run({
input: 'What is machine learning?',
output: 'Machine learning is a subset of artificial intelligence...',
runId: 'optional-run-id',
requestContext: {/* optional context */},
})
.run() inputDirect link to run-input
input:
output:
runId:
requestContext:
groundTruth:
.run() returnsDirect link to run-returns
runId:
score:
reason:
preprocessStepResult:
analyzeStepResult:
preprocessPrompt:
analyzePrompt:
generateScorePrompt:
generateReasonPrompt:
judge:
Judge resultsDirect link to Judge results
The optional judge record contains details about the judge model calls made by prompt-based scorer steps. Its known keys are preprocess, analyze, generateScore, and generateReason. Each key contains an ordered executions array.
interface ScorerJudgeExecutionSuccess {
status: 'success'
prompt: string
output: JSONValue
judgeModelId: string
judgeProvider?: string
usage: ScorerJudgeUsage
attemptCount: number
modelCallCount: number
durationMs: number
cost?: {
amount: number
unit: string
source: string
}
}
type ScorerJudgeExecution = ScorerJudgeExecutionSuccess
interface ScorerJudgeUsage {
inputTokens?: number
outputTokens?: number
totalTokens?: number
reasoningTokens?: number
cachedInputTokens?: number
cacheCreationInputTokens?: number
}
type ScorerJudgeResults = Partial<
Record<
'preprocess' | 'analyze' | 'generateScore' | 'generateReason',
{ executions: ScorerJudgeExecution[] }
>
>
Use the step key to access its judge execution details:
const execution = result.judge?.generateScore?.executions[0]
console.log(execution?.status)
console.log(execution?.judgeModelId)
console.log(execution?.usage.totalTokens)
console.log(execution?.durationMs)
The status value describes the outcome of the logical prompt-step execution, not the quality of the evaluated response. A structured-output fallback that eventually succeeds creates one success execution with an attemptCount greater than one.
attemptCount counts judge invocations, including a structured-output fallback. modelCallCount counts the completed model steps across those attempts. durationMs covers the full prompt-step execution.
Function steps don't create judge entries. Usage in this record belongs to the scorer's judge model, not the agent or workflow being evaluated. The optional cost field is present only when the execution directly reports an authoritative cost, source, and unit.
Use Mastra metrics to query aggregate usage, latency, and estimated cost across scorer runs. The judge record describes one scorer run and doesn't query metrics or traces.
Step execution flowDirect link to Step execution flow
When you call .run(), the MastraScorer executes the defined steps in this order:
- preprocess (optional): Extracts or transforms data
- analyze (optional): Processes the input/output and preprocessed data
- generateScore (required): Computes the numerical score
- generateReason (optional): Provides explanation for the score
Each step receives the results from previous steps, allowing you to build complex evaluation pipelines.
Usage exampleDirect link to Usage example
const scorer = createScorer({
name: 'Quality Scorer',
description: 'Evaluates response quality',
})
.preprocess(({ run }) => {
// Extract key information
return { wordCount: run.output.split(' ').length }
})
.analyze(({ run, results }) => {
// Analyze the response
const hasSubstance = results.preprocessStepResult.wordCount > 10
return { hasSubstance }
})
.generateScore(({ results }) => {
// Calculate score
return results.analyzeStepResult.hasSubstance ? 1.0 : 0.0
})
.generateReason(({ score, results }) => {
// Explain the score
const wordCount = results.preprocessStepResult.wordCount
return `Score: ${score}. Response has ${wordCount} words.`
})
// Use the scorer
const result = await scorer.run({
input: 'What is machine learning?',
output: 'Machine learning is a subset of artificial intelligence...',
})
console.log(result.score) // 1.0
console.log(result.reason) // "Score: 1.0. Response has 12 words."
IntegrationDirect link to Integration
MastraScorer instances can be used for agents and workflow steps
See the createScorer reference for detailed information on defining custom scoring logic.