Tone consistency scorer
The createToneScorer() function evaluates the text's emotional tone and sentiment consistency. It can operate in two modes: comparing the agent output's sentiment against a configured referenceTone, or analyzing tone stability across the sentences of the agent output.
ParametersDirect link to Parameters
The createToneScorer() function accepts an optional config object with the following properties:
referenceTone:
This function returns an instance of the MastraScorer class. See the MastraScorer reference for details on the .run() method and its input/output.
.run() returnsDirect link to run-returns
runId:
preprocessStepResult:
score:
.run() returns a result in the following shape:
{
runId: string,
preprocessStepResult: {
score: number,
responseSentiment?: number,
referenceSentiment?: number,
difference?: number,
avgSentiment?: number,
sentimentVariance?: number,
},
score: number
}
Scoring detailsDirect link to Scoring details
The scorer evaluates sentiment consistency through tone pattern analysis and mode-specific scoring.
Scoring ProcessDirect link to Scoring Process
- Analyzes tone patterns:
- Extracts sentiment features
- Computes sentiment scores
- Measures tone variations
- Calculates mode-specific score:
Tone Consistency (
referenceToneset):- Compares the output's sentiment with the
referenceTonesentiment - Calculates the absolute sentiment difference
- Score = max(0, 1 - sentiment_difference)
Tone Stability (no
referenceTone): - Analyzes sentiment across the output's sentences
- Calculates sentiment variance
- Score = max(0, 1 - sentiment_variance)
- Compares the output's sentiment with the
Final score: mode_specific_score
Score interpretationDirect link to Score interpretation
(0-1)
- 1.0: Perfect tone consistency/stability
- 0.7-0.9: Strong consistency with minor variations
- 0.4-0.6: Moderate consistency with noticeable shifts
- 0.1-0.3: Poor consistency with major tone changes
- 0.0: No consistency - completely different tones
preprocessStepResultDirect link to preprocessstepresult
Object with tone metrics:
- responseSentiment: Sentiment score for the response (comparison mode).
- referenceSentiment: Sentiment score for the
referenceTone(comparison mode). - difference: Absolute difference between sentiment scores (comparison mode).
- avgSentiment: Average sentiment across sentences (stability mode).
- sentimentVariance: Variance of sentiment across sentences (stability mode).
ExampleDirect link to Example
Evaluate whether agent responses match a positive reference tone:
import { runEvals } from '@mastra/core/evals'
import { createToneScorer } from '@mastra/evals/scorers/prebuilt'
import { myAgent } from './agent'
const scorer = createToneScorer({
referenceTone: 'We are happy to help and glad you had a great experience!',
})
const result = await runEvals({
data: [
{ input: 'How was your experience with our service?' },
{ input: 'Tell me about the customer support' },
],
scorers: [scorer],
target: myAgent,
onItemComplete: ({ scorerResults }) => {
console.log({
score: scorerResults[scorer.id].score,
})
},
})
console.log(result.scores)
For more details on runEvals, see the runEvals reference.
To add this scorer to an agent, see the Scorers overview guide.