Jev is an evaluation model from TypeSafe that returns structured decisions and probabilities instead of generated text. You provide the context and possible answers. LLMs can classify too, but Jev is built for bounded decisions that don't require multistep reasoning or generated text.
At Mastra, our initial explorations find that Jev is a good fit for some - but not every - classification problem. Having a fixed set of answers helps, but the model still needs enough context to choose between them.
In this post, we'll share an overview of Jev with use cases and working code examples using classifiers in Mastra.
Jev vs LLMs
Many chat models are post-trained with Reinforcement Learning from Human Feedback (RLHF), which optimizes for responses people prefer. A response that sounds convincing, however, is not necessarily a reliable decision for software.
Jev uses Reinforcement Learning for Calibrated Decisions (RLCD), which trains it to return decisions and calibrated probabilities.
Across many predictions from a well-calibrated model, outcomes assigned an 80% probability should occur about 80% of the time.
You give Jev a state, the information it needs to make the decision. That state can be plain text or structured JSON, such as a message, logs, or user information from your application.
In Mastra, classifier support exposes three question types:
- Choice: Selects one of the options you define.
- Score: Returns a position on an ordered scale, which can be fractional.
- Boolean: Returns the estimated probability that the answer is true.
Jev evaluates the state and returns a typed answer for each question.
Jev use cases and examples
Jev got a lot of attention from the community quickly. A week after its release, we ran a workshop on Building a Classifier in Mastra with Jev, building classifiers for several common use cases. The examples we walked through follow the same scenario: a VP of Engineering asks for pricing on 200 seats. Jev evaluates the lead or the draft. The agent writes the reply.
Classifier-as-judge (Jev-as-a-judge)
Classifier-as-judge uses an evaluation model to grade AI responses against defined criteria. Jev fits checks that need a category or score, so you can compare outputs without generating a written critique.
Our scorer example grades six sample replies: do the prices and seat counts match the supplied account data, does the reply ask for a next step, and does it concede on price?
Mastra's createClassifierScorer maps a classifier answer to an eval score. In our example, Jev identifies the next step proposed by the draft, and our scoring code assigns a value to each choice:
const nextStep = createClassifierScorer({
id: 'next-step',
name: 'Asks for a next step',
description: 'A meeting scores 1. Asking them to reply scores 0.4. Ending on the quote scores 0.',
classifier: draftScorer,
question: 'nextStep',
scores: { meeting: 1, reply: 0.4, none: 0 },
state,
});In this example, the state function supplies the lead, account figures, and draft. Each scorer measures a separate criterion. A high score from concedes means the draft gives ground on price.
Our scorers grade the sample drafts. They don't block a response. For judgments that require research, several reasoning steps, or a written explanation, an LLM judge may be a better fit.
Full scorer code · Workshop walkthrough
Jev for tool approval
Tool approval pauses an agent's proposed action for human review before it runs. Jev can help decide which actions need that review when the choice depends on interpreting the tool's arguments and context.
In our tool approval example, the tool's requireApproval callback asks Jev whether the quote needs human review. Jev evaluates the quote arguments and returns a probability. Our callback compares it with LINE, which we set to 0.7:
const result = await approvalCheck.evaluate({ state: input });
const chance = result.answers.needsReview.probability;
const ask = chance >= LINE;When our callback returns true, Mastra's tool approval flow suspends the tool call.
This example demonstrates Mastra's approval integration. Its seat-count and discount rules can also be enforced directly in code.
Our demo simulates the send and approves the paused call automatically. In an application, that resume step would follow the reviewer's approval.
Full approval code · Workshop walkthrough
Jev for workflow branching
Workflow branching uses conditions to choose which paths run next. Jev can turn the meaning of a message, such as buying intent, into typed answers your routing code can use.
Our lead classifier asks three separate questions:
| Question | Type | What it evaluates |
|---|---|---|
fit | Choice | Whether the person's title suggests a technical buying role |
intent | Score | Where the message falls between browsing and ready to buy |
icp | Boolean | The probability that the company fits the ideal customer profile |
In our lead-routing example, we disqualify a poor-fit lead before the agent runs. Other leads go to sales when intent is at least 2. Of those left, leads with an ICP probability above 0.8 go to nurture. The rest go to review.
Mastra's .classifier() step runs the evaluation, and .branch() uses the answers. This is the condition for the sales branch:
async ({ inputData }) =>
inputData.answers.fit.choice !== 'poor' &&
inputData.answers.intent.score >= 2We made the conditions mutually exclusive because .branch() runs every matching branch. This keeps a poor-fit lead from also entering the sales branch.
Full workflow code · Workshop walkthrough
Jev for content guardrails
Content guardrails check what an agent receives and returns against application policies. Jev is useful for checks that depend on meaning, returning a probability your code can use to flag or block content.
Our processor example screens out poor-fit leads, then uses a separate classifier to check whether the reply offers a discount or invites price negotiation.
The input classifier gets the lead and the output classifier gets the draft. We attach these checks with Mastra's ClassifierProcessor:
const result = await agent.generate(prompt(value), {
maxSteps: 3,
inputProcessors: [screenLead(value)],
outputProcessors: [checkDraft()],
});Jev returns the probability of a price concession. Our output processor calls context.abort() when that probability is at least 0.8. We set errorStrategy: 'warn' in this example, so the agent continues if a classifier call fails.
Full processor code · Workshop walkthrough
Where Jev is less useful (so far)
We tried to classify 78 open-source Mastra GitHub issues as urgent, high, or low using Jev, giving it only each issue's title and body as its state. Roughly 10 of the 78 classifications matched Abhi's judgment.
At first glance, issue triage sounds like a perfect decision-model problem. There are only three possible answers. But Jev doesn't investigate the codebase or product roadmap for you. The title and body are only a small representation of a much larger system, and priority can depend on who is affected, the severity of the failure, and what the team is building next.
That gives us a useful way to think about where Jev fits. If the task requires gathering more evidence or several steps of reasoning, a classifier alone may be the wrong tool. An agent can gather that evidence first, and a classifier can evaluate the resulting state.
If you already have the relevant data and need to quickly classify it so another part of your software can act on the result, Jev becomes much more compelling. Even then, compare its answers with expected labels and test the cases where it should defer to a human. TypeSafe's Jev 1.13 docs identify multistep reasoning, numerical precision, and irrelevant context as sources of errors.
How to get started
Classifier support is live in Mastra. We shipped support soon after Jev launched so you can use Jev and other compatible evaluation models across workflows, agent guardrails, tool approval checks, and evals.
To try it, clone the workshop repo, follow the setup instructions, and run pnpm score in examples/15-classifier with your TypeSafe API key configured. That runs our lead-scoring example without an agent. From there, try the examples above, or use the Classifier reference to add it to an existing Mastra project.
Pick one decision in your agent and try it with Jev. Define the possible answers, supply the relevant state, and compare the results with what you expect. Keep the questions focused and let your application decide what to do with the answers.
