Realtime voice
QuickstartDirect link to Quickstart
Realtime voice turns a Mastra agent into a live call a user can talk over, in the browser or over the phone. Mastra builds it on LiveKit, an open source WebRTC platform for realtime audio and video.
The @mastra/livekit package connects Mastra agents to the LiveKit Agents framework: LiveKit owns the audio loop like voice activity detection, streaming speech-to-text, semantic turn detection, barge-in, and text-to-speech. Your Mastra agent generates every reply with its own model, tools, and memory.
Use realtime voice when you need low-latency, interruptible voice conversations. For provider-based speech-to-speech without LiveKit, see Speech to Speech.
These steps take you from an empty project to a voice agent you can talk to. A voice session has two moving parts you set up here: an API route on your Mastra server that hands out access tokens, and a separate worker process that runs the audio pipeline and calls your agent each turn.
Install the integration package along with the LiveKit plugins for voice activity detection and turn detection:
- npm
- pnpm
- Yarn
- Bun
npm install @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekitpnpm add @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekityarn add @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekitbun add @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekitSet your LiveKit credentials inside an
.envfile. Create a free project on LiveKit Cloud, or run a local server withlivekit-server --dev:.envLIVEKIT_URL=wss://your-project.livekit.cloudLIVEKIT_API_KEY=your-api-keyLIVEKIT_API_SECRET=your-api-secretAdd a voice agent to your Mastra instance and expose a connection route. The
liveKitConnectionRoute()helper adds aPOST /voice/livekit/connection-detailsendpoint that mints a LiveKit token and dispatches your agent into a room:src/mastra/index.tsimport { Mastra } from '@mastra/core/mastra'import { Agent } from '@mastra/core/agent'import { liveKitConnectionRoute } from '@mastra/livekit'const supportAgent = new Agent({id: 'support',name: 'Support',instructions: 'You are a friendly phone support agent. Keep replies short and conversational.',model: 'openai/gpt-5-mini',})export const mastra = new Mastra({agents: { support: supportAgent },server: {apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],},})Create the worker. It runs as a separate process, answers LiveKit sessions, and calls your agent each turn. Worker APIs live on the
@mastra/livekit/workerentry point, so the Mastra server never loads the LiveKit agents runtime. This example uses LiveKit Inference model strings for speech-to-text and text-to-speech, so no provider plugins are required:src/mastra/voice-worker.tsimport { fileURLToPath } from 'node:url'import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker'import { mastra } from './index'export default createLiveKitWorker({mastra,agent: 'support',stt: 'deepgram/nova-3',tts: 'cartesia/sonic-3',turnDetection: 'multilingual',greeting: 'Hi! How can I help you today?',})if (process.argv[1] === fileURLToPath(import.meta.url)) {runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })}The
agentoption selects which Mastra agent answers each session. Pass a fixed key as shown, or omit it to use theagentIdfrom the dispatch metadata, so one worker can serve every agent on your Mastra instance.Download the turn detection and voice activity detection models once. Then run the worker in one terminal and your Mastra server in another:
npx livekit-agents download-filesnpx tsx src/mastra/voice-worker.ts dev- npm
- pnpm
- Yarn
- Bun
npm run devpnpm run devyarn devbun run devThe worker registers with your LiveKit server and waits for sessions, while
mastra devserves the connection route.Talk to your agent. Open the hosted LiveKit Agents Playground and connect it to your project to start a call without building a frontend.
To wire up your own app instead, call the connection route for a token.
POST /voice/livekit/connection-detailsaccepts optionalagentId,threadId, andresourceIdfields in the request body and returns:{"serverUrl": "wss://your-project.livekit.cloud","roomName": "mastra-voice-a1b2c3d4","participantName": "user-1","participantToken": "eyJhbGci..."}This response matches the contract used by LiveKit's frontend starters, so apps built from agent-starter-react or the LiveKit React components work without changes.
Turn detection and interruptionsDirect link to Turn detection and interruptions
LiveKit decides when the user finished speaking and when the agent was interrupted. The defaults work well; tune them with turnHandling:
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
turnHandling: {
endpointing: { mode: 'dynamic', minDelay: 300, maxDelay: 3000 },
interruption: { minDuration: 500, resumeFalseInterruption: true },
},
})
turnDetection: 'multilingual': Runs LiveKit's semantic end-of-turn model locally on CPU. It reads the live transcript to avoid cutting users off mid-thought. Use'vad'or'stt'for silence-based endpointing instead.endpointing: Bounds how long the agent waits after the user stops speaking.interruption: Controls barge-in. When the user speaks over the agent, LiveKit stops playback and cancels the in-flight Mastra stream, so token generation stops too.preemptiveGeneration: Starts the Mastra agent's reply while the user is still finishing, hiding time-to-first-token. The worker disables it by default: each preemptive attempt runs the Mastra agent on an interim transcript, and every run persists the user message, which duplicates messages in the thread. Re-enable it withpreemptiveGeneration: { enabled: true }if latency matters more than exact thread history.
See the LiveKit turn detection docs for all options.
Per-call voices and transcriptionDirect link to Per-call voices and transcription
The top-level stt and tts options apply to every call. To pick them per call, one voice or language per tenant, set the configuration.stt and configuration.tts resolvers instead. Each resolver runs once per call with the dispatch metadata, request context, room name, and job context, and returns a value accepted by the matching top-level option. The value is either a plugin instance or an inference model string. Return undefined to fall back to the top-level option.
The following example gives each tenant its own text-to-speech voice, keyed off the tenant entry in the dispatch metadata:
import * as cartesia from '@livekit/agents-plugin-cartesia'
// One voice id per tenant, resolved from the dispatch metadata on each call.
const tenantVoices: Record<string, string> = {
meridian: 'your-cartesia-voice-id-1',
coastal: 'your-cartesia-voice-id-2',
}
// The resolver runs during call setup, so cache plugin instances across calls.
const ttsByVoice = new Map<string, cartesia.TTS>()
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
configuration: {
tts: ({ requestContext }) => {
const voice = tenantVoices[requestContext?.tenant as string]
if (!voice) return undefined // fall back to the top-level `tts`
let tts = ttsByVoice.get(voice)
if (!tts) {
tts = new cartesia.TTS({ voice })
ttsByVoice.set(voice, tts)
}
return tts
},
},
})
configuration.stt works the same way for per-call transcription, for example a different transcription model or language per tenant. The greeting has a matching per-call form: configuration.greeting.text accepts a resolver with the same call context, so one worker can open with each tenant's own phrasing.
Memory and threadsDirect link to Memory and threads
When the resolved Mastra agent has memory configured, each call becomes one memory thread:
threaddefaults to thethreadIdfrom dispatch metadata, then to the room name.resourcedefaults to theresourceIdfrom dispatch metadata, then to the thread. Send your end user's id here so calls group under the right user. Mastra Studio sends the agent id, matching how its sidebar lists threads.- When the thread doesn't exist yet, the worker creates it titled "Voice call" with metadata
{ source: 'livekit' }, and the spoken greeting is saved as the first assistant message so the thread reads as a full call transcript (disable withpersistGreeting: false).
Each turn sends only the new user input; Mastra Memory supplies history, semantic recall, and working memory. Pin a session to an existing thread by passing threadId in the connection request body, which is useful for continuing a text conversation by voice. In Studio, starting a call from an open chat binds the call to that thread, and the transcript fills into the chat after each exchange.
When a user interrupts the agent, the in-flight generation aborts and nothing from that turn is persisted at that moment. LiveKit keeps the part the user actually heard in its transcript, and on the next turn the worker re-sends that heard-only fragment so the thread backfills to match the call. A user who hangs up right after interrupting leaves that final fragment unrecorded. See interrupted turns for the details and a reconciliation recipe.
Speak while tools runDirect link to Speak while tools run
Voice conversations can't go silent while a slow tool runs. Use toolFeedback to speak a short phrase when the Mastra agent starts a tool call:
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
toolFeedback: ({ toolName }) =>
toolName === 'searchOrders' ? 'Let me look that up.' : undefined,
})
The phrase is spoken as part of the reply and recorded in the transcript.
Generate replies with a workflowDirect link to Generate replies with a workflow
By default the worker generates each reply with a Mastra agent. To run multi-step logic per turn (for example classify intent, route, call tools in sequence, then compose a reply), generate replies with a Mastra workflow instead. Set workflow in place of agent.
LiveKit still owns the audio loop and calls into Mastra once per turn, so the workflow runs to completion each turn. The workflow can't suspend or resume, and no conversation state carries between turns. Pass the transcript in through workflowInput so the workflow stays stateless:
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra,
workflow: 'phoneConversation',
workflowInput: ({ chatCtx }) => ({ history: chatContextToMessages(chatCtx) }),
replyStep: 'generateResponse',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
})
A workflow streams structured step events, not text. To speak tokens as they generate, the reply step pipes its agent's text into the step writer:
const generateResponse = createStep({
id: 'generateResponse',
// input and output schemas omitted
execute: async ({ inputData, mastra, writer, abortSignal }) => {
const stream = await mastra.getAgent('voice').stream(inputData.history, { abortSignal })
await stream.textStream.pipeTo(writer)
return { assistantMessage: await stream.text }
},
})
replyStep: Restricts spoken output to one step. Omit it to speak every step that writes to itswriter.resultText: A fallback that derives the reply from the final run result when no step streams text. Streaming throughwritergives lower time-to-first-token, so prefer it.abortSignal: Forward the step'sabortSignalintoagent.stream()so barge-in stops generation promptly. When the user interrupts, the worker cancels the run.generate: For full control, pass ageneratefunction instead. It can be any reply generator that turns a turn into a text stream.
With a workflow, the worker doesn't persist turns automatically the way an agent's stream() does. Persist conversation history inside the workflow, or keep the LiveKit transcript as the source of truth and pass it in each turn.
Use Mastra as the LLM componentDirect link to Use Mastra as the LLM component
createLiveKitWorker() owns the LiveKit session for you. To own the session yourself, use MastraLLM instead: a standard LiveKit LLM plugin that puts a Mastra agent in the llm slot of your own voice.AgentSession. The Mastra app, agent loop, tools, memory, observability, runs on your Mastra server, and the worker reaches it over HTTP. The worker process needs no Mastra app, database, or model provider keys.
import { fileURLToPath } from 'node:url'
import { defineAgent, voice } from '@livekit/agents'
import * as silero from '@livekit/agents-plugin-silero'
import { MastraLLM } from '@mastra/livekit/plugin'
import { runLiveKitWorker } from '@mastra/livekit/worker'
export default defineAgent({
entry: async ctx => {
await ctx.connect()
const session = new voice.AgentSession({
llm: new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
memory: { thread: ctx.room.name!, resource: 'user-7' },
}),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
vad: await silero.VAD.load(),
// Required with `memory`: LiveKit enables preemptive generation by default.
turnHandling: { preemptiveGeneration: { enabled: false } },
})
await session.start({
// These instructions never reach the Mastra agent; its own instructions apply.
agent: new voice.Agent({ instructions: 'Replies come from the Mastra agent.' }),
room: ctx.room,
})
session.say('Hi! How can I help you today?')
},
})
if (process.argv[1] === fileURLToPath(import.meta.url)) {
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
}
Both paths share the same reply pipeline underneath; choose by who should own the session:
createLiveKitWorker() | MastraLLM | |
|---|---|---|
| Session ownership | The worker helper builds and manages the AgentSession | Your code builds the session; every LiveKit option and hook is yours |
| Where the Mastra app runs | In the worker process | On your Mastra server, reached over HTTP (or in-process via agent) |
| Worker process needs | Your Mastra app, storage, and model provider keys | Only the LiveKit SDK and network access to your server |
| Built-in conveniences | Greeting, consent gating, agent-initiated hang-up, thread bootstrap, observability roll-up | Rebuild what you need with the session helpers |
| Best for | Fastest path to a working voice agent; Studio voice mode | Existing LiveKit apps and full control over the session |
Tools stay on the Mastra agent and execute on the server. LiveKit-side tools passed to the session are ignored. Tool activity reaches the worker through toolFeedback (spoken filler), onToolCall (fires as each tool call starts), and onTurnComplete (fires after each reply with the text, tool calls, and token usage). Agent-initiated hang-up takes a few lines: pair onToolCall with runEndCall().
Don't combine the memory option with LiveKit's preemptiveGeneration, which LiveKit enables by default in sessions you build yourself. A speculative turn that completes before LiveKit discards it persists a user message and a never-spoken reply to the thread. Set turnHandling: { preemptiveGeneration: { enabled: false } }, or run without memory and pass the full transcript each turn.
MastraLLM also accepts an in-process Mastra agent instance, session ownership without a second deployment, or a custom generate function. The remote transport is available standalone as createRemoteAgentReplyGenerator(), which also plugs into createLiveKitWorker's generate option to run the batteries-included worker against a remote server.
Server-initiated sessionsDirect link to Server-initiated sessions
Use dispatchVoiceSession() to add a voice agent to a room from your own code, for example to join an existing room or to drive an outbound SIP call:
import { dispatchVoiceSession } from '@mastra/livekit'
await dispatchVoiceSession({
roomName: 'support-call-42',
agentName: 'mastra-voice',
metadata: { agentId: 'support', threadId: 'thread-42', resourceId: 'user-7' },
})
ObservabilityDirect link to Observability
When the Mastra instance has observability configured, the worker traces each call. It opens one voice call span per session and nests everything under it:
- Every turn's Mastra agent run, with model generation, tool calls, and memory operations, exactly as a text chat records them.
- A child span for each LiveKit pipeline metric: speech-to-text, text-to-speech, end-of-utterance (turn detection), voice activity detection, and the model's time-to-first-token. These carry the latency and audio measurements that text traces can't show.
- A per-model usage roll-up (token, character, and audio totals for the whole call) written to the span when the session ends.
The worker is a separate process, so point storage at a backend that accepts concurrent writes from both the server and the worker. SQLite-backed LibSQL works. Single-writer stores don't. Traces, memory, and threads can share one store:
import { Mastra } from '@mastra/core/mastra'
import { LibSQLStore } from '@mastra/libsql'
import { Observability, MastraStorageExporter } from '@mastra/observability'
export const mastra = new Mastra({
storage: new LibSQLStore({ id: 'voice-agent-storage', url: 'file:./voice-agent.db' }),
observability: new Observability({
configs: {
default: {
serviceName: 'voice-agent',
exporters: [new MastraStorageExporter()],
},
},
}),
})
Tracing is on by default. Pass observability: false to createLiveKitWorker to turn it off.
DeploymentDirect link to Deployment
The worker is a separate process from your Mastra server. Deploy it as a long-running Node service with the production command:
node dist/voice-worker.js start
LiveKit's guidance on sizing, graceful shutdown, and hosting applies unchanged. See Deploying agents. Workers connect outbound to LiveKit, so they don't need inbound ports.
How it worksDirect link to How it works
A LiveKit voice session involves three pieces:
- Your Mastra server mints a LiveKit access token and dispatches your agent into a room. The dispatch carries metadata such as the Mastra agent id, memory thread, and resource.
- A LiveKit agent worker (a separate long-running process) receives the job and runs the audio pipeline. Audio flows between the browser and the worker over WebRTC and never passes through your Mastra HTTP server.
- Each time the user finishes a turn, the worker calls the Mastra agent's
stream()with the new input and speaks the streamed text. When the user interrupts, LiveKit cancels the stream and Mastra stops generating.
Conversation history lives in Mastra Memory, so voice sessions and text chat can share one thread.
RelatedDirect link to Related
API referenceDirect link to API reference
The @mastra/livekit package connects Mastra agents to the LiveKit Agents framework. LiveKit runs the audio pipeline (voice activity detection, speech-to-text, turn detection, text-to-speech, barge-in) and the package bridges reply generation to a Mastra agent's stream() call.
See Realtime voice for setup and concepts.
The package has three entry points:
@mastra/livekit: server-side APIs,liveKitConnectionRoute(),dispatchVoiceSession(),pipeAgentReplyToWriter(),serializeSessionMetadata(), andcreateEndCallTool(). Import these from Mastra server code. This entry never loads the LiveKit agents runtime.@mastra/livekit/worker: the worker runtime,createLiveKitWorker(),runLiveKitWorker(),chatContextToMessages(), and the session helpersspeakGreeting(),waitForAgentDoneSpeaking(), andrunEndCall(). Import it only from the worker entry file.@mastra/livekit/plugin: the LLM-component plugin,MastraLLMandcreateRemoteAgentReplyGenerator(). Import it in workers that build their ownvoice.AgentSession.createRemoteAgentReplyGenerator()is also exported from@mastra/livekit/workerbecause it plugs intocreateLiveKitWorker()'sgenerateoption.MastraLLMis plugin-only.
createLiveKitWorker()Direct link to createlivekitworker
Builds a LiveKit agent definition that answers voice sessions with Mastra agents. Use it as the default export of your worker entry file.
import { fileURLToPath } from 'node:url'
import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
})
if (process.argv[1] === fileURLToPath(import.meta.url)) {
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
}
OptionsDirect link to Options
mastra:
agent?:
workflow?:
workflowInput?:
replyStep?:
resultText?:
generate?:
stt?:
tts?:
vad?:
turnDetection?:
turnHandling?:
sessionOptions?:
memory?:
toolFeedback?:
onTurnComplete?:
configuration?:
greeting?:
consentPolicy?:
endCall?:
stt?:
tts?:
greeting?:
persistGreeting?:
observability?:
voice call span per session: every turn's agent run nests under it, LiveKit's STT, TTS, end-of-utterance, VAD, and LLM latency metrics become child spans, and the span closes with a per-model usage roll-up. Pass false to disable.inputOptions?:
outputOptions?:
onSessionStart?:
runLiveKitWorker()Direct link to runlivekitworker
Starts the LiveKit worker CLI (dev, start, and connect subcommands) for a worker entry file. Call it from the file that default-exports the worker definition, guarded so it only runs when executed directly (the worker spawns a child process per session that re-imports the same file). Using this helper instead of cli.runApp from @livekit/agents guarantees the worker runtime and the bridge share one copy of the LiveKit SDK.
OptionsDirect link to Options
entry:
agentName?:
serverOptions?:
pipeAgentReplyToWriter()Direct link to pipeagentreplytowriter
Streams a Mastra agent's reply into a workflow step's writer on the workflow reply path. It forwards the agent's text deltas, so text-to-speech starts before the full reply is ready, and its tool-call chunks, so toolFeedback fires and onTurnComplete sees the tool list. Piping only stream.textStream silently drops tool calls. Pass the step's abortSignal to agent.stream() so barge-in stops generation promptly.
import { pipeAgentReplyToWriter } from '@mastra/livekit'
const generateResponse = createStep({
id: 'generateResponse',
// input and output schemas omitted
execute: async ({ inputData, mastra, writer, abortSignal }) => {
const stream = await mastra.getAgent('support').stream(inputData.turn, { abortSignal })
const reply = await pipeAgentReplyToWriter(stream, writer)
return { reply }
},
})
Returns: Promise<string>, the accumulated reply text.
ParametersDirect link to Parameters
agentStream:
writer:
chatContextToMessages()Direct link to chatcontexttomessages
Converts a LiveKit chat context into plain messages accepted by agent.stream(), excluding instructions and function calls. Use it in workflowInput to pass the full transcript into a stateless workflow.
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker'
export default createLiveKitWorker({
mastra,
workflow: 'phoneConversation',
workflowInput: ({ chatCtx }) => ({ history: chatContextToMessages(chatCtx) }),
})
Returns: VoiceTurnMessage[], where each entry is { role: 'system' | 'user' | 'assistant'; content: string; id?: string }.
MastraLLMDirect link to mastrallm
A standard LiveKit LLM plugin (llm.LLM) backed by a Mastra agent. Use it when you build the voice.AgentSession yourself and want Mastra in the llm slot. createLiveKitWorker() is the managed alternative. See Use Mastra as the LLM component for how to choose.
With remote, the plugin streams each turn from your Mastra server over HTTP using Server-Sent Events (SSE). The agent loop, tools, and memory run server-side, and interrupting the agent aborts the server-side generation.
import { voice } from '@livekit/agents'
import { MastraLLM } from '@mastra/livekit/plugin'
const session = new voice.AgentSession({
llm: new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
memory: { thread: callId, resource: userId },
}),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
// Required with `memory`: LiveKit enables preemptive generation by default.
turnHandling: { preemptiveGeneration: { enabled: false } },
})
The plugin reports provider as mastra and model as the agent id, so LiveKit metrics and fallback adapters identify it like any other LLM.
Constructor optionsDirect link to Constructor options
Provide exactly one reply source: remote, agent, or generate.
remote?:
agent?:
generate?:
memory?:
requestContext?:
toolFeedback?:
onToolCall?:
onTurnComplete?:
Don't combine memory with the session's preemptiveGeneration option, which LiveKit enables by default in sessions you build yourself. A speculative turn that completes before LiveKit discards it persists a user message and a never-spoken reply to the thread. Set turnHandling: { preemptiveGeneration: { enabled: false } } on the session. Stateless mode (no memory) works with preemptive generation.
Tools run on the Mastra agentDirect link to Tools run on the Mastra agent
Tools are defined and executed server-side on the Mastra agent. The plugin never forwards LiveKit tool definitions: if the session passes a non-empty toolCtx, it logs a one-time warning naming the ignored tools. Every tool must complete server-side: a tool that requires approval or client-side execution fails the turn with a descriptive error instead of hanging the call.
Tool activity reaches the worker through toolFeedback, onToolCall, and onTurnComplete.
InstructionsDirect link to Instructions
LiveKit injects your voice.Agent's instructions into the chat context of every request. The plugin drops them because the server-side Mastra agent's own instructions are authoritative. To change the prompt, change the Mastra agent.
Interrupted turnsDirect link to Interrupted turns
When the user interrupts a reply:
- The plugin cancels the stream. The server aborts generation and persists nothing from that turn.
- LiveKit records the part the user actually heard in its chat context, flagged as interrupted.
- On the next turn, the plugin re-sends that heard-only fragment, ordered before the new user message, so the memory thread backfills to match the call. Messages carry LiveKit's message ids and the server deduplicates by id, so retries and re-sends stay idempotent.
A user who hangs up immediately after interrupting leaves that final fragment unrecorded. When the transcript must capture it, reconcile immediately from the session event; the shared message id means the next turn's re-send upserts instead of duplicating:
import { voice } from '@livekit/agents'
import { MastraClient } from '@mastra/client-js'
const client = new MastraClient({ baseUrl: process.env.MASTRA_URL! })
session.on(voice.AgentSessionEventTypes.ConversationItemAdded, ({ item }) => {
if (item.type !== 'message' || item.role !== 'assistant' || !item.interrupted) return
void client.saveMessageToMemory({
agentId: 'support',
messages: [
{
id: item.id,
threadId: callId,
resourceId: userId,
role: 'assistant',
content: item.textContent ?? '',
type: 'text',
createdAt: new Date(),
},
],
})
})
Usage metricsDirect link to Usage metrics
When the server reports token usage for a turn, the plugin feeds it to LiveKit, so the session's metrics_collected events carry time-to-first-token, duration, and token counts like any LLM plugin. The same usage object (promptTokens, completionTokens, promptCachedTokens, totalTokens) arrives on onTurnComplete as result.usage.
Errors and timeoutsDirect link to Errors and timeouts
The transport throws LiveKit's APIError types (APIStatusError, APIConnectionError, APITimeoutError), so the session's retry policy (connOptions.maxRetry) and FallbackAdapter failover work unchanged. A turn is never retried after its first token: a voice reply is better failed fast than replayed half-heard.
A connect and first-token watchdog uses the session's connOptions.timeoutMs (10 seconds by default), so a server that accepts the connection but never streams can't cause indefinite dead air.
If the Mastra server goes down mid-call, each reply attempt fails with a typed error after its retries, and LiveKit closes the session after several consecutive failed replies. Restore the server before that budget runs out and the call recovers on the next turn.
Message contentDirect link to Message content
Message extraction is text-only: image content is dropped, and audio content is included only through its transcript. Voice pipelines aren't affected, but items you inject into the chat context yourself must carry text.
createRemoteAgentReplyGenerator()Direct link to createremoteagentreplygenerator
Builds a reply generator that runs the agent loop on a remote Mastra server over HTTP/SSE. MastraLLM's remote mode uses it internally. Use it directly through createLiveKitWorker's generate option to run the batteries-included worker against a remote server:
import { createLiveKitWorker, createRemoteAgentReplyGenerator } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra, // local instance for logger and worker config; replies come from the remote server
generate: createRemoteAgentReplyGenerator({
baseUrl: process.env.MASTRA_URL!,
agentId: 'support',
}),
memory: ({ metadata, roomName }) => ({ thread: metadata.threadId ?? roomName }),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
})
On the generate path the worker-level toolFeedback and onTurnComplete options don't apply, and the worker's end-call detection doesn't fire; pass the hooks to the generator instead.
Cancelling a turn (barge-in) tears down the HTTP request, which aborts generation on the server. Errors are thrown as LiveKit APIError types. The retries option applies only to initial connection attempts. A turn is never retried after its first chunk.
Returns: VoiceReplyGenerator.
OptionsDirect link to Options
baseUrl:
agentId:
apiPrefix?:
headers?:
fetch?:
timeoutMs?:
retries?:
body?:
toolFeedback?:
onToolCall?:
onTurnComplete?:
speakGreeting()Direct link to speakgreeting
Speaks an opening greeting on a session you own, honoring interruption and playout options. Returns the LiveKit SpeechHandle, or undefined when there's no greeting text. createLiveKitWorker() uses it internally for its greeting configuration.
import { speakGreeting } from '@mastra/livekit/worker'
await speakGreeting(session, {
text: "You've reached support. You're speaking with an AI assistant.",
allowInterruptions: false,
awaitPlayout: true,
})
ParametersDirect link to Parameters
session:
greeting:
waitForAgentDoneSpeaking()Direct link to waitforagentdonespeaking
Resolves once the agent is no longer producing or playing a reply: its state has left thinking and speaking. Resolves immediately when the agent is already idle, and always resolves within maxWaitMs (30 seconds by default) as a safety cap. Use it before tearing a session down so closing words play out instead of being cut off.
import { waitForAgentDoneSpeaking } from '@mastra/livekit/worker'
await waitForAgentDoneSpeaking(session)
runEndCall()Direct link to runendcall
Ends the call after the agent asks to hang up. It waits for the agent's closing words and speaks an optional final message without interruption. It then deletes the room and hangs up the caller, including SIP callers. The job shuts down with its registered callbacks.
Pair it with MastraLLM's onToolCall and an end-call tool on the server-side agent to rebuild agent-initiated hang-up on a session you own:
import { MastraLLM } from '@mastra/livekit/plugin'
import { DEFAULT_END_CALL_TOOL, runEndCall } from '@mastra/livekit/worker'
let ending = false
const llm = new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
onToolCall: ({ toolName }) => {
if (toolName !== DEFAULT_END_CALL_TOOL || ending) return
ending = true
void runEndCall(session, ctx, {}, console)
},
})
The exported constants DEFAULT_END_CALL_TOOL ('endCall'), DEFAULT_END_CALL_REASON, and DEFAULT_END_CALL_MAX_WAIT_MS (30000) hold the defaults.
ParametersDirect link to Parameters
session:
ctx:
config:
logger:
createEndCallTool()Direct link to createendcalltool
Builds the Mastra tool an agent calls when it wants to end the call. The tool signals intent and can run optional bookkeeping. The worker performs the actual hang-up. The tool lives on the server-safe root entry. Add it to agents defined in server code.
import { Agent } from '@mastra/core/agent'
import { createEndCallTool } from '@mastra/livekit'
const supportAgent = new Agent({
id: 'support',
name: 'Support',
instructions:
'Help the caller. When everything is wrapped up, say goodbye and call endCall as your final action.',
model: 'openai/gpt-5-mini',
tools: { endCall: createEndCallTool() },
})
With createLiveKitWorker(), set configuration: { endCall: {} } and the worker watches for the tool and hangs up. On a session you own, rebuild the hang-up with runEndCall().
OptionsDirect link to Options
id?:
description?:
onEndCall?:
liveKitConnectionRoute()Direct link to livekitconnectionroute
Returns an API route that mints a LiveKit access token with the voice agent dispatched into the room. Frontends call it to join a session.
import { Mastra } from '@mastra/core/mastra'
import { liveKitConnectionRoute } from '@mastra/livekit'
export const mastra = new Mastra({
server: {
apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],
},
})
The route accepts a JSON body with optional agentId, threadId, and resourceId fields and responds with { serverUrl, roomName, participantName, participantToken }. The threadId defaults to the generated room name.
OptionsDirect link to Options
path?:
serverUrl?:
apiKey?:
apiSecret?:
agentName?:
ttl?:
requiresAuth?:
roomName?:
participantIdentity?:
metadata?:
dispatchVoiceSession()Direct link to dispatchvoicesession
Dispatches a Mastra voice agent into a LiveKit room programmatically: for server-initiated sessions such as outbound calls.
import { dispatchVoiceSession } from '@mastra/livekit'
await dispatchVoiceSession({
roomName: 'support-call-42',
agentName: 'mastra-voice',
metadata: { agentId: 'support', threadId: 'thread-42' },
})
OptionsDirect link to Options
roomName:
agentName?:
metadata?:
serverUrl?:
apiKey?:
apiSecret?:
LiveKitSessionMetadataDirect link to livekitsessionmetadata
The metadata passed from the Mastra server to the worker through LiveKit job dispatch.
agentId?:
threadId?:
resourceId?:
requestContext?:
The metadata travels as a JSON string. liveKitConnectionRoute() and dispatchVoiceSession() serialize it for you; use serializeSessionMetadata(metadata) when dispatching through your own code, or write the JSON directly in LiveKit-side configuration such as a SIP dispatch rule. Entries in requestContext reach the agent's runtime-defined instructions, tools, and input processors on every turn of the call.