LiveKit
QuickstartDirect link to Quickstart
Realtime voice turns a Mastra agent into a live call a user can talk over, in the browser or over the phone. Mastra builds it on LiveKit, an open source WebRTC platform for realtime audio and video.
The @mastra/livekit package connects Mastra agents to the LiveKit Agents framework: LiveKit owns the audio loop like voice activity detection, streaming speech-to-text, semantic turn detection, barge-in, and text-to-speech. Your Mastra agent generates every reply with its own model, tools, and memory.
Use realtime voice when you need low-latency, interruptible voice conversations. For provider-based speech-to-speech without LiveKit, see Speech to Speech.
These steps take you from an empty project to a voice agent you can talk to. A voice session has two moving parts you set up here: an API route on your Mastra server that hands out access tokens, and a separate worker process that runs the audio pipeline and calls your agent each turn.
Install the integration package along with the LiveKit plugins for voice activity detection and turn detection:
- npm
- pnpm
- Yarn
- Bun
npm install @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekitpnpm add @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekityarn add @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekitbun add @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekitSet your LiveKit credentials inside an
.envfile. Create a free project on LiveKit Cloud, or run a local server withlivekit-server --dev:.envLIVEKIT_URL=wss://your-project.livekit.cloudLIVEKIT_API_KEY=your-api-keyLIVEKIT_API_SECRET=your-api-secretAdd a voice agent to your Mastra instance and expose a connection route. The
liveKitConnectionRoute()helper adds aPOST /voice/livekit/connection-detailsendpoint that mints a LiveKit token and dispatches your agent into a room:src/mastra/index.tsimport { Mastra } from '@mastra/core/mastra'import { Agent } from '@mastra/core/agent'import { liveKitConnectionRoute } from '@mastra/livekit'const supportAgent = new Agent({id: 'support',name: 'Support',instructions: 'You are a friendly phone support agent. Keep replies short and conversational.',model: 'openai/gpt-5-mini',})export const mastra = new Mastra({agents: { support: supportAgent },server: {apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],},})Create the worker. It runs as a separate process, answers LiveKit sessions, and calls your agent each turn. Worker APIs live on the
@mastra/livekit/workerentry point, so the Mastra server never loads the LiveKit agents runtime. This example uses LiveKit Inference model strings for speech-to-text and text-to-speech, so no provider plugins are required:src/mastra/voice-worker.tsimport { fileURLToPath } from 'node:url'import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker'import { mastra } from './index'export default createLiveKitWorker({mastra,agent: 'support',stt: 'deepgram/nova-3',tts: 'cartesia/sonic-3',turnDetection: 'multilingual',greeting: 'Hi! How can I help you today?',})if (process.argv[1] === fileURLToPath(import.meta.url)) {runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })}The
agentoption selects which Mastra agent answers each session. Pass a fixed key as shown, or omit it to use theagentIdfrom the dispatch metadata, so one worker can serve every agent on your Mastra instance.Download the turn detection and voice activity detection models once. Then run the worker in one terminal and your Mastra server in another:
npx livekit-agents download-filesnpx tsx src/mastra/voice-worker.ts dev- npm
- pnpm
- Yarn
- Bun
npm run devpnpm run devyarn devbun run devThe worker registers with your LiveKit server and waits for sessions, while
mastra devserves the connection route.Talk to your agent. Open the hosted LiveKit Agents Playground and connect it to your project to start a call without building a frontend.
To wire up your own app instead, call the connection route for a token.
POST /voice/livekit/connection-detailsaccepts optionalagentId,threadId, andresourceIdfields in the request body and returns:{"serverUrl": "wss://your-project.livekit.cloud","roomName": "mastra-voice-a1b2c3d4","participantName": "user-1","participantToken": "eyJhbGci..."}This response matches the contract used by LiveKit's frontend starters, so apps built from agent-starter-react or the LiveKit React components work without changes.
Turn detection and interruptionsDirect link to Turn detection and interruptions
LiveKit decides when the user finished speaking and when the agent was interrupted. The defaults work well; tune them with turnHandling:
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
turnHandling: {
endpointing: { mode: 'dynamic', minDelay: 300, maxDelay: 3000 },
interruption: { minDuration: 500, resumeFalseInterruption: true },
},
})
turnDetection: 'multilingual': Runs LiveKit's semantic end-of-turn model locally on CPU. It reads the live transcript to avoid cutting users off mid-thought. Use'vad'or'stt'for silence-based endpointing instead.endpointing: Bounds how long the agent waits after the user stops speaking.interruption: Controls barge-in. When the user speaks over the agent, LiveKit stops playback and cancels the in-flight Mastra stream, so token generation stops too.preemptiveGeneration: Starts the Mastra agent's reply while the user is still finishing, hiding time-to-first-token. The worker disables it by default: each preemptive attempt runs the Mastra agent on an interim transcript, and a run that LiveKit later discards has already persisted a partial user message and a partial, never-spoken reply to the thread. Re-enable it withpreemptiveGeneration: { enabled: true }if latency matters more than exact thread history, or keep both by running turns read-only; see preemptive generation with memory.
See the LiveKit turn detection docs for all options.
To disable voice activity detection, set vad: false on createLiveKitWorker(). The worker passes null to LiveKit, which otherwise enables its bundled detector when vad is omitted. Choose a turn detection mode that doesn't require VAD, such as 'manual', when disabling it. Values in sessionOptions override the worker's generated options, including vad.
Per-call voices and transcriptionDirect link to Per-call voices and transcription
The top-level stt and tts options apply to every call. To pick them per call, one voice or language per tenant, set the configuration.stt and configuration.tts resolvers instead. Each resolver runs once per call with the dispatch metadata, request context, room name, and job context, and returns a value accepted by the matching top-level option. The value is either a plugin instance or an inference model string. Return undefined to fall back to the top-level option.
The following example gives each tenant its own text-to-speech voice, keyed off the tenant entry in the dispatch metadata:
import * as cartesia from '@livekit/agents-plugin-cartesia'
// One voice id per tenant, resolved from the dispatch metadata on each call.
const tenantVoices: Record<string, string> = {
meridian: 'your-cartesia-voice-id-1',
coastal: 'your-cartesia-voice-id-2',
}
// The resolver runs during call setup, so cache plugin instances across calls.
const ttsByVoice = new Map<string, cartesia.TTS>()
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
configuration: {
tts: ({ requestContext }) => {
const voice = tenantVoices[requestContext?.tenant as string]
if (!voice) return undefined // fall back to the top-level `tts`
let tts = ttsByVoice.get(voice)
if (!tts) {
tts = new cartesia.TTS({ voice })
ttsByVoice.set(voice, tts)
}
return tts
},
},
})
configuration.stt works the same way for per-call transcription, for example a different transcription model or language per tenant. The greeting has a matching per-call form: configuration.greeting.text accepts a resolver with the same call context, so one worker can open with each tenant's own phrasing.
Per-call turn detectionDirect link to Per-call turn detection
LiveKit's TurnDetector classes read the job's inference executor when constructed, so they can only be created inside a LiveKit job, not at module scope where the worker options live. To use one, set the configuration.turnDetection resolver. It runs once per call with the same call context as configuration.stt and returns anything the top-level turnDetection option accepts. Return undefined to fall back to the top-level option.
import { turnDetector } from '@livekit/agents-plugin-livekit'
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
configuration: {
// Constructed inside the job, where the inference executor is available.
turnDetection: () => new turnDetector.MultilingualModel(0.2),
},
})
The semantic model's inference runners must be registered before the agent server boots, so when this resolver is set the worker imports @livekit/agents-plugin-livekit up front. Keep the top-level turnDetection set to 'multilingual' or 'english' to pre-register only that model; otherwise both stay available.
Memory and threadsDirect link to Memory and threads
When the resolved Mastra agent has memory configured, each call becomes one memory thread:
threaddefaults to thethreadIdfrom dispatch metadata, then to the room name.resourcedefaults to theresourceIdfrom dispatch metadata, then to the thread. Send your end user's id here so calls group under the right user. Mastra Studio sends the agent id, matching how its sidebar lists threads.- When the thread doesn't exist yet, the worker creates it titled "Voice call" with metadata
{ source: 'livekit' }, and the spoken greeting is saved as the first assistant message so the thread reads as a full call transcript (disable withpersistGreeting: false).
Each turn sends only the new user input; Mastra Memory supplies history, semantic recall, and working memory. Pin a session to an existing thread by passing threadId in the connection request body, which is useful for continuing a text conversation by voice. In Studio, starting a call from an open chat binds the call to that thread, and the transcript fills into the chat after each exchange.
When a user interrupts the agent, the in-flight generation aborts and nothing from that turn is persisted at that moment. LiveKit keeps the part the user actually heard in its transcript, and on the next turn the worker re-sends that heard-only fragment so the thread backfills to match the call. A user who hangs up right after interrupting leaves that final fragment unrecorded. See interrupted turns for the details and a reconciliation recipe.
Preemptive generation with memoryDirect link to Preemptive generation with memory
LiveKit's preemptive generation calls the Mastra agent on interim transcripts and discards runs whose transcript changed. The plugin can't tell a speculative run from a real turn, so with memory set every run persists, including discarded ones. To keep preemptive generation on without corrupting the thread, pass options: { readOnly: true } in the memory mapping. The agent still reads history, semantic recall, and working memory from the thread but writes nothing, so speculative runs leave no trace. Persistence of committed turns then belongs to you: save them from LiveKit's ConversationItemAdded event, which fires only for items the session committed. Messages keep LiveKit's ids, so saves stay idempotent across retries.
import { voice } from '@livekit/agents'
import { createLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra,
agent: 'support',
memory: ({ metadata, roomName }) => ({
thread: metadata.threadId ?? roomName,
resource: metadata.resourceId ?? roomName,
options: { readOnly: true },
}),
turnHandling: { preemptiveGeneration: { enabled: true } },
onSessionStart: async ({ session, ctx, agent }) => {
const mapping = agent.memory
const memory = await mastra.getAgent('support').getMemory()
if (!mapping || !memory) return
let shuttingDown = false
const maxRetries = 5
const retryTimers = new Set<ReturnType<typeof setTimeout>>()
ctx.addShutdownCallback(async () => {
shuttingDown = true
for (const timer of retryTimers) clearTimeout(timer)
retryTimers.clear()
})
session.on(voice.AgentSessionEventTypes.ConversationItemAdded, ({ item }) => {
if (item.type !== 'message' || (item.role !== 'user' && item.role !== 'assistant')) return
const persist = async (attempt = 0): Promise<void> => {
try {
await memory.saveMessages({
messages: [
{
id: item.id,
threadId: mapping.thread,
resourceId: mapping.resource ?? mapping.thread,
role: item.role,
content: {
format: 2,
parts: [{ type: 'text', text: item.textContent ?? '' }],
},
type: 'text',
createdAt: new Date(),
},
],
})
} catch (error) {
if (shuttingDown) return
if (attempt >= maxRetries) {
console.error(`Failed to persist committed voice item ${item.id}; giving up`, error)
return
}
console.error(`Failed to persist committed voice item ${item.id}; retrying`, error)
const delay = Math.min(1_000 * 2 ** attempt, 30_000)
const timer = setTimeout(() => {
retryTimers.delete(timer)
void persist(attempt + 1)
}, delay)
retryTimers.add(timer)
}
}
void persist()
})
},
})
The same options field works on MastraVoiceAgent and MastraLLM; on the remote transport it's forwarded in the request body as memory.options.
Speak while tools runDirect link to Speak while tools run
Voice conversations can't go silent while a slow tool runs. Use toolFeedback to speak a short phrase when the Mastra agent starts a tool call:
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
toolFeedback: ({ toolName }) =>
toolName === 'searchOrders' ? 'Let me look that up.' : undefined,
})
The phrase is spoken as part of the reply and recorded in the transcript.
Generate replies with a workflowDirect link to Generate replies with a workflow
By default the worker generates each reply with a Mastra agent. To run multi-step logic per turn (for example classify intent, route, call tools in sequence, then compose a reply), generate replies with a Mastra workflow instead. Set workflow in place of agent.
LiveKit still owns the audio loop and calls into Mastra once per turn, so the workflow runs to completion each turn. The workflow can't suspend or resume, and no conversation state carries between turns. Pass the transcript in through workflowInput so the workflow stays stateless:
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra,
workflow: 'phoneConversation',
workflowInput: ({ chatCtx }) => ({ history: chatContextToMessages(chatCtx) }),
replyStep: 'generateResponse',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
})
A workflow streams structured step events, not text. To speak tokens as they generate, the reply step pipes its agent's text into the step writer:
const generateResponse = createStep({
id: 'generateResponse',
// input and output schemas omitted
execute: async ({ inputData, mastra, writer, abortSignal }) => {
const stream = await mastra.getAgent('voice').stream(inputData.history, { abortSignal })
await stream.textStream.pipeTo(writer)
return { assistantMessage: await stream.text }
},
})
replyStep: Restricts spoken output to one step. Omit it to speak every step that writes to itswriter.resultText: A fallback that derives the reply from the final run result when no step streams text. Streaming throughwritergives lower time-to-first-token, so prefer it.abortSignal: Forward the step'sabortSignalintoagent.stream()so barge-in stops generation promptly. When the user interrupts, the worker cancels the run.generate: For full control, pass ageneratefunction instead. It can be any reply generator that turns a turn into a text stream.
With a workflow, the worker doesn't persist turns automatically the way an agent's stream() does. Persist conversation history inside the workflow, or keep the LiveKit transcript as the source of truth and pass it in each turn.
Use Mastra as the LLM componentDirect link to Use Mastra as the LLM component
createLiveKitWorker() owns the LiveKit session for you. To own the session yourself, use MastraLLM instead: a standard LiveKit LLM plugin that puts a Mastra agent in the llm slot of your own voice.AgentSession. The Mastra app, agent loop, tools, memory, observability, runs on your Mastra server, and the worker reaches it over HTTP. The worker process needs no Mastra app, database, or model provider keys.
import { fileURLToPath } from 'node:url'
import { defineAgent, voice } from '@livekit/agents'
import * as silero from '@livekit/agents-plugin-silero'
import { MastraLLM } from '@mastra/livekit/plugin'
import { runLiveKitWorker } from '@mastra/livekit/worker'
export default defineAgent({
entry: async ctx => {
await ctx.connect()
const session = new voice.AgentSession({
llm: new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
memory: { thread: ctx.room.name!, resource: 'user-7' },
}),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
vad: await silero.VAD.load(),
// Required with `memory` unless memory.options.readOnly is set: LiveKit enables
// preemptive generation by default.
turnHandling: { preemptiveGeneration: { enabled: false } },
})
await session.start({
// These instructions never reach the Mastra agent; its own instructions apply.
agent: new voice.Agent({ instructions: 'Replies come from the Mastra agent.' }),
room: ctx.room,
})
session.say('Hi! How can I help you today?')
},
})
if (process.argv[1] === fileURLToPath(import.meta.url)) {
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
}
Both paths share the same reply pipeline underneath; choose by who should own the session:
createLiveKitWorker() | MastraLLM | |
|---|---|---|
| Session ownership | The worker helper builds and manages the AgentSession | Your code builds the session; every LiveKit option and hook is yours |
| Where the Mastra app runs | In the worker process | On your Mastra server, reached over HTTP (or in-process via agent) |
| Worker process needs | Your Mastra app, storage, and model provider keys | Only the LiveKit SDK and network access to your server |
| Built-in conveniences | Greeting, consent gating, agent-initiated hang-up, thread bootstrap, observability roll-up | Rebuild what you need with the session helpers |
| Best for | Fastest path to a working voice agent; Studio voice mode | Existing LiveKit apps and full control over the session |
Tools stay on the Mastra agent and execute on the server. LiveKit-side tools passed to the session are ignored. Tool activity reaches the worker through toolFeedback (spoken filler), onToolCall (fires as each tool call starts), and onTurnComplete (fires after each reply with the text, tool calls, and token usage). Agent-initiated hang-up takes a few lines: pair onToolCall with runEndCall().
Don't combine the memory option with LiveKit's preemptiveGeneration, which LiveKit enables by default in sessions you build yourself. A speculative turn persists a partial user message and a partial, never-spoken reply to the thread before LiveKit discards it. Set turnHandling: { preemptiveGeneration: { enabled: false } }, run without memory and pass the full transcript each turn, or set memory.options.readOnly and persist committed turns yourself; see preemptive generation with memory.
MastraLLM also accepts an in-process Mastra agent instance, session ownership without a second deployment, or a custom generate function. The remote transport is available standalone as createRemoteAgentReplyGenerator(), which also plugs into createLiveKitWorker's generate option to run the batteries-included worker against a remote server.
Server-initiated sessionsDirect link to Server-initiated sessions
Use dispatchVoiceSession() to add a voice agent to a room from your own code, for example to join an existing room or to drive an outbound SIP call:
import { dispatchVoiceSession } from '@mastra/livekit'
await dispatchVoiceSession({
roomName: 'support-call-42',
agentName: 'mastra-voice',
metadata: { agentId: 'support', threadId: 'thread-42', resourceId: 'user-7' },
})
Record callsDirect link to Record calls
Pass recording to dispatchVoiceSession() or liveKitConnectionRoute() to create a room with LiveKit auto egress. Mastra creates the room with recording configured, then dispatches the voice agent. Omitting recording preserves the existing behavior and return values.
Recording requires LiveKit Cloud or a self-hosted egress service, plus an output destination. The example below records the caller and agent into a single mixed OGG audio file in an S3 bucket. Set RECORDINGS_S3_BUCKET, RECORDINGS_S3_REGION, RECORDINGS_S3_ACCESS_KEY, and RECORDINGS_S3_SECRET on your server, alongside the LiveKit credentials.
Install the server SDK to use its output format enum:
- npm
- pnpm
- Yarn
- Bun
npm install livekit-server-sdk
pnpm add livekit-server-sdk
yarn add livekit-server-sdk
bun add livekit-server-sdk
Define recording settings in a server-only module:
import type { LiveKitRecordingOptions } from '@mastra/livekit'
import { EncodedFileType } from 'livekit-server-sdk'
export const recording: LiveKitRecordingOptions = {
room: {
audioOnly: true,
fileOutputs: [
{
fileType: EncodedFileType.OGG,
filepath: 'calls/{room_name}.ogg',
output: {
case: 's3',
value: {
bucket: process.env.RECORDINGS_S3_BUCKET!,
region: process.env.RECORDINGS_S3_REGION!,
accessKey: process.env.RECORDINGS_S3_ACCESS_KEY!,
secret: process.env.RECORDINGS_S3_SECRET!,
},
},
},
],
},
}
LiveKitRecordingOptions accepts the settings used to construct LiveKit's RoomEgress, as either a plain object or an SDK instance. Set room.audioOnly: true for mixed audio, or use tracks or participant for separate recordings. Mastra passes these settings to LiveKit without choosing a format or storage provider for you. See LiveKit output options for other destinations.
For an outbound phone call, configure recording before adding the phone participant:
import { randomUUID } from 'node:crypto'
import { dispatchVoiceSession } from '@mastra/livekit'
import { recording } from './mastra/recording'
const roomName = `support-call-${randomUUID()}`
await dispatchVoiceSession({
roomName,
metadata: { agentId: 'support', resourceId: 'caller-7' },
recording,
})
// Add the phone participant to roomName through your existing LiveKit SIP code.
Use a fresh room name for each recorded session and await dispatch before dialing. dispatchVoiceSession() throws LiveKitRecordingRoomConflictError if the room already exists. Import this error from @mastra/livekit to handle the conflict. The existence check and room creation are separate requests, so your application must prevent another process from creating the same room concurrently. This option doesn't start recording in an existing room or capture earlier audio.
For browser sessions, pass the same settings to the connection route:
import { Mastra } from '@mastra/core/mastra'
import { liveKitConnectionRoute } from '@mastra/livekit'
import { recording } from './recording'
export const mastra = new Mastra({
server: {
apiRoutes: [liveKitConnectionRoute({ recording })],
},
})
With recording enabled, the route creates the room and dispatches the agent when connection details are requested, before the browser joins. The response remains { serverUrl, roomName, participantName, participantToken }. Recording settings and storage credentials stay on the server and aren't included in the join token or agent metadata.
The route also accepts a synchronous or asynchronous recording callback. It receives { body, context, roomName }, including the resolved room name, so your server can choose settings for each session. Return undefined to skip recording for that request. Recording configuration isn't read directly from the request body.
An existing recording room returns HTTP 409 before agent dispatch or token issuance. Other room creation and dispatch errors reject the operation; the route doesn't issue a token on failure. If dispatch fails after room creation, the room remains in LiveKit. Successful setup doesn't guarantee that the recording finishes or uploads successfully. Use LiveKit's egress events to check completion and retrieve file details before making a recording available to users. To verify the setup, make a test call, end the room, and listen to the completed file under calls/ in your bucket for audio from both participants.
Recording lifecycle and consentDirect link to Recording lifecycle and consent
Enabling recording starts recording during room creation, before a browser connects or a phone participant joins. An abandoned connection request can still start the agent and consume recording and storage resources. Configure authentication and application limits on the connection route before exposing it to users.
createConsentTool() and configuration.consentPolicy don't control LiveKit egress. They don't delay, pause, or stop an automatic recording. If your application needs consent during the call before capturing audio, obtain it first and use LiveKit's explicit egress APIs to start recording in the existing room. A spoken greeting alone doesn't gate this feature.
Recordings follow the LiveKit room lifecycle. Disconnecting only the agent doesn't necessarily end the room or finalize the file. Wait for a successful egress completion status and file results before marking a recording ready. Treat failed setup and abandoned rooms as application cleanup responsibilities.
The audio file lives in the configured output storage, separately from Mastra memory, transcripts, and traces. Deleting a Mastra thread doesn't delete the recording. Your application manages storage access, retention, and download links for end users. Keep storage credentials in server configuration, including when choosing recording settings per session.
Review recordings in StudioDirect link to Review recordings in Studio
Register liveKitRecordingRoute() to add Review Audio to a voice call's trace panel in Studio. The button opens an audio player for that call. Recording and trace storage must both be configured: the route reads the stored voice call span and passes its roomName to your server-side resolver. The browser supplies only the trace ID.
The resolver returns { url, expiresAt? }, or undefined when no file is available. Use a short-lived signed URL for private storage. The route doesn't persist playback URLs in traces, and Studio requests a fresh URL each time the review opens. Missing files and playback errors offer Refresh recording so you can retry after the upload finishes or a link expires.
For the S3 configuration above, use the same calls/{room_name}.ogg object key for recording and playback. Room names must remain unique across calls, including after a room has ended. Install the AWS SDK in your application:
- npm
- pnpm
- Yarn
- Bun
npm install @aws-sdk/client-s3 @aws-sdk/s3-request-presigner
pnpm add @aws-sdk/client-s3 @aws-sdk/s3-request-presigner
yarn add @aws-sdk/client-s3 @aws-sdk/s3-request-presigner
bun add @aws-sdk/client-s3 @aws-sdk/s3-request-presigner
import {
GetObjectCommand,
HeadObjectCommand,
S3Client,
S3ServiceException,
} from '@aws-sdk/client-s3'
import { getSignedUrl } from '@aws-sdk/s3-request-presigner'
import { liveKitRecordingRoute } from '@mastra/livekit'
import { authorizeRecording } from './recording-access'
const s3 = new S3Client({
region: process.env.RECORDINGS_S3_REGION!,
credentials: {
accessKeyId: process.env.RECORDINGS_S3_ACCESS_KEY!,
secretAccessKey: process.env.RECORDINGS_S3_SECRET!,
},
})
export const recordingReviewRoute = liveKitRecordingRoute({
authorize: authorizeRecording,
resolveRecording: async ({ roomName }) => {
const object = {
Bucket: process.env.RECORDINGS_S3_BUCKET!,
Key: `calls/${roomName}.ogg`,
}
try {
await s3.send(new HeadObjectCommand(object))
} catch (error) {
if (error instanceof S3ServiceException && error.$metadata.httpStatusCode === 404)
return undefined
throw error
}
const expiresIn = 15 * 60
return {
url: await getSignedUrl(
s3,
new GetObjectCommand({ ...object, ResponseContentType: 'audio/ogg' }),
{ expiresIn },
),
expiresAt: new Date(Date.now() + expiresIn * 1000).toISOString(),
}
},
})
Add recordingReviewRoute to server.apiRoutes alongside liveKitConnectionRoute({ recording }). Keep the default path, GET /voice/livekit/recordings/:traceId, for Studio discovery. Both the connection route and server-initiated dispatch recordings can be reviewed, as long as the worker persisted the call trace and your resolver can locate its file.
The review route requires authentication by default and an explicit authorize({ traceId, context }) callback. Implement the authorizeRecording helper in your application to check the authenticated user's access to the requested trace against your trusted ownership or tenant records. Return true only when access is allowed. Authentication alone doesn't establish recording ownership.
The route rejects missing authorization callbacks at registration, including when requiresAuth is false. A denied request returns 403 before trace storage is read or a playback URL is generated. Policy errors also stop resolution. The resolver receives the stored span and request context only after authorization succeeds.
Use requiresAuth: false only for a local demo. It doesn't bypass the required recording authorization callback. If a demo policy allows all recordings, keep that policy explicitly limited to local development and replace it before deployment.
Recording upload permissions alone don't grant playback access. The signing credentials need s3:GetObject for the recording objects, including the HEAD check. S3 also requires s3:ListBucket to return a missing-object 404 instead of 403. Your application can use separate read credentials for playback. The browser receives a temporary URL that authorizes access to that file, so treat it as sensitive. See S3 object permissions and presigned URLs.
liveKitRecordingRoute() is storage-independent. You can resolve URLs from another provider or a recording index populated by LiveKit completion events. It doesn't depend on LiveKit retaining historical egress jobs. If you use timestamped filenames or change buckets, retain the mapping from room to object in your application and resolve that mapping instead of constructing a key.
ObservabilityDirect link to Observability
When the Mastra instance has observability configured, the worker traces each call. It opens one voice call span per session and nests everything under it:
- Every turn's Mastra agent run, with model generation, tool calls, and memory operations, exactly as a text chat records them.
- A child span for each LiveKit pipeline metric: speech-to-text, text-to-speech, end-of-utterance (turn detection), voice activity detection, and the model's time-to-first-token. These carry the latency and audio measurements that text traces can't show.
- A per-model usage roll-up (token, character, and audio totals for the whole call) written to the span when the session ends.
The worker is a separate process, so point storage at a backend that accepts concurrent writes from both the server and the worker. SQLite-backed LibSQL works. Single-writer stores don't. Traces, memory, and threads can share one store:
import { Mastra } from '@mastra/core/mastra'
import { LibSQLStore } from '@mastra/libsql'
import { Observability, MastraStorageExporter } from '@mastra/observability'
export const mastra = new Mastra({
storage: new LibSQLStore({ id: 'voice-agent-storage', url: 'file:./voice-agent.db' }),
observability: new Observability({
configs: {
default: {
serviceName: 'voice-agent',
exporters: [new MastraStorageExporter()],
},
},
}),
})
Tracing is on by default. Pass observability: false to createLiveKitWorker to turn it off.
Request recording playback from your applicationDirect link to Request recording playback from your application
Import the browser client helper from @mastra/livekit/client. Install @mastra/client-js 1.51.2 or newer within 1.x to use this entry point:
- npm
- pnpm
- Yarn
- Bun
npm install @mastra/livekit @mastra/client-js
pnpm add @mastra/livekit @mastra/client-js
yarn add @mastra/livekit @mastra/client-js
bun add @mastra/livekit @mastra/client-js
import { MastraClient } from '@mastra/client-js'
import { getLiveKitRecording } from '@mastra/livekit/client'
import type { LiveKitRecordingResponse } from '@mastra/livekit/client'
const client = new MastraClient({
baseUrl: 'https://mastra.example.com',
credentials: 'include',
})
const recording: LiveKitRecordingResponse = await getLiveKitRecording(client, traceId)
if (recording.status === 'ready') {
audioElement.src = recording.url
}
Register liveKitRecordingRoute() on the server first. The helper requests its default /voice/livekit/recordings/:traceId path, independently of the client's API prefix. It preserves the client's authentication headers, cookies, custom fetch function, and abort signal. Pass { signal } as a third argument to cancel an individual request. Failed requests use MastraClientError and aren't retried automatically.
@mastra/client-js@1.51.2 declares support for Zod 3 and 4, but its @ai-sdk/ui-utils@1.2.11 dependency declares a Zod 3 peer dependency. This mismatch can cause strict npm installs with Zod 4 to fail. Use Zod ^3.25.76 with that SDK to satisfy both peer ranges. Server and worker consumers without the optional client SDK also support Zod 4.
The /client entry point excludes LiveKit server and worker imports. Other integration entry points don't require @mastra/client-js. Installing @mastra/livekit still installs its package-level dependencies; the browser entry point only controls which code enters the browser bundle. Storage credentials and recording authorization remain on the server.
Turn metrics and playback completionDirect link to Turn metrics and playback completion
createLiveKitWorker() accepts two observers for evaluating voice calls:
import { createLiveKitWorker } from '@mastra/livekit/worker'
export default createLiveKitWorker({
mastra,
agent: 'support',
onTurnMetrics: metrics => {
console.info('voice metric', metrics)
},
onSpeechComplete: speech => {
console.info('speech outcome', speech.speechId, speech.outcome)
},
})
Observers run asynchronously and aren't awaited by the audio pipeline. Errors are logged. Keep their work small. Neither observer writes or reconciles memory.
onTurnMetrics receives versioned records with phase: 'generation' or phase: 'speech'. Generation records contain a turnId, a unique attemptId, a terminal outcome, durationMs, optional firstTextMs, and tool timings. Attempts for the same user message share its turn ID. Tool durations run from the bridge receiving a tool call to receiving its result; remote durations include transport and aren't pure server execution times. A tool without a result has no duration.
Speech records also contain LiveKit's speechId. The bridge carries turn and attempt IDs in assistant message metadata so they can be joined after playback. Speech without a correlated assistant message, such as cancelled speech that produced no output, uses its speech ID as the turn ID and omits the attempt ID. Don't infer a correlation by matching event arrival order.
| Measurement | Meaning |
|---|---|
Generation firstTextMs | Time from starting generation to the first text forwarded toward speech synthesis. Can include filler. |
Speech firstAudioMs | Time from caller speech ending to the first output audio, as reported by LiveKit. Can include filler. |
Speech completionMs | Time from caller speech ending to completed output playback. Omitted for interrupted, cancelled, or failed speech. |
speechStartedAt / speechEndedAt | Playback timestamps in Unix epoch milliseconds when available. |
Missing measurements are undefined, not zero. Local generation durations use a monotonic clock. LiveKit's speech timing values are converted to milliseconds. When observability is configured, these records appear as events in the existing voice call trace. Pipeline metric events retain LiveKit's speech ID when supplied.
onSpeechComplete runs once when a speech handle finishes, or when the session closes with that speech pending. It supplies the speech metrics plus optional playedText and transcriptSource. Outcomes are completed, interrupted, cancelled, or failed. A completed speech lifecycle alone doesn't prove that audio was produced; check whether the relevant audio measurements are present.
The measurement source is server-playout. LiveKit's committed transcript can be partial or estimated, depending on the output's synchronization support. It doesn't prove that the remote browser or phone listener heard every word. transcriptSource: 'unavailable' means no committed assistant transcript was available.
Keep using onTurnComplete for its existing generation contract. It can fire before speech playback finishes, with interrupted: false, even when the caller interrupts the audio later. Use onSpeechComplete for playback outcomes. Existing agent memory may already contain generated text; adding this hook doesn't change that persistence behavior.
For a session you own, call observeVoiceSession(session, { onTurnMetrics, onSpeechComplete }) from @mastra/livekit/worker or @mastra/livekit/plugin before session.start(). Its return value detaches the observer and finalizes pending speech as cancelled. The worker attaches its observer automatically.
Speech segment boundariesDirect link to Speech segment boundaries
The built-in agent, remote, and workflow reply generators preserve text-segment and tool-step boundaries. Tool feedback is flushed after its text is queued, so an acknowledgment can reach text-to-speech while a tool is still running. A flush requests a segment boundary; it doesn't guarantee immediate audio from every speech provider.
Custom generators can continue returning string streams. To emit an explicit boundary, return a ReadableStream<VoiceReplyChunk> and enqueue VOICE_TEXT_FLUSH, both exported from @mastra/livekit/worker and @mastra/livekit/plugin. Consumers that read a reply generator directly must distinguish strings from boundary objects. Never concatenate a boundary object into a transcript.
The worker translates boundaries into LiveKit's FlushSentinel. With the MastraLLM plugin, add the node adapter to the LiveKit agent you own:
import { voice } from '@livekit/agents'
import { MastraLLM, mastraLLMNode } from '@mastra/livekit/plugin'
class SupportAgent extends voice.Agent {
override llmNode(...args: Parameters<voice.Agent['llmNode']>) {
return mastraLLMNode(this, ...args)
}
}
const session = new voice.AgentSession({
llm: new MastraLLM({ agent: supportAgent }),
// Configure STT and TTS for your application.
})
await session.start({ agent: new SupportAgent({ instructions: 'Support assistant' }), room })
Without the adapter, the plugin still emits text, but LiveKit's default node doesn't translate the plugin's boundary metadata into flush markers. Use pipeAgentReplyToWriter() in a workflow reply step to preserve the boundaries and tool results used for timing.
DeploymentDirect link to Deployment
The worker is a separate process from your Mastra server. Deploy it as a long-running Node service with the production command:
node dist/voice-worker.js start
LiveKit's guidance on sizing, graceful shutdown, and hosting applies unchanged. See Deploying agents. Workers connect outbound to LiveKit, so they don't need inbound ports.
How it worksDirect link to How it works
A LiveKit voice session involves three pieces:
- Your Mastra server mints a LiveKit access token and dispatches your agent into a room. The dispatch carries metadata such as the Mastra agent id, memory thread, and resource.
- A LiveKit agent worker (a separate long-running process) receives the job and runs the audio pipeline. Audio flows between the browser and the worker over WebRTC and never passes through your Mastra HTTP server.
- Each time the user finishes a turn, the worker calls the Mastra agent's
stream()with the new input and speaks the streamed text. When the user interrupts, LiveKit cancels the stream and Mastra stops generating.
Conversation history lives in Mastra Memory, so voice sessions and text chat can share one thread.
RelatedDirect link to Related
Upgrade to 0.5.0Direct link to Upgrade to 0.5.0
The 0.5.0 prerelease requires LiveKit Agents 1.7.1 or newer within 1.x. The speech helpers use FlushSentinel and SpeechHandle.exception(), which are absent in Agents 1.4.0. Test the supplied prerelease version or package archive before deploying it. Updating a dependency range doesn't update an existing lockfile: check the installed versions after installation.
If you used client.getLiveKitRecording(traceId) in an earlier snapshot, switch to getLiveKitRecording(client, traceId) from @mastra/livekit/client. Import LiveKitRecordingResponse from the same entry point instead of GetLiveKitRecordingResponse from client-js.
Authorize recording accessDirect link to Authorize recording access
Before upgrading an application that registers liveKitRecordingRoute(), add an authorize({ traceId, context }) callback that checks the authenticated user's access to the requested trace. The callback is required even when requiresAuth: false; omitting it throws during route registration. Return true only to allow access, before trace lookup or playback URL resolution. See Review recordings in Studio for the setup example.
Install compatible packagesDirect link to Install compatible packages
Use the same exact release for Agents and its plugins. The tested baseline is:
| Package | Version |
|---|---|
@livekit/agents | 1.7.1 |
@livekit/agents-plugin-livekit | 1.7.1 |
@livekit/agents-plugin-silero | 1.7.1 |
@livekit/rtc-node | 0.13.34 |
livekit-server-sdk | 2.16.0 |
For a local prerelease archive, install it with the matching LiveKit packages:
- npm
- pnpm
- Yarn
- Bun
npm install ./mastra-livekit-0.5.0-next.0.tgz @livekit/agents@1.7.1 @livekit/agents-plugin-livekit@1.7.1 @livekit/agents-plugin-silero@1.7.1 @livekit/rtc-node@0.13.34 livekit-server-sdk@2.16.0
pnpm add ./mastra-livekit-0.5.0-next.0.tgz @livekit/agents@1.7.1 @livekit/agents-plugin-livekit@1.7.1 @livekit/agents-plugin-silero@1.7.1 @livekit/rtc-node@0.13.34 livekit-server-sdk@2.16.0
yarn add ./mastra-livekit-0.5.0-next.0.tgz @livekit/agents@1.7.1 @livekit/agents-plugin-livekit@1.7.1 @livekit/agents-plugin-silero@1.7.1 @livekit/rtc-node@0.13.34 livekit-server-sdk@2.16.0
bun add ./mastra-livekit-0.5.0-next.0.tgz @livekit/agents@1.7.1 @livekit/agents-plugin-livekit@1.7.1 @livekit/agents-plugin-silero@1.7.1 @livekit/rtc-node@0.13.34 livekit-server-sdk@2.16.0
Replace the archive path with the supplied prerelease. Install any other @livekit/agents-plugin-* packages at the matching release. Agents 1.7.1 and Silero 1.7.1 require RTC ^0.13.34; RTC 1.x doesn't satisfy that requirement. Don't override peer dependency errors to install it.
Keep Node.js 22.13 or newer. Both @mastra/livekit and Agents require Zod ^3.25.76 || ^4.1.8. Upgrade older Zod versions before installing 0.5.0. Verify the installed dependency tree and initialize the worker's models:
- npm
- pnpm
- Yarn
- Bun
npm ls @mastra/livekit @livekit/agents @livekit/agents-plugin-livekit @livekit/agents-plugin-silero @livekit/rtc-node
npx tsx src/mastra/voice-worker.ts download-files
pnpm ls @mastra/livekit @livekit/agents @livekit/agents-plugin-livekit @livekit/agents-plugin-silero @livekit/rtc-node
npx tsx src/mastra/voice-worker.ts download-files
yarn list --pattern "@mastra/livekit|@livekit/agents|@livekit/agents-plugin-livekit|@livekit/agents-plugin-silero|@livekit/rtc-node"
npx tsx src/mastra/voice-worker.ts download-files
bun pm ls @mastra/livekit @livekit/agents @livekit/agents-plugin-livekit @livekit/agents-plugin-silero @livekit/rtc-node
npx tsx src/mastra/voice-worker.ts download-files
Use a single installed copy of Agents for both the integration and a custom voice.AgentSession. Multiple copies can fail LiveKit's runtime class checks.
Check VAD and turn detectionDirect link to Check VAD and turn detection
vad: false now disables VAD at session creation. Previously, the worker skipped plugin prewarming but passed undefined to LiveKit, which could enable its bundled VAD. If an application relied on that behavior, configure a detector explicitly.
import { createLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra,
agent: 'support',
vad: false,
turnDetection: 'manual',
})
The default still loads the Silero plugin during prewarm. The 'english' and 'multilingual' turn detection options still select their existing text models. LiveKit's newer inference.VAD and inference.TurnDetector APIs are available through the existing instance and resolver options; switching to them is a separate choice that can change endpointing, model downloads, and cloud usage.
Review custom LiveKit toolsDirect link to Review custom LiveKit tools
Mastra tools continue to run through the Mastra agent. If you also implement LiveKit tools or a custom LiveKit plugin, account for the changes since Agents 1.4:
ToolContextis a class. Use its accessors andflatten()instead of enumerating or mutating it as a plain object.- Named tool lists and duplicate-name checks are supported. Object shorthand with anonymous tools remains supported; converting every tool map isn't required.
- Provider tools use subclasses of
ProviderToolinstead of the olderProviderDefinedToolAPI.
See the LiveKit tool API for custom plugin migration details.
Review shutdownDirect link to Review shutdown
The onCallEnd hook still uses LiveKit's job shutdown callbacks. Keep cleanup within the configured shutdown window, and verify that it completes when a caller disconnects or a worker receives a shutdown signal.
Verify recordings and callsDirect link to Verify recordings and calls
The integration's room-level recording configuration and S3 permissions stay the same. LiveKit Agents' AgentSession.start({ record }) controls separate session recording and telemetry. Disabling one doesn't disable the other.
Before promoting the prerelease, test both a browser call and a phone call through a configured SIP trunk. Verify both speakers, interruptions, tools, greeting and consent behavior, hangup, the completed trace, and onCallEnd. When recording is enabled, wait for the S3 object and play it with Review Audio. Also test recording disabled and rejection of an existing room name. Run these checks on the deployment image with its model files and native libraries installed.
API referenceDirect link to API reference
The @mastra/livekit package connects Mastra agents to the LiveKit Agents framework. LiveKit runs the audio pipeline (voice activity detection, speech-to-text, turn detection, text-to-speech, barge-in) and the package bridges reply generation to a Mastra agent's stream() call.
See Realtime voice for setup and concepts.
The package has three entry points:
@mastra/livekit: server-side APIs,liveKitConnectionRoute(),dispatchVoiceSession(),pipeAgentReplyToWriter(),serializeSessionMetadata(), andcreateEndCallTool(). Import these from Mastra server code. This entry never loads the LiveKit agents runtime.@mastra/livekit/worker: the worker runtime,createLiveKitWorker(),runLiveKitWorker(),chatContextToMessages(), the per-session agent classMastraVoiceAgent, and the session helpersspeakGreeting(),waitForAgentDoneSpeaking(), andrunEndCall(). Import it only from the worker entry file.@mastra/livekit/plugin: the LLM-component plugin,MastraLLMandcreateRemoteAgentReplyGenerator(). Import it in workers that build their ownvoice.AgentSession.createRemoteAgentReplyGenerator()is also exported from@mastra/livekit/workerbecause it plugs intocreateLiveKitWorker()'sgenerateoption.MastraLLMis plugin-only.
createLiveKitWorker()Direct link to createlivekitworker
Builds a LiveKit agent definition that answers voice sessions with Mastra agents. Use it as the default export of your worker entry file.
import { fileURLToPath } from 'node:url'
import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
})
if (process.argv[1] === fileURLToPath(import.meta.url)) {
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
}
OptionsDirect link to Options
mastra:
agent?:
workflow?:
workflowInput?:
replyStep?:
resultText?:
generate?:
stt?:
tts?:
vad?:
turnDetection?:
turnHandling?:
sessionOptions?:
memory?:
options is forwarded to the agent as per-call memory config; { readOnly: true } keeps speculative turns off the thread.toolFeedback?:
onTurnComplete?:
configuration?:
greeting?:
consentPolicy?:
endCall?:
stt?:
tts?:
turnDetection?:
greeting?:
persistGreeting?:
observability?:
voice call span per session: every turn's agent run nests under it, LiveKit's STT, TTS, end-of-utterance, VAD, and LLM latency metrics become child spans, and the span closes with a per-model usage roll-up. Pass false to disable.inputOptions?:
outputOptions?:
onSessionStart?:
runLiveKitWorker()Direct link to runlivekitworker
Starts the LiveKit worker CLI (dev, start, and connect subcommands) for a worker entry file. Call it from the file that default-exports the worker definition, guarded so it only runs when executed directly (the worker spawns a child process per session that re-imports the same file). Using this helper instead of cli.runApp from @livekit/agents guarantees the worker runtime and the bridge share one copy of the LiveKit SDK.
OptionsDirect link to Options
entry:
agentName?:
serverOptions?:
pipeAgentReplyToWriter()Direct link to pipeagentreplytowriter
Streams a Mastra agent's reply into a workflow step's writer on the workflow reply path. It forwards the agent's text deltas, so text-to-speech starts before the full reply is ready, and its tool-call chunks, so toolFeedback fires and onTurnComplete sees the tool list. Piping only stream.textStream silently drops tool calls. Pass the step's abortSignal to agent.stream() so barge-in stops generation promptly.
import { pipeAgentReplyToWriter } from '@mastra/livekit'
const generateResponse = createStep({
id: 'generateResponse',
// input and output schemas omitted
execute: async ({ inputData, mastra, writer, abortSignal }) => {
const stream = await mastra.getAgent('support').stream(inputData.turn, { abortSignal })
const reply = await pipeAgentReplyToWriter(stream, writer)
return { reply }
},
})
Returns: Promise<string>, the accumulated reply text.
ParametersDirect link to Parameters
agentStream:
writer:
chatContextToMessages()Direct link to chatcontexttomessages
Converts a LiveKit chat context into plain messages accepted by agent.stream(), excluding instructions and function calls. Use it in workflowInput to pass the full transcript into a stateless workflow.
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker'
export default createLiveKitWorker({
mastra,
workflow: 'phoneConversation',
workflowInput: ({ chatCtx }) => ({ history: chatContextToMessages(chatCtx) }),
})
Returns: VoiceTurnMessage[], where each entry is { role: 'system' | 'user' | 'assistant'; content: string; id?: string }.
MastraVoiceAgentDirect link to mastravoiceagent
The LiveKit voice.Agent subclass that createLiveKitWorker() builds for every session. Replies come from a Mastra agent (or a custom generate source) through the agent's llmNode; LiveKit keeps the audio loop, turn detection, and barge-in. Construct it yourself when you own the voice.AgentSession, for example to test a Mastra-backed agent with @livekit/agents' voice.testing harness without speech-to-text, text-to-speech, or a live worker. createMastraVoiceAgent(options) is an equivalent factory.
import { initializeLogger, voice } from '@livekit/agents'
import { MastraVoiceAgent } from '@mastra/livekit/worker'
import { supportAgent } from './agents/support'
// Required outside a LiveKit worker: AgentSession needs the LiveKit logger initialized.
initializeLogger({ level: 'silent', pretty: false })
const session = new voice.AgentSession()
await session.start({ agent: new MastraVoiceAgent({ agent: supportAgent, memory: false }) })
// run() returns a RunResult, not a promise; wait() resolves when the turn completes.
const result = session.run({ userInput: 'What are your opening hours?' })
await result.wait()
result.expect.nextEvent().isMessage({ role: 'assistant' })
result.expect.noMoreEvents()
The agent carries its own placeholder llm.LLM so LiveKit runs the reply pipeline; generation always goes through llmNode, so a FakeLLM in the session's llm slot is ignored and calling the placeholder's chat() throws. To stub the model in tests, give the Mastra agent a mock model or pass a custom generate function.
OptionsDirect link to Options
Provide exactly one reply source: agent or generate.
agent?:
generate?:
memory?:
options is forwarded to the agent as per-call memory config, e.g. { readOnly: true }. When false, the full in-session LiveKit context is sent every turn.requestContext?:
toolFeedback?:
onToolCall?:
onTurnComplete?:
greetingReminder?:
streamOptions?:
instructions?:
id / stt / vad / tts / turnHandling?:
MastraLLMDirect link to mastrallm
A standard LiveKit LLM plugin (llm.LLM) backed by a Mastra agent. Use it when you build the voice.AgentSession yourself and want Mastra in the llm slot. createLiveKitWorker() is the managed alternative. See Use Mastra as the LLM component for how to choose.
With remote, the plugin streams each turn from your Mastra server over HTTP using Server-Sent Events (SSE). The agent loop, tools, and memory run server-side, and interrupting the agent aborts the server-side generation.
import { voice } from '@livekit/agents'
import { MastraLLM } from '@mastra/livekit/plugin'
const session = new voice.AgentSession({
llm: new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
memory: { thread: callId, resource: userId },
}),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
// Required with `memory` unless memory.options.readOnly is set: LiveKit enables
// preemptive generation by default.
turnHandling: { preemptiveGeneration: { enabled: false } },
})
The plugin reports provider as mastra and model as the agent id, so LiveKit metrics and fallback adapters identify it like any other LLM.
Constructor optionsDirect link to Constructor options
Provide exactly one reply source: remote, agent, or generate.
remote?:
agent?:
generate?:
memory?:
options is forwarded in the request body as memory.options, e.g. { readOnly: true }. When omitted, the full LiveKit chat context is sent every turn.requestContext?:
toolFeedback?:
onToolCall?:
onTurnComplete?:
Don't combine memory with the session's preemptiveGeneration option, which LiveKit enables by default in sessions you build yourself. A speculative turn persists a partial user message and a partial, never-spoken reply to the thread before LiveKit discards it. Set turnHandling: { preemptiveGeneration: { enabled: false } } on the session, or set memory.options.readOnly and persist committed turns yourself; see preemptive generation with memory. Stateless mode (no memory) works with preemptive generation.
Tools run on the Mastra agentDirect link to Tools run on the Mastra agent
Tools are defined and executed server-side on the Mastra agent. The plugin never forwards LiveKit tool definitions: if the session passes a non-empty toolCtx, it logs a one-time warning naming the ignored tools. Every tool must complete server-side: a tool that requires approval or client-side execution fails the turn with a descriptive error instead of hanging the call.
Tool activity reaches the worker through toolFeedback, onToolCall, and onTurnComplete.
InstructionsDirect link to Instructions
LiveKit injects your voice.Agent's instructions into the chat context of every request. The plugin drops them because the server-side Mastra agent's own instructions are authoritative. To change the prompt, change the Mastra agent.
Interrupted turnsDirect link to Interrupted turns
When the user interrupts a reply:
- The plugin cancels the stream. The server aborts generation and persists nothing from that turn.
- LiveKit records the part the user actually heard in its chat context, flagged as interrupted.
- On the next turn, the plugin re-sends that heard-only fragment, ordered before the new user message, so the memory thread backfills to match the call. Messages carry LiveKit's message ids and the server deduplicates by id, so retries and re-sends stay idempotent.
A user who hangs up immediately after interrupting leaves that final fragment unrecorded. When the transcript must capture it, reconcile immediately from the session event; the shared message id means the next turn's re-send upserts instead of duplicating:
import { voice } from '@livekit/agents'
import { MastraClient } from '@mastra/client-js'
const client = new MastraClient({ baseUrl: process.env.MASTRA_URL! })
session.on(voice.AgentSessionEventTypes.ConversationItemAdded, ({ item }) => {
if (item.type !== 'message' || item.role !== 'assistant' || !item.interrupted) return
void client.saveMessageToMemory({
agentId: 'support',
messages: [
{
id: item.id,
threadId: callId,
resourceId: userId,
role: 'assistant',
content: item.textContent ?? '',
type: 'text',
createdAt: new Date(),
},
],
})
})
Usage metricsDirect link to Usage metrics
When the server reports token usage for a turn, the plugin feeds it to LiveKit, so the session's metrics_collected events carry time-to-first-token, duration, and token counts like any LLM plugin. The same usage object (promptTokens, completionTokens, promptCachedTokens, totalTokens) arrives on onTurnComplete as result.usage.
Errors and timeoutsDirect link to Errors and timeouts
The transport throws LiveKit's APIError types (APIStatusError, APIConnectionError, APITimeoutError), so the session's retry policy (connOptions.maxRetry) and FallbackAdapter failover work unchanged. A turn is never retried after its first token: a voice reply is better failed fast than replayed half-heard.
A connect and first-token watchdog uses the session's connOptions.timeoutMs (10 seconds by default), so a server that accepts the connection but never streams can't cause indefinite dead air.
If the Mastra server goes down mid-call, each reply attempt fails with a typed error after its retries, and LiveKit closes the session after several consecutive failed replies. Restore the server before that budget runs out and the call recovers on the next turn.
Message contentDirect link to Message content
Message extraction is text-only: image content is dropped, and audio content is included only through its transcript. Voice pipelines aren't affected, but items you inject into the chat context yourself must carry text.
createRemoteAgentReplyGenerator()Direct link to createremoteagentreplygenerator
Builds a reply generator that runs the agent loop on a remote Mastra server over HTTP/SSE. MastraLLM's remote mode uses it internally. Use it directly through createLiveKitWorker's generate option to run the batteries-included worker against a remote server:
import { createLiveKitWorker, createRemoteAgentReplyGenerator } from '@mastra/livekit/worker'
import { mastra } from './index'
export default createLiveKitWorker({
mastra, // local instance for logger and worker config; replies come from the remote server
generate: createRemoteAgentReplyGenerator({
baseUrl: process.env.MASTRA_URL!,
agentId: 'support',
}),
memory: ({ metadata, roomName }) => ({ thread: metadata.threadId ?? roomName }),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
})
On the generate path the worker-level toolFeedback and onTurnComplete options don't apply, and the worker's end-call detection doesn't fire; pass the hooks to the generator instead.
Cancelling a turn (barge-in) tears down the HTTP request, which aborts generation on the server. Errors are thrown as LiveKit APIError types. The retries option applies only to initial connection attempts. A turn is never retried after its first chunk.
Returns: VoiceReplyGenerator.
OptionsDirect link to Options
baseUrl:
agentId:
apiPrefix?:
headers?:
fetch?:
timeoutMs?:
retries?:
body?:
toolFeedback?:
onToolCall?:
onTurnComplete?:
speakGreeting()Direct link to speakgreeting
Speaks an opening greeting on a session you own, honoring interruption and playout options. Returns the LiveKit SpeechHandle, or undefined when there's no greeting text. createLiveKitWorker() uses it internally for its greeting configuration.
import { speakGreeting } from '@mastra/livekit/worker'
await speakGreeting(session, {
text: "You've reached support. You're speaking with an AI assistant.",
allowInterruptions: false,
awaitPlayout: true,
})
ParametersDirect link to Parameters
session:
greeting:
waitForAgentDoneSpeaking()Direct link to waitforagentdonespeaking
Resolves once the agent is no longer producing or playing a reply: its state has left thinking and speaking. Resolves immediately when the agent is already idle, and always resolves within maxWaitMs (30 seconds by default) as a safety cap. Use it before tearing a session down so closing words play out instead of being cut off.
import { waitForAgentDoneSpeaking } from '@mastra/livekit/worker'
await waitForAgentDoneSpeaking(session)
runEndCall()Direct link to runendcall
Ends the call after the agent asks to hang up. It waits for the agent's closing words and speaks an optional final message without interruption. It then deletes the room and hangs up the caller, including SIP callers. The job shuts down with its registered callbacks.
Pair it with MastraLLM's onToolCall and an end-call tool on the server-side agent to rebuild agent-initiated hang-up on a session you own:
import { MastraLLM } from '@mastra/livekit/plugin'
import { DEFAULT_END_CALL_TOOL, runEndCall } from '@mastra/livekit/worker'
let ending = false
const llm = new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
onToolCall: ({ toolName }) => {
if (toolName !== DEFAULT_END_CALL_TOOL || ending) return
ending = true
void runEndCall(session, ctx, {}, console)
},
})
The exported constants DEFAULT_END_CALL_TOOL ('endCall'), DEFAULT_END_CALL_REASON, and DEFAULT_END_CALL_MAX_WAIT_MS (30000) hold the defaults.
ParametersDirect link to Parameters
session:
ctx:
config:
logger:
createEndCallTool()Direct link to createendcalltool
Builds the Mastra tool an agent calls when it wants to end the call. The tool signals intent and can run optional bookkeeping. The worker performs the actual hang-up. The tool lives on the server-safe root entry. Add it to agents defined in server code.
import { Agent } from '@mastra/core/agent'
import { createEndCallTool } from '@mastra/livekit'
const supportAgent = new Agent({
id: 'support',
name: 'Support',
instructions:
'Help the caller. When everything is wrapped up, say goodbye and call endCall as your final action.',
model: 'openai/gpt-5-mini',
tools: { endCall: createEndCallTool() },
})
With createLiveKitWorker(), set configuration: { endCall: {} } and the worker watches for the tool and hangs up. On a session you own, rebuild the hang-up with runEndCall().
OptionsDirect link to Options
id?:
description?:
onEndCall?:
liveKitConnectionRoute()Direct link to livekitconnectionroute
Returns an API route that mints a LiveKit access token with the voice agent dispatched into the room. Frontends call it to join a session.
import { Mastra } from '@mastra/core/mastra'
import { liveKitConnectionRoute } from '@mastra/livekit'
export const mastra = new Mastra({
server: {
apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],
},
})
The route accepts a JSON body with optional agentId, threadId, and resourceId fields and responds with { serverUrl, roomName, participantName, participantToken }. The threadId defaults to the generated room name.
OptionsDirect link to Options
path?:
serverUrl?:
apiKey?:
apiSecret?:
agentName?:
ttl?:
requiresAuth?:
roomName?:
participantIdentity?:
metadata?:
recording?:
dispatchVoiceSession()Direct link to dispatchvoicesession
Dispatches a Mastra voice agent into a LiveKit room programmatically: for server-initiated sessions such as outbound calls.
import { dispatchVoiceSession } from '@mastra/livekit'
await dispatchVoiceSession({
roomName: 'support-call-42',
agentName: 'mastra-voice',
metadata: { agentId: 'support', threadId: 'thread-42' },
})
OptionsDirect link to Options
roomName:
agentName?:
metadata?:
recording?:
serverUrl?:
apiKey?:
apiSecret?:
LiveKitSessionMetadataDirect link to livekitsessionmetadata
The metadata passed from the Mastra server to the worker through LiveKit job dispatch.
agentId?:
threadId?:
resourceId?:
requestContext?:
The metadata travels as a JSON string. liveKitConnectionRoute() and dispatchVoiceSession() serialize it for you; use serializeSessionMetadata(metadata) when dispatching through your own code, or write the JSON directly in LiveKit-side configuration such as a SIP dispatch rule. Entries in requestContext reach the agent's runtime-defined instructions, tools, and input processors on every turn of the call.