Skip to main content

LiveKit

Quickstart
Direct link to Quickstart

Realtime voice turns a Mastra agent into a live call a user can talk over, in the browser or over the phone. Mastra builds it on LiveKit, an open source WebRTC platform for realtime audio and video.

The @mastra/livekit package connects Mastra agents to the LiveKit Agents framework: LiveKit owns the audio loop like voice activity detection, streaming speech-to-text, semantic turn detection, barge-in, and text-to-speech. Your Mastra agent generates every reply with its own model, tools, and memory.

Use realtime voice when you need low-latency, interruptible voice conversations. For provider-based speech-to-speech without LiveKit, see Speech to Speech.

These steps take you from an empty project to a voice agent you can talk to. A voice session has two moving parts you set up here: an API route on your Mastra server that hands out access tokens, and a separate worker process that runs the audio pipeline and calls your agent each turn.

  1. Install the integration package along with the LiveKit plugins for voice activity detection and turn detection:

    npm install @mastra/livekit @livekit/agents @livekit/agents-plugin-silero @livekit/agents-plugin-livekit
  2. Set your LiveKit credentials inside an .env file. Create a free project on LiveKit Cloud, or run a local server with livekit-server --dev:

    .env
    LIVEKIT_URL=wss://your-project.livekit.cloud
    LIVEKIT_API_KEY=your-api-key
    LIVEKIT_API_SECRET=your-api-secret
  3. Add a voice agent to your Mastra instance and expose a connection route. The liveKitConnectionRoute() helper adds a POST /voice/livekit/connection-details endpoint that mints a LiveKit token and dispatches your agent into a room:

    src/mastra/index.ts
    import { Mastra } from '@mastra/core/mastra'
    import { Agent } from '@mastra/core/agent'
    import { liveKitConnectionRoute } from '@mastra/livekit'

    const supportAgent = new Agent({
    id: 'support',
    name: 'Support',
    instructions: 'You are a friendly phone support agent. Keep replies short and conversational.',
    model: 'openai/gpt-5-mini',
    })

    export const mastra = new Mastra({
    agents: { support: supportAgent },
    server: {
    apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],
    },
    })
  4. Create the worker. It runs as a separate process, answers LiveKit sessions, and calls your agent each turn. Worker APIs live on the @mastra/livekit/worker entry point, so the Mastra server never loads the LiveKit agents runtime. This example uses LiveKit Inference model strings for speech-to-text and text-to-speech, so no provider plugins are required:

    src/mastra/voice-worker.ts
    import { fileURLToPath } from 'node:url'
    import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker'
    import { mastra } from './index'

    export default createLiveKitWorker({
    mastra,
    agent: 'support',
    stt: 'deepgram/nova-3',
    tts: 'cartesia/sonic-3',
    turnDetection: 'multilingual',
    greeting: 'Hi! How can I help you today?',
    })

    if (process.argv[1] === fileURLToPath(import.meta.url)) {
    runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
    }

    The agent option selects which Mastra agent answers each session. Pass a fixed key as shown, or omit it to use the agentId from the dispatch metadata, so one worker can serve every agent on your Mastra instance.

  5. Download the turn detection and voice activity detection models once. Then run the worker in one terminal and your Mastra server in another:

    npx livekit-agents download-files
    npx tsx src/mastra/voice-worker.ts dev
    npm run dev

    The worker registers with your LiveKit server and waits for sessions, while mastra dev serves the connection route.

  6. Talk to your agent. Open the hosted LiveKit Agents Playground and connect it to your project to start a call without building a frontend.

    To wire up your own app instead, call the connection route for a token. POST /voice/livekit/connection-details accepts optional agentId, threadId, and resourceId fields in the request body and returns:

    {
    "serverUrl": "wss://your-project.livekit.cloud",
    "roomName": "mastra-voice-a1b2c3d4",
    "participantName": "user-1",
    "participantToken": "eyJhbGci..."
    }

    This response matches the contract used by LiveKit's frontend starters, so apps built from agent-starter-react or the LiveKit React components work without changes.

Turn detection and interruptions
Direct link to Turn detection and interruptions

LiveKit decides when the user finished speaking and when the agent was interrupted. The defaults work well; tune them with turnHandling:

src/mastra/voice-worker.ts
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
turnHandling: {
endpointing: { mode: 'dynamic', minDelay: 300, maxDelay: 3000 },
interruption: { minDuration: 500, resumeFalseInterruption: true },
},
})
  • turnDetection: 'multilingual': Runs LiveKit's semantic end-of-turn model locally on CPU. It reads the live transcript to avoid cutting users off mid-thought. Use 'vad' or 'stt' for silence-based endpointing instead.
  • endpointing: Bounds how long the agent waits after the user stops speaking.
  • interruption: Controls barge-in. When the user speaks over the agent, LiveKit stops playback and cancels the in-flight Mastra stream, so token generation stops too.
  • preemptiveGeneration: Starts the Mastra agent's reply while the user is still finishing, hiding time-to-first-token. The worker disables it by default: each preemptive attempt runs the Mastra agent on an interim transcript, and a run that LiveKit later discards has already persisted a partial user message and a partial, never-spoken reply to the thread. Re-enable it with preemptiveGeneration: { enabled: true } if latency matters more than exact thread history, or keep both by running turns read-only; see preemptive generation with memory.

See the LiveKit turn detection docs for all options.

To disable voice activity detection, set vad: false on createLiveKitWorker(). The worker passes null to LiveKit, which otherwise enables its bundled detector when vad is omitted. Choose a turn detection mode that doesn't require VAD, such as 'manual', when disabling it. Values in sessionOptions override the worker's generated options, including vad.

Per-call voices and transcription
Direct link to Per-call voices and transcription

The top-level stt and tts options apply to every call. To pick them per call, one voice or language per tenant, set the configuration.stt and configuration.tts resolvers instead. Each resolver runs once per call with the dispatch metadata, request context, room name, and job context, and returns a value accepted by the matching top-level option. The value is either a plugin instance or an inference model string. Return undefined to fall back to the top-level option.

The following example gives each tenant its own text-to-speech voice, keyed off the tenant entry in the dispatch metadata:

src/mastra/voice-worker.ts
import * as cartesia from '@livekit/agents-plugin-cartesia'

// One voice id per tenant, resolved from the dispatch metadata on each call.
const tenantVoices: Record<string, string> = {
meridian: 'your-cartesia-voice-id-1',
coastal: 'your-cartesia-voice-id-2',
}
// The resolver runs during call setup, so cache plugin instances across calls.
const ttsByVoice = new Map<string, cartesia.TTS>()

export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
configuration: {
tts: ({ requestContext }) => {
const voice = tenantVoices[requestContext?.tenant as string]
if (!voice) return undefined // fall back to the top-level `tts`
let tts = ttsByVoice.get(voice)
if (!tts) {
tts = new cartesia.TTS({ voice })
ttsByVoice.set(voice, tts)
}
return tts
},
},
})

configuration.stt works the same way for per-call transcription, for example a different transcription model or language per tenant. The greeting has a matching per-call form: configuration.greeting.text accepts a resolver with the same call context, so one worker can open with each tenant's own phrasing.

Per-call turn detection
Direct link to Per-call turn detection

LiveKit's TurnDetector classes read the job's inference executor when constructed, so they can only be created inside a LiveKit job, not at module scope where the worker options live. To use one, set the configuration.turnDetection resolver. It runs once per call with the same call context as configuration.stt and returns anything the top-level turnDetection option accepts. Return undefined to fall back to the top-level option.

src/mastra/voice-worker.ts
import { turnDetector } from '@livekit/agents-plugin-livekit'

export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
configuration: {
// Constructed inside the job, where the inference executor is available.
turnDetection: () => new turnDetector.MultilingualModel(0.2),
},
})

The semantic model's inference runners must be registered before the agent server boots, so when this resolver is set the worker imports @livekit/agents-plugin-livekit up front. Keep the top-level turnDetection set to 'multilingual' or 'english' to pre-register only that model; otherwise both stay available.

Memory and threads
Direct link to Memory and threads

When the resolved Mastra agent has memory configured, each call becomes one memory thread:

  • thread defaults to the threadId from dispatch metadata, then to the room name.
  • resource defaults to the resourceId from dispatch metadata, then to the thread. Send your end user's id here so calls group under the right user. Mastra Studio sends the agent id, matching how its sidebar lists threads.
  • When the thread doesn't exist yet, the worker creates it titled "Voice call" with metadata { source: 'livekit' }, and the spoken greeting is saved as the first assistant message so the thread reads as a full call transcript (disable with persistGreeting: false).

Each turn sends only the new user input; Mastra Memory supplies history, semantic recall, and working memory. Pin a session to an existing thread by passing threadId in the connection request body, which is useful for continuing a text conversation by voice. In Studio, starting a call from an open chat binds the call to that thread, and the transcript fills into the chat after each exchange.

When a user interrupts the agent, the in-flight generation aborts and nothing from that turn is persisted at that moment. LiveKit keeps the part the user actually heard in its transcript, and on the next turn the worker re-sends that heard-only fragment so the thread backfills to match the call. A user who hangs up right after interrupting leaves that final fragment unrecorded. See interrupted turns for the details and a reconciliation recipe.

Preemptive generation with memory
Direct link to Preemptive generation with memory

LiveKit's preemptive generation calls the Mastra agent on interim transcripts and discards runs whose transcript changed. The plugin can't tell a speculative run from a real turn, so with memory set every run persists, including discarded ones. To keep preemptive generation on without corrupting the thread, pass options: { readOnly: true } in the memory mapping. The agent still reads history, semantic recall, and working memory from the thread but writes nothing, so speculative runs leave no trace. Persistence of committed turns then belongs to you: save them from LiveKit's ConversationItemAdded event, which fires only for items the session committed. Messages keep LiveKit's ids, so saves stay idempotent across retries.

src/mastra/voice-worker.ts
import { voice } from '@livekit/agents'
import { createLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'

export default createLiveKitWorker({
mastra,
agent: 'support',
memory: ({ metadata, roomName }) => ({
thread: metadata.threadId ?? roomName,
resource: metadata.resourceId ?? roomName,
options: { readOnly: true },
}),
turnHandling: { preemptiveGeneration: { enabled: true } },
onSessionStart: async ({ session, ctx, agent }) => {
const mapping = agent.memory
const memory = await mastra.getAgent('support').getMemory()
if (!mapping || !memory) return

let shuttingDown = false
const maxRetries = 5
const retryTimers = new Set<ReturnType<typeof setTimeout>>()
ctx.addShutdownCallback(async () => {
shuttingDown = true
for (const timer of retryTimers) clearTimeout(timer)
retryTimers.clear()
})

session.on(voice.AgentSessionEventTypes.ConversationItemAdded, ({ item }) => {
if (item.type !== 'message' || (item.role !== 'user' && item.role !== 'assistant')) return

const persist = async (attempt = 0): Promise<void> => {
try {
await memory.saveMessages({
messages: [
{
id: item.id,
threadId: mapping.thread,
resourceId: mapping.resource ?? mapping.thread,
role: item.role,
content: {
format: 2,
parts: [{ type: 'text', text: item.textContent ?? '' }],
},
type: 'text',
createdAt: new Date(),
},
],
})
} catch (error) {
if (shuttingDown) return
if (attempt >= maxRetries) {
console.error(`Failed to persist committed voice item ${item.id}; giving up`, error)
return
}
console.error(`Failed to persist committed voice item ${item.id}; retrying`, error)
const delay = Math.min(1_000 * 2 ** attempt, 30_000)
const timer = setTimeout(() => {
retryTimers.delete(timer)
void persist(attempt + 1)
}, delay)
retryTimers.add(timer)
}
}

void persist()
})
},
})

The same options field works on MastraVoiceAgent and MastraLLM; on the remote transport it's forwarded in the request body as memory.options.

Speak while tools run
Direct link to Speak while tools run

Voice conversations can't go silent while a slow tool runs. Use toolFeedback to speak a short phrase when the Mastra agent starts a tool call:

src/mastra/voice-worker.ts
export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
toolFeedback: ({ toolName }) =>
toolName === 'searchOrders' ? 'Let me look that up.' : undefined,
})

The phrase is spoken as part of the reply and recorded in the transcript.

Generate replies with a workflow
Direct link to Generate replies with a workflow

By default the worker generates each reply with a Mastra agent. To run multi-step logic per turn (for example classify intent, route, call tools in sequence, then compose a reply), generate replies with a Mastra workflow instead. Set workflow in place of agent.

LiveKit still owns the audio loop and calls into Mastra once per turn, so the workflow runs to completion each turn. The workflow can't suspend or resume, and no conversation state carries between turns. Pass the transcript in through workflowInput so the workflow stays stateless:

src/mastra/voice-worker.ts
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker'
import { mastra } from './index'

export default createLiveKitWorker({
mastra,
workflow: 'phoneConversation',
workflowInput: ({ chatCtx }) => ({ history: chatContextToMessages(chatCtx) }),
replyStep: 'generateResponse',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
})

A workflow streams structured step events, not text. To speak tokens as they generate, the reply step pipes its agent's text into the step writer:

src/mastra/workflows/phone-conversation.ts
const generateResponse = createStep({
id: 'generateResponse',
// input and output schemas omitted
execute: async ({ inputData, mastra, writer, abortSignal }) => {
const stream = await mastra.getAgent('voice').stream(inputData.history, { abortSignal })
await stream.textStream.pipeTo(writer)
return { assistantMessage: await stream.text }
},
})
  • replyStep: Restricts spoken output to one step. Omit it to speak every step that writes to its writer.
  • resultText: A fallback that derives the reply from the final run result when no step streams text. Streaming through writer gives lower time-to-first-token, so prefer it.
  • abortSignal: Forward the step's abortSignal into agent.stream() so barge-in stops generation promptly. When the user interrupts, the worker cancels the run.
  • generate: For full control, pass a generate function instead. It can be any reply generator that turns a turn into a text stream.

With a workflow, the worker doesn't persist turns automatically the way an agent's stream() does. Persist conversation history inside the workflow, or keep the LiveKit transcript as the source of truth and pass it in each turn.

Use Mastra as the LLM component
Direct link to Use Mastra as the LLM component

createLiveKitWorker() owns the LiveKit session for you. To own the session yourself, use MastraLLM instead: a standard LiveKit LLM plugin that puts a Mastra agent in the llm slot of your own voice.AgentSession. The Mastra app, agent loop, tools, memory, observability, runs on your Mastra server, and the worker reaches it over HTTP. The worker process needs no Mastra app, database, or model provider keys.

src/mastra/voice-worker-plugin.ts
import { fileURLToPath } from 'node:url'
import { defineAgent, voice } from '@livekit/agents'
import * as silero from '@livekit/agents-plugin-silero'
import { MastraLLM } from '@mastra/livekit/plugin'
import { runLiveKitWorker } from '@mastra/livekit/worker'

export default defineAgent({
entry: async ctx => {
await ctx.connect()

const session = new voice.AgentSession({
llm: new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
memory: { thread: ctx.room.name!, resource: 'user-7' },
}),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
vad: await silero.VAD.load(),
// Required with `memory` unless memory.options.readOnly is set: LiveKit enables
// preemptive generation by default.
turnHandling: { preemptiveGeneration: { enabled: false } },
})

await session.start({
// These instructions never reach the Mastra agent; its own instructions apply.
agent: new voice.Agent({ instructions: 'Replies come from the Mastra agent.' }),
room: ctx.room,
})

session.say('Hi! How can I help you today?')
},
})

if (process.argv[1] === fileURLToPath(import.meta.url)) {
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
}

Both paths share the same reply pipeline underneath; choose by who should own the session:

createLiveKitWorker()MastraLLM
Session ownershipThe worker helper builds and manages the AgentSessionYour code builds the session; every LiveKit option and hook is yours
Where the Mastra app runsIn the worker processOn your Mastra server, reached over HTTP (or in-process via agent)
Worker process needsYour Mastra app, storage, and model provider keysOnly the LiveKit SDK and network access to your server
Built-in conveniencesGreeting, consent gating, agent-initiated hang-up, thread bootstrap, observability roll-upRebuild what you need with the session helpers
Best forFastest path to a working voice agent; Studio voice modeExisting LiveKit apps and full control over the session

Tools stay on the Mastra agent and execute on the server. LiveKit-side tools passed to the session are ignored. Tool activity reaches the worker through toolFeedback (spoken filler), onToolCall (fires as each tool call starts), and onTurnComplete (fires after each reply with the text, tool calls, and token usage). Agent-initiated hang-up takes a few lines: pair onToolCall with runEndCall().

warning

Don't combine the memory option with LiveKit's preemptiveGeneration, which LiveKit enables by default in sessions you build yourself. A speculative turn persists a partial user message and a partial, never-spoken reply to the thread before LiveKit discards it. Set turnHandling: { preemptiveGeneration: { enabled: false } }, run without memory and pass the full transcript each turn, or set memory.options.readOnly and persist committed turns yourself; see preemptive generation with memory.

MastraLLM also accepts an in-process Mastra agent instance, session ownership without a second deployment, or a custom generate function. The remote transport is available standalone as createRemoteAgentReplyGenerator(), which also plugs into createLiveKitWorker's generate option to run the batteries-included worker against a remote server.

Server-initiated sessions
Direct link to Server-initiated sessions

Use dispatchVoiceSession() to add a voice agent to a room from your own code, for example to join an existing room or to drive an outbound SIP call:

import { dispatchVoiceSession } from '@mastra/livekit'

await dispatchVoiceSession({
roomName: 'support-call-42',
agentName: 'mastra-voice',
metadata: { agentId: 'support', threadId: 'thread-42', resourceId: 'user-7' },
})

Record calls
Direct link to Record calls

Pass recording to dispatchVoiceSession() or liveKitConnectionRoute() to create a room with LiveKit auto egress. Mastra creates the room with recording configured, then dispatches the voice agent. Omitting recording preserves the existing behavior and return values.

Recording requires LiveKit Cloud or a self-hosted egress service, plus an output destination. The example below records the caller and agent into a single mixed OGG audio file in an S3 bucket. Set RECORDINGS_S3_BUCKET, RECORDINGS_S3_REGION, RECORDINGS_S3_ACCESS_KEY, and RECORDINGS_S3_SECRET on your server, alongside the LiveKit credentials.

Install the server SDK to use its output format enum:

npm install livekit-server-sdk

Define recording settings in a server-only module:

src/mastra/recording.ts
import type { LiveKitRecordingOptions } from '@mastra/livekit'
import { EncodedFileType } from 'livekit-server-sdk'

export const recording: LiveKitRecordingOptions = {
room: {
audioOnly: true,
fileOutputs: [
{
fileType: EncodedFileType.OGG,
filepath: 'calls/{room_name}.ogg',
output: {
case: 's3',
value: {
bucket: process.env.RECORDINGS_S3_BUCKET!,
region: process.env.RECORDINGS_S3_REGION!,
accessKey: process.env.RECORDINGS_S3_ACCESS_KEY!,
secret: process.env.RECORDINGS_S3_SECRET!,
},
},
},
],
},
}

LiveKitRecordingOptions accepts the settings used to construct LiveKit's RoomEgress, as either a plain object or an SDK instance. Set room.audioOnly: true for mixed audio, or use tracks or participant for separate recordings. Mastra passes these settings to LiveKit without choosing a format or storage provider for you. See LiveKit output options for other destinations.

For an outbound phone call, configure recording before adding the phone participant:

src/start-call.ts
import { randomUUID } from 'node:crypto'
import { dispatchVoiceSession } from '@mastra/livekit'
import { recording } from './mastra/recording'

const roomName = `support-call-${randomUUID()}`

await dispatchVoiceSession({
roomName,
metadata: { agentId: 'support', resourceId: 'caller-7' },
recording,
})

// Add the phone participant to roomName through your existing LiveKit SIP code.

Use a fresh room name for each recorded session and await dispatch before dialing. dispatchVoiceSession() throws LiveKitRecordingRoomConflictError if the room already exists. Import this error from @mastra/livekit to handle the conflict. The existence check and room creation are separate requests, so your application must prevent another process from creating the same room concurrently. This option doesn't start recording in an existing room or capture earlier audio.

For browser sessions, pass the same settings to the connection route:

src/mastra/index.ts
import { Mastra } from '@mastra/core/mastra'
import { liveKitConnectionRoute } from '@mastra/livekit'
import { recording } from './recording'

export const mastra = new Mastra({
server: {
apiRoutes: [liveKitConnectionRoute({ recording })],
},
})

With recording enabled, the route creates the room and dispatches the agent when connection details are requested, before the browser joins. The response remains { serverUrl, roomName, participantName, participantToken }. Recording settings and storage credentials stay on the server and aren't included in the join token or agent metadata.

The route also accepts a synchronous or asynchronous recording callback. It receives { body, context, roomName }, including the resolved room name, so your server can choose settings for each session. Return undefined to skip recording for that request. Recording configuration isn't read directly from the request body.

An existing recording room returns HTTP 409 before agent dispatch or token issuance. Other room creation and dispatch errors reject the operation; the route doesn't issue a token on failure. If dispatch fails after room creation, the room remains in LiveKit. Successful setup doesn't guarantee that the recording finishes or uploads successfully. Use LiveKit's egress events to check completion and retrieve file details before making a recording available to users. To verify the setup, make a test call, end the room, and listen to the completed file under calls/ in your bucket for audio from both participants.

Enabling recording starts recording during room creation, before a browser connects or a phone participant joins. An abandoned connection request can still start the agent and consume recording and storage resources. Configure authentication and application limits on the connection route before exposing it to users.

createConsentTool() and configuration.consentPolicy don't control LiveKit egress. They don't delay, pause, or stop an automatic recording. If your application needs consent during the call before capturing audio, obtain it first and use LiveKit's explicit egress APIs to start recording in the existing room. A spoken greeting alone doesn't gate this feature.

Recordings follow the LiveKit room lifecycle. Disconnecting only the agent doesn't necessarily end the room or finalize the file. Wait for a successful egress completion status and file results before marking a recording ready. Treat failed setup and abandoned rooms as application cleanup responsibilities.

The audio file lives in the configured output storage, separately from Mastra memory, transcripts, and traces. Deleting a Mastra thread doesn't delete the recording. Your application manages storage access, retention, and download links for end users. Keep storage credentials in server configuration, including when choosing recording settings per session.

Review recordings in Studio
Direct link to Review recordings in Studio

Register liveKitRecordingRoute() to add Review Audio to a voice call's trace panel in Studio. The button opens an audio player for that call. Recording and trace storage must both be configured: the route reads the stored voice call span and passes its roomName to your server-side resolver. The browser supplies only the trace ID.

The resolver returns { url, expiresAt? }, or undefined when no file is available. Use a short-lived signed URL for private storage. The route doesn't persist playback URLs in traces, and Studio requests a fresh URL each time the review opens. Missing files and playback errors offer Refresh recording so you can retry after the upload finishes or a link expires.

For the S3 configuration above, use the same calls/{room_name}.ogg object key for recording and playback. Room names must remain unique across calls, including after a room has ended. Install the AWS SDK in your application:

npm install @aws-sdk/client-s3 @aws-sdk/s3-request-presigner
src/mastra/recording-playback.ts
import {
GetObjectCommand,
HeadObjectCommand,
S3Client,
S3ServiceException,
} from '@aws-sdk/client-s3'
import { getSignedUrl } from '@aws-sdk/s3-request-presigner'
import { liveKitRecordingRoute } from '@mastra/livekit'
import { authorizeRecording } from './recording-access'

const s3 = new S3Client({
region: process.env.RECORDINGS_S3_REGION!,
credentials: {
accessKeyId: process.env.RECORDINGS_S3_ACCESS_KEY!,
secretAccessKey: process.env.RECORDINGS_S3_SECRET!,
},
})

export const recordingReviewRoute = liveKitRecordingRoute({
authorize: authorizeRecording,
resolveRecording: async ({ roomName }) => {
const object = {
Bucket: process.env.RECORDINGS_S3_BUCKET!,
Key: `calls/${roomName}.ogg`,
}
try {
await s3.send(new HeadObjectCommand(object))
} catch (error) {
if (error instanceof S3ServiceException && error.$metadata.httpStatusCode === 404)
return undefined
throw error
}
const expiresIn = 15 * 60
return {
url: await getSignedUrl(
s3,
new GetObjectCommand({ ...object, ResponseContentType: 'audio/ogg' }),
{ expiresIn },
),
expiresAt: new Date(Date.now() + expiresIn * 1000).toISOString(),
}
},
})

Add recordingReviewRoute to server.apiRoutes alongside liveKitConnectionRoute({ recording }). Keep the default path, GET /voice/livekit/recordings/:traceId, for Studio discovery. Both the connection route and server-initiated dispatch recordings can be reviewed, as long as the worker persisted the call trace and your resolver can locate its file.

The review route requires authentication by default and an explicit authorize({ traceId, context }) callback. Implement the authorizeRecording helper in your application to check the authenticated user's access to the requested trace against your trusted ownership or tenant records. Return true only when access is allowed. Authentication alone doesn't establish recording ownership.

The route rejects missing authorization callbacks at registration, including when requiresAuth is false. A denied request returns 403 before trace storage is read or a playback URL is generated. Policy errors also stop resolution. The resolver receives the stored span and request context only after authorization succeeds.

warning

Use requiresAuth: false only for a local demo. It doesn't bypass the required recording authorization callback. If a demo policy allows all recordings, keep that policy explicitly limited to local development and replace it before deployment.

Recording upload permissions alone don't grant playback access. The signing credentials need s3:GetObject for the recording objects, including the HEAD check. S3 also requires s3:ListBucket to return a missing-object 404 instead of 403. Your application can use separate read credentials for playback. The browser receives a temporary URL that authorizes access to that file, so treat it as sensitive. See S3 object permissions and presigned URLs.

liveKitRecordingRoute() is storage-independent. You can resolve URLs from another provider or a recording index populated by LiveKit completion events. It doesn't depend on LiveKit retaining historical egress jobs. If you use timestamped filenames or change buckets, retain the mapping from room to object in your application and resolve that mapping instead of constructing a key.

Observability
Direct link to Observability

When the Mastra instance has observability configured, the worker traces each call. It opens one voice call span per session and nests everything under it:

  • Every turn's Mastra agent run, with model generation, tool calls, and memory operations, exactly as a text chat records them.
  • A child span for each LiveKit pipeline metric: speech-to-text, text-to-speech, end-of-utterance (turn detection), voice activity detection, and the model's time-to-first-token. These carry the latency and audio measurements that text traces can't show.
  • A per-model usage roll-up (token, character, and audio totals for the whole call) written to the span when the session ends.

The worker is a separate process, so point storage at a backend that accepts concurrent writes from both the server and the worker. SQLite-backed LibSQL works. Single-writer stores don't. Traces, memory, and threads can share one store:

src/mastra/index.ts
import { Mastra } from '@mastra/core/mastra'
import { LibSQLStore } from '@mastra/libsql'
import { Observability, MastraStorageExporter } from '@mastra/observability'

export const mastra = new Mastra({
storage: new LibSQLStore({ id: 'voice-agent-storage', url: 'file:./voice-agent.db' }),
observability: new Observability({
configs: {
default: {
serviceName: 'voice-agent',
exporters: [new MastraStorageExporter()],
},
},
}),
})

Tracing is on by default. Pass observability: false to createLiveKitWorker to turn it off.

Request recording playback from your application
Direct link to Request recording playback from your application

Import the browser client helper from @mastra/livekit/client. Install @mastra/client-js 1.51.2 or newer within 1.x to use this entry point:

npm install @mastra/livekit @mastra/client-js
import { MastraClient } from '@mastra/client-js'
import { getLiveKitRecording } from '@mastra/livekit/client'
import type { LiveKitRecordingResponse } from '@mastra/livekit/client'

const client = new MastraClient({
baseUrl: 'https://mastra.example.com',
credentials: 'include',
})

const recording: LiveKitRecordingResponse = await getLiveKitRecording(client, traceId)
if (recording.status === 'ready') {
audioElement.src = recording.url
}

Register liveKitRecordingRoute() on the server first. The helper requests its default /voice/livekit/recordings/:traceId path, independently of the client's API prefix. It preserves the client's authentication headers, cookies, custom fetch function, and abort signal. Pass { signal } as a third argument to cancel an individual request. Failed requests use MastraClientError and aren't retried automatically.

@mastra/client-js@1.51.2 declares support for Zod 3 and 4, but its @ai-sdk/ui-utils@1.2.11 dependency declares a Zod 3 peer dependency. This mismatch can cause strict npm installs with Zod 4 to fail. Use Zod ^3.25.76 with that SDK to satisfy both peer ranges. Server and worker consumers without the optional client SDK also support Zod 4.

The /client entry point excludes LiveKit server and worker imports. Other integration entry points don't require @mastra/client-js. Installing @mastra/livekit still installs its package-level dependencies; the browser entry point only controls which code enters the browser bundle. Storage credentials and recording authorization remain on the server.

Turn metrics and playback completion
Direct link to Turn metrics and playback completion

createLiveKitWorker() accepts two observers for evaluating voice calls:

src/mastra/voice-worker.ts
import { createLiveKitWorker } from '@mastra/livekit/worker'

export default createLiveKitWorker({
mastra,
agent: 'support',
onTurnMetrics: metrics => {
console.info('voice metric', metrics)
},
onSpeechComplete: speech => {
console.info('speech outcome', speech.speechId, speech.outcome)
},
})

Observers run asynchronously and aren't awaited by the audio pipeline. Errors are logged. Keep their work small. Neither observer writes or reconciles memory.

onTurnMetrics receives versioned records with phase: 'generation' or phase: 'speech'. Generation records contain a turnId, a unique attemptId, a terminal outcome, durationMs, optional firstTextMs, and tool timings. Attempts for the same user message share its turn ID. Tool durations run from the bridge receiving a tool call to receiving its result; remote durations include transport and aren't pure server execution times. A tool without a result has no duration.

Speech records also contain LiveKit's speechId. The bridge carries turn and attempt IDs in assistant message metadata so they can be joined after playback. Speech without a correlated assistant message, such as cancelled speech that produced no output, uses its speech ID as the turn ID and omits the attempt ID. Don't infer a correlation by matching event arrival order.

MeasurementMeaning
Generation firstTextMsTime from starting generation to the first text forwarded toward speech synthesis. Can include filler.
Speech firstAudioMsTime from caller speech ending to the first output audio, as reported by LiveKit. Can include filler.
Speech completionMsTime from caller speech ending to completed output playback. Omitted for interrupted, cancelled, or failed speech.
speechStartedAt / speechEndedAtPlayback timestamps in Unix epoch milliseconds when available.

Missing measurements are undefined, not zero. Local generation durations use a monotonic clock. LiveKit's speech timing values are converted to milliseconds. When observability is configured, these records appear as events in the existing voice call trace. Pipeline metric events retain LiveKit's speech ID when supplied.

onSpeechComplete runs once when a speech handle finishes, or when the session closes with that speech pending. It supplies the speech metrics plus optional playedText and transcriptSource. Outcomes are completed, interrupted, cancelled, or failed. A completed speech lifecycle alone doesn't prove that audio was produced; check whether the relevant audio measurements are present.

The measurement source is server-playout. LiveKit's committed transcript can be partial or estimated, depending on the output's synchronization support. It doesn't prove that the remote browser or phone listener heard every word. transcriptSource: 'unavailable' means no committed assistant transcript was available.

Keep using onTurnComplete for its existing generation contract. It can fire before speech playback finishes, with interrupted: false, even when the caller interrupts the audio later. Use onSpeechComplete for playback outcomes. Existing agent memory may already contain generated text; adding this hook doesn't change that persistence behavior.

For a session you own, call observeVoiceSession(session, { onTurnMetrics, onSpeechComplete }) from @mastra/livekit/worker or @mastra/livekit/plugin before session.start(). Its return value detaches the observer and finalizes pending speech as cancelled. The worker attaches its observer automatically.

Speech segment boundaries
Direct link to Speech segment boundaries

The built-in agent, remote, and workflow reply generators preserve text-segment and tool-step boundaries. Tool feedback is flushed after its text is queued, so an acknowledgment can reach text-to-speech while a tool is still running. A flush requests a segment boundary; it doesn't guarantee immediate audio from every speech provider.

Custom generators can continue returning string streams. To emit an explicit boundary, return a ReadableStream<VoiceReplyChunk> and enqueue VOICE_TEXT_FLUSH, both exported from @mastra/livekit/worker and @mastra/livekit/plugin. Consumers that read a reply generator directly must distinguish strings from boundary objects. Never concatenate a boundary object into a transcript.

The worker translates boundaries into LiveKit's FlushSentinel. With the MastraLLM plugin, add the node adapter to the LiveKit agent you own:

import { voice } from '@livekit/agents'
import { MastraLLM, mastraLLMNode } from '@mastra/livekit/plugin'

class SupportAgent extends voice.Agent {
override llmNode(...args: Parameters<voice.Agent['llmNode']>) {
return mastraLLMNode(this, ...args)
}
}

const session = new voice.AgentSession({
llm: new MastraLLM({ agent: supportAgent }),
// Configure STT and TTS for your application.
})
await session.start({ agent: new SupportAgent({ instructions: 'Support assistant' }), room })

Without the adapter, the plugin still emits text, but LiveKit's default node doesn't translate the plugin's boundary metadata into flush markers. Use pipeAgentReplyToWriter() in a workflow reply step to preserve the boundaries and tool results used for timing.

Deployment
Direct link to Deployment

The worker is a separate process from your Mastra server. Deploy it as a long-running Node service with the production command:

node dist/voice-worker.js start

LiveKit's guidance on sizing, graceful shutdown, and hosting applies unchanged. See Deploying agents. Workers connect outbound to LiveKit, so they don't need inbound ports.

How it works
Direct link to How it works

A LiveKit voice session involves three pieces:

  1. Your Mastra server mints a LiveKit access token and dispatches your agent into a room. The dispatch carries metadata such as the Mastra agent id, memory thread, and resource.
  2. A LiveKit agent worker (a separate long-running process) receives the job and runs the audio pipeline. Audio flows between the browser and the worker over WebRTC and never passes through your Mastra HTTP server.
  3. Each time the user finishes a turn, the worker calls the Mastra agent's stream() with the new input and speaks the streamed text. When the user interrupts, LiveKit cancels the stream and Mastra stops generating.

Conversation history lives in Mastra Memory, so voice sessions and text chat can share one thread.

Upgrade to 0.5.0
Direct link to Upgrade to 0.5.0

The 0.5.0 prerelease requires LiveKit Agents 1.7.1 or newer within 1.x. The speech helpers use FlushSentinel and SpeechHandle.exception(), which are absent in Agents 1.4.0. Test the supplied prerelease version or package archive before deploying it. Updating a dependency range doesn't update an existing lockfile: check the installed versions after installation.

If you used client.getLiveKitRecording(traceId) in an earlier snapshot, switch to getLiveKitRecording(client, traceId) from @mastra/livekit/client. Import LiveKitRecordingResponse from the same entry point instead of GetLiveKitRecordingResponse from client-js.

Authorize recording access
Direct link to Authorize recording access

Before upgrading an application that registers liveKitRecordingRoute(), add an authorize({ traceId, context }) callback that checks the authenticated user's access to the requested trace. The callback is required even when requiresAuth: false; omitting it throws during route registration. Return true only to allow access, before trace lookup or playback URL resolution. See Review recordings in Studio for the setup example.

Install compatible packages
Direct link to Install compatible packages

Use the same exact release for Agents and its plugins. The tested baseline is:

PackageVersion
@livekit/agents1.7.1
@livekit/agents-plugin-livekit1.7.1
@livekit/agents-plugin-silero1.7.1
@livekit/rtc-node0.13.34
livekit-server-sdk2.16.0

For a local prerelease archive, install it with the matching LiveKit packages:

npm install ./mastra-livekit-0.5.0-next.0.tgz @livekit/agents@1.7.1 @livekit/agents-plugin-livekit@1.7.1 @livekit/agents-plugin-silero@1.7.1 @livekit/rtc-node@0.13.34 livekit-server-sdk@2.16.0

Replace the archive path with the supplied prerelease. Install any other @livekit/agents-plugin-* packages at the matching release. Agents 1.7.1 and Silero 1.7.1 require RTC ^0.13.34; RTC 1.x doesn't satisfy that requirement. Don't override peer dependency errors to install it.

Keep Node.js 22.13 or newer. Both @mastra/livekit and Agents require Zod ^3.25.76 || ^4.1.8. Upgrade older Zod versions before installing 0.5.0. Verify the installed dependency tree and initialize the worker's models:

npm ls @mastra/livekit @livekit/agents @livekit/agents-plugin-livekit @livekit/agents-plugin-silero @livekit/rtc-node
npx tsx src/mastra/voice-worker.ts download-files

Use a single installed copy of Agents for both the integration and a custom voice.AgentSession. Multiple copies can fail LiveKit's runtime class checks.

Check VAD and turn detection
Direct link to Check VAD and turn detection

vad: false now disables VAD at session creation. Previously, the worker skipped plugin prewarming but passed undefined to LiveKit, which could enable its bundled VAD. If an application relied on that behavior, configure a detector explicitly.

src/mastra/voice-worker.ts
import { createLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'

export default createLiveKitWorker({
mastra,
agent: 'support',
vad: false,
turnDetection: 'manual',
})

The default still loads the Silero plugin during prewarm. The 'english' and 'multilingual' turn detection options still select their existing text models. LiveKit's newer inference.VAD and inference.TurnDetector APIs are available through the existing instance and resolver options; switching to them is a separate choice that can change endpointing, model downloads, and cloud usage.

Review custom LiveKit tools
Direct link to Review custom LiveKit tools

Mastra tools continue to run through the Mastra agent. If you also implement LiveKit tools or a custom LiveKit plugin, account for the changes since Agents 1.4:

  • ToolContext is a class. Use its accessors and flatten() instead of enumerating or mutating it as a plain object.
  • Named tool lists and duplicate-name checks are supported. Object shorthand with anonymous tools remains supported; converting every tool map isn't required.
  • Provider tools use subclasses of ProviderTool instead of the older ProviderDefinedTool API.

See the LiveKit tool API for custom plugin migration details.

Review shutdown
Direct link to Review shutdown

The onCallEnd hook still uses LiveKit's job shutdown callbacks. Keep cleanup within the configured shutdown window, and verify that it completes when a caller disconnects or a worker receives a shutdown signal.

Verify recordings and calls
Direct link to Verify recordings and calls

The integration's room-level recording configuration and S3 permissions stay the same. LiveKit Agents' AgentSession.start({ record }) controls separate session recording and telemetry. Disabling one doesn't disable the other.

Before promoting the prerelease, test both a browser call and a phone call through a configured SIP trunk. Verify both speakers, interruptions, tools, greeting and consent behavior, hangup, the completed trace, and onCallEnd. When recording is enabled, wait for the S3 object and play it with Review Audio. Also test recording disabled and rejection of an existing room name. Run these checks on the deployment image with its model files and native libraries installed.

API reference
Direct link to API reference

The @mastra/livekit package connects Mastra agents to the LiveKit Agents framework. LiveKit runs the audio pipeline (voice activity detection, speech-to-text, turn detection, text-to-speech, barge-in) and the package bridges reply generation to a Mastra agent's stream() call.

See Realtime voice for setup and concepts.

The package has three entry points:

createLiveKitWorker()
Direct link to createlivekitworker

Builds a LiveKit agent definition that answers voice sessions with Mastra agents. Use it as the default export of your worker entry file.

src/mastra/voice-worker.ts
import { fileURLToPath } from 'node:url'
import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker'
import { mastra } from './index'

export default createLiveKitWorker({
mastra,
agent: 'support',
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
turnDetection: 'multilingual',
})

if (process.argv[1] === fileURLToPath(import.meta.url)) {
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' })
}

Options
Direct link to Options

mastra:

Mastra
The Mastra instance whose agents handle voice sessions.

agent?:

string | (args) => string | Agent | Promise<string | Agent>
Which Mastra agent answers each session: a fixed agent key or id, or a resolver called per session with the dispatch metadata and job context. Defaults to the agentId from the dispatch metadata.

workflow?:

string | Workflow | (args) => string | Promise<string>
Generate each turn's reply with a Mastra workflow instead of an agent: a Workflow instance, a fixed workflow key or id, or a resolver that returns a workflow id per session. The workflow runs once to completion per turn (no suspend or resume). Mutually exclusive with agent; requires workflowInput.

workflowInput?:

(args: VoiceTurnContext & { metadata }) => unknown | Promise<unknown>
Maps a turn into the workflow inputData. Required when workflow is set. A stateless mapping that passes the full transcript each turn avoids carrying conversation state in the workflow.

replyStep?:

string
Only stream text from this workflow step id. Defaults to every step that writes to its writer.

resultText?:

(result: unknown) => string | undefined
Fallback when the workflow streams no text via writer: derive the spoken reply from the final run result.

generate?:

VoiceReplyGenerator
Lowest-level escape hatch: supply any reply generator directly (a custom workflow, remote bridge, and so on).

stt?:

STT | string
Speech-to-text: a LiveKit plugin instance or an inference model string such as 'deepgram/nova-3'. For per-call selection, set the configuration.stt resolver — it takes precedence, with this option as the fallback.

tts?:

TTS | string
Text-to-speech: a LiveKit plugin instance or an inference model string such as 'cartesia/sonic-3'. For per-call selection, set the configuration.tts resolver — it takes precedence, with this option as the fallback.

vad?:

VAD | 'silero' | false
= 'silero'
Voice activity detection. 'silero' loads the Silero VAD from @livekit/agents-plugin-silero during prewarm. Pass an instance to bring your own, or false to disable.

turnDetection?:

'multilingual' | 'english' | TurnDetectionMode
End-of-turn detection. 'multilingual' and 'english' load LiveKit's semantic turn detector from @livekit/agents-plugin-livekit. Other values such as 'vad', 'stt', or 'manual' pass through. To construct a TurnDetector instance per call, set the configuration.turnDetection resolver — it takes precedence, with this option as the fallback.

turnHandling?:

Partial<TurnHandlingOptions>
Turn handling tuning: endpointing delays, interruption sensitivity, preemptive generation. The worker disables preemptiveGeneration unless set here — each preemptive attempt re-runs the Mastra agent and persists partial user and assistant messages unless memory.options.readOnly is set.

sessionOptions?:

Partial<AgentSessionOptions>
Extra LiveKit AgentSession options merged over what this helper builds.

memory?:

false | ((args) => { thread, resource, options? } | false)
Memory mapping. Defaults to { thread: metadata.threadId ?? room name, resource: metadata.resourceId ?? thread } when the resolved agent has memory configured. Pass false to disable, or a function to customize. options is forwarded to the agent as per-call memory config; { readOnly: true } keeps speculative turns off the thread.

toolFeedback?:

(toolCall) => string | undefined
Called when the Mastra agent starts a tool call mid-reply. Return a short phrase to speak while the tool runs.

onTurnComplete?:

(ctx: VoiceTurnCompleteContext) => void | Promise<void>
Called once per turn after the reply finished streaming to text-to-speech. Runs off the audio path and is not awaited. The context carries the produced reply (text, toolCalls, interrupted, usage) and the resolved memory mapping.

configuration?:

LiveKitWorkerConfiguration
Grouped conversation and compliance configuration: the opening greeting and AI disclosure, consent requirements, agent-initiated hang-up, and per-call STT/TTS/turn detection selection.
LiveKitWorkerConfiguration

greeting?:

GreetingConfiguration
The opening greeting and AI disclosure: text (a fixed string or a per-call resolver for per-tenant greetings), allowInterruptions, awaitPlayout, persist, and periodic re-disclosure via repeatEvery and repeatText.

consentPolicy?:

ConsentConfiguration
The call's consent policy, as named requirements (starting with summaryStorage). Declarative only — the worker blocks nothing by itself. Capture grants at runtime with createConsentTool and enforce them in your own code; the declared policy surfaces on onCallEnd for cross-checking.

endCall?:

EndCallConfiguration
Agent-initiated hang-up: the worker watches each turn for the end-call tool (pair with createEndCallTool), waits for the agent's closing words to play out, then disconnects — running onCallEnd on the way out.

stt?:

(context: VoiceCallContext) => STT | string | undefined
Per-call speech-to-text: a resolver invoked once per call (post-connect) with { metadata, requestContext, roomName, ctx }, returning anything the top-level stt option accepts. Return undefined to fall back to the top-level stt. Cache plugin instances across calls — the resolver runs during call setup.

tts?:

(context: VoiceCallContext) => TTS | string | undefined
Per-call text-to-speech: a resolver invoked once per call (post-connect) with { metadata, requestContext, roomName, ctx }, returning anything the top-level tts option accepts — one voice or language per tenant. Return undefined to fall back to the top-level tts. Cache plugin instances across calls.

turnDetection?:

(context: VoiceCallContext) => TurnDetectionMode | undefined
Per-call end-of-turn detection: a resolver invoked once per call (post-connect, inside the LiveKit job) with { metadata, requestContext, roomName, ctx }, returning anything the top-level turnDetection option accepts. Use it to construct LiveKit TurnDetector instances, which need the job's inference executor. Return undefined to fall back to the top-level turnDetection.

greeting?:

string
Static greeting spoken when the session starts. Deprecated: prefer configuration.greeting.text.

persistGreeting?:

boolean
= true
Save the spoken greeting to the memory thread as an assistant message, making the saved thread a faithful call transcript. Only applies when a greeting is set and memory is enabled. Deprecated: prefer configuration.greeting.persist.

observability?:

boolean
= true
Trace each call when the Mastra instance has observability configured. Opens a voice call span per session: every turn's agent run nests under it, LiveKit's STT, TTS, end-of-utterance, VAD, and LLM latency metrics become child spans, and the span closes with a per-model usage roll-up. Pass false to disable.

inputOptions?:

Partial<RoomInputOptions>
LiveKit room input options passed to session.start().

outputOptions?:

Partial<RoomOutputOptions>
LiveKit room output options passed to session.start().

onSessionStart?:

(args: { session, ctx, agent, metadata }) => void | Promise<void>
Called after the session starts. Attach event listeners or trigger replies here.

runLiveKitWorker()
Direct link to runlivekitworker

Starts the LiveKit worker CLI (dev, start, and connect subcommands) for a worker entry file. Call it from the file that default-exports the worker definition, guarded so it only runs when executed directly (the worker spawns a child process per session that re-imports the same file). Using this helper instead of cli.runApp from @livekit/agents guarantees the worker runtime and the bridge share one copy of the LiveKit SDK.

Options
Direct link to Options

entry:

string | URL
The worker entry module whose default export is the agent definition. Pass import.meta.url.

agentName?:

string
= 'mastra-voice'
LiveKit agent name for explicit dispatch.

serverOptions?:

Partial<ServerOptions>
Extra LiveKit ServerOptions merged over what this helper builds.

pipeAgentReplyToWriter()
Direct link to pipeagentreplytowriter

Streams a Mastra agent's reply into a workflow step's writer on the workflow reply path. It forwards the agent's text deltas, so text-to-speech starts before the full reply is ready, and its tool-call chunks, so toolFeedback fires and onTurnComplete sees the tool list. Piping only stream.textStream silently drops tool calls. Pass the step's abortSignal to agent.stream() so barge-in stops generation promptly.

src/mastra/workflows/phone-conversation.ts
import { pipeAgentReplyToWriter } from '@mastra/livekit'

const generateResponse = createStep({
id: 'generateResponse',
// input and output schemas omitted
execute: async ({ inputData, mastra, writer, abortSignal }) => {
const stream = await mastra.getAgent('support').stream(inputData.turn, { abortSignal })
const reply = await pipeAgentReplyToWriter(stream, writer)
return { reply }
},
})

Returns: Promise<string>, the accumulated reply text.

Parameters
Direct link to Parameters

agentStream:

AgentReplyStreamLike
The stream returned by agent.stream() — anything exposing a fullStream async iterable.

writer:

WritableStream<unknown>
The workflow step's writer.

chatContextToMessages()
Direct link to chatcontexttomessages

Converts a LiveKit chat context into plain messages accepted by agent.stream(), excluding instructions and function calls. Use it in workflowInput to pass the full transcript into a stateless workflow.

src/mastra/voice-worker.ts
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker'

export default createLiveKitWorker({
mastra,
workflow: 'phoneConversation',
workflowInput: ({ chatCtx }) => ({ history: chatContextToMessages(chatCtx) }),
})

Returns: VoiceTurnMessage[], where each entry is { role: 'system' | 'user' | 'assistant'; content: string; id?: string }.

MastraVoiceAgent
Direct link to mastravoiceagent

The LiveKit voice.Agent subclass that createLiveKitWorker() builds for every session. Replies come from a Mastra agent (or a custom generate source) through the agent's llmNode; LiveKit keeps the audio loop, turn detection, and barge-in. Construct it yourself when you own the voice.AgentSession, for example to test a Mastra-backed agent with @livekit/agents' voice.testing harness without speech-to-text, text-to-speech, or a live worker. createMastraVoiceAgent(options) is an equivalent factory.

src/mastra/voice-agent.test.ts
import { initializeLogger, voice } from '@livekit/agents'
import { MastraVoiceAgent } from '@mastra/livekit/worker'
import { supportAgent } from './agents/support'

// Required outside a LiveKit worker: AgentSession needs the LiveKit logger initialized.
initializeLogger({ level: 'silent', pretty: false })

const session = new voice.AgentSession()
await session.start({ agent: new MastraVoiceAgent({ agent: supportAgent, memory: false }) })

// run() returns a RunResult, not a promise; wait() resolves when the turn completes.
const result = session.run({ userInput: 'What are your opening hours?' })
await result.wait()
result.expect.nextEvent().isMessage({ role: 'assistant' })
result.expect.noMoreEvents()

The agent carries its own placeholder llm.LLM so LiveKit runs the reply pipeline; generation always goes through llmNode, so a FakeLLM in the session's llm slot is ignored and calling the placeholder's chat() throws. To stub the model in tests, give the Mastra agent a mock model or pass a custom generate function.

Options
Direct link to Options

Provide exactly one reply source: agent or generate.

agent?:

Agent
In-process Mastra agent. Tools and memory run inside it.

generate?:

VoiceReplyGenerator
Custom reply source, for example from createRemoteAgentReplyGenerator(). A generate source owns its own hooks; toolFeedback, onToolCall, onTurnComplete, and streamOptions only apply to the agent source.

memory?:

MastraVoiceAgentMemory | false
= false
Conversation persistence as { thread, resource?, options? }. When set, only messages new since the agent last spoke are sent each turn and Mastra Memory supplies history. options is forwarded to the agent as per-call memory config, e.g. { readOnly: true }. When false, the full in-session LiveKit context is sent every turn.

requestContext?:

RequestContext | Record<string, unknown>
Request context entries forwarded to every generation.

toolFeedback?:

(toolCall: VoiceToolCall) => string | undefined | void
Return a short phrase to speak while a tool runs. Agent source only.

onToolCall?:

(toolCall: VoiceToolCall) => void
Called as each tool call starts, before its result is known. Keep it cheap and non-throwing. Agent source only.

onTurnComplete?:

VoiceTurnCompleteHook
Called once per turn after the reply finished streaming to text-to-speech. Fire-and-forget; errors are logged. Agent source only.

greetingReminder?:

{ everyMs: number; text?: string }
Periodic AI re-disclosure: once everyMs has elapsed, the next reply is prefixed with text (spoken at the turn boundary). The worker derives this from configuration.greeting.repeatEvery / repeatText.

streamOptions?:

MastraStreamOptions
Extra options merged into every agent.stream() call. Agent source only.

instructions?:

string
LiveKit agent instructions. Not used for reply generation; the Mastra agent applies its own.

id / stt / vad / tts / turnHandling?:

voice.AgentOptions['id' | 'stt' | 'vad' | 'tts' | 'turnHandling']
Passed through to the LiveKit voice.Agent constructor. Use them to set per-agent speech components or turn handling.

MastraLLM
Direct link to mastrallm

A standard LiveKit LLM plugin (llm.LLM) backed by a Mastra agent. Use it when you build the voice.AgentSession yourself and want Mastra in the llm slot. createLiveKitWorker() is the managed alternative. See Use Mastra as the LLM component for how to choose.

With remote, the plugin streams each turn from your Mastra server over HTTP using Server-Sent Events (SSE). The agent loop, tools, and memory run server-side, and interrupting the agent aborts the server-side generation.

src/mastra/voice-worker-plugin.ts
import { voice } from '@livekit/agents'
import { MastraLLM } from '@mastra/livekit/plugin'

const session = new voice.AgentSession({
llm: new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
memory: { thread: callId, resource: userId },
}),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
// Required with `memory` unless memory.options.readOnly is set: LiveKit enables
// preemptive generation by default.
turnHandling: { preemptiveGeneration: { enabled: false } },
})

The plugin reports provider as mastra and model as the agent id, so LiveKit metrics and fallback adapters identify it like any other LLM.

Constructor options
Direct link to Constructor options

Provide exactly one reply source: remote, agent, or generate.

remote?:

RemoteMastraAgentOptions
Remote Mastra server reached over HTTP. Takes the same connection options as createRemoteAgentReplyGenerator(): baseUrl, agentId, apiPrefix, headers, fetch, timeoutMs, retries, body.

agent?:

Agent
In-process Mastra agent. Session ownership without a second deployment.

generate?:

VoiceReplyGenerator
Custom reply source. A generate source owns its own hooks; toolFeedback, onToolCall, and onTurnComplete below only apply to the remote and agent sources.

memory?:

{ thread: string; resource?: string; options?: MemoryConfig } | false
= false
Conversation persistence, resolved per call (for example from the SIP caller identity). When set, only messages new since the agent last spoke are sent each turn and Mastra Memory supplies history. options is forwarded in the request body as memory.options, e.g. { readOnly: true }. When omitted, the full LiveKit chat context is sent every turn.

requestContext?:

RequestContext | Record<string, unknown>
Request context forwarded to generation (tenant, dialed number, and so on).

toolFeedback?:

(toolCall: VoiceToolCall) => string | undefined
Return a short phrase to speak while a server-side tool runs.

onToolCall?:

(toolCall: VoiceToolCall) => void
Called as each tool call starts, mid-stream. Pair with runEndCall() to implement your own agent-initiated hang-up flow.

onTurnComplete?:

(ctx: VoiceTurnCompleteContext) => void | Promise<void>
Called once per turn after the reply finished streaming, off the audio path and not awaited. The context carries the produced reply: text, toolCalls, interrupted, and usage.
warning

Don't combine memory with the session's preemptiveGeneration option, which LiveKit enables by default in sessions you build yourself. A speculative turn persists a partial user message and a partial, never-spoken reply to the thread before LiveKit discards it. Set turnHandling: { preemptiveGeneration: { enabled: false } } on the session, or set memory.options.readOnly and persist committed turns yourself; see preemptive generation with memory. Stateless mode (no memory) works with preemptive generation.

Tools run on the Mastra agent
Direct link to Tools run on the Mastra agent

Tools are defined and executed server-side on the Mastra agent. The plugin never forwards LiveKit tool definitions: if the session passes a non-empty toolCtx, it logs a one-time warning naming the ignored tools. Every tool must complete server-side: a tool that requires approval or client-side execution fails the turn with a descriptive error instead of hanging the call.

Tool activity reaches the worker through toolFeedback, onToolCall, and onTurnComplete.

Instructions
Direct link to Instructions

LiveKit injects your voice.Agent's instructions into the chat context of every request. The plugin drops them because the server-side Mastra agent's own instructions are authoritative. To change the prompt, change the Mastra agent.

Interrupted turns
Direct link to Interrupted turns

When the user interrupts a reply:

  1. The plugin cancels the stream. The server aborts generation and persists nothing from that turn.
  2. LiveKit records the part the user actually heard in its chat context, flagged as interrupted.
  3. On the next turn, the plugin re-sends that heard-only fragment, ordered before the new user message, so the memory thread backfills to match the call. Messages carry LiveKit's message ids and the server deduplicates by id, so retries and re-sends stay idempotent.

A user who hangs up immediately after interrupting leaves that final fragment unrecorded. When the transcript must capture it, reconcile immediately from the session event; the shared message id means the next turn's re-send upserts instead of duplicating:

src/mastra/voice-worker-plugin.ts
import { voice } from '@livekit/agents'
import { MastraClient } from '@mastra/client-js'

const client = new MastraClient({ baseUrl: process.env.MASTRA_URL! })

session.on(voice.AgentSessionEventTypes.ConversationItemAdded, ({ item }) => {
if (item.type !== 'message' || item.role !== 'assistant' || !item.interrupted) return
void client.saveMessageToMemory({
agentId: 'support',
messages: [
{
id: item.id,
threadId: callId,
resourceId: userId,
role: 'assistant',
content: item.textContent ?? '',
type: 'text',
createdAt: new Date(),
},
],
})
})

Usage metrics
Direct link to Usage metrics

When the server reports token usage for a turn, the plugin feeds it to LiveKit, so the session's metrics_collected events carry time-to-first-token, duration, and token counts like any LLM plugin. The same usage object (promptTokens, completionTokens, promptCachedTokens, totalTokens) arrives on onTurnComplete as result.usage.

Errors and timeouts
Direct link to Errors and timeouts

The transport throws LiveKit's APIError types (APIStatusError, APIConnectionError, APITimeoutError), so the session's retry policy (connOptions.maxRetry) and FallbackAdapter failover work unchanged. A turn is never retried after its first token: a voice reply is better failed fast than replayed half-heard.

A connect and first-token watchdog uses the session's connOptions.timeoutMs (10 seconds by default), so a server that accepts the connection but never streams can't cause indefinite dead air.

If the Mastra server goes down mid-call, each reply attempt fails with a typed error after its retries, and LiveKit closes the session after several consecutive failed replies. Restore the server before that budget runs out and the call recovers on the next turn.

Message content
Direct link to Message content

Message extraction is text-only: image content is dropped, and audio content is included only through its transcript. Voice pipelines aren't affected, but items you inject into the chat context yourself must carry text.

createRemoteAgentReplyGenerator()
Direct link to createremoteagentreplygenerator

Builds a reply generator that runs the agent loop on a remote Mastra server over HTTP/SSE. MastraLLM's remote mode uses it internally. Use it directly through createLiveKitWorker's generate option to run the batteries-included worker against a remote server:

src/mastra/voice-worker.ts
import { createLiveKitWorker, createRemoteAgentReplyGenerator } from '@mastra/livekit/worker'
import { mastra } from './index'

export default createLiveKitWorker({
mastra, // local instance for logger and worker config; replies come from the remote server
generate: createRemoteAgentReplyGenerator({
baseUrl: process.env.MASTRA_URL!,
agentId: 'support',
}),
memory: ({ metadata, roomName }) => ({ thread: metadata.threadId ?? roomName }),
stt: 'deepgram/nova-3',
tts: 'cartesia/sonic-3',
})

On the generate path the worker-level toolFeedback and onTurnComplete options don't apply, and the worker's end-call detection doesn't fire; pass the hooks to the generator instead.

Cancelling a turn (barge-in) tears down the HTTP request, which aborts generation on the server. Errors are thrown as LiveKit APIError types. The retries option applies only to initial connection attempts. A turn is never retried after its first chunk.

Returns: VoiceReplyGenerator.

Options
Direct link to Options

baseUrl:

string
Base URL of the remote Mastra server, for example https://my-app.example.com.

agentId:

string
The agent's registered key or id on the remote Mastra instance.

apiPrefix?:

string
= '/api'
Path prefix for the Mastra API.

headers?:

Record<string, string> | () => Record<string, string> | Promise<Record<string, string>>
Static headers, or a resolver invoked per turn — for example to mint a fresh authorization token.

fetch?:

typeof fetch
= globalThis.fetch
Injectable fetch implementation for tests or proxies.

timeoutMs?:

number
= 10000
Connect and first-token timeout in milliseconds. When used through MastraLLM, defaults to the session's connOptions.timeoutMs instead.

retries?:

number
= 2
Initial-connection retry attempts, before the first chunk only. When used through MastraLLM, the LiveKit session owns retries and this is forced to 0.

body?:

Record<string, unknown>
Extra fields merged into each stream request body.

toolFeedback?:

(toolCall: VoiceToolCall) => string | undefined
Return a short phrase to speak while a server-side tool runs.

onToolCall?:

(toolCall: VoiceToolCall) => void
Called as each tool call starts, mid-stream.

onTurnComplete?:

(ctx: VoiceTurnCompleteContext) => void | Promise<void>
Called once per turn after the reply finished streaming, off the audio path.

speakGreeting()
Direct link to speakgreeting

Speaks an opening greeting on a session you own, honoring interruption and playout options. Returns the LiveKit SpeechHandle, or undefined when there's no greeting text. createLiveKitWorker() uses it internally for its greeting configuration.

import { speakGreeting } from '@mastra/livekit/worker'

await speakGreeting(session, {
text: "You've reached support. You're speaking with an AI assistant.",
allowInterruptions: false,
awaitPlayout: true,
})

Parameters
Direct link to Parameters

session:

voice.AgentSession
The session to speak on.

greeting:

{ text?: string; allowInterruptions?: boolean; awaitPlayout?: boolean }
The greeting text and playout options. When awaitPlayout is true, the returned promise resolves after the greeting finished playing (or was interrupted).

waitForAgentDoneSpeaking()
Direct link to waitforagentdonespeaking

Resolves once the agent is no longer producing or playing a reply: its state has left thinking and speaking. Resolves immediately when the agent is already idle, and always resolves within maxWaitMs (30 seconds by default) as a safety cap. Use it before tearing a session down so closing words play out instead of being cut off.

import { waitForAgentDoneSpeaking } from '@mastra/livekit/worker'

await waitForAgentDoneSpeaking(session)

runEndCall()
Direct link to runendcall

Ends the call after the agent asks to hang up. It waits for the agent's closing words and speaks an optional final message without interruption. It then deletes the room and hangs up the caller, including SIP callers. The job shuts down with its registered callbacks.

Pair it with MastraLLM's onToolCall and an end-call tool on the server-side agent to rebuild agent-initiated hang-up on a session you own:

src/mastra/voice-worker-plugin.ts
import { MastraLLM } from '@mastra/livekit/plugin'
import { DEFAULT_END_CALL_TOOL, runEndCall } from '@mastra/livekit/worker'

let ending = false

const llm = new MastraLLM({
remote: { baseUrl: process.env.MASTRA_URL!, agentId: 'support' },
onToolCall: ({ toolName }) => {
if (toolName !== DEFAULT_END_CALL_TOOL || ending) return
ending = true
void runEndCall(session, ctx, {}, console)
},
})

The exported constants DEFAULT_END_CALL_TOOL ('endCall'), DEFAULT_END_CALL_REASON, and DEFAULT_END_CALL_MAX_WAIT_MS (30000) hold the defaults.

Parameters
Direct link to Parameters

session:

voice.AgentSession
The session whose agent is finishing its closing words.

ctx:

JobContext
The LiveKit job context used to delete the room and shut down.

config:

{ message?: string; reason?: string; maxWaitMs?: number; drainMs?: number }
Optional final message spoken before hang-up, the shutdown reason to record, the safety cap on waiting for closing words, and the post-playout drain (default 800ms) that lets audio buffered at the caller finish playing before the room is deleted — LiveKit's playout accounting is worker-local, so hanging up the instant it clears clips the goodbye.

logger:

{ warn: (message: string, ...args: unknown[]) => void }
Receives warnings when teardown steps fail. Pass your logger or console.

createEndCallTool()
Direct link to createendcalltool

Builds the Mastra tool an agent calls when it wants to end the call. The tool signals intent and can run optional bookkeeping. The worker performs the actual hang-up. The tool lives on the server-safe root entry. Add it to agents defined in server code.

src/mastra/agents/support-agent.ts
import { Agent } from '@mastra/core/agent'
import { createEndCallTool } from '@mastra/livekit'

const supportAgent = new Agent({
id: 'support',
name: 'Support',
instructions:
'Help the caller. When everything is wrapped up, say goodbye and call endCall as your final action.',
model: 'openai/gpt-5-mini',
tools: { endCall: createEndCallTool() },
})

With createLiveKitWorker(), set configuration: { endCall: {} } and the worker watches for the tool and hangs up. On a session you own, rebuild the hang-up with runEndCall().

Options
Direct link to Options

id?:

string
= 'endCall'
Tool id the agent calls to end the call. Must match the name the worker watches for (the worker's configuration.endCall.tool, or your own onToolCall check).

description?:

string
Override the description the model sees when deciding to call the tool.

onEndCall?:

(request: { reason?: string; resourceId?: string; threadId?: string }) => void | Promise<void>
Bookkeeping hook called when the agent invokes the tool — record the reason or mark the call resolved. Runs inside the turn; keep it quick. It does not hang up the call.

liveKitConnectionRoute()
Direct link to livekitconnectionroute

Returns an API route that mints a LiveKit access token with the voice agent dispatched into the room. Frontends call it to join a session.

src/mastra/index.ts
import { Mastra } from '@mastra/core/mastra'
import { liveKitConnectionRoute } from '@mastra/livekit'

export const mastra = new Mastra({
server: {
apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],
},
})

The route accepts a JSON body with optional agentId, threadId, and resourceId fields and responds with { serverUrl, roomName, participantName, participantToken }. The threadId defaults to the generated room name.

Options
Direct link to Options

path?:

string
= '/voice/livekit/connection-details'
Route path.

serverUrl?:

string
= process.env.LIVEKIT_URL
LiveKit server URL.

apiKey?:

string
= process.env.LIVEKIT_API_KEY
LiveKit API key.

apiSecret?:

string
= process.env.LIVEKIT_API_SECRET
LiveKit API secret.

agentName?:

string
= 'mastra-voice'
LiveKit agent name for explicit dispatch. Must match the worker's agentName.

ttl?:

string | number
= '15m'
Token time-to-live.

requiresAuth?:

boolean
= true
Whether the route requires authentication.

roomName?:

string | (args) => string
Room name or a function that derives one from the request. With recording enabled, use a unique name for each call. An existing room returns HTTP 409 before agent dispatch or token issuance.

participantIdentity?:

string | (args) => string
Participant identity or a function that derives one from the request.

metadata?:

(args) => LiveKitSessionMetadata | Promise<LiveKitSessionMetadata>
Builds the session metadata delivered to the worker. Defaults to passing through agentId, threadId, and resourceId from the request body.

recording?:

LiveKitRecordingOptions | (args) => LiveKitRecordingOptions | undefined | Promise<LiveKitRecordingOptions | undefined>
Creates a fresh room with automatic recording and dispatches the agent before returning connection details. The server-side callback receives body, context, and the resolved roomName. Return undefined to skip recording. Existing rooms return HTTP 409. Other setup failures reject the request.

dispatchVoiceSession()
Direct link to dispatchvoicesession

Dispatches a Mastra voice agent into a LiveKit room programmatically: for server-initiated sessions such as outbound calls.

import { dispatchVoiceSession } from '@mastra/livekit'

await dispatchVoiceSession({
roomName: 'support-call-42',
agentName: 'mastra-voice',
metadata: { agentId: 'support', threadId: 'thread-42' },
})

Options
Direct link to Options

roomName:

string
Room to dispatch the agent into. Created on demand. With recording enabled, use a unique name for each call. An existing room throws LiveKitRecordingRoomConflictError.

agentName?:

string
= 'mastra-voice'
Must match the worker's agentName.

metadata?:

LiveKitSessionMetadata
Session metadata: agentId, threadId, resourceId, requestContext.

recording?:

LiveKitRecordingOptions
Creates a fresh room with automatic recording before dispatching. Accepts LiveKit RoomEgress settings as a plain object or SDK instance. Existing rooms and setup failures reject the operation. Omit to preserve dispatch-only behavior.

serverUrl?:

string
= process.env.LIVEKIT_URL
LiveKit server URL.

apiKey?:

string
= process.env.LIVEKIT_API_KEY
LiveKit API key.

apiSecret?:

string
= process.env.LIVEKIT_API_SECRET
LiveKit API secret.

LiveKitSessionMetadata
Direct link to livekitsessionmetadata

The metadata passed from the Mastra server to the worker through LiveKit job dispatch.

agentId?:

string
Mastra agent to run, by registered key or agent id.

threadId?:

string
Memory thread id. Defaults to the LiveKit room name.

resourceId?:

string
Memory resource id, typically the end user id.

requestContext?:

Record<string, unknown>
Plain-object entries restored into a RequestContext for agent execution.

The metadata travels as a JSON string. liveKitConnectionRoute() and dispatchVoiceSession() serialize it for you; use serializeSessionMetadata(metadata) when dispatching through your own code, or write the JSON directly in LiveKit-side configuration such as a SIP dispatch rule. Entries in requestContext reach the agent's runtime-defined instructions, tools, and input processors on every turn of the call.