LLM
session.llm({
vendor: 'openai',
model: 'gpt-realtime',
auth: { apiKey },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
events: [
'conversation.item.*',
'response.output_audio_transcript.done',
'input_audio_buffer.committed'
],
llmOptions: {
response_create: {
output_modalities: ['audio'],
instructions: 'Greet the caller warmly in English and ask how you can help today.',
audio: {
output: {
voice: 'alloy',
format: { type: 'audio/pcm', rate: 24000 }
}
},
max_output_tokens: 4096
},
session_update: {
type: 'realtime',
instructions:
'You are a friendly, helpful voice assistant on a phone call. ' +
'Always respond in English unless the caller explicitly speaks another language. ' +
'Keep responses concise and natural for spoken conversation. ' +
'If asked about weather, call the get_weather function.',
tools: [
{
name: 'get_weather',
type: 'function',
description: 'Get the weather at a given location',
parameters: {
type: 'object',
properties: {
location: {
type: 'string',
description: 'Location to get the weather from',
},
scale: {
type: 'string',
enum: ['fahrenheit', 'celsius'],
},
},
required: ['location', 'scale'],
},
},
],
tool_choice: 'auto',
audio: {
input: {
format: { type: 'audio/pcm', rate: 24000 },
transcription: { model: 'whisper-1' },
turn_detection: {
type: 'server_vad',
threshold: 0.8,
prefix_padding_ms: 300,
silence_duration_ms: 500
}
},
output: {
format: { type: 'audio/pcm', rate: 24000 },
voice: 'alloy'
}
}
}
}
});
Parameters
modelstringName of the LLM model.
vendorstringName of the LLM vendor.
actionHookstringWebhook that will be called when the LLM session ends.
authobjectObject containing authentication credentials; format according to the model.
connectOptionsobjectObject containing information such as the URI to connect to.
eventHookstringWebhook that will be called when a requested LLM event happens (e.g., transcript).
eventsarrayArray of event names listing the events requested (wildcards allowed).
handoffobjectDeclarative transfer-to-human configuration. When present, the runtime injects a
transfer_to_human tool and runs the packaged transfer choreography
when the model calls it. See Transfer-to-human handoff.
hangupobjectEnable the built-in hangup tool. When present, the runtime injects a hangup tool into the model's toolset; when the model calls it, the call is ended. Works across all supported s2s vendors. See Built-in Hangup Tool below.
hangup.reasonstringDefault reason placed in the X-Reason SIP header on the outbound BYE. Used as a fallback when the model does not supply its own reason at call time.
llmOptionsobjectObject containing instructions for the LLM; format dependent on the LLM model.
toolHookstringWebhook that will be called when the LLM wants to call a function.
The following LLMs are currently supported:
- OpenAI Realtime API
- OpenAI GPT Live API (limited-access alpha)
- Deepgram Voice Agent
- Ultravox
- ElevenLabs
- Google Gemini Live API (via Gemini Developer API or Vertex AI)
- AssemblyAI Voice Agent
- xAI Voice Agent
- Azure Voice Live
OpenAI Realtime
Set vendor: 'openai' and supply an OpenAI API key via auth.apiKey. llmOptions carries two payloads — response_create and session_update — that are forwarded to OpenAI as the corresponding response.create and session.update client events. The example at the top of this page uses the GA format.
Model selection
modelstringGA realtime model name. Common choices:
gpt-realtime— the standard GA realtime conversation model. Recommended starting point.gpt-realtime-2— reasoning-capable variant. Emits aphase: "commentary"message beforephase: "final_answer"in its output items; both phases produce audio, so the caller hears the model "think out loud" before the actual answer. Useful for some assistants, surprising for typical greeting-first flows.gpt-realtime-whisper— streaming transcription model. Emitsresponse.output_audio_transcript.deltaevents but does not endpoint turns server-side (no.completedin current GA); not suitable for normal speech-to-speech or gather flows.
OpenAI deprecated the Realtime preview models (gpt-4o-realtime-preview-*) on 2026-05-12; use a GA model.
Instructions
session_update.instructions sets the persona that persists for the whole session; response_create.instructions overrides it for a single response only. GA realtime models do not assume English by default — if Wipple CPaaS speaks first (which it does when response.create fires before any caller audio), anchor the language in instructions (e.g. "Always respond in English unless the caller speaks another language") or you may get a greeting in an arbitrary language.
Audio format
The audio format on the wire to OpenAI is fixed at {type: 'audio/pcm', rate: 24000}. Wipple CPaaS resamples to and from the channel's native rate (typically G.711) automatically. Any audio.input.format or audio.output.format set in the application is overridden — they are present in the example only for completeness.
Legacy preview-format compatibility
Apps written against the preview API still work without changes. Wipple CPaaS detects the legacy flat format — top-level modalities, voice, input_audio_format, output_audio_format, input_audio_transcription, turn_detection, temperature — inside session_update or response_create, converts it to the GA format on the wire, and aliases GA event names back to their preview names before forwarding to your eventHook. For example, an app that subscribed to response.audio_transcript.done continues to receive events with that exact type even though the server now emits response.output_audio_transcript.done. A one-time deprecation warning is logged when the legacy format is detected.
GA-invalid fields are stripped silently with a WARN log:
temperatureinresponse_create— no longer accepted in GA.output_audio_formatas a flat string — replaced byaudio.output.format: {type, rate}(which Wipple CPaaS overrides topcm24 kHz anyway).
The compatibility layer is intended as a transitional shim — plan to update applications to the GA format before it is removed.
OpenAI-specific llmOptions fields
response_createobjectThe initial response.create.response payload sent to OpenAI when the session opens. Drives the first response from the assistant.
response_create.output_modalitiesarrayArray of output modalities. Typically ['audio'] for voice. (GA renamed this from preview's modalities.)
response_create.instructionsstringPer-response instruction override. Replaces session-level instructions for this one response only — useful for the opening greeting. Anchor the language explicitly when the model speaks first.
response_create.audioobjectAudio configuration object. Use audio.output.voice to choose a voice (e.g. alloy, marin) and audio.output.format (overridden to pcm 24 kHz by Wipple CPaaS).
response_create.max_output_tokensnumberMaximum tokens for this response.
session_updateobjectThe initial session.update.session payload. Sets persona, tools, audio config, and VAD for the whole session.
session_update.typestringMust be 'realtime' in GA. Required when present.
session_update.instructionsstringSession-wide persona prompt. Persists for every response in the session unless overridden by response_create.instructions.
session_update.audio.input.transcription.modelstringAuxiliary transcription model for converting caller audio into text in the conversation log. Common values: whisper-1, gpt-4o-transcribe, gpt-realtime-whisper. This is separate from the conversation model.
session_update.audio.input.turn_detectionobjectServer-side VAD configuration: {type: 'server_vad', threshold, prefix_padding_ms, silence_duration_ms}. Defaults are server-supplied if omitted. type: 'semantic_vad' with an eagerness field is also supported on models that accept it.
session_update.toolsarrayArray of function-tool definitions exposed to the model. Each entry needs name, type: 'function', description, and JSON Schema parameters. The model calls a tool by emitting response.output_item.done with item.type: 'function_call'; Wipple CPaaS routes that to your toolHook.
session_update.tool_choicestring'auto' (default), 'none', or a specific tool name to force.
OpenAI GPT Live
Set vendor: 'gptlive' and supply an OpenAI API key in auth.apiKey to talk to OpenAI's GPT Live API.
Alpha API
GPT Live requires enrollment in OpenAI's Early Access Program. An ordinary OpenAI key completes the connection and is then refused by OpenAI at the first server event with
Voice session access denied, which Wipple CPaaS reports ascompletion_reason: 'server error'.Expect event names and fields to change while the API is in alpha.
Not a drop-in for
vendor: 'openai'GPT Live is a different protocol from the OpenAI Realtime API, not a newer model for it. If you are moving an app over from
vendor: 'openai', read Migrating from the Realtime API first — the payloads are not interchangeable.
A minimum configuration — note that on its own this produces a silent call, because nothing has asked the model to speak first (see Making the agent speak first):
session.llm({
vendor: 'gptlive',
model: 'gpt-live-1-boulder-alpha',
auth: { apiKey: process.env.OPENAI_API_KEY },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
llmOptions: {
session_update: {
instructions:
'You are a friendly, helpful voice assistant on a phone call. ' +
'Always respond in English unless the caller explicitly speaks another language. ' +
'Keep responses concise and natural for spoken conversation.',
audio: {
output: { voice: 'marin' }
}
}
}
});
Model selection
modelstringGPT Live model name; defaults to gpt-live-1-boulder-alpha. Set it on the Function. It must not appear inside session_update at all — Wipple CPaaS rejects the Function if it does, even if you set it in only that one place.
Audio format
There is no audio format to configure. GPT Live fixes it at 24 kHz mono PCM, and Wipple CPaaS converts to and from the caller's codec for you.
session_update is required
llmOptions.session_update is not optional here: GPT Live will not accept the caller's audio until it has your configuration. Wipple CPaaS sends it first and waits for session.started — the event you should hang your greeting off, described next.
There is no response_create. Once the session starts, the model drives the conversation itself.
Making the agent speak first
Because there is no response_create, nothing tells the model to take the first turn — and putting the greeting in instructions is not enough. The model waits for the caller, who hears silence.
To open the call, send a session.context.append as soon as you receive session.started. Give it the wording you want and tell it when to speak:
session.on('/event', (evt) => {
if (evt.type === 'session.started') {
session.updateLlm({
type: 'session.context.append',
content: [{
type: 'input_text',
text: 'Immediately greet the caller using the exact text below. Do not wait for the '
+ 'caller to speak first. After the greeting, pause and listen.\n\n'
+ 'Hi, I am your support assistant. How can I help you today?',
}],
});
}
});
Use the same pattern any time you need the agent to say something specific mid-call — a disclosure, a transfer notice, a closing.
Greetings are requested, not guaranteed
A context append guides the model. Per OpenAI, it may paraphrase your wording, or occasionally stay silent. Omit the wording and the model will use its own.
If the exact words matter — a legal disclosure, a brand greeting — do not ask the model. Play it yourself with a
sayFunction before thellmFunction.Note
Requesting a greeting needs WebSocket transport. The
llm:updatecommand is not accepted by the RESTupdateCallAPI, so a webhook-only application has no way to send one.
Delegation: how the model asks for work
Delegation is GPT Live's distinguishing concept, and it has no direct analogue in the other vendors on this page. Anything the model needs from outside the spoken conversation — a fact it does not know, a function it wants run — arrives as a delegation. Your choice of kind determines whether you can use tools at all.
session_update.delegationobjectOmit it entirely and the model simply never asks for anything — no tools, no context requests. That is the simplest way to start.
session_update.delegation.typestring'client' — the model asks your application for background information in prose. Simple to implement, but there is no function calling.
'responses' — the model runs a turn on OpenAI's Responses API that can call your functions. Required if you want tools, mcpServers, handoff or hangup.
Responses delegation (function calling)
Declare your tools inside the delegation. Note that responses.model is a second, separate model that runs the delegated turn — it is not the voice model on the Function:
delegation: {
type: 'responses',
responses: {
model: 'gpt-5.5', // required
tools: [
{
type: 'function',
name: 'get_weather',
description: 'Get the current weather for a city.',
parameters: {
type: 'object',
properties: { location: { type: 'string' } },
required: ['location']
}
}
]
}
}
Warning
Tools go in
delegation.responses.tools— notdelegation.tools— and both the nestedresponsesobject and itsmodelare required.Wipple CPaaS rejects the Function before the call connects if you configure
handoff,hangupormcpServerswithout them. It cannot check your own tool declarations, so if you omitresponses.modelthere, OpenAI refuses the session at startup and you getcompletion_reason: 'server error'.
Client delegation (supplying context)
The model emits delegation.created with item.target: 'client'; you answer with the item's id and up to 500 tokens of prose in a single input_text part:
session.on('/event', (evt) => {
if (evt.type === 'delegation.created' && evt.item.target === 'client') {
session.updateLlm({
type: 'delegation.context.append',
delegation_item_id: evt.item.id,
content: [{ type: 'input_text', text: 'It is 62 degrees and raining in Seattle.' }]
});
}
});
Always answer a client delegation. There is no way to cancel one, so if you leave it unanswered the model waits and the caller hears silence.
Handling tool calls
With a responses delegation, Wipple CPaaS calls your toolHook with {name, args, tool_call_id} just as it does for every other vendor. What differs is the envelope you send the result back in:
session.on('/toolCall', async(evt) => {
const { tool_call_id, name, args } = evt;
session.sendToolOutput(tool_call_id, {
type: 'delegation.function_call_output.create',
item: {
type: 'function_call_output',
call_id: tool_call_id,
output: JSON.stringify({ temperature: 62, conditions: 'rain' })
}
});
});
output must be a string. Send one of these per call if the model asked for several, and do not follow it with a response.create — unlike the Realtime API, the server resumes on its own.
The built-in handoff and hangup tools, and any mcpServers you configure, all work here too — but only with a responses delegation, since a client delegation has no way to call a function.
Following the conversation
GPT Live is in alpha and has no public event reference, so this is the list. Name the ones you want in the Function's events array — wildcards such as response.* work, as do 'all' plus -eventName exclusions.
| Event | What it gives you |
|---|---|
session.started |
The session is live. Send your greeting request here |
session.updated, session.context.appended |
Acknowledgements of your client events |
turn.created, turn.delta, turn.done |
Utterance-level view of the conversation. turn.done carries turn.role ('user' or 'assistant') and turn.transcript |
input_transcript.added, output_transcript.added |
Live partial transcript fragments, caller and agent |
delegation.created, delegation.context.appended, delegation.function_call_output.created |
Delegation lifecycle |
response.* |
The delegated Responses turn: response.created, response.output_text.delta, response.completed, response.failed, response.done, and others |
output_audio.playback_started, output_audio.playback_stopped |
Emitted by Wipple CPaaS, not OpenAI: when agent audio actually starts and stops reaching the caller |
session.usage.updated |
Cumulative token usage; final totals arrive on session.closed |
error |
A rejected client event or a server-side problem |
Warning
If you omit
eventsentirely, Wipple CPaaS forwards everything — including transcript fragments, which are high volume. Name the events you need.
Prefer turn.done over the *_transcript.added fragments for following the conversation: fragment boundaries follow speech cadence rather than complete thoughts, so one sentence may arrive in several pieces.
Barge-in needs no work on your part — Wipple CPaaS detects it and flushes the queued agent audio; the underlying signal is not forwarded to your eventHook.
Updating the session mid-call
An llm:update command (session.updateLlm() in the Node SDK) accepts these five events:
session.update— change configuration. Sparse: fields you omit keep their values. Replacingdelegationrequires the whole object.session.context.append— add context or request a spoken messagedelegation.context.append— answer a client delegationdelegation.function_call_output.create— return a tool resultsession.close— end the session gracefully
Anything else is discarded silently — you will not get an error back, so check the type if an update appears to do nothing.
How the session ends
The actionHook fires with a completion_reason:
completion_reason |
Meaning |
|---|---|
normal conversation end |
The session closed cleanly |
session closed: <reason> |
OpenAI closed the session for its own reason, which is included |
disconnect from remote end |
OpenAI dropped the connection |
server error |
GPT Live refused the session at startup. The payload also carries an error object with OpenAI's reason — check it first |
connection failure |
Wipple CPaaS could not open the connection at all: network, DNS or proxy |
hangup |
The model ended the call using the built-in hangup tool |
A handoff that bridges reports the transfer outcome instead. Unlike some other vendors, GPT Live never reports server failure.
Not every problem ends the call. Once the session is running, a rejected client event — a stale delegation_item_id, an over-long context append — or a failed delegation is reported to your eventHook and the conversation continues. Handle those events if you want to recover; ignore them and the model continues without the context or tool result you were sending.
Note
cancelOnBargeInandcancelOnResponseTimeouthave no effect with this vendor, because GPT Live has nothing to cancel.responseTimeoutMsdoes work, but it is disabled by default — set it to a non-zero value and you get aresponse.timeoutevent if a delegation stalls, which you can recover from with anllm:update.
Migrating from the Realtime API
If you are porting an app from vendor: 'openai':
Realtime (vendor: 'openai') |
GPT Live (vendor: 'gptlive') |
|---|---|
llmOptions.response_create |
Not supported — remove it. Greet with session.context.append instead |
session_update.tools |
session_update.delegation.responses.tools |
session_update.turn_detection |
Not supported — turn detection is handled for you |
session_update.audio input/output formats |
Not supported — audio is fixed at 24 kHz mono PCM |
Tool result via conversation.item.create |
delegation.function_call_output.create |
A follow-on response.create after a tool result |
Nothing — the server resumes by itself |
response.create / response.cancel mid-call |
Not supported |
input_audio_buffer.speech_started for barge-in |
Wipple CPaaS detects barge-in itself; the signal is not forwarded to your eventHook |
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| The agent never speaks and the caller hears silence | No session.context.append on session.started — see Making the agent speak first. Note that event traffic looks normal in this case, so don't take playback events as proof the agent spoke |
| The caller goes silent mid-call, after the agent had been talking | An unanswered client delegation. There is no way to cancel one — always reply with delegation.context.append |
Function ends immediately with server error |
Usually a key not enrolled in the Early Access Program (OpenAI refuses with Voice session access denied). Check the error object on the actionHook for the real reason |
| The agent greets the caller but uses its own words | You asked for a greeting without supplying the wording, or the model paraphrased anyway. Use a say Function if the wording is fixed |
| The model never calls your functions | Tools declared outside delegation.responses.tools, or delegation.type is 'client' |
| The Function is rejected before the call connects | model set inside session_update, or handoff/hangup/mcpServers configured without a responses delegation |
A session.updateLlm() call appears to do nothing |
The event type is not one of the five accepted ones; unrecognized types are discarded without an error |
xAI Voice Agent
Set vendor: 'xai' and supply an xAI API key via auth.apiKey. xAI's Voice Agent speaks the same OpenAI Realtime GA dialect described in the OpenAI Realtime section above — llmOptions carries the same response_create and session_update payloads, in the same GA format, forwarded as the response.create and session.update client events. Wipple CPaaS connects to wss://api.x.ai/v1/realtime.
session.llm({
vendor: 'xai',
model: 'grok-voice-latest',
auth: { apiKey: process.env.XAI_API_KEY },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
llmOptions: {
session_update: {
type: 'realtime',
instructions:
'You are a friendly, helpful voice assistant on a phone call. ' +
'Always respond in English unless the caller explicitly speaks another language. ' +
'Keep responses concise and natural for spoken conversation.',
turn_detection: { type: 'server_vad' },
audio: {
output: { voice: 'eve' }
}
},
response_create: {
output_modalities: ['audio'],
instructions: 'Greet the caller warmly in English and ask how you can help today.'
}
}
});
Model selection
modelstringxAI Voice Agent model name.
grok-voice-latest— default; currently aliases togrok-voice-think-fast-1.0.grok-voice-think-fast-1.0— flagship model.
Voices
Set the voice via session_update.audio.output.voice (or response_create.audio.output.voice), same field as OpenAI. Available voices: eve (default), ara, rex, sal, leo.
Required session_update
Unlike OpenAI, llmOptions.session_update is required for xai — audio does not begin flowing until after the first session.updated server event is received.
Audio format
Audio on the wire to xAI is pcm16 at 24 kHz. Wipple CPaaS forces this rate/format regardless of what is set in session_update.audio; do not attempt to select a different format or sample rate.
Turn detection
Unlike OpenAI GA, which nests turn detection under session_update.audio.input.turn_detection, xAI expects turn_detection top-level in session_update — i.e. session_update.turn_detection, not session_update.audio.input.turn_detection. { type: 'server_vad' } is the common setting. If turn_detection is omitted or set to null (manual turn mode), the application is responsible for turn-taking and can send input_audio_buffer.commit / input_audio_buffer.clear client events via an llm:update command.
Tool calls
Tool/function calling uses the same fields as OpenAI's (session_update.tools, tool_choice). The completed tool call arrives on the response.function_call_arguments.done event rather than OpenAI's response.output_item.done; Wipple CPaaS routes it to your toolHook the same way.
Input transcription
When session_update.audio.input.transcription is configured, caller-speech transcripts arrive on conversation.item.input_audio_transcription.updated. Unlike OpenAI's .completed event, this event is cumulative — each update contains the full transcript so far, not just a delta.
Azure Voice Live
Set vendor: 'voicelive' and point connectOptions.host at your Microsoft Foundry or Azure Speech resource to use the Azure Voice Live API. Wipple CPaaS connects to wss://{host}/voice-live/realtime.
Voice Live is not the same thing as vendor: 'microsoft'. That vendor targets the older Azure OpenAI Realtime deployment endpoint (openai/realtime) and shares its payload shapes with vendor: 'openai'. Voice Live is a separate service that keeps the Realtime event vocabulary but uses its own flat session shape and adds Azure-only capabilities: Azure Speech voices, semantic VAD, server-side noise suppression and echo cancellation, word timestamps and visemes, and non-realtime chat models (e.g. gpt-4.1) paired with Azure speech to text.
session.llm({
vendor: 'voicelive',
model: 'gpt-realtime',
auth: { apiKey: process.env.AZURE_VOICELIVE_API_KEY },
connectOptions: { host: 'my-resource.services.ai.azure.com' },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
llmOptions: {
session_update: {
modalities: ['text', 'audio'],
instructions: 'You are a friendly, helpful voice assistant on a phone call.',
voice: { name: 'en-US-Ava:DragonHDLatestNeural', type: 'azure-standard' },
turn_detection: { type: 'azure_semantic_vad', silence_duration_ms: 500, remove_filler_words: true },
input_audio_noise_reduction: { type: 'azure_deep_noise_suppression' }
},
response_create: {
instructions: 'Greet the caller warmly and ask how you can help today.'
}
}
});
Connection options
connectOptions.hoststringYour resource hostname, e.g. my-resource.services.ai.azure.com (or my-resource.cognitiveservices.azure.com for older resources).
connectOptions.apiVersionstringVoice Live api-version. Defaults to 2026-04-10.
connectOptions.agentIdstringFoundry Agent Service agent id. When set, the session is addressed by agent (agent_id, plus projectId as project_id) instead of by model.
Authentication
Supply either auth.apiKey, which Wipple CPaaS sends as the api-key query parameter, or auth.accessToken, a Microsoft Entra ID token sent as an Authorization: Bearer header. Entra tokens are short-lived, so mint one per call in your webhook; the token must be issued for the https://ai.azure.com/.default scope.
Required session_update
llmOptions.session_update is required — caller audio does not flow until the first session.updated server event arrives.
It is sent verbatim, in Voice Live's flat shape: voice, turn_detection, modalities and input_audio_transcription sit at the top level of session, not nested under audio.input / audio.output the way OpenAI's GA Realtime format has them. Do not port an openai_s2s session_update across unchanged.
Voices
Set voice as an object, not a string: { name, type }, where type is azure-standard (Azure Neural, HD and MAI voices), azure-custom, or azure-realtime-native (the curated voice set for the azure-realtime model). HD voices accept an optional temperature, and any Azure voice accepts rate between 0.5 and 1.5.
Turn detection
Beyond server_vad, Voice Live offers azure_semantic_vad and azure_semantic_vad_multilingual, which decide end-of-turn from meaning rather than volume and can drop filler words (remove_filler_words: true) so an "umm" does not trigger barge-in. These work with every model, unlike OpenAI's semantic_vad.
Audio format
Audio on the wire is pcm16 at 24 kHz, which matches Voice Live's default input_audio_sampling_rate. Wipple CPaaS forces this rate; do not set input_audio_sampling_rate: 16000.
Input transcription
Azure speech to text is active automatically only with a non-multimodal model (e.g. gpt-4.1). A native-audio model such as gpt-realtime-2.1 emits no conversation.item.input_audio_transcription.completed events at all unless you ask for transcription explicitly:
session_update: {
input_audio_transcription: { model: 'azure-speech' }
}
With a non-multimodal model, Azure speech to text is active automatically. Set input_audio_transcription.model to choose another: azure-speech, mai-transcribe, or — with gpt-realtime / gpt-realtime-mini — whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize.
Voice Live-only events
response.audio_timestamp.delta / .done (word timestamps, enabled with output_audio_timestamp_types: ['word']) and response.animation_viseme.delta / .done (enabled with animation: {outputs: ['viseme_id']}) are forwarded to your eventHook alongside the standard Realtime events.
Avatar output is not supported
Voice Live's text to speech avatar requires a separate WebRTC SDP exchange with the service, which has no place in a SIP call. Setting
avatarinsession_updatewill not produce video.
Google Gemini Live
Set vendor: 'google' to use Gemini's Live API. The same vendor: 'google' integration reaches the Live API through two access paths:
| Access path | Host | Auth | When to use |
|---|---|---|---|
| Gemini Developer API | generativelanguage.googleapis.com |
API key (AIza…) from Google AI Studio |
Dev / prototype / hobby. No SLA, shared quotas. |
| Vertex AI Live API | {LOCATION}-aiplatform.googleapis.com |
OAuth 2 Bearer token (ya29.…) minted from a GCP service-account JSON |
Production. SLA, IAM, audit logs, project-level quotas. |
Both speak the same JSON wire protocol — setup, realtimeInput, serverContent, modelTurn, toolCall, sessionResumptionUpdate are identical. Only the URL, auth, and model resource format differ. llmOptions.setup is forwarded verbatim to Google's BidiGenerateContentSetup message after the WebSocket connects.
Gemini Developer API (API key)
Default access path. Provide an API key via auth.apiKey; the key starts with AIza. Model name uses the short form models/<id>. connectOptions is optional — defaults to the Gemini Developer endpoint.
session.llm({
vendor: 'google',
model: 'models/gemini-2.0-flash-live-001',
auth: { apiKey: process.env.GEMINI_API_KEY },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
llmOptions: {
setup: {
generationConfig: {
speechConfig: {
voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Aoede' } }
}
},
systemInstruction: {
parts: [{ text: 'You are a helpful assistant named Barbara.' }]
}
},
greeting: 'Greet the caller warmly and ask how you can help.',
sessionResumption: {}
}
});
Vertex AI Live API (OAuth Bearer)
Production access path. The application mints an OAuth access token from a Google Cloud service-account JSON key and passes the token string (starting with ya29.) as auth.apiKey. Wipple CPaaS detects the ya29. prefix and sends it as an Authorization: Bearer … header on the WebSocket upgrade. Model name must be the full Vertex resource path.
const { JWT } = require('google-auth-library');
const fs = require('fs');
const saKey = JSON.parse(fs.readFileSync(process.env.GOOGLE_SERVICE_ACCOUNT_KEY_PATH, 'utf8'));
const jwtClient = new JWT({
email: saKey.client_email,
key: saKey.private_key,
scopes: ['https://www.googleapis.com/auth/cloud-platform']
});
const { token } = await jwtClient.getAccessToken(); // "ya29...."
const projectId = saKey.project_id;
const location = 'us-central1';
session.llm({
vendor: 'google',
model: `projects/${projectId}/locations/${location}/publishers/google/models/gemini-live-2.5-flash-native-audio`,
auth: { apiKey: token },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
connectOptions: {
host: `${location}-aiplatform.googleapis.com`,
path: '/ws/google.cloud.aiplatform.v1beta1.LlmBidiService/BidiGenerateContent'
},
llmOptions: {
setup: {
generationConfig: {
speechConfig: {
voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Aoede' } }
}
},
systemInstruction: {
parts: [{ text: 'You are a helpful assistant named Barbara.' }]
}
},
greeting: 'Greet the caller warmly and ask how you can help.'
}
});
Setup checklist for Vertex AI:
- GCP project with
aiplatform.googleapis.comAPI enabled and billing attached. - Service account with the
roles/aiplatform.userIAM role on the project. - Service-account JSON key file accessible to the application (store outside repo, reference via env var).
- Host region prefix in
connectOptions.hostmust match thelocations/<region>segment inmodel. e.g.us-central1-aiplatform.googleapis.comwithprojects/.../locations/us-central1/.... Mismatched regions returnPublisher Model not found.
OAuth token lifetime: tokens issued from a service-account JWT expire after \~1 hour. Wipple CPaaS uses whatever token the application provides at session start; it is not refreshed mid-session. The WebSocket stays open across expiry, but a reconnect with a stale token will fail. For typical call durations this is fine; mint a fresh token per call.
Google-specific connectOptions fields
hoststringWebSocket host. Default: generativelanguage.googleapis.com (Gemini Developer). For Vertex AI use {LOCATION}-aiplatform.googleapis.com matching the region in your model resource path.
pathstringWebSocket path. Default: /ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent (Gemini Developer). For Vertex AI use /ws/google.cloud.aiplatform.v1beta1.LlmBidiService/BidiGenerateContent.
versionstringLegacy — Gemini Developer API version string (v1, v1beta, v1alpha). Ignored when path is set. Kept for backward compatibility.
Google-specific llmOptions fields
setupobjectThe BidiGenerateContentSetup object sent to Gemini right after the websocket connects. The model field is populated automatically from the Function's model parameter. generationConfig.responseModalities is forced to audio.
greetingstring | objectOptional proactive greeting. When set, Wipple CPaaS sends a text message to Gemini immediately after setup so the agent speaks first without waiting for the caller to speak. Accepts either a string or an object with a text field. The value is an instruction to the model, not the literal words — for example "Greet the caller warmly" rather than "Hello, how can I help?".
Implemented using realtimeInput.text so it works on both the 2.0 Live models and gemini-3.1-flash-live-preview. (On 3.1, clientContent is reserved for seeding history and does not trigger a model response, which is why realtimeInput.text is used.)
sessionResumptionobjectEnable session resumption. Pass {} to opt in, or { handle: "..." } to resume a previous session. Resumption handles are delivered back to the application via llm_event sessionResumptionUpdate messages.
AssemblyAI Voice Agent
Set vendor: 'assemblyai' and supply your AssemblyAI API key via auth.api_key. llmOptions is the AssemblyAI Voice Agent session payload passed through verbatim — there is no Wipple CPaaS-specific wrapper. Wipple CPaaS wraps it as {type: 'session.update', session: <llmOptions>} and sends it as the first client message after the websocket connects to wss://agents.assemblyai.com/v1/ws. See the AssemblyAI Voice Agent product page and Voice Agent API docs for an overview.
The audio format is not configurable. AssemblyAI Voice Agent only accepts audio/pcm at 24 kHz, which Wipple CPaaS uses unconditionally — session.input.format / session.output.format set by the application are overridden. Wipple CPaaS resamples to/from the channel's native rate automatically.
session.llm({
vendor: 'assemblyai',
auth: { api_key: process.env.ASSEMBLYAI_API_KEY },
actionHook: '/final',
eventHook: '/event',
toolHook: '/toolCall',
events: ['all'],
llmOptions: {
system_prompt: 'You are a helpful voice agent.',
greeting: 'Hello, how can I help you today?',
output: { voice: 'ivy' },
input: {
keyterms: ['weather', 'temperature'],
turn_detection: {
vad_threshold: 0.5,
min_silence: 1000,
max_silence: 3000,
interrupt_response: true
}
},
tools: [
{
type: 'function',
name: 'getWeather',
description: 'Get current weather for a given city',
parameters: {
type: 'object',
properties: {
location: { type: 'string', description: 'City name' },
scale: { type: 'string', enum: ['celsius', 'fahrenheit'] }
},
required: ['location']
}
}
]
}
});
AssemblyAI-specific auth fields
api_keystringYour AssemblyAI API key. Sent as Authorization: Bearer <api_key> on the WebSocket handshake.
AssemblyAI-specific llmOptions fields
AssemblyAI's protocol requires a session.update message, but every field inside is optional — pass llmOptions: {} to start with all server defaults.
system_promptstringSystem prompt for the agent.
greetingstringInitial greeting the agent will speak when the session opens.
outputobjectOutput audio configuration. Supports voice — see the AssemblyAI voices reference for available IDs. The format sub-field is overridden by Wipple CPaaS.
inputobjectInput audio configuration. Supports keyterms (array of biasing terms) and turn_detection (vad_threshold, min_silence, max_silence, interrupt_response). The format sub-field is overridden by Wipple CPaaS.
toolsarrayArray of tool definitions. Each entry must include type: "function", name, description, and parameters (JSON Schema). Wipple CPaaS auto-fills type: "function" if omitted.
Tool calls
The agent invokes a tool by emitting a tool.call server event. Wipple CPaaS routes it to the application's toolHook with {name, args, tool_call_id}. The application replies via session.sendToolOutput(tool_call_id, {type: 'tool.result', tool_call_id, result}). The result should be a string (JSON-stringify objects before sending) — Wipple CPaaS JSON-stringifies non-string result values automatically.
Transfer-to-human handoff
Add a handoff block to let the realtime model transfer the caller to a human. The runtime
injects a transfer_to_human tool — no toolHook is needed for it. When the caller asks
for a human and the model calls the tool, Wipple CPaaS runs the packaged
transfer choreography to the configured destination.
session
.llm({
vendor: 'openai',
model: 'gpt-realtime',
auth: { apiKey: process.env.OPENAI_API_KEY },
llmOptions: {
response_create: { instructions: 'You are a helpful support agent.' },
session_update: { type: 'realtime', instructions: 'You are a helpful support agent.' },
},
handoff: {
mode: 'blind',
blindMethod: 'dial',
target: [{ type: 'user', name: 'agent-desk@sip.example.com' }],
},
actionHook: '/llm-complete',
})
.send();
The handoff block accepts every transfer option (mode,
target, blindMethod, disposition, confirm, etc.), plus brief ('auto' | 'none'
| { template }), briefSynthesizer, toolName, and toolDescription — see
the agent Function's handoff section for the details,
which apply identically here. When the human leg bridges, the actionHook reports the
transfer outcome.
Note
For
vendor: 'gptlive'the injectedtransfer_to_humantool needs a function-calling channel, sollmOptions.session_update.delegation.typemust be'responses'; the Function is rejected otherwise, rather than silently leaving the caller with no way to reach a human — see OpenAI GPT Live.
Built-in Hangup Tool
Set hangup to let the model end the call on its own. The runtime injects a hangup tool into the model's toolset; you do not define it in llmOptions and you do not handle it in your toolHook — the runtime intercepts the call, hangs up, and ends the llm Function. This works the same way across every supported s2s vendor (OpenAI, Deepgram Voice Agent, Ultravox, ElevenLabs, Google Gemini Live, AssemblyAI, xAI Voice Agent, Azure Voice Live, OpenAI GPT Live — the last requires delegation.type: 'responses').
session.llm({
vendor: 'openai',
model: 'gpt-realtime',
auth: { apiKey },
hangup: { reason: 'conversation complete' },
actionHook: '/final',
llmOptions: {
session_update: {
type: 'realtime',
instructions:
'You are a helpful voice assistant. When the caller says goodbye, call the hangup tool.',
},
},
});
The injected tool accepts an optional reason argument that the model may fill in when it decides to end the call. The reason placed in the X-Reason SIP header on the outbound BYE is resolved as follows:
- The model-supplied
reasonargument, if present. - Otherwise the app-supplied
hangup.reasondefault, if configured. - Otherwise no
X-Reasonheader is sent.
Pass an empty object (hangup: {}) to enable the tool with no default reason.
Note
When the model calls the hangup tool, the call is released immediately. The
actionHookstill fires, but any follow-on Functions it returns are discarded because the call is already being torn down.Note
For ElevenLabs, tools are configured on the ElevenLabs agent rather than injected per-session; configure a client tool named
hangupon your agent and Wipple CPaaS will intercept it and end the call.