Wipple CPaaS
Menu
Functions / LLM

LLM

session.llm({
  vendor: 'openai',
  model: 'gpt-realtime',
  auth: { apiKey },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  events: [
    'conversation.item.*',
    'response.output_audio_transcript.done',
    'input_audio_buffer.committed'
  ],
  llmOptions: {
    response_create: {
      output_modalities: ['audio'],
      instructions: 'Greet the caller warmly in English and ask how you can help today.',
      audio: {
        output: {
          voice: 'alloy',
          format: { type: 'audio/pcm', rate: 24000 }
        }
      },
      max_output_tokens: 4096
    },
    session_update: {
      type: 'realtime',
      instructions:
        'You are a friendly, helpful voice assistant on a phone call. ' +
        'Always respond in English unless the caller explicitly speaks another language. ' +
        'Keep responses concise and natural for spoken conversation. ' +
        'If asked about weather, call the get_weather function.',
      tools: [
        {
          name: 'get_weather',
          type: 'function',
          description: 'Get the weather at a given location',
          parameters: {
            type: 'object',
            properties: {
              location: {
                type: 'string',
                description: 'Location to get the weather from',
              },
              scale: {
                type: 'string',
                enum: ['fahrenheit', 'celsius'],
              },
            },
            required: ['location', 'scale'],
          },
        },
      ],
      tool_choice: 'auto',
      audio: {
        input: {
          format: { type: 'audio/pcm', rate: 24000 },
          transcription: { model: 'whisper-1' },
          turn_detection: {
            type: 'server_vad',
            threshold: 0.8,
            prefix_padding_ms: 300,
            silence_duration_ms: 500
          }
        },
        output: {
          format: { type: 'audio/pcm', rate: 24000 },
          voice: 'alloy'
        }
      }
    }
  }
});

Parameters

modelstringrequired

Name of the LLM model.


vendorstringrequired

Name of the LLM vendor.


actionHookstring

Webhook that will be called when the LLM session ends.


authobject

Object containing authentication credentials; format according to the model.


connectOptionsobject

Object containing information such as the URI to connect to.


eventHookstring

Webhook that will be called when a requested LLM event happens (e.g., transcript).


eventsarray

Array of event names listing the events requested (wildcards allowed).


handoffobject

Declarative transfer-to-human configuration. When present, the runtime injects a transfer_to_human tool and runs the packaged transfer choreography when the model calls it. See Transfer-to-human handoff.


hangupobject

Enable the built-in hangup tool. When present, the runtime injects a hangup tool into the model's toolset; when the model calls it, the call is ended. Works across all supported s2s vendors. See Built-in Hangup Tool below.


hangup.reasonstring

Default reason placed in the X-Reason SIP header on the outbound BYE. Used as a fallback when the model does not supply its own reason at call time.


llmOptionsobject

Object containing instructions for the LLM; format dependent on the LLM model.


toolHookstring

Webhook that will be called when the LLM wants to call a function.


The following LLMs are currently supported:

  • OpenAI Realtime API
  • OpenAI GPT Live API (limited-access alpha)
  • Deepgram Voice Agent
  • Ultravox
  • ElevenLabs
  • Google Gemini Live API (via Gemini Developer API or Vertex AI)
  • AssemblyAI Voice Agent
  • xAI Voice Agent
  • Azure Voice Live

OpenAI Realtime

Set vendor: 'openai' and supply an OpenAI API key via auth.apiKey. llmOptions carries two payloads — response_create and session_update — that are forwarded to OpenAI as the corresponding response.create and session.update client events. The example at the top of this page uses the GA format.

Model selection

modelstringrequired

GA realtime model name. Common choices:

  • gpt-realtime — the standard GA realtime conversation model. Recommended starting point.
  • gpt-realtime-2 — reasoning-capable variant. Emits a phase: "commentary" message before phase: "final_answer" in its output items; both phases produce audio, so the caller hears the model "think out loud" before the actual answer. Useful for some assistants, surprising for typical greeting-first flows.
  • gpt-realtime-whisper — streaming transcription model. Emits response.output_audio_transcript.delta events but does not endpoint turns server-side (no .completed in current GA); not suitable for normal speech-to-speech or gather flows.

OpenAI deprecated the Realtime preview models (gpt-4o-realtime-preview-*) on 2026-05-12; use a GA model.


Instructions

session_update.instructions sets the persona that persists for the whole session; response_create.instructions overrides it for a single response only. GA realtime models do not assume English by default — if Wipple CPaaS speaks first (which it does when response.create fires before any caller audio), anchor the language in instructions (e.g. "Always respond in English unless the caller speaks another language") or you may get a greeting in an arbitrary language.

Audio format

The audio format on the wire to OpenAI is fixed at {type: 'audio/pcm', rate: 24000}. Wipple CPaaS resamples to and from the channel's native rate (typically G.711) automatically. Any audio.input.format or audio.output.format set in the application is overridden — they are present in the example only for completeness.

Legacy preview-format compatibility

Apps written against the preview API still work without changes. Wipple CPaaS detects the legacy flat format — top-level modalities, voice, input_audio_format, output_audio_format, input_audio_transcription, turn_detection, temperature — inside session_update or response_create, converts it to the GA format on the wire, and aliases GA event names back to their preview names before forwarding to your eventHook. For example, an app that subscribed to response.audio_transcript.done continues to receive events with that exact type even though the server now emits response.output_audio_transcript.done. A one-time deprecation warning is logged when the legacy format is detected.

GA-invalid fields are stripped silently with a WARN log:

  • temperature in response_create — no longer accepted in GA.
  • output_audio_format as a flat string — replaced by audio.output.format: {type, rate} (which Wipple CPaaS overrides to pcm 24 kHz anyway).

The compatibility layer is intended as a transitional shim — plan to update applications to the GA format before it is removed.

OpenAI-specific llmOptions fields

response_createobjectrequired

The initial response.create.response payload sent to OpenAI when the session opens. Drives the first response from the assistant.


response_create.output_modalitiesarray

Array of output modalities. Typically ['audio'] for voice. (GA renamed this from preview's modalities.)


response_create.instructionsstring

Per-response instruction override. Replaces session-level instructions for this one response only — useful for the opening greeting. Anchor the language explicitly when the model speaks first.


response_create.audioobject

Audio configuration object. Use audio.output.voice to choose a voice (e.g. alloy, marin) and audio.output.format (overridden to pcm 24 kHz by Wipple CPaaS).


response_create.max_output_tokensnumber

Maximum tokens for this response.


session_updateobject

The initial session.update.session payload. Sets persona, tools, audio config, and VAD for the whole session.


session_update.typestringrequired

Must be 'realtime' in GA. Required when present.


session_update.instructionsstring

Session-wide persona prompt. Persists for every response in the session unless overridden by response_create.instructions.


session_update.audio.input.transcription.modelstring

Auxiliary transcription model for converting caller audio into text in the conversation log. Common values: whisper-1, gpt-4o-transcribe, gpt-realtime-whisper. This is separate from the conversation model.


session_update.audio.input.turn_detectionobject

Server-side VAD configuration: {type: 'server_vad', threshold, prefix_padding_ms, silence_duration_ms}. Defaults are server-supplied if omitted. type: 'semantic_vad' with an eagerness field is also supported on models that accept it.


session_update.toolsarray

Array of function-tool definitions exposed to the model. Each entry needs name, type: 'function', description, and JSON Schema parameters. The model calls a tool by emitting response.output_item.done with item.type: 'function_call'; Wipple CPaaS routes that to your toolHook.


session_update.tool_choicestring

'auto' (default), 'none', or a specific tool name to force.


OpenAI GPT Live

Set vendor: 'gptlive' and supply an OpenAI API key in auth.apiKey to talk to OpenAI's GPT Live API.

Alpha API

GPT Live requires enrollment in OpenAI's Early Access Program. An ordinary OpenAI key completes the connection and is then refused by OpenAI at the first server event with Voice session access denied, which Wipple CPaaS reports as completion_reason: 'server error'.

Expect event names and fields to change while the API is in alpha.

Not a drop-in for vendor: 'openai'

GPT Live is a different protocol from the OpenAI Realtime API, not a newer model for it. If you are moving an app over from vendor: 'openai', read Migrating from the Realtime API first — the payloads are not interchangeable.

A minimum configuration — note that on its own this produces a silent call, because nothing has asked the model to speak first (see Making the agent speak first):

session.llm({
  vendor: 'gptlive',
  model: 'gpt-live-1-boulder-alpha',
  auth: { apiKey: process.env.OPENAI_API_KEY },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  llmOptions: {
    session_update: {
      instructions:
        'You are a friendly, helpful voice assistant on a phone call. ' +
        'Always respond in English unless the caller explicitly speaks another language. ' +
        'Keep responses concise and natural for spoken conversation.',
      audio: {
        output: { voice: 'marin' }
      }
    }
  }
});

Model selection

modelstring

GPT Live model name; defaults to gpt-live-1-boulder-alpha. Set it on the Function. It must not appear inside session_update at all — Wipple CPaaS rejects the Function if it does, even if you set it in only that one place.


Audio format

There is no audio format to configure. GPT Live fixes it at 24 kHz mono PCM, and Wipple CPaaS converts to and from the caller's codec for you.

session_update is required

llmOptions.session_update is not optional here: GPT Live will not accept the caller's audio until it has your configuration. Wipple CPaaS sends it first and waits for session.started — the event you should hang your greeting off, described next.

There is no response_create. Once the session starts, the model drives the conversation itself.

Making the agent speak first

Because there is no response_create, nothing tells the model to take the first turn — and putting the greeting in instructions is not enough. The model waits for the caller, who hears silence.

To open the call, send a session.context.append as soon as you receive session.started. Give it the wording you want and tell it when to speak:

session.on('/event', (evt) => {
  if (evt.type === 'session.started') {
    session.updateLlm({
      type: 'session.context.append',
      content: [{
        type: 'input_text',
        text: 'Immediately greet the caller using the exact text below. Do not wait for the '
          + 'caller to speak first. After the greeting, pause and listen.\n\n'
          + 'Hi, I am your support assistant. How can I help you today?',
      }],
    });
  }
});

Use the same pattern any time you need the agent to say something specific mid-call — a disclosure, a transfer notice, a closing.

Greetings are requested, not guaranteed

A context append guides the model. Per OpenAI, it may paraphrase your wording, or occasionally stay silent. Omit the wording and the model will use its own.

If the exact words matter — a legal disclosure, a brand greeting — do not ask the model. Play it yourself with a say Function before the llm Function.

Note

Requesting a greeting needs WebSocket transport. The llm:update command is not accepted by the REST updateCall API, so a webhook-only application has no way to send one.

Delegation: how the model asks for work

Delegation is GPT Live's distinguishing concept, and it has no direct analogue in the other vendors on this page. Anything the model needs from outside the spoken conversation — a fact it does not know, a function it wants run — arrives as a delegation. Your choice of kind determines whether you can use tools at all.

session_update.delegationobject

Omit it entirely and the model simply never asks for anything — no tools, no context requests. That is the simplest way to start.


session_update.delegation.typestring

'client' — the model asks your application for background information in prose. Simple to implement, but there is no function calling.

'responses' — the model runs a turn on OpenAI's Responses API that can call your functions. Required if you want tools, mcpServers, handoff or hangup.


Responses delegation (function calling)

Declare your tools inside the delegation. Note that responses.model is a second, separate model that runs the delegated turn — it is not the voice model on the Function:

delegation: {
  type: 'responses',
  responses: {
    model: 'gpt-5.5',          // required
    tools: [
      {
        type: 'function',
        name: 'get_weather',
        description: 'Get the current weather for a city.',
        parameters: {
          type: 'object',
          properties: { location: { type: 'string' } },
          required: ['location']
        }
      }
    ]
  }
}

Warning

Tools go in delegation.responses.tools — not delegation.tools — and both the nested responses object and its model are required.

Wipple CPaaS rejects the Function before the call connects if you configure handoff, hangup or mcpServers without them. It cannot check your own tool declarations, so if you omit responses.model there, OpenAI refuses the session at startup and you get completion_reason: 'server error'.

Client delegation (supplying context)

The model emits delegation.created with item.target: 'client'; you answer with the item's id and up to 500 tokens of prose in a single input_text part:

session.on('/event', (evt) => {
  if (evt.type === 'delegation.created' && evt.item.target === 'client') {
    session.updateLlm({
      type: 'delegation.context.append',
      delegation_item_id: evt.item.id,
      content: [{ type: 'input_text', text: 'It is 62 degrees and raining in Seattle.' }]
    });
  }
});

Always answer a client delegation. There is no way to cancel one, so if you leave it unanswered the model waits and the caller hears silence.

Handling tool calls

With a responses delegation, Wipple CPaaS calls your toolHook with {name, args, tool_call_id} just as it does for every other vendor. What differs is the envelope you send the result back in:

session.on('/toolCall', async(evt) => {
  const { tool_call_id, name, args } = evt;

  session.sendToolOutput(tool_call_id, {
    type: 'delegation.function_call_output.create',
    item: {
      type: 'function_call_output',
      call_id: tool_call_id,
      output: JSON.stringify({ temperature: 62, conditions: 'rain' })
    }
  });
});

output must be a string. Send one of these per call if the model asked for several, and do not follow it with a response.create — unlike the Realtime API, the server resumes on its own.

The built-in handoff and hangup tools, and any mcpServers you configure, all work here too — but only with a responses delegation, since a client delegation has no way to call a function.

Following the conversation

GPT Live is in alpha and has no public event reference, so this is the list. Name the ones you want in the Function's events array — wildcards such as response.* work, as do 'all' plus -eventName exclusions.

Event What it gives you
session.started The session is live. Send your greeting request here
session.updated, session.context.appended Acknowledgements of your client events
turn.created, turn.delta, turn.done Utterance-level view of the conversation. turn.done carries turn.role ('user' or 'assistant') and turn.transcript
input_transcript.added, output_transcript.added Live partial transcript fragments, caller and agent
delegation.created, delegation.context.appended, delegation.function_call_output.created Delegation lifecycle
response.* The delegated Responses turn: response.created, response.output_text.delta, response.completed, response.failed, response.done, and others
output_audio.playback_started, output_audio.playback_stopped Emitted by Wipple CPaaS, not OpenAI: when agent audio actually starts and stops reaching the caller
session.usage.updated Cumulative token usage; final totals arrive on session.closed
error A rejected client event or a server-side problem

Warning

If you omit events entirely, Wipple CPaaS forwards everything — including transcript fragments, which are high volume. Name the events you need.

Prefer turn.done over the *_transcript.added fragments for following the conversation: fragment boundaries follow speech cadence rather than complete thoughts, so one sentence may arrive in several pieces.

Barge-in needs no work on your part — Wipple CPaaS detects it and flushes the queued agent audio; the underlying signal is not forwarded to your eventHook.

Updating the session mid-call

An llm:update command (session.updateLlm() in the Node SDK) accepts these five events:

  • session.update — change configuration. Sparse: fields you omit keep their values. Replacing delegation requires the whole object.
  • session.context.append — add context or request a spoken message
  • delegation.context.append — answer a client delegation
  • delegation.function_call_output.create — return a tool result
  • session.close — end the session gracefully

Anything else is discarded silently — you will not get an error back, so check the type if an update appears to do nothing.

How the session ends

The actionHook fires with a completion_reason:

completion_reason Meaning
normal conversation end The session closed cleanly
session closed: <reason> OpenAI closed the session for its own reason, which is included
disconnect from remote end OpenAI dropped the connection
server error GPT Live refused the session at startup. The payload also carries an error object with OpenAI's reason — check it first
connection failure Wipple CPaaS could not open the connection at all: network, DNS or proxy
hangup The model ended the call using the built-in hangup tool

A handoff that bridges reports the transfer outcome instead. Unlike some other vendors, GPT Live never reports server failure.

Not every problem ends the call. Once the session is running, a rejected client event — a stale delegation_item_id, an over-long context append — or a failed delegation is reported to your eventHook and the conversation continues. Handle those events if you want to recover; ignore them and the model continues without the context or tool result you were sending.

Note

cancelOnBargeIn and cancelOnResponseTimeout have no effect with this vendor, because GPT Live has nothing to cancel. responseTimeoutMs does work, but it is disabled by default — set it to a non-zero value and you get a response.timeout event if a delegation stalls, which you can recover from with an llm:update.

Migrating from the Realtime API

If you are porting an app from vendor: 'openai':

Realtime (vendor: 'openai') GPT Live (vendor: 'gptlive')
llmOptions.response_create Not supported — remove it. Greet with session.context.append instead
session_update.tools session_update.delegation.responses.tools
session_update.turn_detection Not supported — turn detection is handled for you
session_update.audio input/output formats Not supported — audio is fixed at 24 kHz mono PCM
Tool result via conversation.item.create delegation.function_call_output.create
A follow-on response.create after a tool result Nothing — the server resumes by itself
response.create / response.cancel mid-call Not supported
input_audio_buffer.speech_started for barge-in Wipple CPaaS detects barge-in itself; the signal is not forwarded to your eventHook

Troubleshooting

Symptom Cause and fix
The agent never speaks and the caller hears silence No session.context.append on session.started — see Making the agent speak first. Note that event traffic looks normal in this case, so don't take playback events as proof the agent spoke
The caller goes silent mid-call, after the agent had been talking An unanswered client delegation. There is no way to cancel one — always reply with delegation.context.append
Function ends immediately with server error Usually a key not enrolled in the Early Access Program (OpenAI refuses with Voice session access denied). Check the error object on the actionHook for the real reason
The agent greets the caller but uses its own words You asked for a greeting without supplying the wording, or the model paraphrased anyway. Use a say Function if the wording is fixed
The model never calls your functions Tools declared outside delegation.responses.tools, or delegation.type is 'client'
The Function is rejected before the call connects model set inside session_update, or handoff/hangup/mcpServers configured without a responses delegation
A session.updateLlm() call appears to do nothing The event type is not one of the five accepted ones; unrecognized types are discarded without an error

xAI Voice Agent

Set vendor: 'xai' and supply an xAI API key via auth.apiKey. xAI's Voice Agent speaks the same OpenAI Realtime GA dialect described in the OpenAI Realtime section above — llmOptions carries the same response_create and session_update payloads, in the same GA format, forwarded as the response.create and session.update client events. Wipple CPaaS connects to wss://api.x.ai/v1/realtime.

session.llm({
  vendor: 'xai',
  model: 'grok-voice-latest',
  auth: { apiKey: process.env.XAI_API_KEY },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  llmOptions: {
    session_update: {
      type: 'realtime',
      instructions:
        'You are a friendly, helpful voice assistant on a phone call. ' +
        'Always respond in English unless the caller explicitly speaks another language. ' +
        'Keep responses concise and natural for spoken conversation.',
      turn_detection: { type: 'server_vad' },
      audio: {
        output: { voice: 'eve' }
      }
    },
    response_create: {
      output_modalities: ['audio'],
      instructions: 'Greet the caller warmly in English and ask how you can help today.'
    }
  }
});

Model selection

modelstringrequired

xAI Voice Agent model name.

  • grok-voice-latest — default; currently aliases to grok-voice-think-fast-1.0.
  • grok-voice-think-fast-1.0 — flagship model.

Voices

Set the voice via session_update.audio.output.voice (or response_create.audio.output.voice), same field as OpenAI. Available voices: eve (default), ara, rex, sal, leo.

Required session_update

Unlike OpenAI, llmOptions.session_update is required for xai — audio does not begin flowing until after the first session.updated server event is received.

Audio format

Audio on the wire to xAI is pcm16 at 24 kHz. Wipple CPaaS forces this rate/format regardless of what is set in session_update.audio; do not attempt to select a different format or sample rate.

Turn detection

Unlike OpenAI GA, which nests turn detection under session_update.audio.input.turn_detection, xAI expects turn_detection top-level in session_update — i.e. session_update.turn_detection, not session_update.audio.input.turn_detection. { type: 'server_vad' } is the common setting. If turn_detection is omitted or set to null (manual turn mode), the application is responsible for turn-taking and can send input_audio_buffer.commit / input_audio_buffer.clear client events via an llm:update command.

Tool calls

Tool/function calling uses the same fields as OpenAI's (session_update.tools, tool_choice). The completed tool call arrives on the response.function_call_arguments.done event rather than OpenAI's response.output_item.done; Wipple CPaaS routes it to your toolHook the same way.

Input transcription

When session_update.audio.input.transcription is configured, caller-speech transcripts arrive on conversation.item.input_audio_transcription.updated. Unlike OpenAI's .completed event, this event is cumulative — each update contains the full transcript so far, not just a delta.

Azure Voice Live

Set vendor: 'voicelive' and point connectOptions.host at your Microsoft Foundry or Azure Speech resource to use the Azure Voice Live API. Wipple CPaaS connects to wss://{host}/voice-live/realtime.

Voice Live is not the same thing as vendor: 'microsoft'. That vendor targets the older Azure OpenAI Realtime deployment endpoint (openai/realtime) and shares its payload shapes with vendor: 'openai'. Voice Live is a separate service that keeps the Realtime event vocabulary but uses its own flat session shape and adds Azure-only capabilities: Azure Speech voices, semantic VAD, server-side noise suppression and echo cancellation, word timestamps and visemes, and non-realtime chat models (e.g. gpt-4.1) paired with Azure speech to text.

session.llm({
  vendor: 'voicelive',
  model: 'gpt-realtime',
  auth: { apiKey: process.env.AZURE_VOICELIVE_API_KEY },
  connectOptions: { host: 'my-resource.services.ai.azure.com' },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  llmOptions: {
    session_update: {
      modalities: ['text', 'audio'],
      instructions: 'You are a friendly, helpful voice assistant on a phone call.',
      voice: { name: 'en-US-Ava:DragonHDLatestNeural', type: 'azure-standard' },
      turn_detection: { type: 'azure_semantic_vad', silence_duration_ms: 500, remove_filler_words: true },
      input_audio_noise_reduction: { type: 'azure_deep_noise_suppression' }
    },
    response_create: {
      instructions: 'Greet the caller warmly and ask how you can help today.'
    }
  }
});

Connection options

connectOptions.hoststringrequired

Your resource hostname, e.g. my-resource.services.ai.azure.com (or my-resource.cognitiveservices.azure.com for older resources).


connectOptions.apiVersionstring

Voice Live api-version. Defaults to 2026-04-10.


connectOptions.agentIdstring

Foundry Agent Service agent id. When set, the session is addressed by agent (agent_id, plus projectId as project_id) instead of by model.


Authentication

Supply either auth.apiKey, which Wipple CPaaS sends as the api-key query parameter, or auth.accessToken, a Microsoft Entra ID token sent as an Authorization: Bearer header. Entra tokens are short-lived, so mint one per call in your webhook; the token must be issued for the https://ai.azure.com/.default scope.

Required session_update

llmOptions.session_update is required — caller audio does not flow until the first session.updated server event arrives.

It is sent verbatim, in Voice Live's flat shape: voice, turn_detection, modalities and input_audio_transcription sit at the top level of session, not nested under audio.input / audio.output the way OpenAI's GA Realtime format has them. Do not port an openai_s2s session_update across unchanged.

Voices

Set voice as an object, not a string: { name, type }, where type is azure-standard (Azure Neural, HD and MAI voices), azure-custom, or azure-realtime-native (the curated voice set for the azure-realtime model). HD voices accept an optional temperature, and any Azure voice accepts rate between 0.5 and 1.5.

Turn detection

Beyond server_vad, Voice Live offers azure_semantic_vad and azure_semantic_vad_multilingual, which decide end-of-turn from meaning rather than volume and can drop filler words (remove_filler_words: true) so an "umm" does not trigger barge-in. These work with every model, unlike OpenAI's semantic_vad.

Audio format

Audio on the wire is pcm16 at 24 kHz, which matches Voice Live's default input_audio_sampling_rate. Wipple CPaaS forces this rate; do not set input_audio_sampling_rate: 16000.

Input transcription

Azure speech to text is active automatically only with a non-multimodal model (e.g. gpt-4.1). A native-audio model such as gpt-realtime-2.1 emits no conversation.item.input_audio_transcription.completed events at all unless you ask for transcription explicitly:

session_update: {
  input_audio_transcription: { model: 'azure-speech' }
}

With a non-multimodal model, Azure speech to text is active automatically. Set input_audio_transcription.model to choose another: azure-speech, mai-transcribe, or — with gpt-realtime / gpt-realtime-mini — whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize.

Voice Live-only events

response.audio_timestamp.delta / .done (word timestamps, enabled with output_audio_timestamp_types: ['word']) and response.animation_viseme.delta / .done (enabled with animation: {outputs: ['viseme_id']}) are forwarded to your eventHook alongside the standard Realtime events.

Avatar output is not supported

Voice Live's text to speech avatar requires a separate WebRTC SDP exchange with the service, which has no place in a SIP call. Setting avatar in session_update will not produce video.

Google Gemini Live

Set vendor: 'google' to use Gemini's Live API. The same vendor: 'google' integration reaches the Live API through two access paths:

Access path Host Auth When to use
Gemini Developer API generativelanguage.googleapis.com API key (AIza…) from Google AI Studio Dev / prototype / hobby. No SLA, shared quotas.
Vertex AI Live API {LOCATION}-aiplatform.googleapis.com OAuth 2 Bearer token (ya29.…) minted from a GCP service-account JSON Production. SLA, IAM, audit logs, project-level quotas.

Both speak the same JSON wire protocol — setup, realtimeInput, serverContent, modelTurn, toolCall, sessionResumptionUpdate are identical. Only the URL, auth, and model resource format differ. llmOptions.setup is forwarded verbatim to Google's BidiGenerateContentSetup message after the WebSocket connects.

Gemini Developer API (API key)

Default access path. Provide an API key via auth.apiKey; the key starts with AIza. Model name uses the short form models/<id>. connectOptions is optional — defaults to the Gemini Developer endpoint.

session.llm({
  vendor: 'google',
  model: 'models/gemini-2.0-flash-live-001',
  auth: { apiKey: process.env.GEMINI_API_KEY },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  llmOptions: {
    setup: {
      generationConfig: {
        speechConfig: {
          voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Aoede' } }
        }
      },
      systemInstruction: {
        parts: [{ text: 'You are a helpful assistant named Barbara.' }]
      }
    },
    greeting: 'Greet the caller warmly and ask how you can help.',
    sessionResumption: {}
  }
});

Vertex AI Live API (OAuth Bearer)

Production access path. The application mints an OAuth access token from a Google Cloud service-account JSON key and passes the token string (starting with ya29.) as auth.apiKey. Wipple CPaaS detects the ya29. prefix and sends it as an Authorization: Bearer … header on the WebSocket upgrade. Model name must be the full Vertex resource path.

const { JWT } = require('google-auth-library');
const fs = require('fs');

const saKey = JSON.parse(fs.readFileSync(process.env.GOOGLE_SERVICE_ACCOUNT_KEY_PATH, 'utf8'));
const jwtClient = new JWT({
  email: saKey.client_email,
  key: saKey.private_key,
  scopes: ['https://www.googleapis.com/auth/cloud-platform']
});

const { token } = await jwtClient.getAccessToken();   // "ya29...."
const projectId = saKey.project_id;
const location = 'us-central1';

session.llm({
  vendor: 'google',
  model: `projects/${projectId}/locations/${location}/publishers/google/models/gemini-live-2.5-flash-native-audio`,
  auth: { apiKey: token },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  connectOptions: {
    host: `${location}-aiplatform.googleapis.com`,
    path: '/ws/google.cloud.aiplatform.v1beta1.LlmBidiService/BidiGenerateContent'
  },
  llmOptions: {
    setup: {
      generationConfig: {
        speechConfig: {
          voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Aoede' } }
        }
      },
      systemInstruction: {
        parts: [{ text: 'You are a helpful assistant named Barbara.' }]
      }
    },
    greeting: 'Greet the caller warmly and ask how you can help.'
  }
});

Setup checklist for Vertex AI:

  1. GCP project with aiplatform.googleapis.com API enabled and billing attached.
  2. Service account with the roles/aiplatform.user IAM role on the project.
  3. Service-account JSON key file accessible to the application (store outside repo, reference via env var).
  4. Host region prefix in connectOptions.host must match the locations/<region> segment in model. e.g. us-central1-aiplatform.googleapis.com with projects/.../locations/us-central1/.... Mismatched regions return Publisher Model not found.

OAuth token lifetime: tokens issued from a service-account JWT expire after \~1 hour. Wipple CPaaS uses whatever token the application provides at session start; it is not refreshed mid-session. The WebSocket stays open across expiry, but a reconnect with a stale token will fail. For typical call durations this is fine; mint a fresh token per call.

Google-specific connectOptions fields

hoststring

WebSocket host. Default: generativelanguage.googleapis.com (Gemini Developer). For Vertex AI use {LOCATION}-aiplatform.googleapis.com matching the region in your model resource path.


pathstring

WebSocket path. Default: /ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent (Gemini Developer). For Vertex AI use /ws/google.cloud.aiplatform.v1beta1.LlmBidiService/BidiGenerateContent.


versionstring

Legacy — Gemini Developer API version string (v1, v1beta, v1alpha). Ignored when path is set. Kept for backward compatibility.


Google-specific llmOptions fields

setupobjectrequired

The BidiGenerateContentSetup object sent to Gemini right after the websocket connects. The model field is populated automatically from the Function's model parameter. generationConfig.responseModalities is forced to audio.


greetingstring | object

Optional proactive greeting. When set, Wipple CPaaS sends a text message to Gemini immediately after setup so the agent speaks first without waiting for the caller to speak. Accepts either a string or an object with a text field. The value is an instruction to the model, not the literal words — for example "Greet the caller warmly" rather than "Hello, how can I help?".

Implemented using realtimeInput.text so it works on both the 2.0 Live models and gemini-3.1-flash-live-preview. (On 3.1, clientContent is reserved for seeding history and does not trigger a model response, which is why realtimeInput.text is used.)


sessionResumptionobject

Enable session resumption. Pass {} to opt in, or { handle: "..." } to resume a previous session. Resumption handles are delivered back to the application via llm_event sessionResumptionUpdate messages.


AssemblyAI Voice Agent

Set vendor: 'assemblyai' and supply your AssemblyAI API key via auth.api_key. llmOptions is the AssemblyAI Voice Agent session payload passed through verbatim — there is no Wipple CPaaS-specific wrapper. Wipple CPaaS wraps it as {type: 'session.update', session: <llmOptions>} and sends it as the first client message after the websocket connects to wss://agents.assemblyai.com/v1/ws. See the AssemblyAI Voice Agent product page and Voice Agent API docs for an overview.

The audio format is not configurable. AssemblyAI Voice Agent only accepts audio/pcm at 24 kHz, which Wipple CPaaS uses unconditionally — session.input.format / session.output.format set by the application are overridden. Wipple CPaaS resamples to/from the channel's native rate automatically.

session.llm({
  vendor: 'assemblyai',
  auth: { api_key: process.env.ASSEMBLYAI_API_KEY },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  events: ['all'],
  llmOptions: {
    system_prompt: 'You are a helpful voice agent.',
    greeting: 'Hello, how can I help you today?',
    output: { voice: 'ivy' },
    input: {
      keyterms: ['weather', 'temperature'],
      turn_detection: {
        vad_threshold: 0.5,
        min_silence: 1000,
        max_silence: 3000,
        interrupt_response: true
      }
    },
    tools: [
      {
        type: 'function',
        name: 'getWeather',
        description: 'Get current weather for a given city',
        parameters: {
          type: 'object',
          properties: {
            location: { type: 'string', description: 'City name' },
            scale: { type: 'string', enum: ['celsius', 'fahrenheit'] }
          },
          required: ['location']
        }
      }
    ]
  }
});

AssemblyAI-specific auth fields

api_keystringrequired

Your AssemblyAI API key. Sent as Authorization: Bearer <api_key> on the WebSocket handshake.


AssemblyAI-specific llmOptions fields

AssemblyAI's protocol requires a session.update message, but every field inside is optional — pass llmOptions: {} to start with all server defaults.

system_promptstring

System prompt for the agent.


greetingstring

Initial greeting the agent will speak when the session opens.


outputobject

Output audio configuration. Supports voice — see the AssemblyAI voices reference for available IDs. The format sub-field is overridden by Wipple CPaaS.


inputobject

Input audio configuration. Supports keyterms (array of biasing terms) and turn_detection (vad_threshold, min_silence, max_silence, interrupt_response). The format sub-field is overridden by Wipple CPaaS.


toolsarray

Array of tool definitions. Each entry must include type: "function", name, description, and parameters (JSON Schema). Wipple CPaaS auto-fills type: "function" if omitted.


Tool calls

The agent invokes a tool by emitting a tool.call server event. Wipple CPaaS routes it to the application's toolHook with {name, args, tool_call_id}. The application replies via session.sendToolOutput(tool_call_id, {type: 'tool.result', tool_call_id, result}). The result should be a string (JSON-stringify objects before sending) — Wipple CPaaS JSON-stringifies non-string result values automatically.

Transfer-to-human handoff

Add a handoff block to let the realtime model transfer the caller to a human. The runtime injects a transfer_to_human tool — no toolHook is needed for it. When the caller asks for a human and the model calls the tool, Wipple CPaaS runs the packaged transfer choreography to the configured destination.

session
  .llm({
    vendor: 'openai',
    model: 'gpt-realtime',
    auth: { apiKey: process.env.OPENAI_API_KEY },
    llmOptions: {
      response_create: { instructions: 'You are a helpful support agent.' },
      session_update: { type: 'realtime', instructions: 'You are a helpful support agent.' },
    },
    handoff: {
      mode: 'blind',
      blindMethod: 'dial',
      target: [{ type: 'user', name: 'agent-desk@sip.example.com' }],
    },
    actionHook: '/llm-complete',
  })
  .send();

The handoff block accepts every transfer option (mode, target, blindMethod, disposition, confirm, etc.), plus brief ('auto' | 'none' | { template }), briefSynthesizer, toolName, and toolDescription — see the agent Function's handoff section for the details, which apply identically here. When the human leg bridges, the actionHook reports the transfer outcome.

Note

For vendor: 'gptlive' the injected transfer_to_human tool needs a function-calling channel, so llmOptions.session_update.delegation.type must be 'responses'; the Function is rejected otherwise, rather than silently leaving the caller with no way to reach a human — see OpenAI GPT Live.

Built-in Hangup Tool

Set hangup to let the model end the call on its own. The runtime injects a hangup tool into the model's toolset; you do not define it in llmOptions and you do not handle it in your toolHook — the runtime intercepts the call, hangs up, and ends the llm Function. This works the same way across every supported s2s vendor (OpenAI, Deepgram Voice Agent, Ultravox, ElevenLabs, Google Gemini Live, AssemblyAI, xAI Voice Agent, Azure Voice Live, OpenAI GPT Live — the last requires delegation.type: 'responses').

session.llm({
  vendor: 'openai',
  model: 'gpt-realtime',
  auth: { apiKey },
  hangup: { reason: 'conversation complete' },
  actionHook: '/final',
  llmOptions: {
    session_update: {
      type: 'realtime',
      instructions:
        'You are a helpful voice assistant. When the caller says goodbye, call the hangup tool.',
    },
  },
});

The injected tool accepts an optional reason argument that the model may fill in when it decides to end the call. The reason placed in the X-Reason SIP header on the outbound BYE is resolved as follows:

  1. The model-supplied reason argument, if present.
  2. Otherwise the app-supplied hangup.reason default, if configured.
  3. Otherwise no X-Reason header is sent.

Pass an empty object (hangup: {}) to enable the tool with no default reason.

Note

When the model calls the hangup tool, the call is released immediately. The actionHook still fires, but any follow-on Functions it returns are discarded because the call is already being torn down.

Note

For ElevenLabs, tools are configured on the ElevenLabs agent rather than injected per-session; configure a client tool named hangup on your agent and Wipple CPaaS will intercept it and end the call.