Agent
session
.agent({
stt: {
vendor: 'deepgram',
language: 'multi',
deepgramOptions: { model: 'nova-3-general' },
},
tts: {
vendor: 'cartesia',
voice: '9626c31c-bec5-4cca-baa8-f8ba9e84c8bc',
},
llm: {
vendor: 'openai',
model: 'gpt-4.1-mini',
llmOptions: {
messages: [
{ role: 'system', content: 'You are a helpful voice assistant.' },
],
tools: [
{
name: 'get_weather',
description: 'Get current weather for a city',
parameters: {
type: 'object',
properties: {
city: { type: 'string' },
},
required: ['city'],
},
},
],
},
},
turnDetection: 'krisp',
earlyGeneration: true,
bargeIn: { enable: true, minSpeechDuration: 0.5 },
noResponseTimeout: 12,
toolHook: '/tool-call',
eventHook: '/agent-event',
actionHook: '/agent-complete',
})
.send();
The agent Function wires together three components — STT, LLM, and TTS — with integrated turn detection to create a complete voice AI agent. Unlike the llm Function (which connects to speech-to-speech APIs like OpenAI Realtime), the agent Function lets you mix and match separate vendors for each component.
The agent Function manages the full conversational turn cycle:
- User speaks → STT produces a transcript
- Turn detection decides the user is done speaking
- Transcript is sent to the LLM
- LLM response tokens stream to TTS
- TTS audio plays back to the caller
- If the user barges in, TTS stops and a new turn begins
Parameters
actionHookstringA webhook invoked when the agent Function ends. The payload includes a completion_reason field indicating why it terminated.
bargeInobjectControls whether and how the user can interrupt the assistant while it is speaking.
bargeIn.enablebooleanAllow the user to interrupt the assistant while it is speaking.
bargeIn.minSpeechDurationnumberSeconds of detected speech required before confirming an interruption. Prevents brief noises from cutting off the assistant.
bargeIn.stickybooleanIf true, once the user interrupts, the assistant does not resume speaking the interrupted response.
earlyGenerationbooleanEnable speculative LLM prompting before end-of-turn is confirmed. When using Krisp turn detection, set this to true to speculatively prompt the LLM before Krisp confirms the turn has ended. If the transcript matches when the turn ends, buffered tokens are released immediately — reducing response latency. Deepgram Flux performs early generation automatically regardless of this setting.
eventHookstringA webhook invoked for agent events. Receives event types: user_transcript, agent_response, user_interruption, turn_end, and history_summarized. See eventHook Events below.
greetingbooleanWhether the LLM should generate an initial greeting before the user speaks. Set to false if you want the agent to wait silently for the user to speak first.
handoffobjectDeclarative transfer-to-human configuration. When present, the runtime injects a
transfer_to_human tool into the LLM's toolset and runs the packaged
transfer choreography when the model calls it. See
Transfer-to-human handoff.
hangupobjectEnable the built-in hangup tool. When present, the runtime injects a hangup tool into the LLM's toolset; when the LLM calls it, the call is ended. See Built-in Hangup Tool below.
hangup.reasonstringDefault reason placed in the X-Reason SIP header on the outbound BYE. Used as a fallback when the LLM does not supply its own reason at call time.
idstringAn optional unique identifier for this Function instance.
llmobjectLLM configuration. This is the only required property. See Supported LLM Vendors for the full list of supported vendors.
llm.vendorstringLLM vendor id. One of: anthropic, azure-openai, baseten, bedrock, deepseek, google, groq, huggingface, moonshot, openai, vertex-gemini, vertex-openai, xai, zai.
llm.modelstringModel name. Format is vendor-specific (e.g., gpt-4o-mini, claude-haiku-4-5-20251001, gemini-2.5-flash, llama-3.3-70b-versatile, meta-llama/Llama-3.3-70B-Instruct). For Azure OpenAI, this is your deployment name, not the model id.
llm.labelstringOnly needed when the account has more than one LLM credential for the same vendor — give each a distinct label when the credential is created and reference it explicitly here. Most accounts have a single credential per vendor and should leave this unset.
llm.authobjectInline authentication credentials. Format depends on the vendor. If not provided, Wipple CPaaS looks up credentials by vendor (and label, if the rare multi-credential case applies) from the LLM credentials configured for your account.
llm.llmOptionsobjectLLM options including system prompt, tools, and generation parameters.
llm.llmOptions.messagesarrayInitial conversation messages. Typically includes a system message with instructions: [{ role: 'system', content: '...' }].
llm.llmOptions.systemPromptstringSystem prompt for the LLM. Alternative to providing a system message in messages.
llm.llmOptions.toolsarrayTool/function definitions available to the LLM. Uses OpenAI function-calling format. See Tool Calling below.
llm.llmOptions.maxTokensnumberMaximum number of tokens in the LLM response.
mcpServersarrayExternal MCP servers that provide tools to the LLM. The agent Function connects to each server at startup, discovers available tools, and makes them callable by the LLM. Each entry requires a url property and optionally auth and roots.
noiseIsolationstring or objectEnable server-side noise isolation to reduce background noise. As a string, pass "krisp" or "rnnoise" for default settings. As an object: { mode: "krisp", level: 80, direction: "read" }. direction can be "read" (filter caller audio, default) or "write" (filter outbound audio). Availability of Krisp depends on your service plan.
noResponseTimeoutnumberSeconds to wait after the assistant finishes speaking before prompting the user to respond. When triggered, the LLM is prompted with a system cue to check if the user is still there. Set to 0 to disable.
sttobjectSpeech-to-text configuration. See recognizer for available properties. Key properties: vendor, language, hints, and vendor-specific options like deepgramOptions.
toolHookstringA webhook invoked when the LLM requests a tool/function call. The payload includes tool_call_id, name, and arguments (already parsed as an object). See Tool Calling below.
ttsobjectText-to-speech configuration. See synthesizer for available properties. Key properties: vendor, voice, language, and vendor-specific options.
turnDetectionstring or objectTurn detection strategy. Controls when the agent Function decides the user has finished speaking.
As a string:
"stt"— Uses the STT vendor's native end-of-utterance signal. For most vendors this is silence-based. Vendors with smarter built-in turn detection (deepgramflux, assemblyai, speechmatics) always use their native detection regardless of this setting."krisp"— Uses the Krisp acoustic end-of-turn model, which analyzes speech patterns rather than just silence.
As an object (Krisp only):
mode—"krisp"(required)threshold— Confidence threshold 0.0–1.0. Lower values trigger earlier turn transitions. Default: 0.5.model— Optional Krisp model name override.
Tool Calling
Define tools in llm.llmOptions.tools and handle calls via toolHook. The tool call payload includes tool_call_id, name, and arguments (already parsed — an object, not a JSON string).
In WebSocket mode, respond with session.sendToolOutput(tool_call_id, result):
session.on('/tool-call', async (evt) => {
const { tool_call_id, name, arguments: args } = evt;
if (name === 'get_weather') {
const weather = await fetchWeather(args.city);
session.sendToolOutput(tool_call_id, weather);
return;
}
session.sendToolOutput(tool_call_id, `Unknown tool: ${name}`);
});
In webhook mode, return the tool result as JSON in the HTTP response body.
Alternatively, connect external MCP servers to provide tools automatically without defining them inline.
MCP Servers
Instead of (or in addition to) defining tools inline, you can connect to external MCP servers. The agent Function connects to each server at startup via SSE or Streamable HTTP transport, discovers available tools, and makes them callable by the LLM.
session
.agent({
llm: { vendor: 'openai', model: 'gpt-4.1', llmOptions: {
messages: [{ role: 'system', content: 'You are a sports assistant.' }],
}},
stt: { vendor: 'deepgram', language: 'en-US' },
tts: { vendor: 'cartesia', voice: 'sonic-english' },
mcpServers: [
{ url: 'https://livescoremcp.com/sse' },
],
actionHook: '/agent-complete',
})
.send();
When the LLM requests a tool call, the agent Function checks MCP servers first. If the tool name matches one discovered from an MCP server, the call is dispatched there directly. If no MCP server provides the tool, it falls through to the toolHook.
Transfer-to-human handoff
Add a handoff block to let the agent transfer the caller to a human. The runtime injects
a transfer_to_human tool into the LLM's toolset — you don't write a toolHook for it.
When the caller asks for a human and the model calls the tool, Wipple CPaaS runs the packaged
transfer choreography to the configured destination.
session
.agent({
stt: { vendor: 'deepgram', language: 'en-US' },
tts: { vendor: 'cartesia', voice: 'sonic-english' },
llm: { vendor: 'openai', model: 'gpt-4.1', llmOptions: {
messages: [{ role: 'system', content: 'You are a helpful support agent.' }],
}},
handoff: {
mode: 'blind',
blindMethod: 'dial',
brief: 'none',
target: [{ type: 'user', name: 'agent-desk@sip.example.com' }],
},
actionHook: '/agent-complete',
})
.send();
The handoff block accepts every transfer option (mode,
target, blindMethod, disposition, confirm, etc.), plus:
handoff.briefstring or objectControls the spoken summary delivered to the human on a warm transfer. 'auto' lets the
LLM write the summary; 'none' sends no spoken brief (and keeps the injected tool
argument-free); { template: '...' } guides the LLM's summary.
handoff.briefSynthesizerobjectOptional voice/vendor (synthesizer object) for the spoken brief. Defaults to the session synthesizer.
handoff.toolNamestringOverride the injected tool name.
handoff.toolDescriptionstringOverride the injected tool description shown to the LLM.
When the human leg bridges, the agent's actionHook reports the transfer outcome (see the
transfer actionHook properties).
Built-in Hangup Tool
Set hangup to let the LLM end the call on its own. The runtime injects a hangup tool into the LLM's toolset; you do not define it in llm.llmOptions.tools and you do not handle it in your toolHook — the runtime intercepts the call, hangs up, and ends the agent Function.
session
.agent({
llm: { vendor: 'openai', model: 'gpt-4.1-mini', llmOptions: {
messages: [{ role: 'system', content: 'You are a helpful assistant. When the caller says goodbye, call the hangup tool.' }],
}},
stt: { vendor: 'deepgram', language: 'en-US' },
tts: { vendor: 'cartesia', voice: '9626c31c-bec5-4cca-baa8-f8ba9e84c8bc' },
hangup: { reason: 'conversation complete' },
actionHook: '/agent-complete',
})
.send();
The injected tool accepts an optional reason argument that the LLM may fill in when it decides to end the call. The reason placed in the X-Reason SIP header on the outbound BYE is resolved as follows:
- The LLM-supplied
reasonargument, if present. - Otherwise the app-supplied
hangup.reasondefault, if configured. - Otherwise no
X-Reasonheader is sent.
Pass an empty object (hangup: {}) to enable the tool with no default reason.
Note
When the LLM calls the hangup tool, the call is released immediately. The
actionHookstill fires (withcompletion_reason: "hangup"), but any follow-on Functions it returns are discarded because the call is already being torn down.
eventHook Events
The eventHook receives real-time events during the conversation. In WebSocket mode, listen with session.on('/your-event-hook', handler).
turn_end
Sent at the end of each conversational turn. The most useful event for observability.
{
"type": "turn_end",
"transcript": "What's the weather in Portland?",
"confidence": 0.998,
"response": "The current temperature in Portland is 52°F with wind speed 12 km/h.",
"interrupted": false,
"latency": {
"stt_ms": 320,
"eot_ms": 180,
"llm_ms": 890,
"tool_ms": 420,
"tts_ms": 210,
"preflight": {
"result": "hit",
"tokens": 12
}
},
"tool_calls": [
{ "name": "get_weather", "rtt_ms": 420 }
]
}
Latency fields (all in milliseconds):
stt_ms— STT processing time (user stops talking → final transcript received)eot_ms— Additional wait for end-of-turn detection after transcriptllm_ms— Pure LLM thinking time (tool RTT subtracted)tool_ms— Total time spent in tool callstts_ms— TTS engine latency (text sent → first audio received)preflight— Early generation metrics:result(hit,miss, orpending) andtokensbuffered on a hit
user_transcript
Sent when the user's final transcript is available.
{
"type": "user_transcript",
"transcript": "What's the weather in Portland?"
}
agent_response
Sent when the LLM finishes generating its response.
{
"type": "agent_response",
"response": "The current temperature in Portland is 52°F."
}
user_interruption
Sent when the user barges in while the assistant is speaking.
{
"type": "user_interruption"
}
history_summarized
Sent when conversation history summarization completes (requires JAMBONES_PIPELINE_SUMMARIZE_TURNS environment variable).
{
"type": "history_summarized",
"turn": 8,
"messages_dropped": 5,
"messages_kept": 6,
"summary": "The user is a software developer looking for a MacBook Pro..."
}
Mid-conversation Updates
The agent Function supports asynchronous updates while a conversation is in progress. Updates can be sent via WebSocket (session.updateAgent(data)) or REST API.
update_instructions
Replace the LLM system prompt mid-conversation.
session.updateAgent({
type: 'update_instructions',
instructions: 'You are now a billing support agent.',
});
inject_context
Append messages to the LLM conversation history. System messages are routed to the system prompt for vendors that don't support inline system messages.
session.updateAgent({
type: 'inject_context',
messages: [
{ role: 'user', content: 'CRM context: Customer name: Sarah Mitchell. Account tier: Gold.' },
],
});
update_tools
Replace the tool set available to the LLM.
session.updateAgent({
type: 'update_tools',
tools: [
{
name: 'transfer_call',
description: 'Transfer the caller to a specialist',
parameters: { type: 'object', properties: { department: { type: 'string' } } },
},
],
});
generate_reply
Prompt the LLM to generate a new response. Use interrupt: true to cancel the current response and generate immediately.
session.updateAgent({
type: 'generate_reply',
interrupt: true,
user_input: 'URGENT: Tell the customer about the flash sale.',
});
Supported LLM Vendors
| Vendor | llm.vendor |
Example models |
|---|---|---|
| Anthropic | anthropic |
claude-haiku-4-5-20251001, claude-sonnet-4-6 |
| AWS Bedrock | bedrock |
amazon.nova-micro-v1:0, us.anthropic.claude-haiku-4-5-20251001-v1:0 |
| Azure OpenAI | azure-openai |
your deployment name |
| Baseten | baseten |
deepseek-ai/DeepSeek-V3.1, zai-org/GLM-5, moonshotai/Kimi-K2.6, openai/gpt-oss-120b |
| DeepSeek | deepseek |
deepseek-v4-flash, deepseek-v4-pro |
| Google AI Studio | google |
gemini-2.5-flash, gemini-2.5-pro |
| Groq | groq |
llama-3.3-70b-versatile, llama-3.1-8b-instant |
| HuggingFace | huggingface |
meta-llama/Llama-3.3-70B-Instruct, …:fastest |
| Moonshot (Kimi) | moonshot |
kimi-k2-0711-preview, kimi-latest, moonshot-v1-8k |
| OpenAI | openai |
gpt-5.4-mini, gpt-5.4, gpt-5.5, gpt-4o-mini |
| Vertex AI — Gemini | vertex-gemini |
gemini-2.5-flash, gemini-2.5-pro |
| Vertex AI — Partner Models | vertex-openai |
meta/llama-3.3-70b-instruct-maas, mistral-large |
| Z.ai (GLM) | zai |
glm-4.6, glm-4.5-air |