Wipple CPaaS
Menu
Functions / Agent

Agent

session
  .agent({
    stt: {
      vendor: 'deepgram',
      language: 'multi',
      deepgramOptions: { model: 'nova-3-general' },
    },
    tts: {
      vendor: 'cartesia',
      voice: '9626c31c-bec5-4cca-baa8-f8ba9e84c8bc',
    },
    llm: {
      vendor: 'openai',
      model: 'gpt-4.1-mini',
      llmOptions: {
        messages: [
          { role: 'system', content: 'You are a helpful voice assistant.' },
        ],
        tools: [
          {
            name: 'get_weather',
            description: 'Get current weather for a city',
            parameters: {
              type: 'object',
              properties: {
                city: { type: 'string' },
              },
              required: ['city'],
            },
          },
        ],
      },
    },
    turnDetection: 'krisp',
    earlyGeneration: true,
    bargeIn: { enable: true, minSpeechDuration: 0.5 },
    noResponseTimeout: 12,
    toolHook: '/tool-call',
    eventHook: '/agent-event',
    actionHook: '/agent-complete',
  })
  .send();

The agent Function wires together three components — STT, LLM, and TTS — with integrated turn detection to create a complete voice AI agent. Unlike the llm Function (which connects to speech-to-speech APIs like OpenAI Realtime), the agent Function lets you mix and match separate vendors for each component.

The agent Function manages the full conversational turn cycle:

  1. User speaks → STT produces a transcript
  2. Turn detection decides the user is done speaking
  3. Transcript is sent to the LLM
  4. LLM response tokens stream to TTS
  5. TTS audio plays back to the caller
  6. If the user barges in, TTS stops and a new turn begins

Parameters

actionHookstring

A webhook invoked when the agent Function ends. The payload includes a completion_reason field indicating why it terminated.


bargeInobject

Controls whether and how the user can interrupt the assistant while it is speaking.


bargeIn.enablebooleandefault: true

Allow the user to interrupt the assistant while it is speaking.


bargeIn.minSpeechDurationnumberdefault: 0.5

Seconds of detected speech required before confirming an interruption. Prevents brief noises from cutting off the assistant.


bargeIn.stickybooleandefault: false

If true, once the user interrupts, the assistant does not resume speaking the interrupted response.


earlyGenerationbooleandefault: false

Enable speculative LLM prompting before end-of-turn is confirmed. When using Krisp turn detection, set this to true to speculatively prompt the LLM before Krisp confirms the turn has ended. If the transcript matches when the turn ends, buffered tokens are released immediately — reducing response latency. Deepgram Flux performs early generation automatically regardless of this setting.


eventHookstring

A webhook invoked for agent events. Receives event types: user_transcript, agent_response, user_interruption, turn_end, and history_summarized. See eventHook Events below.


greetingbooleandefault: true

Whether the LLM should generate an initial greeting before the user speaks. Set to false if you want the agent to wait silently for the user to speak first.


handoffobject

Declarative transfer-to-human configuration. When present, the runtime injects a transfer_to_human tool into the LLM's toolset and runs the packaged transfer choreography when the model calls it. See Transfer-to-human handoff.


hangupobject

Enable the built-in hangup tool. When present, the runtime injects a hangup tool into the LLM's toolset; when the LLM calls it, the call is ended. See Built-in Hangup Tool below.


hangup.reasonstring

Default reason placed in the X-Reason SIP header on the outbound BYE. Used as a fallback when the LLM does not supply its own reason at call time.


idstring

An optional unique identifier for this Function instance.


llmobjectrequired

LLM configuration. This is the only required property. See Supported LLM Vendors for the full list of supported vendors.


llm.vendorstringrequired

LLM vendor id. One of: anthropic, azure-openai, baseten, bedrock, deepseek, google, groq, huggingface, moonshot, openai, vertex-gemini, vertex-openai, xai, zai.


llm.modelstringrequired

Model name. Format is vendor-specific (e.g., gpt-4o-mini, claude-haiku-4-5-20251001, gemini-2.5-flash, llama-3.3-70b-versatile, meta-llama/Llama-3.3-70B-Instruct). For Azure OpenAI, this is your deployment name, not the model id.


llm.labelstring

Only needed when the account has more than one LLM credential for the same vendor — give each a distinct label when the credential is created and reference it explicitly here. Most accounts have a single credential per vendor and should leave this unset.


llm.authobject

Inline authentication credentials. Format depends on the vendor. If not provided, Wipple CPaaS looks up credentials by vendor (and label, if the rare multi-credential case applies) from the LLM credentials configured for your account.


llm.llmOptionsobject

LLM options including system prompt, tools, and generation parameters.


llm.llmOptions.messagesarray

Initial conversation messages. Typically includes a system message with instructions: [{ role: 'system', content: '...' }].


llm.llmOptions.systemPromptstring

System prompt for the LLM. Alternative to providing a system message in messages.


llm.llmOptions.toolsarray

Tool/function definitions available to the LLM. Uses OpenAI function-calling format. See Tool Calling below.


llm.llmOptions.maxTokensnumber

Maximum number of tokens in the LLM response.


mcpServersarray

External MCP servers that provide tools to the LLM. The agent Function connects to each server at startup, discovers available tools, and makes them callable by the LLM. Each entry requires a url property and optionally auth and roots.


noiseIsolationstring or object

Enable server-side noise isolation to reduce background noise. As a string, pass "krisp" or "rnnoise" for default settings. As an object: { mode: "krisp", level: 80, direction: "read" }. direction can be "read" (filter caller audio, default) or "write" (filter outbound audio). Availability of Krisp depends on your service plan.


noResponseTimeoutnumberdefault: 12

Seconds to wait after the assistant finishes speaking before prompting the user to respond. When triggered, the LLM is prompted with a system cue to check if the user is still there. Set to 0 to disable.


sttobject

Speech-to-text configuration. See recognizer for available properties. Key properties: vendor, language, hints, and vendor-specific options like deepgramOptions.


toolHookstring

A webhook invoked when the LLM requests a tool/function call. The payload includes tool_call_id, name, and arguments (already parsed as an object). See Tool Calling below.


ttsobject

Text-to-speech configuration. See synthesizer for available properties. Key properties: vendor, voice, language, and vendor-specific options.


turnDetectionstring or objectdefault: stt

Turn detection strategy. Controls when the agent Function decides the user has finished speaking.

As a string:

  • "stt" — Uses the STT vendor's native end-of-utterance signal. For most vendors this is silence-based. Vendors with smarter built-in turn detection (deepgramflux, assemblyai, speechmatics) always use their native detection regardless of this setting.
  • "krisp" — Uses the Krisp acoustic end-of-turn model, which analyzes speech patterns rather than just silence.

As an object (Krisp only):

  • mode — "krisp" (required)
  • threshold — Confidence threshold 0.0–1.0. Lower values trigger earlier turn transitions. Default: 0.5.
  • model — Optional Krisp model name override.

Tool Calling

Define tools in llm.llmOptions.tools and handle calls via toolHook. The tool call payload includes tool_call_id, name, and arguments (already parsed — an object, not a JSON string).

In WebSocket mode, respond with session.sendToolOutput(tool_call_id, result):

session.on('/tool-call', async (evt) => {
  const { tool_call_id, name, arguments: args } = evt;
  if (name === 'get_weather') {
    const weather = await fetchWeather(args.city);
    session.sendToolOutput(tool_call_id, weather);
    return;
  }
  session.sendToolOutput(tool_call_id, `Unknown tool: ${name}`);
});

In webhook mode, return the tool result as JSON in the HTTP response body.

Alternatively, connect external MCP servers to provide tools automatically without defining them inline.

MCP Servers

Instead of (or in addition to) defining tools inline, you can connect to external MCP servers. The agent Function connects to each server at startup via SSE or Streamable HTTP transport, discovers available tools, and makes them callable by the LLM.

session
  .agent({
    llm: { vendor: 'openai', model: 'gpt-4.1', llmOptions: {
      messages: [{ role: 'system', content: 'You are a sports assistant.' }],
    }},
    stt: { vendor: 'deepgram', language: 'en-US' },
    tts: { vendor: 'cartesia', voice: 'sonic-english' },
    mcpServers: [
      { url: 'https://livescoremcp.com/sse' },
    ],
    actionHook: '/agent-complete',
  })
  .send();

When the LLM requests a tool call, the agent Function checks MCP servers first. If the tool name matches one discovered from an MCP server, the call is dispatched there directly. If no MCP server provides the tool, it falls through to the toolHook.

Transfer-to-human handoff

Add a handoff block to let the agent transfer the caller to a human. The runtime injects a transfer_to_human tool into the LLM's toolset — you don't write a toolHook for it. When the caller asks for a human and the model calls the tool, Wipple CPaaS runs the packaged transfer choreography to the configured destination.

session
  .agent({
    stt: { vendor: 'deepgram', language: 'en-US' },
    tts: { vendor: 'cartesia', voice: 'sonic-english' },
    llm: { vendor: 'openai', model: 'gpt-4.1', llmOptions: {
      messages: [{ role: 'system', content: 'You are a helpful support agent.' }],
    }},
    handoff: {
      mode: 'blind',
      blindMethod: 'dial',
      brief: 'none',
      target: [{ type: 'user', name: 'agent-desk@sip.example.com' }],
    },
    actionHook: '/agent-complete',
  })
  .send();

The handoff block accepts every transfer option (mode, target, blindMethod, disposition, confirm, etc.), plus:

handoff.briefstring or objectdefault: auto

Controls the spoken summary delivered to the human on a warm transfer. 'auto' lets the LLM write the summary; 'none' sends no spoken brief (and keeps the injected tool argument-free); { template: '...' } guides the LLM's summary.


handoff.briefSynthesizerobject

Optional voice/vendor (synthesizer object) for the spoken brief. Defaults to the session synthesizer.


handoff.toolNamestringdefault: transfer_to_human

Override the injected tool name.


handoff.toolDescriptionstring

Override the injected tool description shown to the LLM.


When the human leg bridges, the agent's actionHook reports the transfer outcome (see the transfer actionHook properties).

Built-in Hangup Tool

Set hangup to let the LLM end the call on its own. The runtime injects a hangup tool into the LLM's toolset; you do not define it in llm.llmOptions.tools and you do not handle it in your toolHook — the runtime intercepts the call, hangs up, and ends the agent Function.

session
  .agent({
    llm: { vendor: 'openai', model: 'gpt-4.1-mini', llmOptions: {
      messages: [{ role: 'system', content: 'You are a helpful assistant. When the caller says goodbye, call the hangup tool.' }],
    }},
    stt: { vendor: 'deepgram', language: 'en-US' },
    tts: { vendor: 'cartesia', voice: '9626c31c-bec5-4cca-baa8-f8ba9e84c8bc' },
    hangup: { reason: 'conversation complete' },
    actionHook: '/agent-complete',
  })
  .send();

The injected tool accepts an optional reason argument that the LLM may fill in when it decides to end the call. The reason placed in the X-Reason SIP header on the outbound BYE is resolved as follows:

  1. The LLM-supplied reason argument, if present.
  2. Otherwise the app-supplied hangup.reason default, if configured.
  3. Otherwise no X-Reason header is sent.

Pass an empty object (hangup: {}) to enable the tool with no default reason.

Note

When the LLM calls the hangup tool, the call is released immediately. The actionHook still fires (with completion_reason: "hangup"), but any follow-on Functions it returns are discarded because the call is already being torn down.

eventHook Events

The eventHook receives real-time events during the conversation. In WebSocket mode, listen with session.on('/your-event-hook', handler).

turn_end

Sent at the end of each conversational turn. The most useful event for observability.

{
  "type": "turn_end",
  "transcript": "What's the weather in Portland?",
  "confidence": 0.998,
  "response": "The current temperature in Portland is 52°F with wind speed 12 km/h.",
  "interrupted": false,
  "latency": {
    "stt_ms": 320,
    "eot_ms": 180,
    "llm_ms": 890,
    "tool_ms": 420,
    "tts_ms": 210,
    "preflight": {
      "result": "hit",
      "tokens": 12
    }
  },
  "tool_calls": [
    { "name": "get_weather", "rtt_ms": 420 }
  ]
}

Latency fields (all in milliseconds):

  • stt_ms — STT processing time (user stops talking → final transcript received)
  • eot_ms — Additional wait for end-of-turn detection after transcript
  • llm_ms — Pure LLM thinking time (tool RTT subtracted)
  • tool_ms — Total time spent in tool calls
  • tts_ms — TTS engine latency (text sent → first audio received)
  • preflight — Early generation metrics: result (hit, miss, or pending) and tokens buffered on a hit

user_transcript

Sent when the user's final transcript is available.

{
  "type": "user_transcript",
  "transcript": "What's the weather in Portland?"
}

agent_response

Sent when the LLM finishes generating its response.

{
  "type": "agent_response",
  "response": "The current temperature in Portland is 52°F."
}

user_interruption

Sent when the user barges in while the assistant is speaking.

{
  "type": "user_interruption"
}

history_summarized

Sent when conversation history summarization completes (requires JAMBONES_PIPELINE_SUMMARIZE_TURNS environment variable).

{
  "type": "history_summarized",
  "turn": 8,
  "messages_dropped": 5,
  "messages_kept": 6,
  "summary": "The user is a software developer looking for a MacBook Pro..."
}

Mid-conversation Updates

The agent Function supports asynchronous updates while a conversation is in progress. Updates can be sent via WebSocket (session.updateAgent(data)) or REST API.

update_instructions

Replace the LLM system prompt mid-conversation.

session.updateAgent({
  type: 'update_instructions',
  instructions: 'You are now a billing support agent.',
});

inject_context

Append messages to the LLM conversation history. System messages are routed to the system prompt for vendors that don't support inline system messages.

session.updateAgent({
  type: 'inject_context',
  messages: [
    { role: 'user', content: 'CRM context: Customer name: Sarah Mitchell. Account tier: Gold.' },
  ],
});

update_tools

Replace the tool set available to the LLM.

session.updateAgent({
  type: 'update_tools',
  tools: [
    {
      name: 'transfer_call',
      description: 'Transfer the caller to a specialist',
      parameters: { type: 'object', properties: { department: { type: 'string' } } },
    },
  ],
});

generate_reply

Prompt the LLM to generate a new response. Use interrupt: true to cancel the current response and generate immediately.

session.updateAgent({
  type: 'generate_reply',
  interrupt: true,
  user_input: 'URGENT: Tell the customer about the flash sale.',
});

Supported LLM Vendors

Vendor llm.vendor Example models
Anthropic anthropic claude-haiku-4-5-20251001, claude-sonnet-4-6
AWS Bedrock bedrock amazon.nova-micro-v1:0, us.anthropic.claude-haiku-4-5-20251001-v1:0
Azure OpenAI azure-openai your deployment name
Baseten baseten deepseek-ai/DeepSeek-V3.1, zai-org/GLM-5, moonshotai/Kimi-K2.6, openai/gpt-oss-120b
DeepSeek deepseek deepseek-v4-flash, deepseek-v4-pro
Google AI Studio google gemini-2.5-flash, gemini-2.5-pro
Groq groq llama-3.3-70b-versatile, llama-3.1-8b-instant
HuggingFace huggingface meta-llama/Llama-3.3-70B-Instruct, …:fastest
Moonshot (Kimi) moonshot kimi-k2-0711-preview, kimi-latest, moonshot-v1-8k
OpenAI openai gpt-5.4-mini, gpt-5.4, gpt-5.5, gpt-4o-mini
Vertex AI — Gemini vertex-gemini gemini-2.5-flash, gemini-2.5-pro
Vertex AI — Partner Models vertex-openai meta/llama-3.3-70b-instruct-maas, mistral-large
Z.ai (GLM) zai glm-4.6, glm-4.5-air