Wipple CPaaS
Menu
Functions / Gather

Gather

{
  "verb": "gather",
  "actionHook": "/collect",
  "input": ["digits", "speech"],
  "bargein": true,
  "dtmfBargein": true,
  "finishOnKey": "#",
  "numDigits": 5,
  "timeout": 8,
  "recognizer": {
    "vendor": "google",
    "language": "en-US",
    "hints": ["sales", "support"],
    "hintsBoost": 10
  },
  "say": {
    "text": "To speak to Sales press 1 or say Sales.  To speak to customer support press 2 or say Support"
  }
}

Note

When collecting speech input, the default recognizer for the application will be used unless you overridde it by specifying a recognizer property.

Parameters

actionHookstringrequired

Webhook or websocket URL to send the collected digits or speech to; may be either an absolute or relative URL (see note below).
In the case of webhooks, this will be sent in a POST request.
See below for payload details.


actionHookDelayActionobject

An object describing behaviors to apply if the webhook or websocket application delays in responding to the actionHook.


bargeinboolean

Boolean indicating whether to allow barge-in when collecting speech (i.e., kill audio playback if the caller begins speaking).


dtmfBargeinboolean

Boolean indicating whether to allow barge-in using DTMF keys.


fillerNoiseobjectdeprecated

An object describing audio to play to the caller if the webhook or websocket application delays in responding to the actionHook.
(Deprecated in favor of actionHookDelay).


fillerNoise.enablebooleanrequireddeprecated

Whether to enable or disable filler noise.


fillerNoise.startDelaySecsnumberrequireddeprecated

Integer value specifying the number of seconds to wait for a response from the remote application.


fillerNoise.urlstringrequireddeprecated

HTTP(S) URL to audio to play as filler noise.


finishOnKeystring

A DTMF key that, if received, indicates the end of DTMF input.


inputarraydefault: ['digits']

Array specifying allowed types of input: ['digits'], ['speech'], or ['digits', 'speech'].


interDigitTimeoutnumber

Amount of time in seconds to wait between digits after minDigits have been entered.


listenDuringPromptbooleandefault: true

If false, do not listen for user speech until the nested say or play command has completed.


maxDigitsnumber

Maximum number of DTMF digits expected to gather.


minBargeinWordCountnumberdefault: 1

If bargein is true, only kill speech when this many words are spoken OR a final transcription is returned.


minDigitsnumberdefault: 1

Minimum number of DTMF digits expected to gather.


numDigitsnumber

Exact number of DTMF digits expected to gather.


partialResultHookstring

Webhook to send interim transcription results to.
Partial transcriptions are only generated if this property is set.


playobject

Nested play Function that can be used to prompt the user with an audio file.


recognizerobject

Speech recognition options to override default settings if desired.
See recognizer.


sayobject

Nested say command that can be used to prompt the user with text-to-speech.


Note

The actionHook property may be either an absolute URL or a relative URL. If it is a relative URL, it will be resolved relative to the base URL of the Wipple CPaaS application. Generally speaking, you will typically use a relative URL for the actionHook property, and when using websockets you are strongly encouragedto use a relative URL.

actionHook properties

The payload sent to the actionHook URL will always include a reason property with one of the following values:

  • speechDetected - user speech was collected
  • dtmfDetected - user dtmf entries were collected
  • stt-low-confidence - user speech was collected but the confidence level was below the threshold
  • timeout - neither speech nor dtmf was collected before the timeout expired
  • error - an error occurred during the gather operation, possibly with the speech recognizer

Depending on the reason, additional properties may be included in the payload.

speechDetected

A speech object will be included in the payload. The speech object will include the following properties:

  • language_code - the detected language that was spoken
  • is_final - a boolean indicating whether this was a final or interim transcription
  • channel_tag - a value of 1 if the transcription is from the caller, or 2 if from the callee
  • alternatives - an array of alternative transcriptions, each with a transcript and confidence value
  • vendor - the raw payload returned from the speech vendor

Example:

{
  "language_code": "en",
  "channel_tag": 1,
  "is_final": true,
  "alternatives": [
    {
      "confidence": 0.9848633,
      "transcript": "yes sorry setting up wi fi calling"
    }
  ],
  "vendor": {
    "name": "deepgram",
    "evt": {
      "type": "Results",
      "channel_index": [
        0,
        1
      ],
      "duration": 4.1899986,
      "start": 47.74,
      "is_final": true,
      "speech_final": true,
      "channel": {
        "alternatives": [
          {
            "transcript": "yes sorry setting up a wi fi calling",
            "confidence": 0.9848633,
            "words": [
              {
                  "word": "yes",
                  "start": 48.659542,
                  "end": 48.939404,
                  "confidence": 0.82470703
              },
              {
                  "word": "sorry",
                  "start": 48.939404,
                  "end": 49.259243,
                  "confidence": 0.98535156
              },
              {
                  "word": "setting",
                  "start": 49.259243,
                  "end": 49.61906,
                  "confidence": 0.9941406
              },
              {
                  "word": "up",
                  "start": 49.61906,
                  "end": 49.85894,
                  "confidence": 0.9848633
              },
              {
                  "word": "a",
                  "start": 49.85894,
                  "end": 49.97888,
                  "confidence": 0.94921875
              },
              {
                  "word": "wi",
                  "start": 49.97888,
                  "end": 50.178783,
                  "confidence": 0.80371094
              },
              {
                  "word": "fi",
                  "start": 50.178783,
                  "end": 50.418663,
                  "confidence": 0.99902344
              },
              {
                  "word": "calling",
                  "start": 50.418663,
                  "end": 50.918663,
                  "confidence": 0.9326172
              }
            ]
          }
        ]
      },
      "metadata": {
        "request_id": "b1bf1f58-ce74-42be-8b43-37f2f8331e72",
        "model_info": {
          "name": "phonecall-enhanced",
          "version": "2022-05-12.1",
          "arch": "polaris"
        },
        "model_uuid": "9a15c1cc-65e1-429a-9db6-f4ea6fbf822a"
      },
      "from_finalize": false
    }
  }
}

dtmfDetected

A digits property will be included in the payload containing the digits that were collected.

stt-low-confidence

A speech object will be included, see speechDetectedfor an example. The reason stt-low-confidence indicates to the application that this transcript should be treated with caution as it is likely to be an inaccurate reporting of what the user actually said.

timeout

No additional properties will be included in the payload.

error

A details property will be included in the payload with a description of the error.