Gather
{
"verb": "gather",
"actionHook": "/collect",
"input": ["digits", "speech"],
"bargein": true,
"dtmfBargein": true,
"finishOnKey": "#",
"numDigits": 5,
"timeout": 8,
"recognizer": {
"vendor": "google",
"language": "en-US",
"hints": ["sales", "support"],
"hintsBoost": 10
},
"say": {
"text": "To speak to Sales press 1 or say Sales. To speak to customer support press 2 or say Support"
}
}
Note
When collecting speech input, the default recognizer for the application will be used unless you overridde it by specifying a recognizer property.
Parameters
actionHookstringWebhook or websocket URL to send the collected digits or speech to; may be either an absolute or relative URL (see note below).
In the case of webhooks, this will be sent in a POST request.
See below for payload details.
actionHookDelayActionobjectAn object describing behaviors to apply if the webhook or websocket application delays in responding to the actionHook.
bargeinbooleanBoolean indicating whether to allow barge-in when collecting speech (i.e., kill audio playback if the caller begins speaking).
dtmfBargeinbooleanBoolean indicating whether to allow barge-in using DTMF keys.
fillerNoiseobjectAn object describing audio to play to the caller if the webhook or websocket application delays in responding to the actionHook.
(Deprecated in favor of actionHookDelay).
fillerNoise.enablebooleanWhether to enable or disable filler noise.
fillerNoise.startDelaySecsnumberInteger value specifying the number of seconds to wait for a response from the remote application.
fillerNoise.urlstringHTTP(S) URL to audio to play as filler noise.
finishOnKeystringA DTMF key that, if received, indicates the end of DTMF input.
inputarrayArray specifying allowed types of input: ['digits'], ['speech'], or ['digits', 'speech'].
interDigitTimeoutnumberAmount of time in seconds to wait between digits after minDigits have been entered.
listenDuringPromptbooleanIf false, do not listen for user speech until the nested say or play command has completed.
maxDigitsnumberMaximum number of DTMF digits expected to gather.
minBargeinWordCountnumberIf bargein is true, only kill speech when this many words are spoken OR a final transcription is returned.
minDigitsnumberMinimum number of DTMF digits expected to gather.
numDigitsnumberExact number of DTMF digits expected to gather.
partialResultHookstringWebhook to send interim transcription results to.
Partial transcriptions are only generated if this property is set.
playobjectNested play Function that can be used to prompt the user with an audio file.
recognizerobjectSpeech recognition options to override default settings if desired.
See recognizer.
sayobjectNested say command that can be used to prompt the user with text-to-speech.
Note
The
actionHookproperty may be either an absolute URL or a relative URL. If it is a relative URL, it will be resolved relative to the base URL of the Wipple CPaaS application. Generally speaking, you will typically use a relative URL for theactionHookproperty, and when using websockets you are strongly encouragedto use a relative URL.
actionHook properties
The payload sent to the actionHook URL will always include a reason property with one of the following values:
speechDetected- user speech was collecteddtmfDetected- user dtmf entries were collectedstt-low-confidence- user speech was collected but the confidence level was below the thresholdtimeout- neither speech nor dtmf was collected before the timeout expirederror- an error occurred during the gather operation, possibly with the speech recognizer
Depending on the reason, additional properties may be included in the payload.
speechDetected
A speech object will be included in the payload. The speech object will include the following properties:
language_code- the detected language that was spokenis_final- a boolean indicating whether this was a final or interim transcriptionchannel_tag- a value of 1 if the transcription is from the caller, or 2 if from the calleealternatives- an array of alternative transcriptions, each with atranscriptandconfidencevaluevendor- the raw payload returned from the speech vendor
Example:
{
"language_code": "en",
"channel_tag": 1,
"is_final": true,
"alternatives": [
{
"confidence": 0.9848633,
"transcript": "yes sorry setting up wi fi calling"
}
],
"vendor": {
"name": "deepgram",
"evt": {
"type": "Results",
"channel_index": [
0,
1
],
"duration": 4.1899986,
"start": 47.74,
"is_final": true,
"speech_final": true,
"channel": {
"alternatives": [
{
"transcript": "yes sorry setting up a wi fi calling",
"confidence": 0.9848633,
"words": [
{
"word": "yes",
"start": 48.659542,
"end": 48.939404,
"confidence": 0.82470703
},
{
"word": "sorry",
"start": 48.939404,
"end": 49.259243,
"confidence": 0.98535156
},
{
"word": "setting",
"start": 49.259243,
"end": 49.61906,
"confidence": 0.9941406
},
{
"word": "up",
"start": 49.61906,
"end": 49.85894,
"confidence": 0.9848633
},
{
"word": "a",
"start": 49.85894,
"end": 49.97888,
"confidence": 0.94921875
},
{
"word": "wi",
"start": 49.97888,
"end": 50.178783,
"confidence": 0.80371094
},
{
"word": "fi",
"start": 50.178783,
"end": 50.418663,
"confidence": 0.99902344
},
{
"word": "calling",
"start": 50.418663,
"end": 50.918663,
"confidence": 0.9326172
}
]
}
]
},
"metadata": {
"request_id": "b1bf1f58-ce74-42be-8b43-37f2f8331e72",
"model_info": {
"name": "phonecall-enhanced",
"version": "2022-05-12.1",
"arch": "polaris"
},
"model_uuid": "9a15c1cc-65e1-429a-9db6-f4ea6fbf822a"
},
"from_finalize": false
}
}
}
dtmfDetected
A digits property will be included in the payload containing the digits that were collected.
stt-low-confidence
A speech object will be included, see speechDetectedfor an example. The reason stt-low-confidence
indicates to the application that this transcript should be treated with caution as it is likely to
be an inaccurate reporting of what the user actually said.
timeout
No additional properties will be included in the payload.
error
A details property will be included in the payload with a description of the error.