Configure TTS voices
Every Say Message node in your IVR — and the Customer-first greeting on each queue — uses text-to-speech to read the text you typed. The Voices tab supplies the default voice for each language. Say Message and Gather Input nodes can override that default from the node configuration panel.
Before you start
- A phone number provisioned with at least one IVR flow assigned.
- A decision on which languages your callers speak. Most channels need only one; multi-language operations need a voice per language.
Steps
- Open Settings → Voice.
- Select your phone number in the sidebar, then go to the Voices tab.
- Pick the language you want to configure from the language picker.
- Pick a voice from the catalog for that language. Each language row has a preview button — preview the selected voice before saving.
- If the preview shows an amber warning or notice, the requested voice was coerced to a Twilio-renderable fallback. Preview uses the same compatibility logic as live calls, so the saved voice should match what callers hear.
- (Optional) Adjust speech rate and pitch if the catalog exposes those controls for the selected voice.
- Repeat for each additional language you serve.
- Save.
The IVR’s Say Message nodes, Gather Input prompts, and Customer-first greetings now speak in the chosen default voice when the call’s language matches. To use a different language or voice for a specific Say Message or Gather Input node, open that node’s configuration panel and set Spoken language. Once a language is selected, use Voice to choose the node-specific voice.
How the call’s language is decided
A call’s language is determined in this order:
- For Say Message and Gather Input nodes, if the node has Spoken language selected in its configuration panel, that node uses the selected language.
- If the IVR has a Language Selection node and the caller picked an option, the call’s language is set to that choice.
- Otherwise, the call uses the channel’s default language as configured on the Voices tab.
You can read or set the call’s language mid-flow via a Set Variable step (using the call’s language variable) — useful when an API call earlier in the flow returned the caller’s preferred language.
Verify it worked
Dial your number and walk through an IVR step that includes a Say Message or Gather Input prompt. The message should be spoken in the voice configured for that language, unless the node has its own Voice override.
If the voice sounds wrong, double-check that the call’s language or node-level Spoken language matches the language whose voice you edited — if you only configured an English voice but the call or node ran in French, you’ll hear a fallback.
If a preview or live call uses a fallback, Atender applies the same Twilio compatibility checks in both places. Watch for an amber warning or header-driven notice in preview that explains when the requested voice was coerced to a renderable voice.
Atender preserves regional accent choices when they are compatible, such as using an en-GB voice for an en-US caller or message. If the voice’s base language does not match the caller or message language, Atender falls back to a compatible voice for that language. This applies to live queue announcements and to IVR Say playback where Twilio renders <Say>.
Catalog notes
- The voice catalog covers 33 languages.
- The catalog includes Twilio-renderable Amazon Polly voices plus Google Chirp3-HD voices where Twilio supports them.
- Each language typically has both male and female voice options.
- Some languages have neural or Chirp3-HD voices that sound more natural; others only have standard voices. Preview before committing.
- The Voices tab sets the default voice for each language on the channel. Say Message and Gather Input nodes can use different voices in the same flow by setting Spoken language and then Voice in the node configuration panel.
Troubleshooting
-
Symptom: The voice sounds robotic and you remembered it sounding better in another tenant. Fix: The other tenant likely picked a higher-quality voice. The voice dropdown groups options under Google and Amazon, and the higher-quality Amazon voices are named with a (Generative) or (Neural) suffix — pick one of those if available.
-
Symptom: A Google Chirp3-HD voice is not available for the language or locale you need. Fix: Choose one of the listed Amazon Polly voices for that language instead. Don’t assume every Chirp3-HD voice can be used in every locale — the catalog only shows voices that Twilio can render for that language.
-
Symptom: A Say Message reads numbers wrong (e.g. “twelve oh five” instead of “twelve-oh-five” for 12:05). Fix: TTS engines vary in how they interpret formatted numbers. Spell out tricky values in the Say Message text — “twelve oh five” reads more reliably than “12:05”.
-
Symptom: You can’t pick a voice for a language you need. Fix: That language isn’t in the catalog. Use the closest related language as a workaround (for example, German (DE) for an Austrian or Swiss caller), or upload a pre-recorded Play Audio node for that specific message.