TelonicDocs
English

Integrations

Speech recognition and voice

The speech recognition and voice synthesis providers Telonic works with, how they are chosen for your deployment, and why they always run in your region.

On this page
  1. What speech and voice providers do in your deployment
  2. Providers Telonic works with
  3. Why speech always stays in your region
  4. How providers are chosen
  5. The technical detail
  6. Permissions typical for this category
  7. In practice
  8. What your team controls
  9. Related

On a call, two things decide whether the conversation works: how well the agent hears the caller, and how the agent sounds when it replies. Speech recognition turns the caller's speech into text as they speak, in their language and dialect. Voice synthesis speaks the agent's reply in the voice chosen for your brand.

Telonic works with several specialist providers of each, and chooses them with you for your deployment. Speech recognition and voice synthesis always run in your chosen region.

What speech and voice providers do in your deployment

RoleWhat it means for youExample
Speech recognitionThe caller's words become text as they speak, including Gulf, Egyptian and Levantine Arabic, and Arabic and English mixed in one sentence"أبغى أغير الحجز to Friday, please" transcribed as spoken
Voice synthesisThe reply is spoken in a voice that suits your brand, in each language you useA calm, formal Arabic voice and a matching English voice for a luxury developer
Voice notesWhatsApp voice notes are transcribed before the agent repliesA policyholder describes the damage in a voice note

Providers Telonic works with

Each provider below offers speech recognition, voice synthesis, or both. Which provider fills which role in your deployment is decided with you, by language, dialect, voice and region.

IntegrationWhat the connection covers
DeepgramSpeech recognition, voice synthesis, or both, as chosen for your deployment
ElevenLabsSpeech recognition, voice synthesis, or both, as chosen for your deployment
Fish AudioSpeech recognition, voice synthesis, or both, as chosen for your deployment
RimeSpeech recognition, voice synthesis, or both, as chosen for your deployment
CartesiaSpeech recognition, voice synthesis, or both, as chosen for your deployment
SpeechmaticsSpeech recognition, voice synthesis, or both, as chosen for your deployment
AssemblyAISpeech recognition, voice synthesis, or both, as chosen for your deployment
SonioxSpeech recognition, voice synthesis, or both, as chosen for your deployment

Which providers and models are available in your region is confirmed during implementation, and named in your agreement.

Why speech always stays in your region

Speech recognition and voice synthesis always run on providers deployed in your chosen region. Call audio carries personal information that cannot be redacted (removed or masked) before it is transcribed, and a spoken reply may contain a customer's name or booking details, so both stay in-region. No out-of-region speech or voice provider is offered.

The only component that can be chosen outside your region is the language model, and only by your explicit choice. See Language models.

How providers are chosen

Providers are chosen per deployment, with you, during implementation. Speech recognition is chosen for how well it handles the languages and dialects your customers use, including switching between Arabic and English. Voice synthesis is chosen for the voice your brand needs: your team listens to shortlisted voices and picks one for each language. See Voices.

Telonic's contracts with every provider prohibit using your data to train or improve their models, and zero data retention arrangements (where the provider keeps no copy of what it processes) are used wherever the provider offers them. Every provider used for your deployment is listed in your sub-processors. See Sub-processors.

The technical detail

Speech recognition runs continuously while the caller speaks, so the transcript builds during the call. Providers are reached through a model routing layer (software that sends each request to the provider chosen for that role), so one can be changed without rebuilding the agent. Any change of provider is tested and needs your approval before release.

If a provider has an outage, failover only ever switches to an approved provider in your chosen region, and fallback providers are tested in advance. On-premises deployments run self-hosted speech models. Audio between your phone system and your deployment is encrypted with SRTP (Secure Real-time Transport Protocol) by default.

Permissions typical for this category

Every action in your systems is taken by Telonic's software, under the permissions your team has granted. Speech and voice providers receive only the audio or text for each turn of the conversation, and have no connection of their own to your systems. See Permissions.

In practice

Qamar Stays, a hotel group in Riyadh and Jeddah, sets up voice reservations in the Saudi Arabian region.

  1. Its guests call mostly in Gulf Arabic, often switching into English for dates and room types.
  2. Telonic tests the speech recognition providers deployed in the Saudi Arabian region on recordings of the group's own calls, used with its permission, and recommends one.
  3. The marketing team listens to shortlisted voices and chooses a warm, formal Arabic voice and a matching English voice.
  4. Every provider runs in the Saudi Arabian region, and each is named in the group's agreement.

What your team controls

  • The voice your customers hear, in each language.
  • Approval of every change of provider before it goes live.

Product names and logos are trademarks of their owners. Their mention shows systems Telonic connects to and does not imply partnership or endorsement.