Agents
Models and providers
The speech recognition, language model and voice providers behind every conversation, how they are chosen for your deployment, and how that choice follows your region.
On this page
Every conversation an agent holds relies on AI models from specialist providers: one to turn speech into text, one to understand the conversation and decide the reply, and one to speak that reply. Telonic chooses these providers with you for each deployment, based on the languages and dialects your customers use, the voice that suits your brand, and the region your data must stay in. This page explains what each provider does, how the choice is made, and how it follows the region rule that applies to your deployment.
What each provider does in a conversation
| Provider type | What it does | Used on |
|---|---|---|
| Speech recognition | Turns the caller's speech into text as they speak, in their language and dialect | Voice calls, browser voice and WhatsApp voice notes |
| Language model | Understands the conversation and decides the reply, within the agent's limits, using facts retrieved from your systems and documents | Every channel |
| Voice synthesis | Speaks the agent's reply in the voice chosen for your brand | Voice calls and browser voice |
On WhatsApp, SMS, email and web chat, only the language model is involved in each reply, because the conversation is already in text. The exception is a WhatsApp voice note, which is transcribed first. Images and documents that customers send are read by a model running in your chosen region.
The language model does not work alone. Telonic's own software sits around it: the industry model for your sector (the things your business runs on and the rules attached to each), your limits, the customer record and the connections to your systems.
The language model decides how to express an answer. The facts in the answer come from your systems and approved documents, and the actions available to it are the ones your team has permitted. See How an agent works.
The technical detail
Speech recognition runs continuously while the caller speaks, so the transcript builds during the call rather than after it. The agent starts forming its reply while the caller is finishing their sentence, which keeps the conversation natural rather than delayed.
Telonic uses models through a model routing layer (software that sends each request to the provider chosen for that role), so the provider behind each role can be changed without rebuilding the agent. Your industry model, limits, memory and integrations stay the same when a provider changes.
How providers are chosen for your deployment
Providers are chosen per deployment, with you, during implementation. Three things decide the choice.
| What decides it | Why it matters | Example |
|---|---|---|
| Language and dialect coverage | Speech recognition and voice quality vary by language and dialect. The right provider for Gulf Arabic may differ from the right one for Egyptian Arabic or English | A Riyadh hotel group whose callers mostly speak Gulf Arabic, switching into English for booking terms |
| Voice | The voice is part of your brand. Voices vary by provider in character, gender and how natural they sound in each language | A luxury developer choosing a calm, formal Arabic voice and a matching English voice |
| Region | The provider must be deployed in the region your data stays in | An insurer whose data must remain in Saudi Arabia |
The choice can be changed after go-live (the point at which the agent starts handling real customers). Any change of model or provider is tested against a set of conversations drawn from your industry before it reaches your customers. See Testing before every release.
How the choice follows your region
All processing for a deployment, including speech recognition, the language model and voice synthesis, runs in the region you choose, using models and providers deployed in that region.
A provider that processes outside your chosen region is used only if you explicitly choose it, and only a language model can be chosen this way. If you do, personal information is redacted (removed or masked) before any data reaches it. Redaction covers names in Arabic and Latin script, phone numbers, email addresses, national identity numbers (including Emirates ID and Iqama numbers), passport numbers, payment card numbers, bank account numbers and IBANs, dates of birth and street addresses.
Speech recognition and voice synthesis always run in your chosen region, and no out-of-region speech or voice provider is offered. Audio cannot be redacted before it is transcribed, and the text a voice speaks often contains a customer's name. When text goes to an out-of-region language model, each redacted item is replaced with a placeholder, and the real values are restored in the reply inside your deployment.
| Your choice | Where each provider runs | What leaves your region |
|---|---|---|
| In-region only | Speech recognition, language model and voice synthesis all run in your chosen region | Nothing |
| An out-of-region language model, chosen explicitly by you | Speech recognition and voice synthesis in your region. The language model outside it | The conversation text, the facts retrieved for the reply and the agent's instructions, with personal information redacted first |
The provider choice for your deployment, including any provider outside your region, is recorded in your agreement and in the list of sub-processors (companies that process personal data on Telonic's behalf) for your deployment. See Sub-processors.
The providers Telonic works with
Telonic works with the following providers. Not every provider is deployed in every region, which is one reason the choice is made per deployment. Which providers and models are available in your region is confirmed during implementation, and named in your agreement.
| Provider type | Providers |
|---|---|
| Language models | OpenAI, Anthropic, Google Gemini, Mistral AI, Meta (Llama models), TII (Falcon models) |
| Speech recognition and voice synthesis | Deepgram, ElevenLabs, Fish Audio, Rime, Cartesia, Speechmatics, AssemblyAI, Soniox |
Each speech provider on the list provides speech recognition, voice synthesis, or both.
What providers do with your data
Providers process conversation data only to perform their role in your deployment: transcribing, deciding a reply or speaking it. Telonic's contracts with every provider prohibit using your data to train or improve their models. Zero data retention arrangements, where the provider keeps nothing after processing, are used wherever the provider offers them.
What Telonic carries from one deployment to the next is industry knowledge: workflows, edge cases and the test sets used to check quality. It never carries your customer data.
What Telonic builds, and what it uses
Telonic builds everything that makes foundation models (the large, general-purpose AI models that providers build) useful in your business: the industry model for your sector, the workflows, the limits, the customer record, the connections to your systems, and the tests that check every change. The foundation models themselves come from established providers. Telonic does not train them.
That means improvements in the underlying models reach your deployment without a rebuild. Before any new model is used in your deployment, it is tested against conversations drawn from your industry, and adopted only when it performs at least as well as the one it replaces.
If a provider has an outage, your deployment can switch to another provider you have approved. This is configured during implementation, and the fallback provider is tested in advance, like any other change. Failover only ever switches to an approved provider in your chosen region.
In practice
A hotel group in Riyadh and Jeddah sets up voice reservations with Telonic.
- Its guests call mostly in Gulf Arabic, often switching into English for dates and room types. Some call in Egyptian Arabic or English.
- With the group's IT team, Telonic chooses a speech recognition provider that handles Gulf and Egyptian Arabic and mixed Arabic and English well, deployed in the Saudi Arabian region.
- The marketing team listens to shortlisted voices and chooses a warm, formal voice in Arabic, and a matching voice in English.
- The group keeps every provider in-region. Nothing leaves the Saudi Arabian region.
- Some months later, a newer language model becomes available in the region. Telonic tests it against the group's reservation conversations, confirms it performs at least as well, and the change goes live after the group's approval.
What your team controls
- Whether any provider outside your region may be used, and for which role.
- The voice your customers hear, in each language.
- Approval of any change of model or provider before it goes live. Every such change is approved by you before release.
Related
- How an agent worksAgents
- Hosting and data residencySecurity and data protection
- How data flows through a conversationSecurity and data protection
- VoicesLanguages
- Sub-processorsSecurity and data protection
Product names and logos are trademarks of their owners. Their mention shows systems Telonic connects to and does not imply partnership or endorsement.