TelonicDocs
English

Integrations

Fish Audio

How Fish Audio's speech models can be used in your deployment, always in your chosen region, and what they receive.

On this page
  1. What a speech provider does in your deployment
  2. How it connects
  3. What the provider receives
  4. In practice
  5. Related

Fish Audio is one of the speech providers Telonic works with. Speech providers supply the models behind every spoken conversation: speech recognition (turning a caller's speech into text as they speak) and voice synthesis (speaking the agent's reply in the voice chosen for your brand). Fish Audio provides speech recognition, voice synthesis, or both. Which of its models your deployment uses, and for which role, is decided with you during implementation.

What a speech provider does in your deployment

RoleWhat it means for youExample
Speech recognitionCallers are understood as they speak, and the transcript builds during the call rather than after itA caller's words transcribed as they speak on a voice call or in browser voice, and a WhatsApp voice note transcribed when it arrives
Voice synthesisThe agent's reply is spoken in the voice chosen for your brand, in each languageA calm, formal Arabic voice and a matching English voice
Custom voiceA voice unique to your brand, built from a short recording of a speaker you choose. Configured during implementationA voice recorded by a presenter your marketing team selects

How it connects

Speech recognition and voice synthesis always run in your chosen region. No speech or voice provider outside your region is offered, whatever you choose for the language model. Speech providers are chosen for your deployment on two things: their coverage of the languages and dialects your customers use, and the voices that suit your brand.

Telonic's contract with every provider prohibits using your data to train or improve their models, and zero data retention arrangements (the provider keeps no copy of what it processes) are used wherever the provider offers them. Any change of speech provider or model is tested on conversations drawn from your industry and needs your approval before release. If a provider has an outage, failover switches only to another approved provider in your chosen region, tested in advance and configured during implementation.

What the provider receives

DataNeeded forDefault
Call audio, as the caller speaksSpeech recognitionOnly if the provider is chosen for recognition. Processed in your region
The text of each spoken replyVoice synthesisOnly if the provider is chosen for synthesis. Processed in your region
A recording of the speaker you chooseBuilding a custom voiceOnly if you commission a custom voice, with the speaker's written consent
Access to your systems or the customer recordNot neededNot given
Your data, for training the provider's modelsNot permittedProhibited by contract

Calls are recorded only once the caller has given consent, and nothing is stored without your permission. See Models and providers.

In practice

Nakhla Travel, a tour operator, adds voice to its booking line.

  1. Its customers call in Arabic and English, and many are tourists who speak English as a second language.
  2. Nakhla's marketing team listens to shortlisted voices from providers deployed in its region, Fish Audio among them.
  3. It chooses a warm, clear voice in each language, and the choice is recorded in its agreement.
  4. The voices are tested on booking calls from the travel industry before they reach customers.

Product names and logos are trademarks of their owners. Their mention shows systems Telonic connects to and does not imply partnership or endorsement.