In a voice agent rollout, the knowledge base takes days, the conversation rules are written line by line, the handover thresholds get debated. Voice selection usually happens in the last minute, by ticking the first option in a list. Then the first real calls are played back and someone says: this voice is not us.
The customer does not see your knowledge base in the first thirty seconds. They hear a voice.
Voice is a decision, not a preference
There are three separate components in voice selection, and they are usually compressed into a single "I like this one":
- Perceived gender and age. Who the voice sounds like. Expectations differ by sector; a technical service line and a beauty clinic do not carry the same voice.
- Pace. Words per minute and the pauses between sentences. A fast voice does not sound efficient, it sounds rushed.
- Tone. How sentences rise and fall, where the stress lands. The same sentence can mean "I am helping you" or "I am brushing you off" depending on tone.
These three can be tuned independently. When you dislike a voice, what usually needs changing is not the voice itself but its pace.
What to base the decision on
Not on your taste, but on the caller's state of mind. Three questions that work in practice:
- Is the caller calm or tense? Someone reporting a fault, chasing a delivery or missing an appointment calls tense. A tense caller gets more tense with a fast, cheerful voice.
- Is the conversation giving information or completing a task? An informing voice should be slower and clearer; a transactional voice should use shorter sentences.
- Do you already have a brand voice? If your ads, your on-hold announcement or your in-store messages carry a tone, the agent should not be disconnected from it.
Falling back to the default
A pattern we see often: no voice was ever selected on the account, so the agent speaks with the system default. Nobody did anything wrong — a field was simply left empty. But the customer hears that voice.
The first thing to do after setup is to call your own agent and listen to your own voice. Five minutes of this prevents months of speaking in the wrong tone.
What happens when the voice changes
A voice change is not retroactive: old recordings keep the old voice, new conversations start with the new one. Changing the voice in the middle of a campaign can mean calling the same customer twice with two different voices.
Practical rule: fix the voice before a campaign starts and leave it alone for the duration.
Do you need more than one voice
Usually not. One voice per business is more valuable for recognisability. Two cases justify splitting:
- Different languages. The same voice does not have to sound natural in both Turkish and English; a separate voice can be chosen per language.
- Different functions. If outbound reminders and inbound support carry different personas, the voice can differ too. Bear in mind this means running two agents, which increases management overhead.
Decide by listening, not by reading
Picking from a list by names and labels is misleading. Listen to the same line in three different voices — and use a real sentence from your own knowledge base, not a demo script. "I have booked your appointment for Tuesday at 2pm and I will send you a reminder" tells you far more about fit than any marketing sample.
You can see how the voice agent works on the AI call center page, or hear it directly by having it call you from the demo page.
Tomorrow we look at the first five seconds of a conversation: what the greeting should say, and what it should not.
