What AI Voice Agents Actually Cost Per Call

The short version
An AI voice call has three moving costs: turning speech into text, deciding what to say, and turning the reply back into speech. On the calls we have measured, the third one dominates — and that has practical consequences for how you design an agent.
Where the money actually goes
Measured across real support calls, a typical three-minute conversation breaks down roughly like this:
- Speech-to-text: around 14%. Priced per hour of audio, and the audio is only as long as the call.
- The language model: around 1%. Genuinely negligible at current prices.
- Text-to-speech: around 84%. Priced per character the agent speaks.
That last number surprises most people. The intuition is that the "AI" is the expensive part. It is not — the voice is.
Three consequences worth designing around
1. Using a better language model barely changes the bill
If the model is one percent of the cost, moving to a stronger one costs almost nothing in absolute terms. There is rarely a good reason to run an agent on a weaker model to save money.
2. Short answers are a cost control
Every character the agent speaks is billed. An instruction to keep replies to two or three sentences is not only better conversation design, it is directly cheaper — and it makes the agent feel faster, because callers are not waiting through a paragraph.
3. The voice you choose changes the economics
Faster, lighter speech models can cost roughly half as much per character as premium ones. Whether that trade is worth making depends entirely on the language and the accent you need; for some languages the cheaper model is indistinguishable, for others it is obviously worse. This is worth testing with your own script rather than assuming.
What this means for buying
When a vendor quotes a per-minute rate, the useful questions are which voice model it assumes, whether the price changes if you need a specific accent, and whether long agent responses are billed differently. A quote that cannot answer those is a quote that may move once you are live.
The honest caveat
Cost per minute is the easy number. The one that actually determines whether an agent pays for itself is how many calls it resolves without a human. Published enterprise benchmarks put median containment near 41%, while vendors commonly advertise 70 to 80%. A cheap agent that resolves nothing is not cheap.
If you want to work through what an agent would cost for your call volume and your languages, tell us what you are trying to automate and we will give you real numbers.
Ready to get started with VoxClouds?
Cloud PBX, virtual numbers, and cheap international calls.