
A voice AI call has four separately metered cost components: speech-to-text per audio minute, language model per token, text-to-speech per character or minute, and telephony per call minute. Realistic all-in prices across the market land between roughly $0.13 and $0.50 per minute, depending on vendor and configuration (dated sources below). What you actually pay depends less on the headline rate than on who provides, and who bills, each stage of the pipeline.
This guide breaks down the cost stack, shows why advertised rates understate real bills, and explains how per-stage sourcing changes the math. For PeachDesk's own plans and rates, see PeachDesk pricing.
What are the four costs inside a voice AI minute?
Every minute of an AI-handled call consumes four services, each metered on a different basis. Speech-to-text is billed per audio minute. The language model is billed per token, and a talkative agent burns tokens in both directions. Text-to-speech is billed per character or per generated minute. Telephony, the call leg itself, is billed per call minute by the carrier, and it stays metered in every architecture, including fully self-hosted ones.

- Speech-to-textper audio minute
- Language modelper token, usually the largest share
- Text-to-speechper character or minute
- Telephonyper call minute, always metered
Two consequences follow. First, a single per-minute number always hides a composition; two quotes at the same price can buy very different models. Second, each component can, in principle, be sourced separately, which is the opening that cost control exploits.
Why is the advertised rate never the real rate?
Vendors advertise their most quotable number: a platform fee, a floor, or a starting rate. The realistic bill adds model costs, per-minute add-ons such as knowledge bases or guardrails, and monthly items such as phone numbers and extra concurrency. The gap between the two is large enough that headline rates should be treated as marketing, not budgets.
| Vendor | Advertised headline | Realistic all-in per minute | Source |
|---|---|---|---|
| Vapi | $0.05 (hosting only, excludes model costs) | $0.15 to $0.50 | vapi.ai pricing page; third-party estimates (HappyRobot, Rezora, n8nlab), directional |
| Retell AI | $0.07 (advertised blend floor) | $0.13 to $0.31 | retellai.com pricing page; Cekura third-party analysis |
| Bland AI | $0.11 to $0.14 bundled, tiered, plus telephony | Approximately the headline plus telephony and platform fee | bland.ai pricing page |
| Synthflow | About $0.08 to $0.09 starting (vendor-cited) | $0.15 to $0.24 effective | Secondary reporting (Layer3 Labs), indicative |
Fairness notes. Retell AI publishes a full component rate card, the most transparent pricing surface in this set. Vapi supports BYOK, which can zero the model portion of its bill, though the $0.05 hosting floor still applies. None of these vendors publishes an architectural path to near-zero variable inference cost: every meter has a floor above zero.
How did compliance become a revenue line?
Data controls are sold separately in this market. Vapi charges $2,000 per month for its HIPAA add-on and $1,000 per month for Zero Data Retention (vapi.ai pricing page, accessed 2026-07-30). Bland AI and Synthflow gate their compliance offerings behind enterprise contracts, while Retell AI is the outlier with self-serve BAA signing, though its plan gating is ambiguous across its own materials (retellai.com, accessed 2026-07-30).
PeachDesk takes a different structural approach. In Local mode, audio and text for locally served pipeline stages stay inside your own environment, so data control is a property of the architecture rather than a monthly add-on. That is a deployment fact, not a compliance certification; regulated teams should still run their own review.
How does per-stage sourcing change the math?
PeachDesk classifies every provider request in a call into one of three sourcing modes, set per stage, per agent. The sourcing mode decides both what a stage costs and where its data goes.
- Frontier mode. PeachDesk-managed cloud providers serve the stage. Always metered: you pay per use for the strongest general-purpose models, with no infrastructure to run.
- BYOK. Your own provider keys serve the stage. You pay your providers directly at your negotiated rates, and calls on your own keys never consume PeachDesk credits. Read more about voice AI BYOK.
- Local mode. A full open-model stack on your own GPUs serves the stage. There is no per-minute provider charge for locally served inference; you pay for GPU capacity. Variable per-minute inference cost becomes fixed infrastructure cost, and audio and text stay in your environment.
Because sourcing is per stage, a single agent can run a frontier model where reasoning quality matters most and leading open models where volume and data sensitivity dominate. Two things never change: telephony stays metered by the carrier in every configuration, and Frontier-mode stages stay metered per use. No configuration makes every call free.
The same mechanics decide who owns the margin. An agency reselling metered minutes inherits the vendor's per-minute floor as its cost base, and the markup it can add compresses as volume grows. An agency that controls sourcing owns its unit economics instead, which is the model behind PeachDesk as a white-label voice AI platform.
Compare costs on your own call profile
Bring your minutes, your stages, and your providers. PeachDesk shows an estimated cost per minute for an agent before you run it, and records cost, price, margin, and sourcing mode on every metered call after you do.