AI Voice Agent Platforms Compared: Real Pricing and How to Choose
NextPhone, Kosmo, Fonema, Retell, Bland and Vapi compared at their real per-minute cost, Latin American Spanish quality, integrations, and the decision framework we use at Mintec.
There is no "best AI voice agent platform": there are four routes with very different real prices, and most of the comparisons you find on Google omit exactly the costs that matter. Vapi advertises $0.05 per minute and ends up costing $0.13–0.33 when you add models and telephony. NextPhone charges a flat $199/month with unlimited calls. Fonema has no monthly fee, but its telephony runs $0.015 per minute in Mexico and $0.07 in Colombia or Peru. Choosing wrong is not just overpaying: it's an agent that sounds off to your customers, a platform that won't talk to your CRM, or a project that dies within 90 days. This article compares the six platforms we keep running into in Latin American implementations, with pricing verified in August 2026, and the decision framework we use at Mintec.
Why the per-minute price they advertise is only half the truth
Start with the most common trap. Global platforms sell "from $0.05/min", but that number almost never includes what makes the call actually work:
- The LLM, which decides what to say: $0.02–0.30 per minute depending on the model.
- Speech-to-text (STT): ~$0.01/min.
- Text-to-speech (TTS): $0.05–0.15/min if you want quality voices.
- Telephony (the number and call termination): $0.01–0.05/min in the US — but in several Latin American countries it costs more.
An independent analysis by CloudTalk (August 2026) puts a real Vapi deployment at $0.30–0.33 per minute — six times the advertised rate. Bland admits it on its own pricing page: it says typical stacks run $0.13–0.30 on Vapi and $0.11–0.25 on Retell, and that its flat $0.09 includes the models. Everyone pulls toward their own corner, but the real market range is clear: $0.07–0.35 per minute depending on the route you take.
And in Latin America there's an extra cost no English-language comparison mentions: per-country telephony. Fonema publishes its rates and the spread is brutal: $0.0152/min in Mexico, $0.014 in the US, but $0.0537 in Argentina, $0.0428 in Chile and $0.07 in Colombia and Peru. If you answer calls from multiple countries, your "per-minute cost" can double or triple on termination fees alone.
The 4 routes to an AI voice agent (with real pricing)
In the projects we've seen at Mintec, every platform falls into one of four routes. Each solves a different problem and has a different cost structure.
Route 1: Flat-rate receptionist SaaS — NextPhone, Kosmo
These are the "Netflix of voice agents": you pay a monthly plan, configure it in a day, and never think about minutes or models. NextPhone charges a flat $199/month with unlimited inbound calls, 24/7, according to its 2026 pricing guide — one of the few platforms with no per-minute meter and no overage charges. In Mexico, Kosmo bills in pesos: Basic $1,290 MXN (200 min), Starter $2,800 (900 min), PRO $4,800 (1,900 min) and PYME $6,800 (2,900 min) per month.
The rule we apply: if your business takes fewer than ~500 call minutes a month and nobody on the team wants to touch configuration, this route is correct. The cost is predictable and setup takes hours, not weeks. We covered what a voice agent can do and what it costs you not to have one; this route is the fastest way to start.
Route 2: Usage-based Latin American omnichannel — Fonema
Fonema is the option we like most for businesses that want voice and WhatsApp in the same agent, with genuine Latin American Spanish: 200+ regional voices and sub-1.2-second latency. Its pricing model is "pay only for what you use": no fixed monthly fee, per-minute telephony by country ($0.0152 Mexico, $0.07 Colombia), extra interactions (WhatsApp, email, actions) at $0.01, and additional concurrency lines at $15/month each.
For a Mexican SMB with 1,000 minutes a month, the math is brutally cheap: ~$15/month in telephony. The catch is that the savings depend on your country and how many simultaneous calls you need; the first 5 concurrency lines are included.
Route 3: Global low-code per-minute platforms — Retell, Bland
When you need more control (CRM integrations, warm transfers, outbound campaigns) but don't have a development team, enter the low-code platforms that bundle everything into one per-minute number. Retell charges $0.07+/min with models included in the base plan, unlimited concurrency, and native integrations with n8n, HubSpot and Twilio. Bland charges $0.09/min and includes LLM, STT and TTS in that rate; telephony is billed separately at cost, and they promise a first agent deployed within a day.
These are the platforms we pick when the client already has a technical owner for the project (even a semi-geeky ops person) and the agent needs to talk to the CRM, not just answer calls. Real cost lands between $0.07 and $0.25/min depending on voice quality and the model you choose.
Route 4: Developer orchestrator — Vapi
Vapi is the most powerful and the most expensive in practice: it charges $0.05/min for the platform and passes model costs through to you (you can bring your own API keys). With STT, LLM, TTS and telephony, real cost lands around $0.13–0.33/min, and enterprise deployments typically need $40,000–70,000 annual budgets, per CloudTalk. In exchange you get total control: choose every model, tune latency, custom tools, 100+ languages.
This route only makes sense with a development team that will maintain it. Otherwise it's the most expensive way to learn how to build voice agents. For businesses that do want to build their own stack from scratch, we published the full chatbot architecture that applies the same principle to the voice channel.
What it really costs: the 1,000-minute table
For an honest comparison, here's what an average business (1,000 call minutes per month, August 2026) would pay:
| Platform | Structure | Real cost ~1,000 min/mo | Latin Spanish | Best for |
|---|---|---|---|---|
| NextPhone | Flat $199/mo, unlimited calls | $199 | Yes, configurable | Zero setup, low-mid volume, predictable budget |
| Kosmo (MX) | MXN plans with minutes | $2,800–4,800 MXN per plan | Regional Mexican | SMBs billing in pesos that want a fixed plan |
| Fonema | Usage only (per-country telephony + $0.01 interactions) | ~$15–70 | 200+ regional voices, <1.2s | Voice + WhatsApp in one agent, LatAm |
| Retell | $0.07+/min, models included | $70–250 | Multi, needs local testing | Low-code technical owner, CRM integrations |
| Bland | $0.09/min all-in + telephony | $90–140 | Multi, needs local testing | Scale on a single per-minute rate |
| Vapi | $0.05/min + models + telephony | $130–330+ | Multi (bring your own models) | Development teams, custom use cases |
If you already run a human call center and just want AI on top, the route is different: platforms like CloudTalk ($25–49 per user/month) sell voice agents as an add-on from $99 for 200 minutes. That's the right choice when AI is one more layer of a phone system you already pay for.
The 5-question decision framework
This is the order we follow at Mintec when a client asks for "a voice agent":
- Is there anyone technical on the team? No → Route 1 or 2. Yes, low-code → Route 3. A development team → Route 4.
- How many real call minutes do you have per month? Measure one week and multiply by 4.3; almost nobody does this and everyone overestimates. Under 500 → flat rate. 500–3,000 → usage-based. Over 3,000 → negotiate volume with any platform.
- Do you need voice and WhatsApp in the same agent? In markets like Mexico and Colombia, WhatsApp is the primary channel; if the voice agent can't continue the conversation in chat, you'll duplicate processes. That's where the omnichannel route (Fonema) wins, or a well-built human-machine handover.
- Which countries do you answer calls from? Mexico and the US have cheap telephony on almost any platform. If Colombia, Peru or Argentina are in play, check per-country termination rates before signing: they can triple your real per-minute cost.
- Are you in a regulated industry? Health, finance or personal data changes the rules: you need HIPAA/BAAs or data-residency agreements, and those only exist in enterprise tiers. Per-minute savings are irrelevant next to a fine.
The combination of these answers defines the route. There is no answer "by platform" — only by business profile.
What vendors won't tell you (and clients discover in production)
After implementing these platforms, these are the surprises that keep coming up in our diagnostics:
- Advertised latency is not real latency. Vapi claims <500ms, but users report 6–7-second waits on non-optimized configurations. Fonema and the receptionist SaaS publish more conservative numbers (1.2s) that hold up better in practice.
- Concurrency is the hidden cost. Vapi includes 10 simultaneous lines on pay-as-you-go ($10/line/month extra); Fonema includes 5. If your business has call peaks (lunch hour, campaigns), the concurrency cap matters more than the per-minute price.
- Data retention costs money. Vapi keeps call history for 14 days on the base plan; extended retention or "zero data retention" is a paid add-on ($1,000/month). If your industry regulates what you keep or delete, that's a requirement, not an extra.
- 72% of customers want to know they're talking to an AI (Salesforce, 2026). Hiding it isn't free: trust drops and calls degrade. The best agents introduce themselves as AI in the first second and win on speed, not deception.
- Agentic project cancellation rates are brutal. Gartner predicts over 40% of agentic AI projects will be canceled by 2027, and trust in fully autonomous agents fell from 43% to 27% in a year (Capgemini). It's almost never the model's fault: it's picking the wrong platform for the business profile, or launching and forgetting maintenance (the famous Day 2 problem).
The 20-call test before you sign
Whatever the route, don't sign anything without the field test: record 20 real calls (or have the vendor generate them) where the agent responds to customers with the accent and vocabulary of your region, not the demo. Listen to how the transfer to a human sounds (73.8% of well-handled calls end in correct routing, and that's success, not failure), measure how many calls it resolved alone and how many it dropped, and compare the real cost against your current receptionist bill.
Choosing a voice platform isn't choosing the cheapest per minute: it's choosing the one that survives contact with your real operation. If you want, we can run that diagnostic with you — at Mintec we evaluate volume, countries and processes before touching a single API. And if what you need is the full agent — voice, WhatsApp, CRM and chatbot automation — that evaluation is the first step of the project.
Frequently Asked Questions
How much does an AI voice agent cost per minute in 2026?
Advertised rates run from $0.05 to $0.09 per minute (Vapi, Retell, Bland), but the real cost including models and telephony lands between $0.13 and $0.33 per minute. Flat-rate alternatives like NextPhone charge $199/month for unlimited calls, and in Latin America Fonema charges usage only: from $0.014/minute in the US up to $0.07/minute in Colombia and Peru.
Which AI voice platform is best for Latin American Spanish?
Fonema is purpose-built for Latin American Spanish with 200+ regional voices and sub-1.2-second latency. Kosmo bills in Mexican pesos with plans from $1,290 MXN/month. If you pick a global platform (Retell, Bland, Vapi), test with real accents from your region before committing: neutral Spanish does not sound neutral to your customers.
Vapi, Retell or Bland: which should I choose?
Vapi is an orchestrator for development teams: $0.05/min platform fee plus model costs (STT, LLM, TTS) and telephony, for a real cost of $0.13–0.33/min. Retell ($0.07+/min) and Bland ($0.09/min) bundle the models into one number and are low-code. If you have no technical team, start with a flat-rate SaaS (NextPhone, Kosmo) or an omnichannel platform like Fonema.



