Pricing · one all-in rate

One rate.
Everything in.

Most platforms quote a low headline, then stack the LLM, voice and telephony on top. We own the whole stack, so $0.10/min is your entire AI bill, billed by the second.

fig. 1 — one minute, itemised
USD

$0.10/ min · all-in

Speech recognition
akro
included
Reasoning
LLM
included
Voice
vui
included
Your own API keys
BYOK
none needed
Phone numbers
telephony
BYO or at cost
No per-component billing, and nothing to subscribe to — pay as you go, by the second.

§01 — Comparison

What a minute actually costs.

Not headline rates — the estimated cost of an actual minute of AI, with the parts each vendor leaves off their pricing page added back at the same rates for everyone. The caption under each bar shows the stack, because the gap between a headline and a bill is the whole point. Ours has nothing to add, which is why it is a line and not a range.

fig. 2 — estimated all-in cost per minute of AI

Agent platforms

A product you buy

Fluxions

ASR + reasoning + voice — nothing to add

$0.10 / min

ElevenLabs Agents

$0.08–0.10 voice + $0.02–0.05 for the LLM

$0.10–0.15 / min

Bland

all-in incl. line, plus $299–499/mo platform fee

$0.11–0.14 / min

Retell

$0.07–0.08 voice + up to $0.16 for the LLM

$0.12–0.30 / min

Synthflow

$0.09 voice + $0.02 LLM + $0.04–0.14 of your own keys

$0.15–0.25 / min

Vapi

$0.05 orchestration + 4–6 providers at cost

$0.14–0.30 / min

Model APIs

Parts you assemble yourself — plus the engineering to build and run it

OpenAI Realtime

$0.06–0.18 model + $0.01–0.02 of your compute to host it

$0.07–0.20 / min

Public pricing, checked September 2026, in USD. Add-ons modelled identically for every row: $0.02–0.05 for an LLM where the voice is billed separately, and $0.01–0.02 of your own compute to host a model-API agent. Token rates are converted at OpenAI’s published tokens-per-second, from a well-cached short call to a long one. Telephony is excluded throughout — it is pass-through or bring-your-own almost everywhere (~$0.0085/min US inbound on Twilio, less over SIP), and some deployments have no phone line at all. Engineering time to build and run an agent is not in any of these numbers.

§02 — À la carte

Just need one piece?

The same models, a la carte. Call transcription or speech directly from the API and pay only for what you render.

Transcription

$0.20/ hour of audio

akro-v1 speech-to-text with speaker diarization and non-speech events. Cheaper than OpenAI, Deepgram and AssemblyAI.

$0.01 minimum per request — the floor binds on clips under 3 minutes.

Speech (TTS)

$9/ 1M characters

Expressive VUI text-to-speech (~$0.41 per hour of audio) with voice cloning and non-verbal cues. Undercuts OpenAI; a fraction of ElevenLabs.

§03 — Enterprise

Need more?

High call volume, dedicated capacity or on-prem — let’s build the right plan.

Enterprise

Customtalk to us

Volume rates, dedicated GPU capacity, SLAs and on-prem.

  • Discounted volume per-minute rates
  • Reserved, dedicated GPU capacity
  • Uptime SLAs & priority support
  • On-prem / private-cloud option
Contact us

Two things drive price

Minutes
How long your agent talks.
Parallel streams
How many callers it can handle at once — each is a live conversation on its own line.

Enterprise adds reserved capacity for more parallel streams, and discounted volume rates.

§04 — On-prem

On-prem licensed deployment.

For security and data sovereignty: run the exact same models inside your own infrastructure, fully air-gapped. Annual licence with SLAs and dedicated support — ideal for healthcare, finance and the public sector.

§05 — FAQ

Questions.

What does the per-minute rate include?
The AI, end to end. Speech recognition, the language model and the voice all run on our own stack, so $0.10/min is the whole AI bill — no separate LLM or TTS charge, and no keys to bring. (Phone numbers, if you want us to provide them, are billed separately at cost.)
Does that include phone numbers / telephony?
The rate covers the AI — ASR, reasoning and voice. Phone connectivity is bring-your-own (connect your existing Twilio or SIP number) or we provision numbers at pass-through cost. You can also skip telephony entirely and stream over WebRTC from your web or mobile app. This is also why our rate is directly comparable to the raw model APIs below: none of those include telephony either.
What counts as a minute?
Connected conversation time, billed by the second. Idle or silent time is auto-hung-up so you do not pay for dead air.
What is a concurrent stream?
One live conversation at a time. If five callers talk to your agent at once you need five streams. Minutes are how long you talk; streams are how many talk at once — both affect price.
Do you really run your own models?
Yes. We own ASR, reasoning and voice end to end. That is why we can offer one all-in rate with no component stacking, and deploy the exact same models on-prem.
Can we run it on our own infrastructure?
Yes — licensed on-prem and private-cloud deployment is available for security and data sovereignty. The models run air-gapped inside your network; nothing leaves. Get in touch.

Still deciding?

Hear a real call.

Hear it first