We build the whole voice stack — recognition, reasoning and speech — from the weights up. Small enough to run on a phone. Cheap enough to build a business on.
$0.10/ min
for a whole voice agent — recognition, LLM and voice, all-in
300Mparams
in vui-nano, speaking in real time on the iPhone GPU
$0.20/ hour
of transcription, with speakers and non-speech events
§01 — Phone agents
Agents built on our API take bookings, triage call-outs and fill appointment books for real businesses — you heard three of them in fig. 1.
$0.10 a minute, all-in. Numbers are bring-your-own or provided at cost.
akro — listens
Speech to text — who is speaking, and the breaths, laughs and hesitations between words.
LLM — thinks
Decides what to say, and calls your tools — bookings, lookups, web search.
vui — speaks
Text to an expressive voice, streamed back while it is still being written.
§02 — vui for iPhone
vui runs speech recognition, the language model and a 300M-parameter voice entirely on the iPhone GPU. No server does the thinking, so nothing you say is sent anywhere.
Free during beta · iPhone with iOS 26+ · ~1.2 GB one-time download
§03 — API
We train our own recognition, reasoning and voice, so nothing is stacked on top — no separate LLM bill, no TTS bill, no keys to bring. Billed by the second.
curl -X POST https://api.fluxions.ai/vui/v1/tts \-H "Authorization: YOUR_API_KEY" \-H "Content-Type: application/json" \-d '{"voice": "maeve.h8ff7e07da","input": "[sigh] fine, I will say it one more time."}' \--output speech.wav
§04 — Licensing
The same models, air-gapped in your infrastructure or on your own devices. Annual licence with SLAs and dedicated support.