Voice synthesis in 2026 splits into two camps: the hosted services that bill you per character or per audio token, and the open-weight models you run on your own GPU. The hosted side gets you professional voice clones in a few clicks and a tier that handles your traffic; the local side gets you unlimited generation, full control over the voice, and a licence that does not forbid you from selling the output. Neither camp has swept the other.
This page is the hosted-vs-local comparison. The local-only side of it - VRAM tiers, STT, and the open TTS ranking - lives in the companion Best Local Speech Models by VRAM Tier guide; this page covers pricing, licence, and the catches.
The pricing actually looks like this
Hosted TTS billing varies more than it looks at first glance. ElevenLabs charges by the character of input text; Google bills audio by the audio token; OpenAI mixes both. Below is the lowest published tier from each vendor’s own pricing page, accessed 2026-09-28.
| Service | Lowest paid tier | Price | Unit | What you get |
|---|---|---|---|---|
| ElevenLabs Starter | $5/mo (annual) or $6/mo | $5-$6 | monthly subscription | 30,000 credits; instant voice cloning; commercial use |
| ElevenLabs Creator | $18.33/mo (annual) or $22/mo | $18-$22 | monthly subscription | 121,000 credits; instant + professional cloning |
| ElevenLabs Pro | $82.50-$99/mo | $82-$99 | monthly subscription | 600,000 credits; higher concurrency |
| OpenAI tts-1 | pay-as-you-go | $15 | per 1M characters | Standard TTS, 6 preset voices |
| OpenAI tts-1-hd | pay-as-you-go | $30 | per 1M characters | Higher-fidelity variant |
| OpenAI gpt-4o-mini-tts | pay-as-you-go | $12 | per 1M audio tokens (output) | Controllable speech via prompt |
| Gemini 2.5 Flash Preview TTS | pay-as-you-go | $10 | per 1M audio tokens (output) | Lower latency |
| Gemini 3.1 Flash TTS Preview | pay-as-you-go | $20 | per 1M audio tokens (output) | Lower latency |
| Gemini 3.8 Flash TTS | pay-as-you-go + free tier | $9 (to 2026-12-31) / $18 (from 2027-01-01) | per 1M audio tokens (output) | Studio-grade fidelity |
ElevenLabs is the only one of the three that is subscription-based with rolled-over character quotas; OpenAI and Gemini are pure pay-as-you-go, billed to the second. Gemini 3.8 Flash TTS doubles on 2027-01-01 according to its own pricing page - if you build on it, that is the date your unit economics change.
A second billing wrinkle: 1M audio tokens at Gemini is “25 tokens per second of audio,” so a minute of generated speech is 1,500 tokens. At $9/1M tokens, that is roughly $0.0135 per minute of audio in 2026, before the rise.
Local TTS options and what each one costs you to run
The hosted vendors have one price per output. The local models have a one-time hardware cost and an electricity bill. The catch is that you need a GPU, and not every model fits on every card.
| Model | Params | License | Voice cloning | Footprint |
|---|---|---|---|---|
| Kokoro-82M | 82M | Apache-2.0 | No (54 fixed voices, 8 languages) | Lightweight; footprint not published on card |
| Resemble AI Chatterbox | 500M | MIT | Yes (zero-shot) | VRAM not published on card |
| Voxtral-4B-TTS-2603 | 4B | CC-BY-NC-4.0 (non-commercial) | Yes (20 preset voices + adaptation) | ≥16 GB on a single GPU (per Mistral card) |
| Fish Audio S2 Pro | 4B Slow AR + 400M Fast AR (~5B total) | fish-audio-research-license (non-commercial) | Yes (per Spaces list, separate from main card) | Footprint not published on card |
| Step-Audio-EditX | 3B | Apache-2.0 | Yes (zero-shot, editable delivery) | ~12-15 GB bf16; AWQ 4-bit ~8-10 GB |
Licence column is the part most readers underestimate. The two highest-Elo open-weight TTS models on the August 2026 leaderboard - Fish Audio S2 Pro and Voxtral - both ship non-commercial. Sell a product with their output and you are out of licence. Apache-2.0 and MIT models (Kokoro, Chatterbox, Step-Audio-EditX) do not have that problem.
Footprint column has gaps on purpose. Where the upstream card publishes no VRAM figure, the table says so rather than invent one - see Best Local Speech Models by VRAM Tier for what each tier actually fits.
What the hosted side does that local still cannot match
Three things, in this order:
- Time-to-first-voice-clone. ElevenLabs Starter ($5/mo) lets you upload a short sample and produce a cloned voice in minutes, with no GPU, no Python, no docker. The local equivalent means installing dependencies and running a script, before you think about hardware.
- Long-form stability. Gemini 3.8 Flash TTS markets “long-form stability” explicitly. Local TTS at the 4-12 GB tier is fine for short replies; past several minutes of continuous speech hosted stacks still tend to win, because they have been tuned on audiobook- and podcast-length jobs.
- Studio voices and presets. ElevenLabs’ voice library, OpenAI’s six preset voices with controllable style via prompt, and Gemini’s voice roster are polished assets. Kokoro ships 54 voices across 8 languages; Chatterbox and Step-Audio-EditX clone from your audio, but you supply the voice.
If your project is a chatty in-app assistant, a podcast with a single narrator, or anything under a few thousand characters per request, the local options in the table above will match or beat the hosted quality without the bill.
What the local side does that hosted cannot match
Three things, in the opposite order:
- No per-character cost. A locally run model turns a 10,000-character script into audio for the cost of electricity, which on a 12 GB consumer GPU is roughly a fraction of a cent. The hosted equivalent on ElevenLabs Starter ($5/30k credits) is roughly $1.67 at the same volume. Below a few thousand characters per month, hosted free tiers win.
- Licence freedom. Apache-2.0 (Kokoro, Step-Audio-EditX) and MIT (Chatterbox) permit commercial output without per-seat or per-character fees and without storing your text or voice on someone else’s server. Voxtral and Fish Audio S2 Pro are local but ship under a non-commercial grant, removing most of the licence advantage.
- Audio never leaves your machine. Anything you dictate into a personal voice memo app or generate from a private script stays on your hardware. For TTS the stakes are higher than for chat because voice is biometric, and a cloned voice is a credential.
Where the crossover actually sits
If you are deciding between hosted and local for a single project, the breakpoint is roughly: below ~200,000 characters per month, hosted free or entry tiers are cheaper than the time you would spend installing a local pipeline. Above that, the bill climbs while the local hardware cost stays flat, and local wins on total cost of ownership above about 1-2 million characters per month. Numbers vary with model choice and caching, which is why the broader Local AI vs API Cost Crossover guide treats this as a worked example rather than a single number.
The licence you should pick before you pick the model
A practical order:
- If the output ships in something with a price on it, restrict yourself to Apache-2.0 or MIT local models, or a hosted plan whose terms permit commercial use. ElevenLabs Starter and above permit commercial use; the Free tier does not. OpenAI and Gemini TTS terms permit commercial use under their standard API terms.
- If the output is internal or personal, hosted free tiers and CC-BY-NC-4.0 local models are all fair game; the licence is not the constraint.
- If you need voice cloning, hosted starts at $5/mo (ElevenLabs Starter, instant cloning only) or $18.33/mo (Creator, professional cloning). On the local side, Chatterbox, Step-Audio-EditX, Voxtral and Fish Audio S2 Pro all clone; only the first two can ship commercially.
- If you cannot accept that your audio or voice reference leaves your machine, hosted is out. Local with Apache-2.0 or MIT is the answer: Kokoro (no cloning), Chatterbox (cloning, MIT), or Step-Audio-EditX (cloning, Apache-2.0).
Bottom line
The right answer is “both, depending on the row in the spreadsheet.” ElevenLabs at $5/mo is the cheapest way to get a cloned voice without touching a GPU. Gemini 3.8 Flash TTS at ~$0.0135/min is the cheapest pay-as-you-go TTS at production quality, with a planned price rise on 2027-01-01. Kokoro-82M is the cheapest way to read text aloud and the only one that runs on a CPU; it is not a cloner. For commercial products that clone voices on consumer hardware, Chatterbox (MIT) and Step-Audio-EditX (Apache-2.0) are the two local options that combine cloning, a permissive licence, and an 8 GB card footprint. Fish Audio S2 Pro and Voxtral sit at the top of the open-weight TTS leaderboard and are unusable in commercial products; treat them as research baselines.
For per-tier VRAM and quick-start code for every local model listed here, see the companion Best Local Speech Models by VRAM Tier page. For the broader question of when local beats paying per token, see Local AI vs API Cost Crossover.