AI Voice and TTS Tools: Hosted vs Local (September 2026)

ElevenLabs vs Gemini TTS vs OpenAI TTS vs local open-weight models (Kokoro, Voxtral, Fish Audio S2 Pro). Pricing, licence, and the catches.

Updated September 28, 2026


Voice synthesis in 2026 splits into two camps: the hosted services that bill you per character or per audio token, and the open-weight models you run on your own GPU. The hosted side gets you professional voice clones in a few clicks and a tier that handles your traffic; the local side gets you unlimited generation, full control over the voice, and a licence that does not forbid you from selling the output. Neither camp has swept the other.

This page is the hosted-vs-local comparison. The local-only side of it - VRAM tiers, STT, and the open TTS ranking - lives in the companion Best Local Speech Models by VRAM Tier guide; this page covers pricing, licence, and the catches.

The pricing actually looks like this

Hosted TTS billing varies more than it looks at first glance. ElevenLabs charges by the character of input text; Google bills audio by the audio token; OpenAI mixes both. Below is the lowest published tier from each vendor’s own pricing page, accessed 2026-09-28.

ServiceLowest paid tierPriceUnitWhat you get
ElevenLabs Starter$5/mo (annual) or $6/mo$5-$6monthly subscription30,000 credits; instant voice cloning; commercial use
ElevenLabs Creator$18.33/mo (annual) or $22/mo$18-$22monthly subscription121,000 credits; instant + professional cloning
ElevenLabs Pro$82.50-$99/mo$82-$99monthly subscription600,000 credits; higher concurrency
OpenAI tts-1pay-as-you-go$15per 1M charactersStandard TTS, 6 preset voices
OpenAI tts-1-hdpay-as-you-go$30per 1M charactersHigher-fidelity variant
OpenAI gpt-4o-mini-ttspay-as-you-go$12per 1M audio tokens (output)Controllable speech via prompt
Gemini 2.5 Flash Preview TTSpay-as-you-go$10per 1M audio tokens (output)Lower latency
Gemini 3.1 Flash TTS Previewpay-as-you-go$20per 1M audio tokens (output)Lower latency
Gemini 3.8 Flash TTSpay-as-you-go + free tier$9 (to 2026-12-31) / $18 (from 2027-01-01)per 1M audio tokens (output)Studio-grade fidelity

ElevenLabs is the only one of the three that is subscription-based with rolled-over character quotas; OpenAI and Gemini are pure pay-as-you-go, billed to the second. Gemini 3.8 Flash TTS doubles on 2027-01-01 according to its own pricing page - if you build on it, that is the date your unit economics change.

A second billing wrinkle: 1M audio tokens at Gemini is “25 tokens per second of audio,” so a minute of generated speech is 1,500 tokens. At $9/1M tokens, that is roughly $0.0135 per minute of audio in 2026, before the rise.

Local TTS options and what each one costs you to run

The hosted vendors have one price per output. The local models have a one-time hardware cost and an electricity bill. The catch is that you need a GPU, and not every model fits on every card.

ModelParamsLicenseVoice cloningFootprint
Kokoro-82M82MApache-2.0No (54 fixed voices, 8 languages)Lightweight; footprint not published on card
Resemble AI Chatterbox500MMITYes (zero-shot)VRAM not published on card
Voxtral-4B-TTS-26034BCC-BY-NC-4.0 (non-commercial)Yes (20 preset voices + adaptation)≥16 GB on a single GPU (per Mistral card)
Fish Audio S2 Pro4B Slow AR + 400M Fast AR (~5B total)fish-audio-research-license (non-commercial)Yes (per Spaces list, separate from main card)Footprint not published on card
Step-Audio-EditX3BApache-2.0Yes (zero-shot, editable delivery)~12-15 GB bf16; AWQ 4-bit ~8-10 GB

Licence column is the part most readers underestimate. The two highest-Elo open-weight TTS models on the August 2026 leaderboard - Fish Audio S2 Pro and Voxtral - both ship non-commercial. Sell a product with their output and you are out of licence. Apache-2.0 and MIT models (Kokoro, Chatterbox, Step-Audio-EditX) do not have that problem.

Footprint column has gaps on purpose. Where the upstream card publishes no VRAM figure, the table says so rather than invent one - see Best Local Speech Models by VRAM Tier for what each tier actually fits.

What the hosted side does that local still cannot match

Three things, in this order:

  1. Time-to-first-voice-clone. ElevenLabs Starter ($5/mo) lets you upload a short sample and produce a cloned voice in minutes, with no GPU, no Python, no docker. The local equivalent means installing dependencies and running a script, before you think about hardware.
  2. Long-form stability. Gemini 3.8 Flash TTS markets “long-form stability” explicitly. Local TTS at the 4-12 GB tier is fine for short replies; past several minutes of continuous speech hosted stacks still tend to win, because they have been tuned on audiobook- and podcast-length jobs.
  3. Studio voices and presets. ElevenLabs’ voice library, OpenAI’s six preset voices with controllable style via prompt, and Gemini’s voice roster are polished assets. Kokoro ships 54 voices across 8 languages; Chatterbox and Step-Audio-EditX clone from your audio, but you supply the voice.

If your project is a chatty in-app assistant, a podcast with a single narrator, or anything under a few thousand characters per request, the local options in the table above will match or beat the hosted quality without the bill.

What the local side does that hosted cannot match

Three things, in the opposite order:

  1. No per-character cost. A locally run model turns a 10,000-character script into audio for the cost of electricity, which on a 12 GB consumer GPU is roughly a fraction of a cent. The hosted equivalent on ElevenLabs Starter ($5/30k credits) is roughly $1.67 at the same volume. Below a few thousand characters per month, hosted free tiers win.
  2. Licence freedom. Apache-2.0 (Kokoro, Step-Audio-EditX) and MIT (Chatterbox) permit commercial output without per-seat or per-character fees and without storing your text or voice on someone else’s server. Voxtral and Fish Audio S2 Pro are local but ship under a non-commercial grant, removing most of the licence advantage.
  3. Audio never leaves your machine. Anything you dictate into a personal voice memo app or generate from a private script stays on your hardware. For TTS the stakes are higher than for chat because voice is biometric, and a cloned voice is a credential.

Where the crossover actually sits

If you are deciding between hosted and local for a single project, the breakpoint is roughly: below ~200,000 characters per month, hosted free or entry tiers are cheaper than the time you would spend installing a local pipeline. Above that, the bill climbs while the local hardware cost stays flat, and local wins on total cost of ownership above about 1-2 million characters per month. Numbers vary with model choice and caching, which is why the broader Local AI vs API Cost Crossover guide treats this as a worked example rather than a single number.

The licence you should pick before you pick the model

A practical order:

  1. If the output ships in something with a price on it, restrict yourself to Apache-2.0 or MIT local models, or a hosted plan whose terms permit commercial use. ElevenLabs Starter and above permit commercial use; the Free tier does not. OpenAI and Gemini TTS terms permit commercial use under their standard API terms.
  2. If the output is internal or personal, hosted free tiers and CC-BY-NC-4.0 local models are all fair game; the licence is not the constraint.
  3. If you need voice cloning, hosted starts at $5/mo (ElevenLabs Starter, instant cloning only) or $18.33/mo (Creator, professional cloning). On the local side, Chatterbox, Step-Audio-EditX, Voxtral and Fish Audio S2 Pro all clone; only the first two can ship commercially.
  4. If you cannot accept that your audio or voice reference leaves your machine, hosted is out. Local with Apache-2.0 or MIT is the answer: Kokoro (no cloning), Chatterbox (cloning, MIT), or Step-Audio-EditX (cloning, Apache-2.0).

Bottom line

The right answer is “both, depending on the row in the spreadsheet.” ElevenLabs at $5/mo is the cheapest way to get a cloned voice without touching a GPU. Gemini 3.8 Flash TTS at ~$0.0135/min is the cheapest pay-as-you-go TTS at production quality, with a planned price rise on 2027-01-01. Kokoro-82M is the cheapest way to read text aloud and the only one that runs on a CPU; it is not a cloner. For commercial products that clone voices on consumer hardware, Chatterbox (MIT) and Step-Audio-EditX (Apache-2.0) are the two local options that combine cloning, a permissive licence, and an 8 GB card footprint. Fish Audio S2 Pro and Voxtral sit at the top of the open-weight TTS leaderboard and are unusable in commercial products; treat them as research baselines.

For per-tier VRAM and quick-start code for every local model listed here, see the companion Best Local Speech Models by VRAM Tier page. For the broader question of when local beats paying per token, see Local AI vs API Cost Crossover.