Skip to content
fedi.software

Audio

Text-to-speech models ranked by listeners' blind preference on TTS Arena, priced per million characters or per minute.

Updated · Sources: LiteLLM, TTS Arena

Tier Model Elo Price, USD per 1M characters Context App Compare
S1,519— —
S1,509— —
Compare ()

Tiers are ours: a model's place within the board counting only the models whose whole confidence interval is above it. S — top 10%, A — up to 30%, B — up to 60%, C — up to 85%, D — the rest.

What the audio leaderboard ranks

This page covers text-to-speech models — models that read a text aloud. The ratings come from TTS Arena V2, an open community project by TTS-AGI: listeners hear the same text read by two anonymous voices and pick the one that sounds better. The votes are turned into an Elo rating with an uncertainty range, and the table shows how many votes stand behind each rating.

Prices

Speech models are priced per million characters of input text, or per minute of generated audio, depending on how the provider sells them; the column header names the unit. The lowest current offer comes from models.dev or LiteLLM. The "Low price" chip marks models under 5 USD per million characters.

Tiers and votes

TTS Arena publishes the number of votes per model, so our usual tier rule applies here: a model with fewer than 300 votes gets no tier yet. A model's place is counted from the models whose whole uncertainty range lies above its own, and the share of the board gives S under 10%, A under 30%, B under 60%, C under 85%, D otherwise. The ratings of speech models sit close together and their ranges are wide, so many models end up in tier S at once.

What is not covered

  • Speech recognition (speech-to-text) is not ranked: the source has no arena for it.
  • Voices are compared mostly in English. A voice that wins there may sound less natural in another language, so listen to samples in yours.
  • Latency and streaming support, which matter for voice assistants, are not part of the rating.
  • Some rows are models that are not released yet and are tested under a code name: they have no vendor and no price.

For a narration or an app voice, use the table to shortlist a few models with a good rating and a price that fits your volume, then compare samples at the providers.