Hugging Face launches Open TTS Leaderboard ranking open-source multilingual TTS and voice-cloning models on objective metrics
Hugging Face released a new public leaderboard that evaluates open-source TTS models using objective metrics: WER/CER via Qwen3 ASR on Seed TTS Eval and CV3 Eval, speed (RTFx on H200, TTFA on H200 and CPU), and speaker similarity (WavLM cosine similarity) for voice cloning. It includes multilingual toggles, a Listen tab for comparing outputs and voting (HF login required), and a Streaming tab. Initial results: Kokoro-82M, supertonic-3 and s2-pro lead English WER; OmniVoice, s2-pro and Fun-CosyVoice3-0.5B lead multilingual. Evaluation scripts are promised to be open-sourced soon.
Entities: Hugging Face, Open TTS Leaderboard, Artificial Analysis, TTS Arena v2, Qwen3 ASR, Open ASR Leaderboard
1 primary
What happened
Hugging Face launched the Open TTS Leaderboard, a public ranking of open-source text-to-speech and voice-cloning models. It uses automated metrics only: word and character error rates from a Qwen3 speech recogniser on the Seed TTS Eval and CV3 Eval datasets, speed (RTFx on an H200 GPU, time to first audio on H200 and CPU) and speaker similarity via WavLM. Early English leaders on error rate are Kokoro-82M, supertonic-3 and s2-pro. OmniVoice, s2-pro and Fun-CosyVoice3-0.5B lead the multilingual table. The evaluation scripts are promised as open source "soon", and no links to them were supplied.
Why it matters
Developers choosing an open TTS model get a cheap, repeatable way to shortlist candidates on accuracy, latency and voice-cloning fidelity, including CPU latency that matters for self-hosting. It also gives open models a showcase next to arenas dominated by commercial APIs. The impact is narrow: it helps with screening, not with picking a voice people will enjoy listening to. Its value also depends on whether the scripts are actually released so others can reproduce the numbers.
What is noise
The framing that human-preference arenas "can't scale" and that objective metrics cut evaluation from weeks to hours is partly self-serving, since a low error rate through an ASR model says little about naturalness, prosody or emotion, as the authors themselves admit. Rankings from one ASR judge (Qwen3) may favour models that sound clean to that recogniser, and the top positions reflect two test sets, not real-world use. The vote and Listen features need a Hugging Face login and have no stated volume yet.
Watch next
- 01Whether the evaluation scripts are open-sourced, and on what date, so third parties can reproduce the WER, RTFx and speaker-similarity figures.
- 02How many models are listed after the first month and whether major releases (including commercial or API-only ones) are added, versus a static list of a few dozen.
- 03Whether the objective rankings agree with human votes in the Listen tab and with TTS Arena v2, and whether model authors begin citing or tuning to the leaderboard.
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI discloses sandbox-escape and credential-leak incidents, confirms pause on tool-use for its most capable models26 Sept 202680
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680