Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning TLDR 👉 new TTS leaderboard focused on open-source and multilingual The pace of open-source text-to-speech (TTS) model releases has been incredible.

- TLDR 👉 new TTS leaderboard focused on open-source and multilingual The pace of open-source text-to-speech (TTS) model releases has been incredible.
- On the Hugging Face Hub (as of Sep 30, 2026) there are more than 8K TTS models available 🚀 Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized.
- This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena .
TLDR 👉 new TTS leaderboard focused on open-source and multilingual The pace of open-source text-to-speech (TTS) model releases has been incredible. On the Hugging Face Hub (as of Sep 30, 2026) there are more than 8K TTS models available 🚀 Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. The gold standard is human preference scores such as MOS or MUSHRA (more on metrics ). To this end, several arena-based leaderboards have established themselves as useful reference points for the community: These arenas compare models by presenting users with TTS outputs from two models, and asking them to choose one over the other. After collecting a sufficient number of votes, an Elo score is computed to rank models, typically with the Bradley–Terry model (see Voice Arena methodology ). While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases . This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena . This likely reflects practical factors: adding an API model requires little more than an API key, whereas an open model must be hosted and served by the arena operator, and commercial providers have more reason to seek placement than open-source authors. Another limitation with arena-style evaluation is voter consistency: no arena can ensure that the same voters with the same criteria of “better” can consistently evaluate models over time. Even the preferences of a single person change over time (“A man cannot step into the same river twice” as famously said by Heraclitus). To this end, we've built the Open TTS Leaderboard , which uses objective metrics to evaluate models on complementary aspects of performance: Intelligibility : word/character error rate (WER and CER) between the prompt and the generated audio's transcript, using Qwen3 ASR (top ranking open-source model on the Open ASR Leaderboard ). Speed : inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for quantifying streaming batch size 1 latency on an H200 GPU and CPU. Speaker similarity by computing the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip. By relying on objective metrics evaluating a model drops from a couple weeks (for collecting votes) to a couple hours ⚡ Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations. Our intention with this leaderboard is for it to be shaped by the community ; we want to hear your feedback so the evaluations stay relevant and insightful. The next few sections give an overview of main features of the Open TTS Leaderboard. From the default view of the leaderboard, models are ranked by macro-average WER on the English splits of Seed TTS Eval ( paper ) and CV3 Eval (zero shot) ( paper ). hexgrad/Kokoro-82M , Supertone/supertonic-3 , and fishaudio/s2-pro lead the pack on English WER when averaged on these two splits, while the Pareto plots visualize which models strike a good balance between WER, batched inference (RTFx), and size. English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages. k2-fsa/OmniVoice , fishaudio/s2-pro , and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models. By toggling “Voice cloning”, the models that support this functionality (on the selected languages) can be compared. Moreover, a SIM column for speaker similarity now appears in the table, as well as two more Pareto plots for visualizing the tradeoff between SIM, batched inference, and size. The average WER of some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 , improve under voice cloning, namely when a reference audio is provided. Numbers only tell part of the story, and as mentioned earlier human preference is the ultimate decider . From the “Listen” tab, you can compare the generated outputs that are behind the metrics, to find which model(s) you prefer! Pick the language/dataset you're interested in, whether you want to compare voice cloning , and optionally pick the models or listen to outputs from a random selection. The “Listen” tab fills an important gap in existing TTS leaderboards: a space to explore model outputs of various models. You can even give feedback on the generated outputs. As we collect more votes from the community, we may include this data on the leaderboard. So vote! But please login with your HF account to help us weed out spam/bots. The “Streaming” tab compares the streaming capabilities. Models are ranked by TTFA (time-to-first-audio), which quantifies how long a user waits after probing a model in order to obtain audio that can be played. This is important for voice agents and other interactive apps. For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest. The default view compares performance on an H200 GPU. Results are also available for CPU for a small (but growing) set of models! kyutai/pocket-tts is a great model for streaming on both GPU and CPU! The goal of the Open TTS Leaderboard is not only to keep up with the incredible pace of TTS model releases, but to be shaped by the community; we want to hear your feedback so the evaluations stay relevant and insightful. Let us know which datasets, models, and metrics you want to see! Open-source models , to put forward many great models that have been neglected by arena-style evaluations. Multilingual , since English performance is not a suitable proxy for other languages. We will soon open-source the evaluation scripts, much like the Open ASR Leaderboard repo , so that you can directly provide your feedback and suggestions via GitHub Issues and PRs! Let's shape TTS evaluations together 🤗 The Open ASR Leaderboard Adds Its First Global South Language Introducing Real World VoiceEQ: Measuring the human quality of voice AI very interesting
\n","updatedAt":"2026-09-30T20:40:39.903Z","author":{"_id":"67d8504e837f263bd9d04e0d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/xAez-XvIUIr0j3GpPB55Y.png","fullname":"Eduardo Vela","name":"EducitoEc","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9895163774490356},"editors":["EducitoEc"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/xAez-XvIUIr0j3GpPB55Y.png"],"reactions":[],"isReport":false}},{"id":"6abd8509c7226b846c3cf9d0","author":{"_id":"69f3b7146bbe3ec4c034c3ff","avatarUrl":"/avatars/f47b10ef653f71f3ba5bc7dcc8c085d2.svg","fullname":"Igor Eduardo","name":"Nomad-link-id","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-30T21:54:17.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Glad this board splits multilingual TTS quality from voice-cloning identity instead of one aggregate green.\n\nThe check I'd keep reading: whether the ranking still holds when you hold language/domain fixed versus when cloning is the axis. Otherwise teams optimize the easier subscale and still call the result "SOTA TTS."","html":"Glad this board splits multilingual TTS quality from voice-cloning identity instead of one aggregate green.
\nThe check I'd keep reading: whether the ranking still holds when you hold language/domain fixed versus when cloning is the axis. Otherwise teams optimize the easier subscale and still call the result "SOTA TTS."
\n","updatedAt":"2026-09-30T21:54:17.565Z","author":{"_id":"69f3b7146bbe3ec4c034c3ff","avatarUrl":"/avatars/f47b10ef653f71f3ba5bc7dcc8c085d2.svg","fullname":"Igor Eduardo","name":"Nomad-link-id","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7888889908790588},"editors":["Nomad-link-id"],"editorAvatarUrls":["/avatars/f47b10ef653f71f3ba5bc7dcc8c085d2.svg"],"reactions":[],"isReport":false}},{"id":"6abe0b5d761a4d95dacf1bf0","author":{"_id":"6a3cd0fb3af2b0b0921e329a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/1Yp81w2J8OBsuzrHcQoZu.jpeg","fullname":"Elton Williams","name":"eltonwilliams","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-10-01T07:27:25.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":">English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has <a href="https://moviebox4u.app/\">movie box audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.\n\nOpen TTS Leaderboard is an open evaluation platform designed to make multilingual text-to-speech and voice-cloning models easier to compare at scale. Traditional TTS arenas usually let users listen to outputs from two models and choose the one they prefer. Those votes can then be used to calculate an Elo-style score, often through approaches such as the Bradley–Terry model, giving users a simple way to understand how different systems perform based on human preference.\n\nHuman feedback remains an important part of evaluating speech quality, but relying only on arena-style voting can be difficult as new TTS models are released at a rapid pace. Open-source models can also face practical barriers because they generally need to be hosted and served by the evaluation platform, while API-based commercial models can be added more easily. This can result in open-weight systems being less visible on public leaderboards. Another challenge is consistency: people's preferences and listening criteria can change over time, meaning that votes collected months apart may not always represent exactly the same standards.\n\nOpen TTS Leaderboard aims to address these challenges by providing a scalable and transparent space for evaluating multilingual TTS and voice-cloning systems. The goal is to make comparisons easier to reproduce, expand coverage of open models, and provide useful evaluation signals alongside human preference. By bringing together scalable testing and community feedback, the project can help researchers, developers, and users better understand how speech models perform across different languages, voices, and use cases.","html":"\n\nEnglish performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has <a href="https://moviebox4u.app/\" rel="nofollow">movie box audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
\n
Open TTS Leaderboard is an open evaluation platform designed to make multilingual text-to-speech and voice-cloning models easier to compare at scale. Traditional TTS arenas usually let users listen to outputs from two models and choose the one they prefer. Those votes can then be used to calculate an Elo-style score, often through approaches such as the Bradley–Terry model, giving users a simple way to understand how different systems perform based on human preference.
\nHuman feedback remains an important part of evaluating speech quality, but relying only on arena-style voting can be difficult as new TTS models are released at a rapid pace. Open-source models can also face practical barriers because they generally need to be hosted and served by the evaluation platform, while API-based commercial models can be added more easily. This can result in open-weight systems being less visible on public leaderboards. Another challenge is consistency: people's preferences and listening criteria can change over time, meaning that votes collected months apart may not always represent exactly the same standards.
\nOpen TTS Leaderboard aims to address these challenges by providing a scalable and transparent space for evaluating multilingual TTS and voice-cloning systems. The goal is to make comparisons easier to reproduce, expand coverage of open models, and provide useful evaluation signals alongside human preference. By bringing together scalable testing and community feedback, the project can help researchers, developers, and users better understand how speech models perform across different languages, voices, and use cases.
\n","updatedAt":"2026-10-01T07:27:25.374Z","author":{"_id":"6a3cd0fb3af2b0b0921e329a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/1Yp81w2J8OBsuzrHcQoZu.jpeg","fullname":"Elton Williams","name":"eltonwilliams","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9207108020782471},"editors":["eltonwilliams"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/1Yp81w2J8OBsuzrHcQoZu.jpeg"],"reactions":[],"isReport":false}}],"status":"open","isReport":false,"pinned":false,"locked":false,"collection":"community_blogs"},"contextAuthors":["bezzam","Steveeeeeeen","eustlb","mrfakename"],"primaryEmailConfirmed":false,"discussionRole":0,"acceptLanguages":["*"],"withThread":true,"cardDisplay":false,"repoDiscussionsLocked":false,"hideComments":true}"> Glad this board splits multilingual TTS quality from voice-cloning identity instead of one aggregate green. The check I'd keep reading: whether the ranking still holds when you hold language/domain fixed versus when cloning is the axis. Otherwise teams optimize the easier subscale and still call the result "SOTA TTS." English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has movie box audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages. Open TTS Leaderboard is an open evaluation platform designed to make multilingual text-to-speech and voice-cloning models easier to compare at scale. Traditional TTS arenas usually let users listen to outputs from two models and choose the one they prefer. Those votes can then be used to calculate an Elo-style score, often through approaches such as the Bradley–Terry model, giving users a simple way to understand how different systems perform based on human preference. Human feedback remains an important part of evaluating speech quality, but relying only on arena-style voting can be difficult as new TTS models are released at a rapid pace. Open-source models can also face practical barriers because they generally need to be hosted and served by the evaluation platform, while API-based commercial models can be added more easily. This can result in open-weight systems being less visible on public leaderboards. Another challenge is consistency: people's preferences and listening criteria can change over time, meaning that votes collected months apart may not always represent exactly the same standards. Open TTS Leaderboard aims to address these challenges by providing a scalable and transparent space for evaluating multilingual TTS and voice-cloning systems. The goal is to make comparisons easier to reproduce, expand coverage of open models, and provide useful evaluation signals alongside human preference. By bringing together scalable testing and community feedback, the project can help researchers, developers, and users better understand how speech models perform across different languages, voices, and use cases. Upload images, audio, and videos by dragging in the text input, pasting, or clicking here .Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.