πŸ—£οΈ Open TTS Tracker

February 13, 2025 Β· View on GitHub

A one stop shop to track all open-access/ source TTS models as they come out. Feel free to make a PR for all those that aren't linked here.

This is aimed as a resource to increase awareness for these models and to make it easier for researchers, developers, and enthusiasts to stay informed about the latest advancements in the field.

Note


This repo will only track open source/access codebase TTS models. More motivation for everyone to open-source! πŸ€—

NameGitHubWeightsLicenseFine-tuneLanguagesPaperDemoIssues
AmphionRepoπŸ€— HubMITNoMultilingualPaperπŸ€— Space
AI4BharatRepoπŸ€— HubMITYesIndicPaperDemo
BarkRepoπŸ€— HubMITNoMultilingualPaperπŸ€— Space
EmotiVoiceRepoGDriveApache 2.0YesZH + ENNot AvailableNot AvailableSeparate GUI agreement
Fish-speechRepoπŸ€— HubApache 2.0NoMultilingualPaperπŸ€— Space
Glow-TTSRepoGDriveMITYesEnglishPaperGH Pages
GPT-SoVITSRepoπŸ€— HubMITYesMultilingualNot AvailableNot Available
HierSpeech++RepoGDriveMITNoKR + ENPaperπŸ€— Space
IMS-ToucanRepoGH releaseApache 2.0YesMultilingualPaperπŸ€— Space
KokoroRepoπŸ€— HubApache 2.0YesMultilingualπŸ€— Space
LlasaRepoπŸ€— HubCC-BY-NC 4.0NoMultilingualPaperπŸ€— Space
MahaTTSRepoπŸ€— HubApache 2.0NoEnglish + IndicNot AvailableRecordings, Colab
Matcha-TTSRepoGDriveMITYesEnglishPaperπŸ€— SpaceGPL-licensed phonemizer
MeloTTSRepoπŸ€— HubMITYesMultilingualπŸ€— Space
MetaVoice-1BRepoπŸ€— HubApache 2.0YesMultilingualNot AvailableπŸ€— Space
Neural-HMM TTSRepoGitHubMITYesEnglishPaperGH Pages
OpenVoiceRepoπŸ€— HubCC-BY-NC 4.0NoZH + ENPaperπŸ€— SpaceNon Commercial
OuteTTSRepoπŸ€— HubApache 2.0NoMultilingualπŸ€— Space
OverFlow TTSRepoGitHubMITYesEnglishPaperGH Pages
Parler TTSRepoπŸ€— HubApache 2.0YesEnglishNot AvailableNot Available
pflowTTSUnofficial RepoGDriveMITYesEnglishPaperNot AvailableGPL-licensed phonemizer
PiperRepoπŸ€— HubMITYesMultilingualNot AvailableNot AvailableGPL-licensed phonemizer
PhemeRepoπŸ€— HubCC-BYYesEnglishPaperπŸ€— Space
RAD-MMMRepoGDriveMITYesMultilingualPaperJupyter Notebook, Webpage
RAD-TTSRepoGDriveMITYesEnglishPaperGH Pages
SileroRepoGH linksCC BY-NC-SANoEM + DE + ES + EANot AvailableNot AvailableNon Commercial
StyleTTS 2RepoπŸ€— HubMITYesEnglishPaperπŸ€— SpaceGPL-licensed phonemizer
Tacotron 2Unofficial RepoGDriveBSD-3YesEnglishPaperWebpage
TorToiSe TTSRepoπŸ€— HubApache 2.0YesEnglishTechnical reportπŸ€— Space
TTTSRepoπŸ€— HubMPL 2.0NoZHNot AvailableColab, πŸ€— Space
VALL-EUnofficial RepoNot AvailableMITYesNAPaperNot Available
VITS/ MMS-TTSRepoπŸ€— Hub / MMSApache 2.0YesEnglishPaperπŸ€— SpaceGPL-licensed phonemizer
WhisperSpeechRepoπŸ€— HubMITNoEnglish, PolishNot AvailableπŸ€— Space, Recordings, Colab
XTTSRepoπŸ€— HubCPMLYesMultilingualPaperπŸ€— SpaceNon Commercial
xVASynthRepoπŸ€— HubGPL-3.0YesMultilingualPaperπŸ€— SpaceCopyrighted materials used for training.
ZonosRepoπŸ€— HubApache 2.0NoMultilingualπŸ€— Space

Capability specifics

Click on this to toggle table visibility
NameProcessor
⚑
Phonetic alphabet
πŸ”€
Insta-clone
πŸ‘₯
Emotional control
🎭
Prompting
πŸ“–
Speech control
🎚
Streaming support
🌊
S2S support
🦜
Longform synthesis
AmphionCUDAπŸ‘₯🎭πŸ‘₯❌
BarkCUDA❌🎭 tags❌
EmotiVoice
Fish-speechCUDA❌πŸ‘₯🎭πŸ‘₯❌speed / stability
🎚
🌊🦜Yes
Glow-TTS
GPT-SoVITS
HierSpeech++❌πŸ‘₯🎭πŸ‘₯❌speed / stability
🎚
🦜
IMS-ToucanCUDA❌❌❌❌
KokoroCPU / CUDAπŸ‘₯❌❌❌speed
🎚
🌊❌
LlasaCUDA❌πŸ‘₯🎭❌❌❌❌
MahaTTS
Matcha-TTSIPA❌❌❌speed / stability
🎚
MetaVoice-1BCUDAπŸ‘₯🎭πŸ‘₯❌stability / similarity
🎚
Yes
Neural-HMM TTS
OpenVoiceCUDA❌πŸ‘₯6-type 🎭
πŸ˜‘πŸ˜ƒπŸ˜­πŸ˜―πŸ€«πŸ˜Š
❌
OuteTTSCPU / CUDA❌πŸ‘₯❌❌speed
🎚
🌊❌
OverFlow TTS
pflowTTS
Piper
PhemeCUDA❌πŸ‘₯🎭πŸ‘₯❌stability
🎚
RAD-TTS
Silero
StyleTTS 2CPU / CUDAIPAπŸ‘₯🎭πŸ‘₯❌🌊Yes
Tacotron 2
TorToiSe TTSβŒβŒβŒπŸ“–πŸŒŠ
TTTSCPU/CUDA❌πŸ‘₯
VALL-E
VITS/ MMS-TTSCUDA❌❌❌❌speed
🎚
WhisperSpeechCUDA❌πŸ‘₯🎭πŸ‘₯❌speed
🎚
XTTSCUDA❌πŸ‘₯🎭πŸ‘₯❌speed / stability
🎚
🌊❌
xVASynthCPU / CUDAARPAbet+❌4-type 🎭
πŸ˜‘πŸ˜ƒπŸ˜­πŸ˜―
per‑phoneme
❌speed / pitch / energy / 🎭
🎚
per‑phoneme
❌🦜
ZonosCUDAeSpeakπŸ‘₯🎭❌speed / pitch / quality / emotion
🎚
❌❌
  • Processor - CPU/CUDA/ROCm (single/multi used for inference; Real-time factor should be below 2.0 to qualify for CPU, though some leeway can be given if it supports audio streaming)
  • Phonetic alphabet - None/IPA/ARPAbet (Phonetic transcription that allows to control pronunciation of certain words during inference)
  • Insta-clone - Yes/No (Zero-shot model for quick voice clone)
  • Emotional control - Yes🎭/Strict (Strict, as in has no ability to go in-between states, insta-clone switch/🎭πŸ‘₯)
  • Prompting - Yes/No (A side effect of narrator based datasets and a way to affect the emotional state, ElevenLabs docs)
  • Streaming support - Yes/No (If it is possible to playback audio that is still being generated)
  • Speech control - speed/pitch/ (Ability to change the pitch, duration, energy and/or emotion of generated speech)
  • Speech-To-Speech support - Yes/No (Streaming support implies real-time S2S; S2T=>T2S does not count)

How can you help?

Help make this list more complete. Create demos on the Hugging Face Hub and link them here :) Got any questions? Drop me a DM on Twitter @reach_vb.