Awesome Controllable Speech Synthesis
August 4, 2026 Β· View on GitHub
This is an evolving repo for the survey: Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey.
Abstract
Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has become a rapidly growing research area. This survey provides the first comprehensive review of controllable TTS methods, from traditional control techniques to emerging approaches using natural language prompts. We categorize model architectures, control strategies, and feature representations, while also summarizing challenges, datasets, and evaluations in controllable TTS. This survey aims to guide researchers and practitioners by offering a clear taxonomy and highlighting future directions in this fast-evolving field.
If you find our survey useful for your research, please consider πcitingπ our EMNLP 2025 main conference paper.
ποΈ Video introductions for our paper: YouTube (English), Bilibili (Chinese)
π Letβs make it better together:
- If you find any mistakes, please donβt hesitate to open an issue.
- If you find this project helpful, please consider giving it a β on GitHub to stay updated.
- Pull requests are always welcome if youβd like to contribute papers to this repo.
News
- [2025-09-29] Our paper has been accepted to the EMNLP 2025 Main Conference.
Table of Contents
- Follow-up Papers
- Non-autoregressive Controllable TTS
- Autoregressive Controllable TTS
- Datasets
- Evaluation
- Star History
Follow-up Papers π₯π₯π₯ (Newest First)
2026-07
- Xiang, Bajian, et al. "Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm." arXiv preprint arXiv:2607.23938 (2026). Demo
- Du, Muyang, Shuang Yu, and Junjie Lai. "Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs." arXiv preprint arXiv:2607.21042 (2026). Demo Code
- Luo, Kaicheng, et al. "StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis." arXiv preprint arXiv:2607.19859 (2026). Demo
- Shen, Shengfan, et al. "Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer." arXiv preprint arXiv:2607.17900 (2026).
- Jia, Zhenqi, et al. "AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis." arXiv preprint arXiv:2607.15755 (2026).
- Lou, Haowei, et al. "AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling." arXiv preprint arXiv:2607.12706 (2026).
- Kimura, Kota, Mei Tanaka, and Masakiyo Fujimoto. "Zero-Shot Prosodic Style Control for LLM Agent Speech via Disentangled Conditioning Vectors." Preprint (2026).
- Pamuk, Ahmet Erdem, et al. "FreyaTTS Technical Report." arXiv preprint arXiv:2607.09530 (2026). Code
- Nie, Sihang, et al. "WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS." arXiv preprint arXiv:2607.06461 (2026). Demo & Code
- Chung, Ho-Lam, Kuan-Po Huang, Bo-Ru Lu, and Hung-yi Lee. "FrΓ©chet Distance Loss on Speech Representations for Text-to-Speech Synthesis." arXiv preprint arXiv:2607.06027 (2026).
- Asonitis, Antonis, et al. "GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech." arXiv preprint arXiv:2607.02633 (2026).
- Moon, A. -S., D. Song, and J. Lee. "Effective Mel-Spectrogram Synthesis for Multi-Speaker Text-to-Speech with Band-Specific Acoustic Modeling." IEEE Access (2026).
- Wang, Siyi, James Bailey, and Ting Dang. "A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models." arXiv preprint arXiv:2607.00946 (2026).
- Madha, Wasim, et al. "DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis." Preprint (2026).
- Yu, Zuda, et al. "Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis." arXiv preprint arXiv:2607.00363 (2026).
- Michel, Gaspard, Elena V. Epure, and Christophe Cerisara. "Computational Narrative Understanding for Expressive Text-to-Speech." Findings of the Association for Computational Linguistics: ACL 2026. 2026. Demo Code
- Baral, Anita, et al. "Neural Text-to-Speech for Myaamia: Speech Synthesis for an Indigenous Algonquian Language." Proceedings of the Sixth Workshop on NLP for Indigenous Languages of the Americas. 2026.
2026-06
- Hosseini-Kivanani, Nina, and Sandipana Dowerah. "LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish." arXiv preprint arXiv:2606.31947 (2026).
- Zhu, Chuanbo, et al. "UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling." arXiv preprint arXiv:2606.31128 (2026).
- Li, Xueyi, et al. "Benchmarking Scientific Formula Vocalization in Large Speech Language Models Toward Accessible Learning." In Artificial Intelligence in Education: AIED 2026, pp. 78β93. Springer, 2026. Code & Dataset
- Nie, Sihang, et al. "HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech." arXiv preprint arXiv:2606.28249 (2026). Demo
- Shi, Runwu, Yujin Wang, Hongjin Song, and Chunxiang Jin. "Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS." arXiv preprint arXiv:2606.25672 (2026).
- Liu, Lianbo, et al. "Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis." arXiv preprint arXiv:2606.25369 (2026). Demo Code
- Dong, Xin, Yong Yang, ZhiHao Li, and Qun Yang. "MSE-TTS: Emotion-Controllable Text-to-Speech via Multi-Scale Vision Feature." In 2026 International Joint Conference on Neural Networks (IJCNN). IEEE, 2026. Demo
- Wu, Ziqian. "Phoneme-Level Data Augmentation for Low-Resource TTS." Master's thesis, Uppsala University, 2026.
- Gal, O., and M. Giurgiu. "A Unified Objective Evaluation Framework for Deep Neural Text-to-Speech Synthesis in Romanian." 2026 IEEE International Conference on Automation, Quality and Testing, Robotics (AQTR), Baile Felix, Romania, 2026, pp. 1-6, doi: 10.1109/AQTR70159.2026.11577891.
- Wu, Minghui, et al. "EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis." arXiv preprint arXiv:2606.20650 (2026).
- Annamdevula, Ram, et al. "CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations." arXiv preprint arXiv:2606.25403 (2026).
- Wang, Haoxu, et al. "FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech." arXiv preprint arXiv:2606.23190 (2026).
- Huang, Rongjie, et al. "Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization." Findings of the Association for Computational Linguistics: ACL 2026. 2026.
- Tian, Jinchuan, et al. "Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis." arXiv preprint arXiv:2606.22811 (2026).
- Wang, Dongmei, et al. "AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation." arXiv preprint arXiv:2606.21893 (2026).
- Choi, Jeongsoo, Ji-Hoon Kim, Shujie Hu, and Joon Son Chung. "ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion." arXiv preprint arXiv:2606.21888 (2026).
- Du, Muyang, Jason Roche, and Junjie Lai. "Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead." arXiv preprint arXiv:2606.21882 (2026).
- Kim, Hounsu, and Juhan Nam. "SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion." arXiv preprint arXiv:2606.21157 (2026). Demo & Code
- Akti, Seymanur, and Alexander Waibel. "Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS." arXiv preprint arXiv:2606.23176 (2026).
- Lu, Yao. "DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation." arXiv preprint arXiv:2606.21457 (2026).
- Feng, Fangming, et al. "Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech." Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
- Xie, Hanke, et al. "FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation." arXiv preprint arXiv:2606.09141 (2026).
- Chen, Wenxi, et al. "WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling." arXiv preprint arXiv:2606.03455 (2026).
- Fan, Wei, et al. "BareWave: Waveform-Native Flow-Matching Text-to-Speech." arXiv preprint arXiv:2606.09048 (2026).
- Liang, Qifan, et al. "TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis." Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Demo
- Zhou, Shuoyi, et al. "FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech." arXiv preprint arXiv:2606.19209 (2026).
- Ghosh, Subhankar, et al. "MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data." arXiv preprint arXiv:2606.18485 (2026).
- Eom, SooHwan, et al. "Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning." arXiv preprint arXiv:2606.20266 (2026).
- Singh, Harshit, Ayush Pratap Singh, and Nityanand Mathur. "FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS." arXiv preprint arXiv:2606.20518 (2026).
- Zhang, Fang, Anqi Gou, and Linli Xu. "ExpV2S: Zero-Shot Expressive Video-to-Speech Synthesis via Latent Diffusion Model." Proceedings of the 2026 International Conference on Multimedia Retrieval. 2026.
- Mathur, Nityanand, et al. "How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech." arXiv preprint arXiv:2606.20532 (2026).
- Arigala, Adarsh, et al. "Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech." arXiv preprint arXiv:2606.14750 (2026).
- Ferreira, Alef Iury Siqueira, et al. "Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech." arXiv preprint arXiv:2606.13989 (2026).
- Quang, Vinh Dang, and Huy Ngo Quang. "An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis." arXiv preprint arXiv:2606.14922 (2026).
- Mai, Jialong, et al. "NVMOS: Non-Verbal Vocalization Quality Assessment in Speech." arXiv preprint arXiv:2606.15888 (2026).
- Mou, Zhenwei, et al. "Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity." arXiv preprint arXiv:2606.15267 (2026). Demo
- Mou, Zhenwei, et al. "DuraMark: Duration-Embedded Watermarking in LLM-based TTS." arXiv preprint arXiv:2606.15264 (2026). Demo
- Lin, Yihang, et al. "Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech." arXiv preprint arXiv:2606.13006 (2026).
- Chen, Peijie, et al. "SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations." arXiv preprint arXiv:2606.11611 (2026).
- Zhou, Yixuan, et al. "VoxCPM2 Technical Report." arXiv preprint arXiv:2606.06928 (2026). Code
- Lian, Shi, et al. "dots. tts Technical Report." arXiv preprint arXiv:2606.07080 (2026). Code
- de Brito, Daniel Oliveira, and Arnaldo Candido Junior. "Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech." arXiv preprint arXiv:2606.05367 (2026). Code
2026-05
- Pan, Changhao, et al. "Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios." arXiv preprint arXiv:2605.28618 (2026).
- Yun, Jun-Hak, Seung-Bin Kim, and Seong-Whan Lee. "ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment." arXiv preprint arXiv:2605.30965 (2026). Demo
- Li, Ruiqi, et al. "SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue." arXiv preprint arXiv:2605.30993 (2026). Demo
- Li, Bowen, et al. "PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis." arXiv preprint arXiv:2605.27258 (2026). Code
- Lee, Yoonhyung, et al. "FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations." arXiv preprint arXiv:2605.24618 (2026). Demo
- Lin, Bin, et al. "StepAudio 2.5 Technical Report." arXiv preprint arXiv:2605.23463 (2026).
- Chen, Junyang, et al. "CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS." arXiv preprint arXiv:2605.25930 (2026). Demo
- Kang, Bin, et al. "AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech." arXiv preprint arXiv:2605.17583 (2026). Demo
- Zhu, Yan, Rui Zhou, and Gang Chen. "A Multi-Level Expressive Voice Cloning Method Based on Adaptive Grouped Code Modeling." IEEE Access (2026).
- S. B. M, B. V. Victoria, B. R. L, M. Thirunavukkarasu, K. S and T. Dinesh Kumar, "A Novel Emotion-Aware Generative Adversarial Network for Fine-Grained and Controllable Speech Synthesis," 2026 9th International Conference on Intelligent Computing and Control Systems (ICICCS), Erode, India, 2026, pp. 1252-1257, doi: 10.1109/ICICCS67901.2026.11502912.
2026-04
- Li, Aoduo, et al. "ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis." arXiv preprint arXiv:2604.19055 (2026).
- Wang, Kuang, et al. "Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation." arXiv preprint arXiv:2604.11424 (2026). Demo
- Feng, Tao, et al. "MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora." arXiv preprint arXiv:2604.11552 (2026). Demo
- Zhu, Jian, et al. "AccompGen: Hierarchical Autoregressive Vocal Accompaniment Generation with Dual-Rate Codec Tokenization." arXiv preprint arXiv:2604.09054 (2026). Code
- Cheng, Changhao, et al. "On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation." arXiv preprint arXiv:2604.12383 (2026). Code
- Cong, Gaoxiang, et al. "CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing." arXiv preprint arXiv:2604.12292 (2026).
- Su, Tianhui, et al. "An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding." arXiv preprint arXiv:2604.12438 (2026).
- Jin, Haopeng. "ChronoVoice: Controllable Text-to-Speech with Explicit Time Planning, Speech-Rate Modulation, and Emotion Trajectory Control." (2026). [2026.04]
- Zheng, Qixi, et al. "X-VC: Zero-shot Streaming Voice Conversion in Codec Space." arXiv preprint arXiv:2604.12456 (2026). Code [2026.04]
- H. -J. Guo et al., "QE-XVC: Zero-Shot Cross-Lingual Voice Conversion via Query-Enhancement and Conditional Flow Matching," ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2026, pp. 17342-17346. [2026.04]
- Wang, Xi, et al. "TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis." arXiv preprint arXiv:2604.22225 (2026). Code [2026.04]
- Qiang, Chunyu, et al. "UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions." arXiv preprint arXiv:2604.22209 (2026). Demo [2026.04]
- Cao, Jinglei, et al. "Composer: Learning to Orchestrate Generative Primitives for Controllable Speech Synthesis." [2026.04]
- Tsai, Yun-Shao, et al. "The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation." arXiv preprint arXiv:2604.26347 (2026). [2026.04]
- Halychanskyi, Yurii, Nimet Beyza Bozdag, Mark Hasegawa-Johnson, Dilek Hakkani-TΓΌr, and Volodymyr Kindratenko. "Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing." arXiv preprint arXiv:2604.27273 (2026). Demo
- Ma, Jianbo, and Richard Cartwright. "Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation." arXiv preprint arXiv:2604.19330 (2026). [2026.04]
- Chen, H., Hu, J., Xue, L., Zhan, Q., Li, W., Ma, G., ... & Xie, L. (2026). MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech. arXiv preprint arXiv:2604.17958. [2026.04]
- Tian, Fengping, Peng Bai, Xuanfan Ni, Haoqin Sun, Qingjuan Li, Zhiqiang Qian, Chenyang Lyu et al. "Marco-Voice: A Unified Framework for Expressive Speech Synthesis with Voice Cloning." In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3256-3260. IEEE, 2026. Code [2026.04]
- Fang, Jingyi, Yufei Tang, Zhiyu Wu, Yuanzhong Zheng, Yaoxuan Wang, and Haojun Fei. "QFOCUS: Controllable Synthesis for Automated Speech Stress Editing to Deliver Human-Like Emphatic Intent." In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 17277-17281. IEEE, 2026.
- Chen, Yihang, Hualei Wang, Na Li, and Zhifeng Li. "F5E-TTS: Enhancing Speech Synthesis by Aligning Text with Rich Semantic Representations." In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 18627-18631. IEEE, 2026.
- Lee, Joun Yeop, Heejin Choi, Min-Kyung Kim, Ji-Hyun Lee, and Hoon-Young Cho. "Hierarchical Discrete Flow Matching For Multi-Codebook Codec-Based Text-To-Speech." In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 18347-18351. IEEE, 2026.
- Shi, Yan, Jin Shi, Minchuan Chen, Ziyang Zhuang, Peng Qi, Shaojun Wang, and Jing Xiao. "NCF-TTS: Enhancing Flow Matching Based Text-To-Speech with Neighborhood Consistency Flow." In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 18327-18331. IEEE, 2026.
- Mai, Jialong, Xiaofen Xing, and Xiangmin Xu. "MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control." arXiv preprint arXiv:2604.21164 (2026). [2026.04]
- Ma, Jianbo, and Richard Cartwright. "Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation." arXiv preprint arXiv:2604.19330 (2026). [2026.04]
- Ouyang, Zhicheng, Seong-Gyun Leem, Bach Viet Do, Haibin Wu, Ariya Rastrow, Yuzong Liu, and Florian Metze. "Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning." arXiv preprint arXiv:2604.08709 (2026).
- Su, Xiaosu, Zihan Sun, Peilei Jia, and Jun Gao. "CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation." arXiv preprint arXiv:2604.08363 (2026). Demo [2026.04]
- Arata, Chihiro, and Kiyoshi Kurihara. "T5Gemma-TTS Technical Report." arXiv preprint arXiv:2604.01760 (2026). Code
- Zhu, Han, Lingxuan Ye, Wei Kang, Zengwei Yao, Liyong Guo, Fangjun Kuang, Zhifeng Han, Weiji Zhuang, Long Lin, and Daniel Povey. "OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models." arXiv preprint arXiv:2604.00688 (2026). Code
2026-03
- Xin, Detai, Shujie Hu, Chengzuo Yang, Chen Huang, Guoqiao Yu, Guanglu Wan, and Xunliang Cai. "LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space." arXiv preprint arXiv:2603.29339 (2026). Code
- Fan, Xiaoyu, Huizhi Xie, Wei Zou, and Yunzhang Chen. "LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling." arXiv preprint arXiv:2603.26364 (2026). Demo
- Diwan, Anuj, Eunsol Choi, and David Harwath. "ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining." arXiv preprint arXiv:2603.28737 (2026). Code
- Khamis, Ahmed, and Hesham Ali Ahmed. "LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models." In Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, pp. 47-54. 2026. Dataset
- Huang, Kexin, Liwei Fan, Botian Jiang, Yaozhou Jiang, Qian Tu, Jie Zhu, Yuqian Zhang et al. "MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions." arXiv preprint arXiv:2603.28086 (2026). Code
- Kumar, Lokesh, Nirmesh Shah, Ashishkumar P. Gudmalwar, and Pankaj Wasnik. "Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?." arXiv preprint arXiv:2603.19831 (2026). Demo [2026.03]
- Zhao, Xiutian, et al. "Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models." arXiv preprint arXiv:2603.17231 (2026).
- An, Xiaochun, Xu Zhang, Xiaoge Li, Ercheng Pei, Qingli Yan, and Lang He. "LLM-driven fine-grained emotion parsing and parameterized mapping for conversational TTS." Pattern Recognition (2026): 113544. [2026.03]
- Kim, Seunghee, Bumkyu Park, Kyudan Jung, Joosung Lee, Soyoon Kim, Jeonghoon Kim, Taeuk Kim, and Hwiyeol Jo. "OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models." arXiv preprint arXiv:2603.23938 (2026). [2026.03]
- Zhang, Yuqian, Donghua Yu, Zhengyuan Lin, Botian Jiang, Mingshu Chen, Yaozhou Jiang, Yiwei Zhao et al. "MOSS-TTSD: Text to Spoken Dialogue Generation." arXiv preprint arXiv:2603.19739 (2026). Code [2026.03]
- Hao, Chunbo, Junjie Zheng, Guobin Ma, Yuepeng Jiang, Huakang Chen, Wenjie Tian, Gongyu Chen, Zihao Chen, and Lei Xie. "YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance." arXiv preprint arXiv:2603.24589 (2026). Code
- Bai, Qibing, Yuhan Du, Tom Ko, Shuai Wang, Yannan Wang, and Haizhou Li. "Controllable Accent Normalization via Discrete Diffusion." arXiv preprint arXiv:2603.14275 (2026). Demo [2026.03]
- Ni, Qinke, Huan Liao, Dekun Chen, Yuxiang Wang, and Zhizheng Wu. "NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation." arXiv preprint arXiv:2603.15352 (2026). Demo [2026.03]
- Liu, Changsong, Tianrui Wang, Ye Ni, Yizhou Peng, and Eng Siong Chng. "Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input." arXiv preprint arXiv:2603.06444 (2026). Demo [2026.03]
- Yang, Mu, and John HL Hansen. "Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech." arXiv preprint arXiv:2603.05977 (2026). Demo [2026.03]
- Lertpetchpun, Thanathai, Thanapat Trachu, Jihwan Lee, Tiantian Feng, Dani Byrd, and Shrikanth Narayanan. "Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data." arXiv preprint arXiv:2603.07534 (2026). [2026.03]
2026-02
- Zhou, R., Zhu, Y. & Zhang, B. MTC-VC: A Multi-Task Contrastive Learning Method for Controllable and Efficiency-Balanced Voice Cloning. SIViP 20, 123 (2026). [2026.02]
- J. Kang, Y. Lee and K. Shim, "Encoder-Free Style-Controllable Text-to-Speech with Voice Attribute Vectors," 2026 International Conference on Electronics, Information, and Communication (ICEIC), Macau, China, 2026, pp. 1-4, doi: 10.1109/ICEIC69189.2026.11386084. [2026.02]
- Bai, Qibing, Shuhao Shi, Shuai Wang, Yukai Ju, Yannan Wang, and Haizhou Li. "CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data." arXiv preprint arXiv:2602.19166 (2026). Demo, Code [2026.02]
- Khamis, Ahmed Khaled, and Hesham Ali. "LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models." arXiv preprint arXiv:2602.15675 (2026). Code [2026.02]
- Zhou, Qiangong, and Nagasaka Tomohiro. "UniTAF: A Modular Framework for Joint Text-to-Speech and Audio-to-Face Modeling." arXiv preprint arXiv:2602.15651 (2026). Code [2026.02]
- Ren, Yong, Jiangyan Yi, Jianhua Tao, Zhengqi Wen, and Tao Wang. "Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards." arXiv preprint arXiv:2602.00560 (2026). [2026.02]
- Roy, Rajarshi, Jonathan Raiman, Sang-gil Lee, Teodor-Dumitru Ene, Robert Kirby, Sungwon Kim, Jaehyeon Kim, and Bryan Catanzaro. "PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models." arXiv preprint arXiv:2602.06053 (2026). [2026.02]
- Lu, Q., Bai, B., Xue, J. et al. REA-TTS: Retrieval-Augmented Expressive Audiobook Text-to-Speech Generation with Contrastive Language-Audio Learning. J. Shanghai Jiaotong Univ. (Sci.) (2026). [2026.02]
- Wang, Siyi, Shihong Tan, Siyi Liu, Hong Jia, Gongping Huang, James Bailey, and Ting Dang. "CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering." arXiv preprint arXiv:2602.03420 (2026). [2026.02]
- Zhou, Li, Hao Jiang, Junjie Li, Tianrui Wang, and Haizhou Li. "EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis." arXiv preprint arXiv:2601.22873 (2026). [2026.02]
- Ren, Yong, Jiangyan Yi, Jianhua Tao, Zhengqi Wen, and Tao Wang. "Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards." arXiv preprint arXiv:2602.00560 (2026). Demo [2026.02]
2026-01
- Singh, Ayush Pratap, Harshit Singh, Nityanand Mathur, Akshat Mandloi, and Sudarshan Kamath. "SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS." arXiv preprint arXiv:2601.17086 (2026). [2026.01]
- Pei, Hanchen, Shujie Liu, Yanqing Liu, Jianwei Yu, Yuanhang Qian, Gongping Huang, Sheng Zhao, and Yan Lu. "A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation." arXiv preprint arXiv:2601.12480 (2026). Demo [2026.01]
- Hu, Hangrui, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo et al. "Qwen3-TTS Technical Report." arXiv preprint arXiv:2601.15621 (2026). Code [2026.01]
- Zhang, Leying, Tingxiao Zhou, Haiyang Sun, Mengxiao Bi, and Yanmin Qian. "DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice." arXiv preprint arXiv:2601.15596 (2026). Demo [2026.01]
- Wang, Yao, Lina Yang, Bingzhen Wang, Jinhong Xu, Dongnan Yang, Miao Zhou, and Yuan Yan Tang. "Emo-DiT: Emotional Speech Synthesis With a Diffusion Model Approach to Enhance Naturalness and Emotional Expressiveness." IEEE Transactions on Affective Computing (2026). Code [2026.01]
- Hu, Jingbin, Huakang Chen, Linhan Ma, Dake Guo, Qirui Zhan, Wenhao Li, Haoyu Zhang et al. "VoiceSculptor: Your Voice, Designed By You." arXiv preprint arXiv:2601.10629 (2026). Demo Code [2026.01]
- Chen, Junyang, Yuhang Jia, Hui Wang, Jiaming Zhou, Yaxin Han, Mengying Feng, and Yong Qin. "CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models." arXiv preprint arXiv:2601.05329 (2026). [2026.01]
- Chen, Dekun, Xueyao Zhang, Yuancheng Wang, Kenan Dai, Li Ma, and Zhizheng Wu. "FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions." arXiv preprint arXiv:2601.04656 (2026). Demo [2026.01]
- Li, Haitao, Chunxiang Jin, Chenglin Li, Wenhao Guan, Zhengxing Huang, and Xie Chen. "ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis." arXiv preprint arXiv:2601.03632 (2026). [2026.01]
- Li, Yunpei, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, and Jingchen Shu. "IndexTTS 2.5 Technical Report." arXiv preprint arXiv:2601.03888 (2026). Demo
- Liang, Qifan, Yuansen Liu, Ruixin Wei, Nan Lu, Junchuan Zhao, and Ye Wang. "Segment-Aware Conditioning for Training-Free Intra-Utterance Emotion and Duration Control in Text-to-Speech." arXiv preprint arXiv:2601.03170 (2026). Demo [2026.01]
- Ren, Yong, Jiangyan Yi, Jianhua Tao, Haiyang Sun, Zhengqi Wen, Hao Gu, Le Xu, and Ye Bai. "OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech." arXiv preprint arXiv:2601.01459 (2026). Demo [2026.01]
2025-12
- Li, Weiqin, Qian Chen, Dan Luo, Tianjiao Du, Yafeng Chen, Zhiyong Wu, Xixin Wu, and Helen Meng. "PerTTS: Personalized and Controllable Zero-Shot Spontaneous Style Text-to-Speech Synthesis." IEEE Transactions on Audio, Speech and Language Processing 34 (2025): 545-556. Demo [2025.12]
- Zhou, Haiyang, Zhihua Huang, and Bowen Li. "CL-EDiff: Cross-Lingual Emotional TTS System Based on Diffusion Model." In Man-Machine Speech Communication: 20th National Conference, NCMMSC 2025, Zhenjiang, China, October 16β19, 2025, Proceedings, p. 152. Springer Nature. Demo [2015.12]
- Lan, Tianwei, Yuhang Guo, Mengyuan Deng, Jing Wang, Wenwu Wang, and Chong Feng. "Controllable timbre cloning and style replication with reference speech examples for multimodal human-computer interaction." Neurocomputing (2025): 132529. Demo [2015.12]
- Yu, Fan, Tao Wang, You Wu, Lin Zhu, Wei Deng, Weisheng Han, Wenchao Wang et al. "JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis." arXiv preprint arXiv:2512.19090 (2025). Demo [2015.12]
- Feng, Pengchao, Yao Xiao, Ziyang Ma, Zhikang Niu, Shuai Fan, Yao Li, Sheng Wang, and Xie Chen. "Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis." arXiv preprint arXiv:2512.18699 (2025). Demo, Code [2025.12]
- Zhou, Wangzixi, Bagus Tris Atmaja, and Sakriani Sakti. "TOWARD NATURAL EMOTIONAL TEXT-TO-SPEECH SYSTEM WITH FINE-GRAINED NON-VERBAL EXPRESSION CONTROL." [2025.12]
- Li, Tao, Wengshuo Ge, Zhichao Wang, Zihao Cui, Yong Ma, Yingying Gao, Chao Deng, Shilei Zhang, and Junlan Feng. "DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec." arXiv preprint arXiv:2512.13251 (2025). Demo [2025.12]
- Kojima, Atsushi, Yusuke Fujita, Hao Shi, Tomoya Mizumoto, Mengjie Zhao, and Yui Sudo. "Conversation Context-aware Direct Preference Optimization for Style-Controlled Speech Synthesis." APSIPA ASC25. [2025.12]
- Okamoto, Umi, Sei Ueno, and Akinobu Lee. "Face-conditioned Large-scale Text-to-Speech via Speaker Embedding Prediction from Facial Images." APSIPA ASC25. [2025.12]
- Yin, Kang, Chunyu Qiang, Sirui Zhao, Xiaopeng Wang, Yuzhe Liang, Pengfei Cai, Tong Xu, Chen Zhang and Enhong Chen. βDMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance.β (2025). Demo [2025.12]
- Wang, Cong, Changfeng Gao, Yang Xiang, Zhihao Du, Keyu An, Han Zhao, Qian Chen, Xiangang Li, Yingming Gao, and Ya Li. "RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS." arXiv preprint arXiv:2512.04552 (2025). [2025.12]
2025-11
- R. -G. Bolborici and A. NeacΕu, "Adding Emotion Conditioning in Speech Synthesis via Multi-Term Classifier-Free Guidance," 2025 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), Cluj-Napoca, Romania, 2025, pp. 86-91. Demo [2025.11]
- Qiang, Chunyu, Kang Yin, Xiaopeng Wang, Yuzhe Liang, Jiahui Zhao, Ruibo Fu, Tianrui Wang et al. "InstructAudio: Unified speech and music generation with natural language instruction." arXiv preprint arXiv:2511.18487 (2025). Demo [2025.11]
- Seung-Bin Kim, Jun-Hyeok Cha, Hyung-Seok Oh, Heejin Choi, and Seong-Whan Lee. 2025. FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control. EMNLP 2025. Demo [2025.11]
- Yejin Jeon, Youngjae Kim, Jihyun Lee, Hyounghun Kim, and Gary Lee. 2025. Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis. EMNLP 2025 Findings. [2025.11]
- Yifu Chen, Shengpeng Ji, Ziqing Wang, Hanting Wang, and Zhou Zhao. 2025. InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model. EMNLP 2025 Findings. Demo [2025.11]
- Shehzeen Samarah Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, and Jason Li. 2025. Koel-TTS: Enhancing LLM-based Speech Generation with Preference Alignment and Classifier Free Guidance. EMNLP 2025. Demo [2025.11]
- Zhenqi Jia, Rui Liu, Berrak Sisman, and Haizhou Li. 2025. Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis. EMNLP 2025. Demo & Code [2025.11]
- Jianxing Yu, Gou Zihao, Chen Li, Zhisheng Wang, Peiji Yang, Wenqing Chen, and Jian Yin. 2025. Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of Speech. EMNLP 2025. Demo [2025.11]
- Anuj Diwan, Zhisheng Zheng, David Harwath, and Eunsol Choi. 2025. Scaling Rich Style-Prompted Text-to-Speech Datasets. EMNLP 2025. Dataset & Model & Code [2025.11]
- Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, and David Harwath. 2025. VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing. EMNLP 2025. Demo [2025.11]
- Zhao, Yiwen, Jiatong Shi, Jinchuan Tian, Yuxun Tang, Jiarui Hai, Jionghao Han, and Shinji Watanabe. "Adapting Speech Language Model to Singing Voice Synthesis." In AI for Music Workshop. Demo [2025.11]
- Yan, Chao, Boyong Wu, Peng Yang, Pengfei Tan, Guoqiang Hu, Yuxin Zhang, Fei Tian et al. "Step-Audio-EditX Technical Report." arXiv preprint arXiv:2511.03601 (2025). Code [2025.11]
-
- Yu, Xinyue, Youqing Fang, Pingyu Wu, Guoyang Ye, Wenbo Zhou, Weiming Zhang, and Song Xiao. "MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement." In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 21, pp. 17966-17974. 2026. Demo
- Yu, Xinyue, Youqing Fang, Pingyu Wu, Guoyang Ye, Wenbo Zhou, Weiming Zhang, and Song Xiao. "MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement." In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 21, pp. 17966-17974. 2026. Demo
2025-10
- Tu, Wenming, Guanrou Yang, Ruiqi Yan, Wenxi Chen, Ziyang Ma, Yipeng Kang, Kai Yu, Xie Chen, and Zilong Zheng. "UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models." arXiv preprint arXiv:2510.22588 (2025). Demo, Code [2025.10]
- Lou, Haowei, Hye-Young Paik, Wen Hu, and Lina Yao. "ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation." arXiv preprint arXiv:2510.18308 (2025). Code [2025.10]
- Peng, Yizhou, Yukun Ma, Chong Zhang, Yi-Wen Chao, Chongjia Ni, and Bin Ma. "Mismatch Aware Guidance for Robust Emotion Control in Auto-Regressive TTS Models." arXiv preprint arXiv:2510.13293 (2025).
- Li, Haoxun, Yu Liu, Yuqing Sun, Hanlei Shi, Leyuan Qu, and Taihao Li. "EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS." arXiv preprint arXiv:2510.05758 (2025). Demo [2025.10]
2025-09
- Wang, Yue, Ruotian Ma, Xingyu Chen, Zhengliang Shi, Wanshun Chen, Huang Liu, Jiadi Yao et al. "BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs." arXiv preprint arXiv:2509.26514 (2025). Demo & Code [2025.09]
- Zhang, Ziyu, Hanzhao Li, Jingbin Hu, Wenhao Li, and Lei Xie. "HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis." arXiv preprint arXiv:2509.25842 (2025). Demo [2025.09]
- Wang, Tianrui, Haoyu Wang, Meng Ge, Cheng Gong, Chunyu Qiang, Ziyang Ma, Zikang Huang et al. "Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis." arXiv preprint arXiv:2509.24629 (2025). Demo [2025.09]
- Wang, Sirui, Andong Chen, and Tiejun Zhao. "Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation." arXiv preprint arXiv:2509.20378 (2025). [2025.09]
- Liu, Min, JingJing Yin, Xiang Zhang, Siyu Hao, Yanni Hu, Bin Lin, Yuan Feng, Hongbin Zhou, and Jianhao Ye. "Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook." arXiv preprint arXiv:2509.17516 (2025). Demo [2025.09]
- Lu, Ye-Xin, Yu Gu, Kun Wei, Hui-Peng Du, Yang Ai, and Zhen-Hua Ling. "DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis." arXiv preprint arXiv:2509.14684 (2025). Demo [2025.09]
2025-04 - 2025-08
- Tian, Fengping, Chenyang Lyu, Xuanfan Ni, Haoqin Sun, Qingjuan Li, Zhiqiang Qian, Haijun Li et al. "Marco-voice technical report." arXiv preprint arXiv:2508.02038 (2025). Demo & Code [2025.08]
- Zhang, Xueyao, Junan Zhang, Yuancheng Wang, Chaoren Wang, Yuanzhe Chen, Dongya Jia, Zhuo Chen, and Zhizheng Wu. "Vevo2: Bridging Controllable Speech and Singing Voice Generation via Unified Prosody Learning." arXiv preprint arXiv:2508.16332 (2025). Demo [2025.08]
- Park, Joonyong, and Kenichi Nakamura. "EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens." arXiv preprint arXiv:2508.11273 (2025). [2025.08]
- Bauer, Judith, Frank Zalkow, Meinard MΓΌller, and Christian Dittmar. "Explicit Emphasis Control in Text-to-Speech Synthesis." In Proc. SSW 2025, pp. 21-27. 2025. Demo
- Lemerle, ThΓ©odor, Nicolas Obin, and Axel Roebel. "Lina-Style: Word-Level Style Control in TTS via Interleaved Synthetic Data." In Proc. SSW 2025, pp. 35-39. 2025. [2025.08]
- Zhu, Boyu, Cheng Gong, Muyang Wu, Ruihao Jing, Fan Liu, Xiaolei Zhang, Chi Zhang, and Xuelong Li. ": A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation." arXiv preprint arXiv:2508.09702 (2025). Dataset [2025.08]
- Xie, Tianxin, Shan Yang, Chenxing Li, Dong Yu, and Li Liu. "EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering." arXiv preprint arXiv:2508.03543 (2025). Demo [2025.08]
- Wu, Zhuojun, Dong Liu, Juan Liu, Yechen Wang, Linxi Li, Liwei Jin, Hui Bu, Pengyuan Zhang, and Ming Li. "SMIIP-NV: A Multi-Annotation Non-Verbal Expressive Speech Corpus in Mandarin for LLM-Based Speech Synthesis." ACM Multimedia, 2025. Dataset [2025.07]
- Niu, Rui, Weihao Wu, Jie Chen, Long Ma, and Zhiyong Wu. "A Multi-Stage Framework for Multimodal Controllable Speech Synthesis." arXiv preprint arXiv:2506.20945 (2025). Demo [2025.06]
- Zhou, Siyi, Yiquan Zhou, Yi He, Xun Zhou, Jinchao Wang, Wei Deng, and Jingchen Shu. "IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech." arXiv preprint arXiv:2506.21619 (2025). Demo, Code [2025.06]
- Ju, Zeqian, et al. "EmergentTTS-Eval: A Benchmark for Evaluating Emergent Abilities in Text-to-Speech Systems." Advances in Neural Information Processing Systems 38 (NeurIPS 2025 Datasets and Benchmarks Track). 2025. Code Dataset [2025.05]
- Rong, Yan, Jinting Wang, Guangzhi Lei, Shan Yang, and Li Liu. "AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation." arXiv preprint arXiv:2505.22053 (2025). Demo [2025.05]
- Rong, Yan, Shan Yang, Guangzhi Lei, and Li Liu. "Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation." arXiv preprint arXiv:2504.11002 (2025). Demo [2025.04]
Non-autoregressive Controllable TTS
Below are representative non-autoregressive controllable TTS methods. Each entry follows this format: method name, zero-shot capability, controllability, acoustic model, vocoder, acoustic feature, release date, and code/demo.
NOTE: MelS and LinS represent Mel Spectrogram and Linear Spectrogram, respectively. Among todayβs TTS systems, MelS, latent features (from VAEs, diffusion models, and other flow-based methods), and various types of discrete tokens are the most commonly used acoustic representations.
- ProEmo, Zero-shot (β), Controllability (Pitch, Energy, Emotion, Description), Transformer, HiFi-GAN, MelS, 2025.01, Code
- DrawSpeech, Zero-shot (β), Controllability (Energy, Prosody), Diffusion, HiFi-GAN, MelS, 2025.01, Demo, Code
- DiffStyleTTS, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre), Transformer + Diffusion, HiFi-GAN, MelS, 2025.01, Demo
- HED, Zero-shot (β), Controllability (Emotion), Flow-based Diffusion, Vocos, MelS, 2024.12, Demo
- EmoDubber, Zero-shot (β), Controllability (Prosody, Timbre, Emotion), Transformer + Flow, Flow-based Vocoder, MelS, 2024.12, Demo
- EmoSphere++, Zero-shot (β), Controllability (Prosody, Timbre, Emotion), Transformer + Flow, BigVGAN, MelS, 2024.11, Demo, Code
- MSKU-VTTS, Zero-shot (β), Controllability (Environment, Description), Diffusion, BigvGAN, MelS, 2024.10
- NanoVoice, Zero-shot (β), Controllability (Timbre), Diffusion, BigVGAN, MelS, 2024.09
- NansyTTS, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Description), Transformer, NANSY++, MelS, 2024.09, Demo
- StyleTTS-ZS, Zero-shot (β), Controllability (Timbre), Flow-based Diffusion + GAN, Mel-based Decoder, MelS, 2024.09, Demo
- E1 TTS, Zero-shot (β), Controllability (Timbre), DiT + Flow, BigVGAN, Token + MelS, 2024.09, Demo
- SimpleSpeech 2, Zero-shot (β), Controllability (Speed, Timbre), Flow-based DiT, SQ Codec, Token, 2024.08, Demo, Code
- CCSP, Zero-shot (β), Controllability (Timbre), Diffusion, RVQ-based Codec, Token, 2024.07, Demo
- ArtSpeech, Zero-shot (β), Controllability (Timbre), RNN + CNN, HiFI-GAN, MelS, 2024.07, Demo, Code
- DEX-TTS, Zero-shot (β), Controllability (Timbre), Diffusion, HiFi-GAN, MelS, 2024.06, Code
- MobileSpeech, Zero-shot (β), Controllability (Timbre), Transformer, Vocos, Token, 2024.06, Demo
- E2 TTS, Zero-shot (β), Controllability (Timbre), Transformer + Flow, BigVGAN, MelS, 2024.06, Demo, Code (unofficial)
- DiTTo-TTS, Zero-shot (β), Controllability (Speed, Timbre), DiT + VAE, BigVGAN, MelS, 2024.06, Demo
- SimpleSpeech, Zero-shot (β), Controllability (Timbre), Transformer + Diffusion, SQ Codec, Token, 2024.06, Demo, Code
- AST-LDM, Zero-shot (β), Controllability (Timbre, Environment, Description), Diffusion + VAE, HiFi-GAN, MelS, 2024.06, Demo
- ControlSpeech, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion, Description), Transformer + Diffusion, FACodec, Token, 2024.06, Demo, Code
- InstructTTS, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Emotion, Description), Transformer + Diffusion, HiFi-GAN, Token, 2024.05, Demo
- NaturalSpeech 3, Zero-shot (β), Controllability (Speed, Prosody, Timbre), Transformer + Diffusion, FACodec, Token, 2024.04, Demo
- FlashSpeech, Zero-shot (β), Controllability (Timbre), Latent Consistency Model, EnCodec, Token, 2024.04, Demo, Code
- Audiobox, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Environment, Description), Transformer + Flow, EnCodec, MelS, 2023.12, Demo
- HierSpeech++, Zero-shot (β), Controllability (Timbre), Transformer + VAE + Flow, BigVGAN, MelS, 2023.11, Demo, Code
- E3 TTS, Zero-shot (β), Controllability (Timbre), Diffusion, Not required, Waveform, 2023.11, Demo
- P-Flow, Zero-shot (β), Controllability (Timbre), Transformer + Flow, HiFi-GAN, MelS, 2023.10, Demo, Code (unofficial)
- SpeechFlow, Zero-shot (β), Controllability (Timbre), Transformer + Flow, HiFi-GAN, MelS, 2023.10, Demo
- PromptTTS++, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Emotion, Description), Transformer + Diffusion, BigVGAN, MelS, 2023.09, Demo, Code
- DuIAN-E, Zero-shot (β), Controllability (Pitch, Speed, Prosody), CNN + RNN, HiFi-GAN, MelS, 2023.09, Demo
- VoiceLDM, Zero-shot (β), Controllability (Pitch, Prosody, Timbre, Emotion, Environment, Description), Diffusion, HiFi-GAN, MelS, 2023.09, Demo, Code
- PromptTTS 2, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Description), Diffusion, RVQ-based Codec, Latent Feature, 2023.09, Demo
- MegaTTS 2, Zero-shot (β), Controllability (Prosody, Timbre, Emotion), Decoder-only Transformer + GAN, HiFi-GAN, MelS, 2023.07, Demo, Code (unofficial)
- VoiceBox, Zero-shot (β), Controllability (Timbre), Transformer + Flow, HiFi-GAN, MelS, 2023.06, Demo, Code (unofficial)
- StyleTTS 2, Zero-shot (β), Controllability (Prosody, Timbre, Emotion), Flow-based Diffusion + GAN, HiFi-GAN / iSTFTNet, MelS, 2023.06, Demo, Code
- PromptStyle, Zero-shot (β), Controllability (Pitch, Prosody, Timbre, Emotion, Description), VITS + Flow, HiFi-GAN, MelS, 2023.05, Demo
- NaturalSpeech 2, Zero-shot (β), Controllability (Timbre), Diffusion, RVQ-based Codec, Token, 2023.04, Demo, Code (unofficial)
- Grad-StyleSpeech, Zero-shot (β), Controllability (Timbre), Score-based Diffusion, HiFi-GAN, MelS, 2022.11, Demo
- PromptTTS, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion, Description), Bert + Transformer, HiFi-GAN, MelS, 2022.11, Demo
- CLONE, Zero-shot (β), Controllability (Pitch, Speed, Prosody), Transformer + CNN, WaveNet, MelS + LinS, 2022.07, Demo
- Cauliflow, Zero-shot (β), Controllability (Speed, Prosody), BERT + Flow, UP WaveNet, MelS, 2022.06
- GenerSpeech, Zero-shot (β), Controllability (Timbre), Transformer + Flow, HiFi-GAN, MelS, 2022.05, Demo
- StyleTTS, Zero-shot (β), Controllability (Timbre), CNN + RNN, HiFi-GAN, MelS, 2022.05, Code
- YourTTS, Zero-shot (β), Controllability (Timbre), Transformer + Flow, HiFi-GAN, LinS, 2021.12, Demo & Checkpoint
- DelightfulTTS, Zero-shot (β), Controllability (Pitch, Speed, Prosody), Transformer + CNN, HiFiNet, MelS, 2021.11, Demo
- Meta-StyleSpeech, Zero-shot (β), Controllability (Timbre), Transformer, MelGAN, MelS, 2021.06, Code
- SC-GlowTTS, Zero-shot (β), Controllability (Timbre), Transformer + Flow, HiFi-GAN, MelS, 2021.06, Demo, Code
- StyleTagging-TTS, Zero-shot (β), Controllability (Timbre, Emotion), Transformer + CNN, HiFi-GAN, MelS, 2021.04, Demo
- Parallel Tacotron, Zero-shot (β), Controllability (Prosody), Transformer + CNN, WaveRNN, MelS, 2020.10, Demo
- FastPitch, Zero-shot (β), Controllability (Pitch, Prosody), Transformer, WaveGlow, MelS, 2020.06, Code
- FastSpeech 2, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody), Transformer, Parallel WaveGAN, MelS, 2020.06, Code (unofficial)
- FastSpeech, Zero-shot (β), Controllability (Speed, Prosody), Transformer, WaveGlow, MelS, 2019.05, Code (unofficial)
Autoregressive Controllable TTS
Below are representative non-autoregressive controllable TTS methods. Each entry follows this format: method name, zero-shot capability, controllability, acoustic model, vocoder, acoustic feature, release date, and code/demo.
NOTE: MelS and LinS represent Mel Spectrogram and Linear Spectrogram, respectively. Among todayβs TTS systems, MelS, latent features (from VAEs, diffusion models, and other flow-based methods), and various types of discrete tokens are the most commonly used acoustic representations.
- EmoVoice, Zero-shot (β), Controllability (Emotion, Description), Decoder-only Transformer, HiFi-GAN, Token, 2025.04, Demo
- Spark-TTS, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre), Decoder-only Transformer, BiCodec, Token, 2025.03, Code
- Vevo, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion), Decoder-only Transformer, BigVGAN, Token + MelS, 2025.02, Demo, Code
- Step-Audio, Zero-shot (β), Controllability (Prosody, Timbre, Emotion, Description), Decoder-only Transformer, Flow-based Vocoder, Token, 2025.02, Code
- FleSpeech, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion, Description), Flow-based DiT, WaveGAN, Latent Feature, 2025.01, Demo
- IDEA-TTS, Zero-shot (β), Controllability (Timbre, Environment), Transformer, Flow-based Vocoder, LinS + MelS, 2024.12, Demo, Code
- KALL-E, Zero-shot (β), Controllability (Prosody, Timbre, Emotion), Decoder-only Transformer, WaveVAE, Latent Feature, 2024.12, Demo
- IST-LM, Zero-shot (β), Controllability (Prosody, Timbre), Decoder-only Transformer, HiFi-GAN, Token + MelS, 2024.12
- SLAM-Omni, Zero-shot (β), Controllability (Prosody, Timbre), Decoder-only Transformer, HiFi-GAN, Token + MelS, 2024.12, Demo, Code
- FishSpeech, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, Firefly-GAN,Token, 2024.11, Code
- HALL-E, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, EnCodec, Token, 2024.10
- Takin, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Emotion, Description), Decoder-only Transformer + Flow, HiFi-GAN, Token + MelS, 2024.09, Demo
- Emotional Dimension Control, Zero-shot (β), Controllability (Timbre, Emotion), Decoder-only Transformer + Flow, HiFI-GAN, Token + MelS, 2024.09, Demo
- CoFi-Speech, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, BigVGAN, Token + MelS, 2024.09, Demo
- FireRedTTS, Zero-shot (β), Controllability (Prosody, Timbre), Decoder-only Transformer + Flow, BigVGAN-v2, Token + MelS, 2024.09, Demo, Code
- Emo-DPO, Zero-shot (β), Controllability (Emotion), Decoder-only Transformer, HiFi-GAN, Token + MelS, 2024.09, Demo
- VoxInstruct, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion, Description), Decoder-only Transformer, Vocos, Token, 2024.08, Demo, Code
- MELLE, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, HiFi-GAN, MelS, 2024.07. Demo
- CosyVoice, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Emotion, Description), Decoder-only Transformer + Flow, HiFi-GAN, Token, 2024.07, Demo, Code
- XTTS, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer + GAN, HiFi-GAN-based Vococder, Token + MelS, 2024.06, Demo, Code
- VoiceCraft, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, HiFi-GAN, Token, 2024.06, Code
- Seed-TTS, Zero-shot (β), Controllability (Timbre, Emotion), Decoder-only Transformer + DiT, Unknown Vocoder, Latent Feature, 2024.06, Demo
- VALL-E 2, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, Vocos, Token, 2024.06, Demo, Code (unofficial 1), Code (unofficial 2)
- VALL-E R, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, Vocos, Token, 2024.06, Demo
- ARDiT, Zero-shot (β), Controllability (Speed, Timbre), Decoder-only DiT, BigVGAN, MelS, 2024.06, Demo
- RALL-E, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, SoundStream, Token, 2024.05, Demo
- CLaM-TTS, Zero-shot (β), Controllability (Timbre), Encoder-decoder Transformer, BigVGAN, Token + MelS, 2024.04, Demo
- BaseTTS, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, Speechcode Decoder, Token, 2024.02, Demo
- ELLA-V, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, EnCodec, Token, 2024.01, Demo
- UniAudio, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Timbre, Description), Decoder-only Transformer, UniAudio Codec, Token, 2023.10, Demo, Code
- Salle, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion, Description), Decoder-only Transformer, EnCodec, Token, 2023.08, Demo
- SC VALL-E, Zero-shot (β), Controllability (Pitch, Energy, Speed, Prosody, Timbre, Emotion), Decoder-only Transformer, EnCodec, Token, 2023.07, Demo, Code
- MegaTTS, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer + GAN, HiFi-GAN, MelS, 2023.06, Demo
- TorToise, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer + Diffusion, UnivNet, MelS, 2023.05, Code
- Make-a-voice, Zero-shot (β), Controllability (Timbre), Encoder-decoder Transformer, Unit-based Vocoder, Token, 2023.05, Demo
- VALL-E X, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, EnCodec, Token, 2023.03, Demo, Code (unofficial)
- SpearTTS, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, SoundStream, Token, 2023.02, Demo, Code (unofficial)
- VALL-E, Zero-shot (β), Controllability (Timbre), Decoder-only Transformer, EnCodec, Token, 2023.01, Demo, Code (unofficial 1), Code (unofficial 2)
- MsEmoTTS, Zero-shot (β), Controllability (Pitch, Prosody, Emotion), CNN + RNN, WaveRNN, MelS, 2022.01, Demo
- Flowtron, Zero-shot (β), Controllability (Pitch, Speed, Prosody), CNN + RNN, WaveGlow, MelS, 2020.07, Demo, Code
- DurIAN, Zero-shot (β), Controllability (Pitch, Speed, Prosody), CNN + RNN, MB-WaveRNN, MelS, 2019.09, Demo, Code (unofficial)
- VAE-Tacotron, Zero-shot (β), Controllability (Pitch, Speed, Prosody), VAE, WaveNet, MelS, 2019.02, Code (unoffcial 1), Code (unoffcial 2)
- GMVAE-Tacotron, Zero-shot (β), Controllability (Pitch, Speed, Prosody, Description), VAE, WaveRNN, MelS, 2018.12, Demo, Code (unofficial)
- GST-Tacotron, Zero-shot (β), Controllability (Pitch, Prosody), CNN + RNN, Griffin-Lim, LinS, 2018.03, Demo, Code (unofficial)
- Prosody-Tacotron, Zero-shot (β), Controllability (Pitch, Prosody), RNN, WaveNet, MelS, 2018.03, Demo
Datsets
A summary of open-source datasets for controllable TTS:
| Dataset | Hours | #Speakers | Labels | Lang | Release | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pit. | Ene. | Spe. | Age | Gen. | Emo. | Emp. | Acc. | Top. | Des. | Dia. | |||||
| SpeechCraft | 2,391 | 3,200 | β | β | β | β | β | β | β | β | β | en,zh | 2024 | ||
| Parler-TTS | 50,000 | / | β | β | β | β | β | β | en | 2024 | |||||
| MSceneSpeech | 13 | 13 | β | zh | 2024 | ||||||||||
| VccmDataset | 330 | 1,324 | β | β | β | β | β | β | en | 2024 | |||||
| CLESC | <1 | / | β | β | β | β | en | 2024 | |||||||
| TextrolSpeech | 330 | 1,324 | β | β | β | β | β | β | en | 2023 | |||||
| DailyTalk | 20 | 2 | β | β | β | en | 2023 | ||||||||
| MagicData-RAMC | 180 | 663 | β | β | zh | 2022 | |||||||||
| PromptSpeech | / | / | β | β | β | β | β | en | 2022 | ||||||
| WenetSpeech | 10,000 | / | β | zh | 2021 | ||||||||||
| GigaSpeech | 10,000 | / | β | en | 2021 | ||||||||||
| ESD | 29 | 10 | β | en,zh | 2021 | ||||||||||
| CommonVoice | 2,500 | 50,000 | β | β | β | multi | 2020 | ||||||||
| AISHELL-3 | 85 | 218 | β | β | β | zh | 2020 | ||||||||
| Taskmaster-1 | / | / | β | en | 2019 | ||||||||||
| CMU-MOSEI | 65 | 1,000 | β | en | 2018 | ||||||||||
| RAVDESS | / | 24 | β | β | en | 2018 | |||||||||
| RECOLA | 3.8 | 46 | β | fr | 2013 | ||||||||||
| IEMOCAP | 12 | 10 | β | β | β | β | β | en | 2008 |
Abbreviations: Pit(ch), Ene(rgy)=volume, Spe(ed), Gen(der), Emo(tion), Emp(hasis), Acc(ent), Top(ic), Des(cription), Env(ironment), Dia(logue).
Evaluation
| Metric | Type | Eval Target | GT Required |
|---|---|---|---|
| Mel-Cepstral Distortion (MCD) | Objective | Acoustic similarity | β |
| Frequency Domain Score Difference (FDSD) | Objective | Acoustic similarity | β |
| Word Error Rate (WER) | Objective | Intelligibility | β |
| Cosine Similarity | Objective | Speaker similarity | β |
| Perceptual Evaluation of Speech Quality (PESQ) | Objective | Perceptual quality | β |
| Signal-to-Noise Ratio (SNR) | Objective | Perceptual quality | β |
| Mean Opinion Score (MOS) | Subjective | Preference | |
| Comparison Mean Opinion Score (CMOS) | Subjective | Preference | |
| AB Test | Subjective | Preference | |
| ABX Test | Subjective | Perceptual similarity | β |
GT: Ground truth, : Lower is better, : Higher is better.