DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

November 25, 2025 · View on GitHub

ICASSP 2025 [Paper] [arXiv] [Demo]

Status

This project is currently under active development. We are continuously updating and improving it, with more usage details and features to be released in the future.

Getting started

Download dataset and checkpoints

  1. Download the LJSpeech dataset and place the dataset into data/dataset with structure looks like below:
data/dataset/LJSpeech-1.1
 ┣ metadata.csv
 ┣ wavs
 ┃ ┣ LJ001-0001.wav
 ┃ ┣ LJ001-0002.wav 
 ┃ ┣ ...
 ┣ README
  1. Download the alignments of the LJSpeech dataset LJSpeech.zip. You have to unzip the files in data/dataset/LJSpeech-1.1
  2. Download checkpoints from HuggingFace.
  3. Place the checkpoints into data/checkpoints/

Preprocessing

python preprocessing.py

Training

Train the VAE (Optional)

CUDA_VISIBLE_DEVICES=0 python drawspeech/train/autoencoder.py -c drawspeech/config/vae_ljspeech_22k.yaml

If you don't want to train the VAE, you can just use the VAE checkpoint that we provide.

  • set the variable reload_from_ckpt in drawspeech_ljspeech_22k.yaml to data/checkpoints/vae.ckpt

Train the DrawSpeech

CUDA_VISIBLE_DEVICES=0 python drawspeech/train/latent_diffusion.py -c drawspeech/config/drawspeech_ljspeech_22k.yaml

Inference

If you have trained the model using drawspeech_ljspeech_22k.yaml, use the following syntax:

CUDA_VISIBLE_DEVICES=0 python drawspeech/infer.py --config_yaml drawspeech/config/drawspeech_ljspeech_22k.yaml --list_inference tests/inference.json

If not, please specify the DrawSpeech checkpoint:

CUDA_VISIBLE_DEVICES=0 python drawspeech/infer.py --config_yaml drawspeech/config/drawspeech_ljspeech_22k.yaml --list_inference tests/inference.json --reload_from_ckpt data/checkpoints/drawspeech.ckpt

Acknowledgement

This repository borrows codes from the following repos. Many thanks to the authors for their great work.

Citation

@INPROCEEDINGS{10889767,
  author={Chen, Weidong and Yang, Shan and Li, Guangzhi and Wu, Xixin},
  booktitle={ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, 
  title={DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions}, 
  year={2025},
  volume={},
  number={},
  pages={1-5},
  doi={10.1109/ICASSP49660.2025.10889767}}
}