Speech-Prompts-Adapters

August 4, 2023 · View on GitHub

This Repository surveys the paper focusing on Adapters and Prompting methods for Speech Processing.

NEWS

  • In ICASSP 2023, we will give a tutorial about Paramter-Efficient Learning for speech processing and natural langauge processing. I (Kai-Wei Chang) will cover the topics of adapters and prompts for speech processing.

ICASSP 2023 Tutorial Information

  • Title: Parameter-Efficient Learning for Speech and Language Processing: Adapters, Prompts, and Reprogramming
  • Conference: ICASSP 2023
  • Website: ICASSP 2023 - Tutorials
  • Parameter-Efficient Learning for Speech Processing Slides

Presenters:

  • Pin-Yu Chen (IBM Research)
  • Hung-yi Lee (National Taiwan University)
  • Chao-Han Huck Yang (Georgia Institute of Technology )
  • Kai-Wei Chang (National Taiwan University)
  • Cheng-Han Chiang (National Taiwan University)

Adapters and Prompting for Speech Processing

Adapters for Speech Processing

TitleAuthorsModalityTaskLink
Differentially Private Adapters for Parameter Efficient Acoustic ModelingChun-Wei Ho et al.Speechkeyword SpottingInterspeech 2023
Beyond Universal Transformer: block reusing with adaptor in Transformer for automatic speech recognitionHaoyu Tang et al.SpeechASRarXiv 2023
A Parameter-Efficient Learning Approach to Arabic Dialect Identification with Pre-Trained General-Purpose Speech ModelSrijith Radhakrishnan et al.SpeechDialect IdentificationInterspeech 2023
CHAPTER: Exploiting Convolutional Neural Network Adapters for Self-supervised Speech ModelsZih-Ching Chen et al.Speech[Multiple]arXiv 2022
Parameter Efficient Transfer Learning for Various Speech Processing TasksShinta Otake et al.Speech[Multiple]arXiv 2022
Parameter-efficient transfer learning of pre-trained Transformer models for speaker verification using adaptersJunyi Peng et al.SpeechSpeaker VerificationarXiv 2022
Exploring Efficient-tuning Methods in Self-supervised Speech ModelsZih-Ching Chen et al.Speech[Multiple]SLT 2022
DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised Learning and Its Application to Children’s ASRRuchao Fan, Abeer AlwanSpeechASRInterspeech 2022
Speaker adaptation for Wav2vec2 based dysarthric ASRMurali Karthick Baskar et al.SpeechASRInterspeech 2022
Adaptive multilingual speech recognition with pretrained modelsNgoc-Quan Pham et al.SpeechASRInterspeech 2022
An Adapter Based Pre-Training for Efficient and Scalable Self-Supervised Speech Representation LearningSamuel Kessler et al.SpeechASRICASSP 2022
Efficient Adapter Transfer of Self-Supervised Speech Models for Automatic Speech RecognitionBethan Thomas et al.SpeechASRICASSP 2022
Scaling End-to-End Models for Large-Scale Multilingual ASRBo Li et al.SpeechASRASRU 2021
Meta-Adapter: Efficient Cross-Lingual Adaptation With Meta-LearningWenxin Hou et al.SpeechASRICASSP 2021
Exploiting Adapters for Cross-Lingual Low-Resource Speech RecognitionWenxin Hou et al.SpeechASRTASLP 2021
Lightweight Adapter Tuning for Multilingual Speech TranslationHang Le et al.SpeechSpeech TranslationACL-IJCNLP 2021
Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented SpeechKatrin Tomanek et al.SpeechASREMNLP 2021
Multilingual Speech Recognition with Self-Attention Structured ParameterizationYun Zhu et al.SpeechASRInterspeech 2020
Large-Scale Multilingual Speech Recognition with a Streaming End-to-End ModelAnjuli Kannan et al.SpeechASRInterspeech 2019

Prompting for Speech Processing

TitleAuthorsModalityTaskLink
Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task GeneralizationPuyuan Peng et al.Speech[Multiple]Interspeech 2023
From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech RecognitionChao-Han Huck Yang et al.SpeechASRICASSP 2023
SpeechPrompt v2: Prompt Tuning for Speech Classification TasksKai-Wei Chang et al.Speech[Multiple]arXiv 2023
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal SupervisionEugene Kharitonov et al.Text & SpeechTTSarXiv 2023
Describing emotions with acoustic property prompts for speech emotion recognitionHira Dhamyal et al.Text & SpeechERarXiv 2022
PromptTTS: Controllable Text-to-Speech with Text DescriptionsZhifang Guo et al.Text & SpeechTTSarXiv 2022
Neural Model Reprogramming with Similarity Based Mapping for Low-Resource Spoken Command ClassificationHao Yen et al.SpeechSpoken Command RecognitionarXiv 2022
WAVPROMPT: Towards Few-Shot Spoken Language Understanding with Frozen Language ModelsHeting Gao et al.Text & SpeechSLUInterspeech 2022
An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing TasksKai-Wei Chang et al.Speech[Multiple]Interspeech 2022

Reprogramming and Prompting

For more information about reprogramming and prompting for large pre-trained models, please refer to the "awesome-neural-reprogramming-acoustic-prompting" repository. This topic was also covered in ICASSP 2022 tutorial by Dr. Pin-Yu Chen and Dr. Huck Yang.


Parameter Efficient Learning Methods

TitleAuthorsLink
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-modelsElad Ben Zaken et al.ACL 2022
Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He et al.ICLR 2022
LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu et al.ICLR 2022
Parameter-Efficient Transfer Learning for NLPNeil Houlsby et al.ICML 2019

Acknowledgment

We thank Kuang-Chen Peng, Tzu-Han Lin, and Fabian Ritter for their invaluable contribution to the initial collection.

Contact

This repository is maintained by Kai-Wei Chang (kaiwei.chang.tw@gmail.com) and Zih-Ching Chen. Feel free to contact us or make a pull request