[](https://github.com/FORARTfe/HyMPS#- "AUDIO section") [](https://github.com/FORARTfe/HyMPS/blob/main/Audio/AI-based.md#-- "AI-based category") [](https://github.com/FORARTfe/HyMPS/blob/main/Audio/AI-Voicing.md#--- "Voicing page")
March 10, 2026 · View on GitHub
📁 Cloners - Denoisers - Dubbers - Enhancers - Extractors - Stylers - Speech - TTSers
Warning
Cloners ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| Applio | VITS-based Voice Conversion focused on simplicity, quality and performance | - | |||
| OpenVoice | Versatile Instant Voice Cloning (paper) | - | |||
| Real-Time Voice Cloning | Clone a voice in 5 seconds to generate arbitrary speech in real-time | - | |||
| Real-Time Voice Cloning | An implementation of Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis paper with a vocoder that works in real-time | - |
Denoisers ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| denoise-autoencoder | Denoise audio with convolutional autoencoder | NA | |||
| Denoiser | An AI model to remove noise from the input audio using deep learning model which predicts the type of noise present and filter it out from the audio to give noise-free results | 1 x h5 | |||
| Noise2Noise-denoising | Source code for the "Speech Denoising without Clean Training Data: a Noise2Noise Approach" paper | 24 x pth | |||
| Audio-Denoiser-CNN | - | 1 x h5 | |||
| DNP | A PyTorch implementation for the "Speech Denoising by Accumulating Per-Frequency Modeling Fluctuations" paper | - | |||
| SAB-cnn-audio-denoiser | Tensorflow 2.0 implementation of the paper A Fully Convolutional Neural Network for Speech Enhancement | NA | |||
| Clean Sound with AI | Aims to clean sound recordings from noisy environments using a Convolutional Neural Network (CNN) based on the CleanUNet model | NA | |||
| denoiser | A PyTorch implementation for the "Real Time Speech Enhancement in the Waveform Domain" paper | 1 2 3 4 x th | |||
| Speech-enhancement | Deep learning for audio denoising | 1 x h5 | |||
| DeepFilterNet | Noise supression using deep filtering | 3 x ckpt/best 12 x onnx | |||
| speech-denoiser | denoiser | NA | |||
| DTLN | Tensorflow 2.x implementation of the DTLN real time speech denoising model | 3 x h5 3 x tflite 2 x onnx |
Dubbers ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| Voxella | This app leverages advanced AI algorithms to automatically detect audio and text in the original language, providing effortless translation and native-like voice dubbing | - | |||
| Voice Craft AI | An AI tool to dub videos into multiple regional languages and lip-sync at the same time | - | |||
| DubFlow | It lets you effortlessly dub YouTube videos into any language with high-quality translations and synced audio | - | |||
| Linly-Dubbing | An intelligent multi-language AI dubbing and translation tool that offers diverse and high-quality dubbing options by integrating Linly-Talker’s digital human lip-sync technology, creating a more natural multi-language video experience | - | |||
| voice_ukr_to_eng | Tool to generate English AI Dubbing for a YouTube video | - | |||
| Emotionally-Intelligent-AI-based-movie-dubbing | AI to seamlessly translate and dub content into any language while preserving the original speaker's emotions, characteristics, and authenticity | - | |||
| Multilingual audio visual system with lip synchronization using GAN | This system takes a video in any language and generates a new video with synchronized lip movements speaking in English | - | |||
| Video Dubbing Tool | A fully-featured, multi-language video dubbing tool with a modern Streamlit GUI | - | |||
| Auto Synced & Translated Dubs | Automatically translates the text of a video into chosen languages based on a subtitle file, uses AI voice to dub the video, while keeping it properly synced to the original video using the subtitle's timings | - | |||
| Kara-Audio | Gradio web-ui for vocal remover that uses demucs and MDX-Net + automatic subtitle creation using faster Whisper | - | |||
| Youtube Auto Dubbing | Simple tool for dubbing youtube videos with AI generatied voice (inspired by Auto Synced & Translated Dubs) | - | |||
| pyvideotrans | Translate the video from one language to another and add dubbing | - | |||
| Srt-AI-Voice-Assistant | Subtitle dubbing with multiple AI projects | - | |||
| AutoDub | An advanced AI-powered tool that automatically translates and dubs YouTube videos into different languages while dynamically adjusting video speed | - | |||
| Dubbing AI | Project for dubbing a video in many languages and with many different voices with the power of the AI | - | |||
| Dublaris AI - Video Translator | An automated tool for multilingual video dubbing and subtitling | - | |||
| Pollyduble | Automatic Dubbing with Voice Cloning and Speech Recognition using OpenVoice, MeloTTS, Faster Whisper, VoiceFixer, python-audio-separator and FFmpeg | - | |||
| AI Dubs over Subs | A set of python scripts that take a video file as input and try to create a new video dubbed by AI | - | |||
| AI Video Dubbing | Automates translation and dubbing by transcribing audio with Speech-to-Text, translating it with the Translation API, and generating speech with Text-to-Speech using Google APIs | - | |||
| Google AI Dubbing | Allows you to create localized videos using the same video base and adding translations using Google AI Powered TextToSpeech API | - | |||
| Open dubbing | An AI dubbing system which uses machine learning models to automatically translate and synchronize audio dialogue into different languages | - | |||
| YouDub | An innovative open source tool that focuses on translating and dubbing premium videos from platforms such as YouTube into Chinese | - | |||
| InstantDubing: AI Video Translator | It uses cutting-edge AI technology to transcribe, translate, and then re-voice a video into English in the original speaker's voice | - | |||
| Multivoice | This project uses voice cloning and TTS to deliver natural and engaging dubbed dialogue for a seamless viewing adventure | - | |||
| Open-Source AI Video Dubber | Automatically dub any video into English while keeping the original voice styles | - | |||
| Subdub | A command line Python app offering a video-to-dubbed-video workflow with transcription, translation and synchronisation | - | |||
| WeeaBlind | A program to dub non-english media with modern AI speech synthesis, diarization, and voice cloning | - | |||
| YouTube Auto-Dub | Automated voice dubbing (for YouTube videos) that translates and dubs videos with original voice timbre | - | |||
| AI Voice Translator | A tool that translates audio into another language with the ability to use your own voice | - |
Enhancers ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| Speech Enhancement / Noise Reduction | Demonstrates the process of enhancing and separating mixed audio sources, such as isolating speech from background noise, using a pre-trained model | 4 x ckpt | |||
| AURAL_GAN+predictive_model | Aims to transform low-quality phone recordings into professional-quality audio using a Generative Adversarial Network (GAN) | NA | |||
| Resemble Enhance | AI powered speech denoising and enhancement | 1 x bin 2 x pth 8 x safetensor 12 x pt | |||
| ClearerVoice-Studio | An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction | 8 x pt | |||
| Deep Learning Based Noise Reduction and Speech Enhancement System | Implements two deep learning models, one can classify the type of noise, the other can retain human voice and reduce environmental noise | 4 x h5 | |||
| Audio Enhancement and Denoising using Autoencoders | - | NA |
Extractors ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| deeper-wider-melody | Code for the "Enhancing Vocal Melody Extraction with Multilevel Contexts" paper | 1 x ckpt | |||
| Vocal-Extraction-from-Complex-Audio-Mixtures | A hybrid model combining CNNs and LSTM networks to isolate vocals from complex audio mixtures | NA | |||
| ultimatevocalremovergui | A GUI for a Vocal Remover that uses Deep Neural Networks | 1 x pth |
Stylers ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| AutoPST | Global Rhythm Style Transfer Without Text Transcriptions | - | |||
| AutoVC | Zero-Shot Voice Style Transfer with Only Autoencoder Loss | - | |||
| Voice Conversion with Non-Parallel Data | Deep neural networks for voice conversion (voice style transfer) in Tensorflow | - | |||
| StyleSinger | Code for Style Transfer for Out-of-Domain Singing Voice Synthesis paper | - | |||
| Voice style transfer with random CNN | Audio style transfer with shallow random parameters CNN | - | |||
| Robust-Voice-Style-Transfer | Code for Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion paper | - |
Speech ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| ASRT | A Deep-Learning-Based Chinese Speech Recognition System | - | |||
| whisper.cpp | High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model | - |
TTSers ⌂
| Repository | Short description | Language | License | Weights | Last commit |
|---|---|---|---|---|---|
| OpenTTS | Use Microsoft speech synthesis to generate your own voice package | - | |||
| TorToiSe | A multi-voice TTS system trained with an emphasis on quality | - | |||
| coquiTTS | A deep learning toolkit for Text-to-Speech, battle-tested in research and production | - | |||
| Fish Speech | SOTA Open Source TTS | - |