README.md
July 23, 2026 · View on GitHub
🤗Hugging Face | KVAE GitHub | Habr Article | Project Page | Technical Report (soon)
KVAE-Audio
KVAE-Audio is a continuous, full-band (48 kHz) audio autoencoder. It compresses raw waveforms into compact continuous latents and reconstructs them with high fidelity across speech, music, and general sound. The model is designed not only for faithful reconstruction, but as a latent space for generative models — in our internal text-to-audio pipeline, swapping the autoencoder for KVAE-Audio improves generation quality under a fixed generator.
Inference instruction
Setup
First install PyTorch (adjust for your CUDA setup):
pip install torch==2.11.0 torchvision==0.26.0 torchaudio==2.11.0 --index-url https://download.pytorch.org/whl/cu128
Then clone the repository and install KVAE-Audio from the local package:
git clone https://github.com/kandinskylab/kvae-audio.git
pip install -e ./kvae-audio
Tested with Python 3.14.
Download weights
Checkpoint kvae-audio.pt is hosted on Hugging Face: kandinskylab/KVAE-Audio.
pip install huggingface_hub
hf download kandinskylab/KVAE-Audio kvae-audio.pt --local-dir .
Or from Python:
from huggingface_hub import hf_hub_download
weights_path = hf_hub_download("kandinskylab/KVAE-Audio", "kvae-audio.pt")
Run inference
Place input wavs in input_samples. Reconstructions are written to output_samples.
python scripts/infer.py \
--weights_path kvae-audio.pt \
--input_path input_samples \
--output_path output_samples
Evaluate reconstructions
Compare reference wavs in input_samples against reconstructions in output_samples (same filenames). Writes metrics.csv to output_samples.
python scripts/eval.py \
--input_path input_samples \
--output_path output_samples
Python API
By default encode returns the mean latent (mu); pass sample=True for stochastic sampling. For end-to-end reconstruction, model(waveform, sample_rate)["audio"] is equivalent.
import torch
import soundfile as sf
from kvae_audio import KVAEAudio
model = KVAEAudio.load("kvae-audio.pt", map_location="cpu")
model.eval()
data, sr = sf.read("audio.wav")
waveform = torch.from_numpy(data).float().unsqueeze(0).unsqueeze(0)
length = waveform.shape[-1]
with torch.no_grad():
latents, _, _, _ = model.encode(waveform, sample_rate=sr)
audio = model.decode(latents)[..., :length]
sf.write("output.wav", audio.squeeze(0).T.numpy(), sr)
Evaluation results
Generative quality is established under a fixed generator — same DiT architecture, training data, and number of steps — varying only the autoencoder. We report objective generation metrics and blind human side-by-side below.
Evaluation of latent space qualities for generation
Audio Generation Examples
Below are qualitative examples generated from the same text prompts using four different models.
Example 1 (speech)
Prompt:
Low and gravelly, with a southern Russian accent, a man says slowly, <S>Я говорил тебе, что так и будет<E>. A sharp inhale precedes the line. The space is acoustically dry, with minimal room tone and a slight hiss.
| KVAE-Audio Download | MMAudio Download | DACVAE MovieGen Download | SAME-L Download |
Example 2 (background)
Prompt:
In a home kitchen, oil sizzles in a pan, a knife chops on a board, and a woman hums softly in Russian, a refrigerator humming beneath it all. Small kitchen, light reverberation. High-fidelity.
| KVAE-Audio Download | MMAudio Download | DACVAE MovieGen Download | SAME-L Download |
Example 3 (music)
Prompt:
In a large reverberant hall, a brass band launches into a lively march. Trumpets and cornets carry the melody while trombones drive the harmony, brisk and precise. Clean recording, natural reverb, faint equipment hiss.
| KVAE-Audio Download | MMAudio Download | DACVAE MovieGen Download | SAME-L Download |
AudioCaps test set
| Model      | # Params | Latent dim | CLAP↑   | CE↑    | PQ↑    | FAD (PANNs)↓ | FAD (PASST)↓ | FAD (VGGIsh)↓ |
|---|---|---|---|---|---|---|---|---|
| MMAudio 44.1kHz | 427.6M Â | 40 Â Â Â Â | 0,336 Â | 3,909 Â | 6,192 Â Â | 17,873 Â Â Â | 195,910 Â Â Â | 1,364 Â Â Â Â |
| DACVAE MovieGen | 107.7M Â | 128 Â Â Â Â | 0,313 Â Â | 3,772 Â Â | 6,167 Â Â | 20,558 Â Â Â | 234,312 Â Â Â | 1,700 Â Â Â Â |
| SAME-L Â Â Â Â Â | 852.1M Â | 256 Â Â Â Â | 0,322 Â Â | 3,588 Â Â | 5,756 Â Â | 18,446 Â Â Â | 240,635 Â Â Â | 1,325 Â Â Â |
| KVAE-Audio    | 166.9M  | 64     | 0,344 | 3,982 | 6,242 | 15,381  | 193,760  | 1,210   |
Song Describer
| Model      | # Params | Latent dim | CLAP↑   | CE↑    | PQ↑    | FAD (PANNs)↓ | FAD (PASST)↓ | FAD (VGGIsh)↓ |
|---|---|---|---|---|---|---|---|---|
| MMAudio 44.1kHz | 427.6M Â | 40 Â Â Â Â | 0,356 | 7,136 Â | 7,707 Â | 5,412 Â Â | 158,599 Â | 0,356 Â Â |
| DACVAE MovieGen | 107.7M Â | 128 Â Â Â Â | 0,312 Â Â | 6,953 Â Â | 7,538 Â Â | 10,194 Â Â Â | 214,009 Â Â Â | 1,046 Â Â Â Â |
| SAME-L Â Â Â Â Â | 852.1M Â | 256 Â Â Â Â | 0,345 Â | 7,076 Â Â | 7,465 Â Â | 8,442 Â Â Â Â | 250,668 Â Â Â | 0,987 Â Â Â Â |
| KVAE-Audio    | 166.9M  | 64     | 0,339   | 7,216 | 7,929 | 7,971     | 189,427    | 0,599     |
LibriSpeech test-clean
| Model      | # Params | Latent dim | CLAP↑   | CE↑    | PQ↑    | FAD (PANNs)↓ | FAD (PASST)↓ | FAD (VGGIsh)↓ | WER↓    | CER↓    |
|---|---|---|---|---|---|---|---|---|---|---|
| MMAudio 44.1kHz | 427.6M Â | 40 Â Â Â Â | 0,368 Â Â | 5,704 Â | 6,629 Â Â | 8,305 Â Â Â Â | 105,931 Â | 2,001 Â Â Â | 0,257 Â | 0,593 Â |
| DACVAE MovieGen | 107.7M Â | 128 Â Â Â Â | 0,413 | 5,482 Â Â | 7,052 | 5,008 Â Â Â | 210,478 Â Â Â | 1,501 Â Â | 0,911 Â Â | 1,048 Â Â |
| SAME-L Â Â Â Â Â | 852.1M Â | 256 Â Â Â Â | 0,379 Â Â | 4,617 Â Â | 5,024 Â Â | 10,257 Â Â Â | 301,508 Â Â Â | 2,721 Â Â Â Â | 0,349 Â Â | 0,629 Â Â |
| KVAE-Audio    | 166.9M  | 64     | 0,389  | 5,906 | 6,940  | 4,677   | 185,609   | 2,138     | 0,244 | 0,576 |
Reconstructions
Reconstruction is evaluated on open datasets across domains (the released weights directly substantiate these numbers). Baselines: MMAudio 44.1 kHz VAE, DACVAE from MovieGen Audio, SAME-L (Stable Audio 3 VAE).
AudioSet eval
| Model      | # Params | Latent dim | MEL↓    | STFT↓   | Waveform↓ | SI-SDR↑  | SDR↑    | SNR↑    |
|---|---|---|---|---|---|---|---|---|
| MMAudio 44.1kHz | 427.6M Â | 40 Â Â Â Â | 0,636 Â | 1,938 Â | 0,106 Â Â | -32,080 Â | -2,682 Â Â | -2,686 Â Â |
| DACVAE MovieGen | 107.7M Â | 128 Â Â Â Â | 0,669 Â Â | 2,275 Â Â | 0,029 Â Â | 8,384 Â Â | 9,421 Â Â Â | 9,416 Â Â Â |
| SAME-L Â Â Â Â Â | 852.1M Â | 256 Â Â Â Â | 0,986 Â Â | 2,726 Â Â | 0,027 Â | 9,586 | 10,347 | 10,339 |
| KVAE-Audio    | 166.9M  | 64     | 0,537 | 1,770 | 0,027 | 9,065  | 9,920   | 9,933   |
MUSDB18-HQ
| Model      | # Params | Latent dim | MEL↓    | STFT↓   | Waveform↓ | SI-SDR↑   | SDR↑    | SNR↑    |
|---|---|---|---|---|---|---|---|---|
| MMAudio 44.1kHz | 427.6M Â | 40 Â Â Â Â | 0,681 Â Â | 1,865 Â Â | 0,114 Â Â | -40,204 Â Â | -3,274 Â Â | -3,273 Â Â |
| DACVAE MovieGen | 107.7M Â | 128 Â Â Â Â | 0,519 Â | 1,762 Â | 0,024 Â Â | 9,688 Â Â Â | 10,046 Â Â | 10,047 Â Â |
| SAME-L Â Â Â Â Â | 852.1M Â | 256 Â Â Â Â | 0,668 Â Â | 1,786 Â Â | 0,023 Â | 10,278 Â | 10,648 Â | 10,648 Â |
| KVAE-Audio    | 166.9M  | 64     | 0,516 | 1,725 | 0,022 | 10,390 | 10,675 | 10,677 |
EARS
| Model      | # Params | Latent dim | MEL↓    | STFT↓   | Waveform↓ | SI-SDR↑   | SDR↑    | SNR↑    | PESQ↑   |
|---|---|---|---|---|---|---|---|---|---|
| MMAudio 44.1kHz | 427.6M Â | 40 Â Â Â Â | 0,616 Â Â | 1,395 Â Â | 0,030 Â Â | -29,947 Â Â | -2,728 Â Â | -2,697 Â Â | 2,424 Â Â |
| DACVAE MovieGen | 107.7M Â | 128 Â Â Â Â | 0,453 | 1,310 | 0,006 Â | 10,264 | 10,680 | 10,681 | 4,246 Â |
| SAME-L Â Â Â Â Â | 852.1M Â | 256 Â Â Â Â | 0,774 Â Â | 1,575 Â Â | 0,007 Â Â | 9,939 Â Â Â | 10,374 Â Â | 10,376 Â Â | 2,982 Â Â |
| KVAE-Audio    | 166.9M  | 64     | 0,463  | 1,314  | 0,006 | 9,952   | 10,377  | 10,384  | 4,266 |
Citation
@misc{kvae_audio_2026,
author = {Ivan Kirillov, Denis Parkhomenko, Alexander Ivanov, Azat Saginbaev, Egor Silvestrov, Denis Dimitrov},
title = {KVAE-Audio: a full-band continuous audio tokenizer for generative models},
howpublished = {\url{https://github.com/kandinskylab/kvae-audio}},
year = 2026
}
Package layout
kvae-audio/
├── kvae_audio/
│ ├── model/
│ │ ├── kvae_audio.py
│ │ └── base.py
│ ├── metrics/loss.py
│ └── nn/layers.py
└── scripts/
├── infer.py
└── eval.py