CrispASR

August 16, 2026 · View on GitHub

One C++ binary, 54 ASR backends + 52 TTS engines + multilingual text translation, zero Python dependencies.

CrispASR started as a fork of whisper.cpp and extends that base into a unified speech engine called crispasr, backed by full ggml C++ runtimes for major open-weights ASR and TTS architectures. One build, one binary, one consistent CLI — pick the backend at the command line or let CrispASR auto-detect it from your GGUF file. See Text-to-Speech for the TTS side.

$ crispasr -m ggml-base.en.bin          -f samples/jfk.wav                    # OpenAI Whisper
$ crispasr -m parakeet-tdt-0.6b.gguf    -f samples/jfk.wav                    # NVIDIA Parakeet
$ crispasr -m canary-1b-v2.gguf         -f samples/jfk.wav                    # NVIDIA Canary
$ crispasr -m voxtral-mini-3b-2507.gguf -f samples/jfk.wav                    # Mistral Voxtral
$ crispasr --backend qwen3 -m auto      -f samples/jfk.wav                    # -m auto downloads
$ crispasr --backend kokoro -m auto --tts "Hello world" --tts-output out.wav  # TTS

No Python. No PyTorch. No separate per-model binary. No pip install. Just one C++ binary and a GGUF file.

Browser: All backends compile to WebAssembly (4.3 MB) via build-wasm.sh. Multithreaded, runs entirely client-side with COOP/COEP headers.

Demo: HuggingFace Space — live transcription + TTS + language detection, auto-deployed from hf-space/.

Ecosystem

ProjectWhat it does
CrispASRThis repo — C++ speech engine. 54 ASR + 52 TTS backends, CLI + HTTP server + C-ABI + Python/Rust/Dart/Go/Ruby/Java bindings.
CrisperWeaverCross-platform Flutter transcription app built on CrispASR. Desktop + mobile, model browser with download queue, mic capture, SRT/VTT/JSON export, diarization, batch processing. Fully offline.
CrispEmbedText-related engine via ggml — same philosophy as CrispASR but for embeddings, retrieval, OCR and OMR, Math and Music Notation. Numerous architectures (XLM-R, Qwen3-Embed, Gemma3, ModernBERT, ...), dense + sparse + ColBERT + reranking. PP-OCR, Tesseract, EasyOCR, InternVL2, etc. Python/Rust/Dart bindings.
SusurrusPython ASR GUI with 9 backends (faster-whisper, mlx-whisper, voxtral, insanely-fast-whisper, ...). The Python counterpart to CrispASR's C++ approach.

Table of contents


Supported backends

CrispASR ships 54 ASR backends for transcription/translation and 51 TTS engines for synthesis. It also ships audio-to-audio S2S backends, including Sidon restoration and the VoxCPM2 AudioVAE speech upscaler; see the feature matrix for the complete capability list. Pick at the CLI with --backend NAME, or omit it to let the binary auto-detect from the GGUF metadata. Jump to the TTS table for the synthesis side.

ASR backends

BackendModelArchitectureLanguagesLicense
whisperggml-base.en.bin and all OpenAI Whisper variantsEncoder-decoder transformer99MIT
whisperdistil-whisper/distil-large-v3Distilled Whisper: 32L encoder + 2L decoder (6.3x faster)EnglishMIT
parakeetnvidia/parakeet-tdt-0.6b-v3FastConformer + TDT25 EU (auto-detect)CC-BY-4.0
parakeetnvidia/parakeet-tdt-0.6b-v2FastConformer + TDT, original Open ASR Leaderboard topperen (mixed-case + punct)CC-BY-4.0
parakeetnvidia/parakeet-tdt-1.1b42L FastConformer + TDT, larger English varianten (lowercase)CC-BY-4.0
parakeetnvidia/parakeet-tdt_ctc-110m17L FastConformer + TDT+CTC hybrid; smallest variant, auto-CTC decodeenCC-BY-4.0
parakeetnvidia/parakeet-tdt_ctc-1.1b42L FastConformer + TDT+CTC hybrid; largest, mixed-case + punctenCC-BY-4.0
parakeetnvidia/parakeet-tdt_ctc-0.6b-jaFastConformer-TDT-CTC, xscaling, 80 melsJapaneseCC-BY-4.0
reazonspeechreazon-research/reazonspeech-nemo-v2FastConformer-RNNT, local attn (w=256), 80 mels, 619M paramsJapaneseApache-2.0
fastconformer-ctcnvidia/parakeet-ctc-0.6b24L FastConformer + CTC, 80 mels (same arch as fc-ctc-xlarge)enCC-BY-4.0
fastconformer-ctcnvidia/parakeet-ctc-1.1b42L FastConformer + CTC, 80 melsenCC-BY-4.0
fastconformer-ctcgrider-transwithai/parakeet-ctc-1.1b-ja42L FastConformer + CTC, 80 mels, Japanese fine-tuneJapaneseApache-2.0
canarynvidia/canary-1b-v2FastConformer + Transformer decoder25 EU (explicit -sl/-tl)CC-BY-4.0
canary-qwennvidia/canary-qwen-2.5bFastConformer + Qwen3-1.7B SALMenCC-BY-4.0
lfm2-audioLiquidAI/LFM2.5-Audio-1.5BFastConformer + LFM2 hybrid conv+attention backbone (ASR+TTS)enLFM Open v1.0
lfm2-audioLiquidAI/LFM2.5-Audio-1.5B-JPFastConformer + LFM2 hybrid conv+attention backbone (ASR+TTS)jaLFM Open v1.0
mini-omni2gpt-omni/mini-omni2Whisper-small + Qwen2-0.5B (ASR+TTS+S2S)enMIT
cohereCohereLabs/cohere-transcribe-03-2026Conformer + Transformer13Apache-2.0
cohereefwkjn/cohere-asr-ja-v0.1Japanese fine-tune of cohere-transcribe-03-2026 (TedX/JSUT-tuned)JapaneseApache-2.0
graniteibm-granite/granite-speech-{3.2-8b,3.3-2b,3.3-8b}, granite-4.0-1b-speechConformer + Q-Former + Granite LLM (μP) (more)en fr de es pt jaApache-2.0
granite-4.1ibm-granite/granite-speech-4.1-2b16L Conformer + Q-Former + Granite LLM; single ggml graph (more)en fr de es pt jaApache-2.0
granite-4.1-plusibm-granite/granite-speech-4.1-2b-plus4.1 + hidden-state concat; punctuated output (more)en fr de es ptApache-2.0
granite-4.1-naribm-granite/granite-speech-4.1-2b-narNon-autoregressive: single LLM forward + slot argmax (more)en fr de es ptApache-2.0
fastconformer-ctcnvidia/stt_en_fastconformer_ctc_largeFastConformer + CTC (NeMo family, all sizes)enCC-BY-4.0
voxtralmistralai/Voxtral-Mini-3B-2507Whisper encoder + Mistral 3B LLM8Apache-2.0
voxtral4bmistralai/Voxtral-Mini-4B-Realtime-2602Causal encoder + 3.4B LLM, sliding window13, realtime streamingApache-2.0
qwen3Qwen/Qwen3-ASR-0.6BWhisper-style audio encoder + Qwen3 0.6B LLM30 + 22 Chinese dialectsApache-2.0
qwen3-1.7bQwen/Qwen3-ASR-1.7BWhisper-style audio encoder + Qwen3 1.7B LLM30 + 22 Chinese dialectsApache-2.0
qwen3-ja-animejaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame-hfQwen3-ASR-1.7B fine-tuned for Japanese anime/galgame speechja + 30 langsApache-2.0
mega-asrzhifeixie/Mega-ASRQwen3-ASR-1.7B + merged robustness LoRA; always-on robust pathnoisy / degraded speechApache-2.0
higgs-sttbosonai/higgs-audio-v3-sttWhisper-large-v3 encoder (4 s chunked) + Qwen3-1.7B LLM (more)enApache-2.0
wav2vec2jonatasgrosman/wav2vec2-large-xlsr-53-englishCNN + 24L transformer + CTC head (any Wav2Vec2ForCTC)per-modelApache-2.0
wav2vec2facebook/data2vec-audio-base-960hData2Vec Audio (79 MB Q4_K)EnglishApache-2.0
wav2vec2facebook/hubert-large-ls960-ftHuBERT Large (212 MB Q4_K)EnglishApache-2.0
glm-asrzai-org/GLM-ASR-Nano-2512Whisper encoder + 4-frame projector + Llama 1.5B (GQA)Mandarin (+ Chinese dialects), English, CantoneseMIT
kyutai-sttkyutai/stt-1b-en_frMimi codec (SEANet + RVQ) + 16L causal LMen, frMIT
kyutai-sttkyutai/stt-2.6b-enMimi codec + 48L causal LM (2.6B, English-only; 3.5 s lookahead)enMIT
firered-asrFireRedTeam/FireRedASR2-AEDConformer + CTC + beam search; also LID (120 langs)Mandarin, English, 20+ Chinese dialectsApache-2.0
moonshineUsefulSensors/moonshine-{tiny,base}Conv + 6L enc + 6L dec; multilingual variantsEnglish + 6 langsMIT
moonshine‑defidoriel/moonshine-base-deGerman fine-tune of moonshine-base (6.9% WER CV22)GermanCC‑BY‑NC‑SA‑4.0
moonshine‑tiny‑defidoriel/moonshine-tiny-deGerman fine-tune of moonshine-tiny (11.4% WER CV22)GermanCC‑BY‑NC‑SA‑4.0
moonshine-streamingUsefulSensors/moonshine-streaming-{tiny,small,medium}Streaming: sliding-window encoder + AR decoder (34–245M)EnglishMIT
gemma4-e2bgoogle/gemma-4-E2B-itUSM Conformer 12L + Gemma4 LLM 35L (GQA, PLE)140+ langsApache-2.0
gemma4-e4bgoogle/gemma-4-E4B-itSame USM Conformer 12L + larger Gemma4 LLM 42L (GQA, PLE); runs on --backend gemma4-e2b140+ langsApache-2.0
omniasromniASR-CTC-1B-v2wav2vec2 CNN + 48L transformer + CTC (more)1600+Apache-2.0
omniasr‑300momniASR-CTC-300M-v2Same arch, 24L, ~194 MB Q4_K; auto-chunks >7 s (more)1600+Apache-2.0
omniasr-llmomniASR-LLM-300M-v2Same encoder + 12L LLaMA decoder (more)1600+Apache-2.0
omniasr-llmomniASR-LLM-Unlimited-300M-v2Streaming: 15s segment protocol, unlimited audio (more)1600+Apache-2.0
vibevoicemicrosoft/VibeVoice-ASRσ-VAE ConvNeXt + Qwen2.5-7B (more)50+MIT
vibevoice-bitnetVibeVoice-ASR-BitNetSame arch, TQ2_0 ternary LM (1.6 GB) (more)7+MIT
mimo-asrXiaomiMiMo/MiMo-V2.5-ASR6L transformer + 36L Qwen2 LM + RVQ codec (more)Mandarin + dialects + EnglishMIT
ark-asr ⚠️experimental/WIPcstr/ark-asr-3b-GGUF (base AutoArk-AI/ARK-ASR-3B)Whisper-large-v3 enc (partial RoPE) + Qwen2.5-3B LM (more)19 (zh, en, de, ja, fr, ko, es, pl, it, ro, hu, cs, nl, fi, hr, sk, sl, et, lt)see base
moss-audioOpenMOSS-Team/MOSS-Audio-4B-Instruct32L Whisper encoder + DeepStack 3-tap + 36L Qwen3 LM; audio understanding + ASR (more)zh, enApache-2.0
moss-transcribeOpenMOSS-Team/MOSS-Transcribe-preview-2BQwen3-Omni audio encoder (32L, windowed attn) + GatedMLP adapter + Qwen3-1.7B LM; ASR (more)zh, enApache-2.0
moss-diarizeOpenMOSS-Team/MOSS-Transcribe-Diarize-0.9BStock Whisper encoder (24L, 80 mel) + 4x merge + VQAdaptor + Qwen3-0.6B LM; joint ASR + speaker diarization + timestampsmultiApache-2.0
whisper (tiron) ⚠️experimentalTrelis/tiron (base Trelis/tiron)Whisper large-v3 with an extended vocab: emits inline `<speakerN>` markers + 20 ms timestamps for joint transcription + per-window speaker attribution, via a constrained-decoding grammar; cross-window linking to stable speakers (#295)
funasrFunAudioLLM/Fun-ASR-Nano-251270-block SANM encoder + 2-block Transformer adaptor + Qwen3-0.6B LLMzh, yue, en, ja, koFunASR Model License v1.1 (commercial OK w/ attribution)
fun-asr-mlt-nanoFunAudioLLM/Fun-ASR-MLT-Nano-2512Same architecture, multilingual decoder31 langs incl. de, fr, es, pt, ru, ar, hi, vi, th, koFunASR Model License v1.1
paraformerfunasr/paraformer-zh50-block SANM encoder + CIF predictor + 16-block NAR decoder (single-pass, non-autoregressive); character-level vocab (8404); 220M paramszh, enFunASR Model License (commercial OK w/ attribution)
foxnose (speaker diarization)Wespeaker/wespeaker-voxceleb-resnet34-LMSpeaker diarization via --diarize-method foxnose: WeSpeaker ResNet34-LM 256-d embeddings + GMM/BIC speaker counting + spectral clustering + Viterbi temporal smoothing (more). 3.18 % DER on VoxConverse dev vs the upstream reference implementation's 3.07 %anyweights CC-BY-4.0
gigaamai-sage/GigaAM-v3 (base ai-sage/GigaAM-v3)16-layer rotary Conformer (220M) + CTC or RNN-T head; four revisions — e2e_rnnt / e2e_ctc emit punctuation + casing + ITN from a SentencePiece vocab, rnnt / ctc emit bare lowercase Cyrillic (more)en ruMIT
sensevoiceFunAudioLLM/SenseVoiceSmall70-block SANM encoder + CTC head; emits transcript + language ID + audio-event in one forward pass (non-AR, 15× faster than Whisper-Large); structured C ABI + -oj JSON expose the tags as separate fields. Upstream's emotion classifier is not exposed — see EU AI Act50+ langs; native LID + audio-event tagsFunASR Model License v1.1

Speech-to-speech audio upscaling and restoration

BackendModelArchitectureInput / outputLicense
sidonKevinAHM/Sidon-GGUF (base sarulab-speech/sidon-v0.1)w2v-BERT 2.0 predictor + continuous DAC decoder (more)16 kHz mono → restored 48 kHz monoMIT
voxcpm2-vaeAudioVAE V2 from openbmb/VoxCPM2, converted with --vae-onlyIsolated causal AudioVAE encoder + decoder (more)16 kHz mono → upscaled 48 kHz monoApache-2.0
huggingface-cli download KevinAHM/Sidon-GGUF sidon-v0.1-f16.gguf --local-dir models
crispasr -m models/sidon-v0.1-f16.gguf -f input.wav --s2s --s2s-output restored.wav

python models/convert-voxcpm2-to-gguf.py --input openbmb/VoxCPM2 \
  --output models/voxcpm2-vae-f32.gguf --vae-only
crispasr -m models/voxcpm2-vae-f32.gguf -f input.wav --s2s \
  --s2s-output upscaled.wav

Text-to-Speech models

Synthesis backends, driven by the --tts flag and a --tts-output PATH.wav. See the dedicated Text-to-Speech section below for quick-start commands and engine selection guidance.

BackendModelsArchitectureLanguagesLicense
miottsMioTTS-0.6BQwen3 LLM + MioCodec-v2 FSQ codec (25 Hz, 44.1 kHz output)ja, enApache-2.0
vibevoice-ttsVibeVoice-Realtime-0.5B, VibeVoice-1.5BDPM-Solver++ + σ-VAE decoder; voice presets or cloningen, zhMIT
kugelaudiokugelaudio-0-openQwen2.5-7B LM + 4L DiT diffusion + acoustic VAE decoder; voice cloningmultilingualApache-2.0
qwen3-ttsQwen3-TTS-12Hz-0.6B-Base, 1.7B-Base, 1.7B-VoiceDesignQwen3 talker LM + 12 Hz RVQ (more)multilingualApache-2.0
qwen3-tts-customvoice1.7B-CustomVoiceSame talker + 9 premium built-in speakers (--voice <name>); optional style via --instruct (e.g. "spoke very slowly") (more)multilingualApache-2.0
moss-ttsOpenMOSS-Team/MOSS-TTS-v1.5Qwen3-8B backbone emitting 32 RVQ audio codebooks under a delay pattern, decoded by a 1.6B pure-transformer codec companion; voice cloning via --voice ref.wav; --backend moss-tts -m <backbone> --codec-model <codec>multilingualApache-2.0
moss-tts-localOpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5Qwen3-4B backbone; a 1-layer local/depth transformer autoregressively emits 12 RVQ codebooks per frame (RQ-Transformer, no delay), decoded to 48 kHz by MOSS-Audio-Tokenizer-v2 (downmixed to mono); --backend moss-tts-local -m <backbone> --codec-model <codec>multilingualApache-2.0
omnivoicek2-fsa/OmniVoiceQwen3-0.6B + masked iterative 8-codebook TTS (SoundStorm-style); voice cloning; 600+ languages (more)600+ langsApache-2.0
melottsmyshell-ai/MeloTTS EN_V2VITS2 (6L transformer + SDP/DP + transformer coupling flow + HiFi-GAN); 44.1 kHz, 102 MB + 52 MB BERT Q4_K companion (154 MB total); neural G2P; 4 EN speakers (more)enMIT
piperrhasspy/piper community voicesVITS (6L transformer + SDP + 4-block coupling flow + HiFi-GAN); 22 kHz mono, 30 MB F16 per voice; built-in G2P for EN/DE/FR/ES (--g2p-dict)30+ langs (built-in + espeak dlopen)MIT
kokorohexgrad/Kokoro-82M + German backbonesStyleTTS2 / iSTFTNet (82M); per-voice GGUF (more)en, es, fr, hi, it, ja, pt, zh, deApache-2.0
orpheusOrpheus-3B-FT + SNAC 24 kHzLlama-3.2-3B + SNAC RVQ codec; 8 speakers (more)en, deLlama 3.2 Community License / MIT
chatterboxcstr/chatterbox-GGUF + turbo/kartoffelbox/lahgtna variantsT3 AR + S3Gen flow-matching (more)23 multilingual; separate Arabic (lahgtna-chatterbox) and German turbo (kartoffelbox-turbo) fine-tunesMIT
indexttscstr/indextts-1.5-GGUFGPT-2 AR (24L/1280d) + Conformer conditioning + BigVGAN vocoder; voice cloning via reference audiozh, enApache-2.0
voxcpm2-ttscstr/voxcpm2-GGUFTokenizer-free CFM diffusion AR (TSLM + RALM + LocDiT) at 48 kHz native; zero-shot + voice cloning via --voice <wav>30 languagesApache-2.0
voxtral-ttsmistralai/Voxtral-4B-TTS-2603Ministral-3B AR (26L GQA) + 3L FM acoustic transformer (7-step Euler ODE) + Voxtral codec decoder at 24 kHz; 20 preset voices; SOTA French technical texten, fr, de, es, it, pt, nl, ar, hiCC‑BY‑NC‑4.0
cosyvoice3-ttscstr/cosyvoice3-0.5b-2512-GGUFQwen2-0.5B AR speech-token LM + DiT-CFM (10-step Euler) + HiFT (NSF + iSTFT) at 24 kHz; baked-voice zero-shot cloning via --voice <name>, or any WAV via --voice ref.wav --ref-text "<exact transcript>". --backend cosyvoice3-tts-rl selects upstream's RL-tuned talker (same companions)9 langs + 18 zh dialectsApache-2.0
csmcstr/csm-1b-GGUFSesame CSM-1B conversational TTS: Llama-3.2 1B backbone + 100M depth decoder (32-codebook RVQ) + Kyutai Mimi codec at 24 kHz (more)enApache-2.0
lfm2-audiocstr/lfm2-audio-1.5b-GGUF + jpLFM2.5-Audio ASR+TTS+S2S: FastConformer enc + LFM2 hybrid backbone + depthformer (8-codebook Mimi) + ISTFT detokenizer at 24 kHz; interleaved text+audio generationen, jaLFM Open v1.0
dianari-labs/Dia-1.6BByte-level text encoder (12L) + AR audio decoder (18L GQA + CFG) → 9 delayed DAC codebooks + 44.1 kHz DAC codec; dialogue style with [S1]/[S2] tags (use >100-char prompts)enApache-2.0
zonos-ttscstr/zonos-v0.1-transformer-GGUF + cstr/dac-44khz-GGUFZyphra Zonos-v0.1: 26L GQA AR transformer (2B) + 9-codebook DAC @ 44.1 kHz; CFG-guided; voice cloning via reference WAV (more)enApache-2.0
barkcstr/bark-small-GGUFSuno Bark 3-stage GPT-2 TTS: text→semantic (12L) → coarse EnCodec (12L, 2 codebooks) → fine (12L, 8 codebooks) → EnCodec 24 kHz decoder; speaker conditioning via .npz prompts (--voice <file.npz>)multilingualMIT
speecht5cstr/speecht5-tts-GGUFSpeechT5 80M: char-level encoder (12L) + AR mel decoder (6L) + 5-layer conv postnet + HiFi-GAN at 16 kHz; speaker via 512-d x-vector (--voice <xvector.bin>)enMIT
fastpitchcstr/fastpitch-en-GGUFNVIDIA FastPitch 60M: non-autoregressive parallel TTS — 6L encoder + duration/pitch predictors + 6L decoder + HiFi-GAN at 22 kHz; deterministic, single forward pass (more)enCC-BY-4.0
bananamind-ttsBanaxi-Tech/BananaMind-TTS-V2.1-PreviewBananaMind-TTS 13M: Tacotron-lite char-level encoder (Conv+BN+BiLSTM) + AR GRU decoder with location-sensitive attention + postnet + HiFi-GAN at 22 kHz; fixed voice per locale (more)en, deApache-2.0
parler-ttscstr/parler-tts-mini-v1.1-GGUFParler TTS Mini v1.1 (~900M): T5 encoder + MusicGen decoder + DAC 44.1 kHz; prompt-conditioned (describe voice in text via --instruct)enApache-2.0
outettscstr/outetts-0.3-1b-GGUFOLMo-1B talker + WavTokenizer single-codebook VQ-GAN at 24 kHz; voice cloning via speaker profile JSON (--voice <speaker.json>)enCC-BY-NC-SA-4.0
pocket-ttscstr/pocket-tts-GGUFKyutai Pocket TTS 100M: continuous-latent AR at 12.5 Hz + one-step LSD flow + Mimi VAE 24 kHz; voice cloning via ref audio (more)enCC-BY-4.0 + gated-use conditions
tadacstr/tada-tts-1b-GGUF + HumeAI/tada-3b-mlLlama-3.2 1B/3B backbone + per-token FM diffusion head + TADA codec at 24 kHz; 1:1 text-to-acoustic alignment; default prompt via tada-ref.gguf, custom voices via --voice <tada-ref.gguf> built with models/convert-tada-ref-to-gguf.py (more)enLlama 3.2 Community License
TTS feature matrix
BackendVoice cloningSamplingkHzAuto-downloadFlash attn
vibevoice-ttsyestemp24yesyes
qwen3-ttsyes*temp24yesyes
omnivoiceyestemp24
kokoro24yes
orpheustemp24yesyes
chatterboxyestemp24yesyes
outettsyes (JSON)temp24yesyes
indexttsyestemp24yesyes
voxcpm2-ttsyes48yes
cosyvoice3-ttsyestemp24yesyes
f5-ttsyes24yes
irodori-ttsyes (WAV)VoiceDesign: --instruct48yes
csmtemp24yes
diatemp44yes
barkyes (.npz)temp24yes
speecht5yes (xvec)16yes
parler-ttstemp44yes
fastpitch22
piper22
pocket-ttsyestemp24yes
tadayestemp24yes
dots-ttsyes (--voice ref.wav)16-step CFG Euler48yes

* CustomVoice variant only; Base uses baked speakers via --voice <name>.

Output language. -tl <lang> (or -l) selects the language to speak; cosyvoice3-tts, qwen3-tts and moss-tts act on it natively. For cross-lingual cloning — an English reference clip speaking German, the subtitle-dubbing case — also pass -sl <lang> for the language the reference is spoken in, so cosyvoice3 drops the reference transcript instead of carrying its accent. Over HTTP: "language" + "source_lang" on POST /v1/audio/speech. See docs/tts.md.

Translation

Text-to-text translation, distinct from the audio-side --translate flag (which routes audio → English text on whisper / canary / etc.). Driven by --text "..." -sl <src> -tl <tgt>.

BackendModelsArchitectureLanguagesLicense
m2m100facebook/m2m100_418M12L enc + 12L dec transformer, SentencePiece 128K (more)100 langs, any-to-anyMIT
m2m100-wmt21facebook/wmt21-dense-24-wide-en-x + facebook/wmt21-dense-24-wide-x-enSame as m2m100, scaled to 4.7B (24L enc) (more)English ↔ 7 langs (separate en-x / x-en checkpoints)MIT
madladgoogle/madlad400-3b-mtT5 enc-dec (12L+12L, d=2048, gated-GELU, RMSNorm) (more)419 languagesApache-2.0
# m2m100 base (production-ready)
./build/bin/crispasr --backend m2m100 -m auto \
    --text "Hello world, how are you today?" \
    -sl en -tl de
# → Hallo Welt, wie bist du heute?

# WMT21 dense (English ↔ X, 4.7B — auto-downloads ~2.5 GB).
# Two separate checkpoints: en-x for English-source, x-en for
# English-target. Pick the one matching your `-sl`/`-tl` direction
# (or pass an explicit `-m <path>` to load the other manually).
./build/bin/crispasr --backend m2m100-wmt21 -m auto \
    --text "The president said he would not attend." \
    -sl en -tl de   # uses wmt21-dense-24-wide-en-x

./build/bin/crispasr --backend m2m100-wmt21 \
    -m models/wmt21-dense-24-wide-x-en-q4_k.gguf \
    --text "Le président a dit qu'il ne serait pas présent." \
    -sl fr -tl en   # uses wmt21-dense-24-wide-x-en

# MADLAD-400 3B (419 languages, bit-token-identical to Python SP)
./build/bin/crispasr --backend madlad -m auto \
    --text "Hello world." \
    -sl en -tl ta

For 2-stage pipelines (e.g., ASR → m2m100), use the dedicated --tr-sl / --tr-tl flags; they fall back to -sl / -tl when unset, so single-stage standalone usage is just -sl/-tl.

Post-processing models

Work with all backends.

ModelTaskArchitectureLanguagesLicenseHuggingFace
FireRedPuncPunctuation restorationBERT-base (12L, d=768), 5 classesChinese + EnglishApache-2.0cstr/fireredpunc-GGUF
fullstop-puncPunctuation restorationXLM-RoBERTa-large (24L, d=1024), 6 classesEN, DE, FR, ITMITcstr/fullstop-punc-multilang-GGUF
punctuate-allPunctuation restorationXLM-RoBERTa-base (12L, d=768), 6 classes12 languagesMITcstr/punctuate-all-GGUF
PCSPunc + truecase + SBDXLM-RoBERTa-base (12L), 4 heads47 languagesApache-2.0--punc-model pcs
truecaser‑lstmGerman truecasing (best)BiLSTM char-level (2×150, 3.2 MB, 97.9% F1)GermanApache-2.0--truecase-model lstm
truecaser‑crfGerman truecasingCRF + context features (8.5 MB)GermanMIT--truecase-model crf
truecaser‑deGerman truecasing (simple)Statistical word-frequency (71K entries, 1.7 MB)GermanMIT--truecase-model auto
CLD3Text language IDEmbedding-bag → FC + ReLU → softmax (~1.5 MB F32)109 ISO 639-1Apache-2.0cstr/cld3-GGUF
GlotLID-V3Text language IDfastText supervised, flat softmax2102 ISO 639-3 + scriptApache-2.0cstr/glotlid-GGUF
LID-176Text language IDfastText supervised, hierarchical softmax176 ISO 639-1CC-BY-NC-4.0cstr/fasttext-lid176-GGUF

Audio codecs

Shared codec modules used by TTS backends. Also available standalone for encode/decode.

ModelArchitectureSample RateToken RateLicenseHuggingFace
MioCodec v2WavLM encoder → FSQ(12800) → Transformer decoder + AdaLN-Zero + SnakeBeta upsampler + iSTFT44.1 kHz25 Hz (341 bps)MITcstr/miocodec-v2-44k-GGUF
SNAC 24 kHz3-codebook RVQ + decoder blocks (stride 8/8/4/2)24 kHz3×12.5 HzMITcstr/snac-24khz-GGUF

All runtimes share ggml-based inference. The speech-LLM backends (qwen3, voxtral, voxtral4b, granite, glm-asr, kyutai-stt) inject audio encoder frames directly into an autoregressive language model's input embeddings, instead of using a dedicated CTC/transducer/seq2seq decoder. The fastconformer-ctc backend hosts the NeMo FastConformer-CTC standalone ASR family — stt_en_fastconformer_ctc_{large,xlarge,xxlarge} and the architecturally-identical parakeet-ctc-{0.6b,1.1b} (different training data + tokenizer, same encoder + head shape) — with greedy CTC decoding. Same C++ runtime as the canary-ctc aligner.

Music & audio analysis

Beyond speech, CrispASR runs several music/audio analysis tasks — each a small GGUF with the architecture auto-detected, no Python. See docs/cli.md for the per-task flags and output formats.

  • Source separation (--separate) — split a mix into stems (<input>_<stem>.wav) via mel-band-roformer (vocal/instrumental, MIT) or htdemucs (4-stem). --stems vocals,drums selects a subset; --sep-output-dir sets the output location.
  • Piano transcription (--backend piano-transcription) — piano audio → MIDI note events (88 keys @ 100 fps, ByteDance/Kong CRNN; F16 GGUF ≈ 77 MB).
  • Guitar tablature (--tab) — per-frame fret-per-string grid via TabCNN (Wiggins & Kim, ISMIR 2019; CC BY 4.0 weights). The backend emits per-string emission scores, not a decided tablature — run your own constrained Viterbi via crispasr_session_tab_emissions() for playable output.
  • Beat / downbeat tracking (--beats) — beat grid via Beat This! (CPJKU, ISMIR 2024; MIT for code and weights, no patent-encumbered DBN).
  • Chord recognition (--chords) — chord timeline (.lab) via BTC (ISMIR 2019). Weights are CC-BY-NC-SA, gated behind --accept-license cc-by-nc-sa-4.0.
  • Pitch / F0 estimation (--pitch) — monophonic pitch track via CREPE (MIT).

Feature matrix

Run crispasr --list-backends to see it live. Each backend declares capabilities at runtime; if you ask for a feature the selected backend does not support, CrispASR prints a warning and silently ignores the flag.

Sortable / filterable view: docs/feature-matrix.html — click any column header to sort, type to filter rows, click cap pills to require a capability. Generated from crispasr --list-backends-json (single source of truth — drift impossible). Regenerate via python tools/gen-feature-matrix.py. A Markdown twin lives at docs/feature-matrix.md.

The static table below is a curated subset focusing on the ASR backends and the cross-cutting features that matter for ASR pipelines. The full 105-backend × 27-cap surface is in the generated views.

Featurewhisperparakeetcanarycoheregranitegranite‑4.1voxtralvoxtral4bqwen3fc‑ctcwav2vec2glm‑asrkyutai‑sttfireredmoonshinemoon‑streamomniasromniasr‑llmvibevoicegemma4‑e2bmimo‑asrfunasrparaformersensevoice
Native timestamps
CTC timestamps
Word-level timing-am✔†-am-am-am-am-am-am-am-am-am-am-am-am-am-am-am
Per-token confidence
Language auto-detectLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLIDLID
Speech translation
Speaker diarization
Grammar (GBNF)
Temperature sampling
Beam search
Flash attention
Punctuation toggle
Punc restorationpppppppppppppppppppppppppppppppppppppppppppppppp
Source / target language
Audio Q&A (--ask)******
Streaming
Auto-download (-m auto)
KV quant (CRISPASR_KV_QUANT, plus per-half _K / _V)
mmap weights (CRISPASR_GGUF_MMAP)
TTS

The matrix above covers 24 ASR backends. Additional ASR backends not shown: nemotron (39-lang streaming ASR with cache-aware FastConformer + RNN-T), lfm2-audio (ASR + TTS + S2S in one model), moss-audio (audio understanding + ASR), moss-transcribe (Qwen3-Omni encoder + Qwen3-1.7B ASR), mini-omni2 (ASR + TTS + S2S), kugelaudio (7B audio understanding). See docs/feature-matrix.md for the full 106-backend matrix. TTS-only backends (kokoro, qwen3-tts + variants, vibevoice-tts, orpheus + DE variants, chatterbox / chatterbox-turbo / kartoffelbox-turbo / lahgtna-chatterbox, dia, bark, outetts, zonos, csm, f5-tts, irodori-tts, parler-tts, speecht5, piper, fastpitch, pocket-tts, melotts, cosyvoice3, voxcpm2, tada-tts) all carry the TTS, AUTO_DOWNLOAD, TEMPERATURE, and FLASH_ATTN caps; per-backend cloning + voice-pack support is documented in the Text-to-Speech models table above and docs/tts.md. The vibevoice and lfm2-audio columns mark dual-mode (ASR + TTS) backends.

Key: ✔ = native/built-in, -am = via CTC forced aligner (-am canary-ctc-aligner.gguf or -am qwen3-forced-aligner.gguf), LID = via external language identification pre-step (-l auto), pp = via --punc-model post-processor (FireRedPunc or fullstop-punc), * = experimental or partial support, † = PLUS variant only (native [T:N] word timestamps with -owts; base uses -am). granite-4.1 covers both the regular and -plus variants; granite-4.1-nar is a non-autoregressive variant with encoder+projector only (no LLM decode features). The KV quant row marks backends that honor CRISPASR_KV_QUANT={f16,q8_0,q4_0} — CTC-style backends without a KV cache (parakeet, fc-ctc, wav2vec2, kyutai-stt, firered, moonshine variants, omniasr-CTC) don't apply. The same backends also honor the per-half CRISPASR_KV_QUANT_K / CRISPASR_KV_QUANT_V overrides (llama.cpp --cache-type-k / --cache-type-v parity) for asymmetric K-vs-V precision; common recipe K=q8_0 V=q4_0 saves ~40 % more KV memory than symmetric Q8_0. The mmap weights row marks backends consuming core_gguf::load_weights() and therefore honoring CRISPASR_GGUF_MMAP=1; whisper itself uses upstream's loader and is unaffected. See docs/cli.md Memory footprint for usage + recommended combos.

Speaker diarization as a post-processing step via --diarize:

  • energy / xcorr — stereo-only, no extra deps
  • foxnosebest accuracy, no external deps: WeSpeaker ResNet34-LM embeddings + GMM/BIC speaker counting + spectral clustering + Viterbi smoothing. Estimates the speaker count rather than needing it up front; --diarize-embedder auto fetches the GGUF (24 MB, CC-BY-4.0). 7.3 % DER on VoxConverse dev where pyannote + TitaNet scores 7.8 %, and 3.18 % vs the upstream reference's 3.07 % when scored on turns (more)
  • pyannote — native GGUF (no Python, no sherpa-onnx); add --diarize-embedder auto (TitaNet) or --diarize-embedder indextts (ECAPA-TDNN) for globally stable speaker IDs across long files
  • sherpa / ecapa — external sherpa-onnx subprocess; runs once globally on full audio for consistent speaker IDs (#110)
  • vad-turns — mono-friendly gap-based proxy

The server endpoint supports response_format=diarized_json for structured speaker-labelled output with normalised speaker letters (A, B, C …) — see docs/server.md.

Full reference + tuning knobs (cluster threshold, max speakers, pluggable embedder adapters): see docs/cli.md#diarization.

Language identification for backends without native LID: --lid-backend whisper (default, 75 MB ggml-tiny.bin), --lid-backend silero (native GGUF, 16 MB, 95 languages), or --lid-backend firered (FireRedLID, 1.7 GB, 120 languages — Conformer encoder + Transformer decoder).

Voice activity detection: --vad uses the default Silero VAD (~885 KB, auto-downloaded). Each VAD segment is transcribed independently, producing separate SRT/VTT entries with correct timestamps. Use --vad --split-on-punct for best subtitle output. Four VAD backends: Silero (default), FireRedVAD (-vm firered, recommended), MarbleNet (-vm marblenet, 439 KB, 6 languages), Whisper-VAD-EncDec (-vm whisper-vad, experimental).

Punctuation restoration (--punc-model): CTC-based backends output lowercase without punctuation. Named shortcuts: auto/firered (Chinese+English), fullstop (EN/DE/FR/IT, XLM-R-large), punctuate-all (12 languages, XLM-R-base), pcs (47 languages, punc + truecasing + sentence boundary detection in one model). Or pass a GGUF path directly. Also available via Python/Rust/Dart wrappers (crispasr.PuncModel).

Truecasing (--truecase-model): Restore German noun/name capitalization in lowercase ASR output. Three options in ascending quality: auto (statistical, 9 MB), crf (CRF with context, 24 MB), lstm (BiLSTM char-level, 3.2 MB, recommended — 97.9% F1, handles adjective/noun distinction and formal "Ihnen"). All auto-download from cstr/truecaser-de. Or use --punc-model pcs for neural punc + truecasing in one pass (47 languages).

Which backends produce punctuation natively?
BackendPunctuationCapitalizationNotes
whisperFull punctuation and casing
parakeet
canary
cohereToggleable via --no-punctuation
graniteLLM output
voxtralLLM output
voxtral4bLLM output
qwen3LLM output
funasrLLM output (Qwen3-0.6B decoder). Chinese chars carry full-width period; mlt-nano variant adds Latin-script casing + punctuation.
sensevoiceCTC output with native ITN — toggle via --punctuation / --no-punctuation, controls Arabic-digit vs spelled-out numerals + comma/period emission.
paraformernonoNAR character-level output — add --punc-model
gigaam✔ (e2e_*)✔ (e2e_*)The e2e_* revisions carry punctuation + casing + inverse text normalization in the SentencePiece vocabulary. The charwise ctc / rnnt revisions emit lowercase Cyrillic with no punctuation — but auto-restoration is still suppressed for them, because the auto-enabled FireRedPunc is a Chinese/English model and injects full-width CJK punctuation into Russian. Use an e2e_* revision for punctuated output, or pass an explicit --punc-model.
glm-asrLLM output
kyutai-sttLLM output
moonshineEncoder-decoder output
fastconformer-ctcnonoCTC — add --punc-model
wav2vec2nonoCTC — add --punc-model
firered-asrnonoCTC — add --punc-model
omniasr (CTC)nonoCTC — add --punc-model
omniasr (LLM)Autoregressive decoder

Other freely-licensed alternatives that could be added: felflare/bert-restore-punctuation (MIT, English, includes truecasing), xashru/punctuation-restoration (Apache-2.0, 40+ languages, BiLSTM-CRF).

Progressive subtitle output (--flush-after): By default, non-whisper backends buffer all segments and print output at the end. For real-time subtitle consumption (PotPlayer, custom media players), use --flush-after 1 to print each SRT entry to stdout immediately after its VAD segment is transcribed:

crispasr --backend parakeet -m parakeet.gguf --vad --flush-after 1 -osrt -f long_audio.wav
# SRT entries appear progressively as each segment finishes

JSON output with language detection: When using -l auto -oj, the JSON output includes detected language info:

{
  "crispasr": {
    "backend": "cohere",
    "language": "en",
    "language_detected": "en",
    "language_confidence": 0.977,
    "language_source": "ecapa"
  },
  "transcription": [...]
}

Which backend should I pick?

NeedPick
Battle-tested, all features exposedwhisper
Lowest English WERcohere
Fastest (16x realtime on CPU)moonshine (tiny), fc-ctc (10x)
Multilingual + word timestamps + fastparakeet (2.9x RT)
Multilingual with explicit language controlcanary
Speech translation (X→en or en→X)canary, voxtral, qwen3
30 languages + Chinese dialectsqwen3
1600+ languagesomniasr (CTC or LLM)
Realtime streaming ASR (native incremental encoder, ~2× RT feed; sub-second-token target deferred to phase 2)voxtral4b
Highest-quality offline speech-LLMvoxtral
Apache-licensed speech-LLMgranite, voxtral, qwen3, omniasr-llm
Lightweight CTC-only (fast, no decoder)wav2vec2, fc-ctc, data2vec, omniasr
Russiangigaam (e2e_rnnt — 8.4 % avg WER, punctuation + ITN), whisper, qwen3
Mandarin + Chinese dialectsfirered-asr, qwen3, glm-asr, funasr, paraformer, sensevoice
Multilingual (31 langs) speech-LLMfun-asr-mlt-nano, qwen3, omniasr-llm, gemma4-e2b
Multilingual (50+ langs) + LID + audio-event in one passsensevoice (encoder-only CTC, non-AR, 15× faster than Whisper-Large)

CPU performance tips

Audio-LLM backends (qwen3, voxtral, granite, glm-asr, etc.) run full transformer decoder stacks (28+ layers, 2048-dim) and are dramatically slower on CPU than encoder-only backends. On older dual-core hardware they can drop below 0.01× realtime. If you're on CPU-only hardware:

  • Prefer moonshine (16× RT), fc-ctc (10× RT), parakeet (2.9× RT), or whisper for usable speeds.
  • Use --flush-after 1 to see results as each VAD slice completes instead of waiting for the entire file.
  • Use -pp / --print-progress for per-slice progress indicators on all backends (unified backends show slice-level progress; whisper shows encoder-level progress).
  • Quantize models to Q4_K or Q5_K to reduce memory and compute.

Language detection for backends that don't do it natively

Cohere, canary, granite, voxtral and voxtral4b need an explicit language code up front. If you don't know the language, pass -l auto and crispasr runs an optional LID pre-step before the main transcribe() call:

# Downloads ggml-tiny.bin (75 MB, 99 languages) on first use
crispasr --backend cohere -m $TC/cohere-transcribe-q5_0.gguf \
         -f unknown.wav -l auto
# crispasr[lid]: detected 'en' (p=0.977) via whisper-tiny
# crispasr: LID -> language = 'en' (whisper, p=0.977)

These LID providers are available:

  • --lid-backend whisper (default) — uses a small multilingual ggml-*.bin model via the crispasr C API. Auto-downloads ~75 MB on first use. 99 languages.

  • --lid-backend silero — native GGUF port of Silero's 95-language classifier. 16 MB F32. Runs as a ggml graph (multi-threaded SIMD on CPU, GPU offload on Metal/CUDA; on Vulkan the graph is routed to CPU pending an upstream kernel fix). Analyzes the first 30 s of audio (CRISPASR_SILERO_LID_MAX_S overrides); CRISPASR_SILERO_LID_LEGACY=1 restores the old scalar path.

  • --lid-backend ecaparecommended: ECAPA-TDNN (Apache-2.0). Purpose-built for language ID. Very high accuracy on TTS benchmark. Two variants via --lid-model:

  • --lid-backend firered — FireRedLID (Conformer encoder + Transformer decoder). Q4_K (544 MB), 120 languages including Chinese dialects. Slower but covers more languages.

  • --lid-backend probe — no second model at all: ask the ASR model itself. Transcribes a 20 s clip once per language the model declares and keeps the best-scoring candidate (length × text-LID agreement × distinct-token ratio², the last term catching the repetitive output a wrong-language prompt produces). Currently implemented by cohere. The reason to prefer it is correctness, not just the saved download: an external detector knows 99 languages while Cohere Transcribe accepts 14 — and its Arabic finetune only en/ar — so external LID regularly returns a language the model was never trained on, and Cohere answers a wrong language fluently rather than failing. The probe cannot. Cost is one encode + one short decode per candidate, so it runs automatically only for a model with ≤ 4 languages (CRISPASR_COHERE_PROBE_MAX_LANGS); CRISPASR_COHERE_PROBE_TEXTLID=0 drops the text-LID agreement term.

    The ceiling is about cost, not accuracy. Measured on the real models: the two-language Arabic finetune picks ar for an Arabic clip (p=0.675) and en for samples/jfk.wav (p=0.647); the 14-language base model, probed across all 14, also gets both right (en p=0.169, ar p=0.254) — it is simply slower than an external detector. The encoder output is language-independent, so the probe encodes once and decodes per candidate (encode is ~87 % of a pass); a 14-candidate probe measured 12 s → 4-5 s against one-encode-per-candidate, byte-identical output. CRISPASR_COHERE_PROBE_REUSE_ENC=0 restores the naive path.

    The one soft spot worth knowing: asking the model for a language it was not trained on can yield a clean translation rather than garbage, which a text LID then confirms — "fluent French out" is not evidence of French in. That is what makes a mismatched whitelist dangerous, and it is measurable: force a 14-language list onto the two-language Arabic finetune and its fr probe returns real French and wins. The real base model's fr probe instead code-switches ("Et so, my fellow Americans…", agreement 0.00) and loses, as it should.

These VAD providers are available:

  • Silero VAD (default) — ~885 KB, auto-downloaded via --vad. Industry-standard, well-tested.
  • FireRedVAD — DFSMN-based, 2.4 MB, F1=97.57%. Pass --vad -vm firered to auto-download. Recommended.
  • MarbleNet — NVIDIA 1D separable CNN, 439 KB, 6 languages (EN/DE/FR/ES/RU/ZH). Pass --vad -vm marblenet to auto-download. Smallest model. (cstr/marblenet-vad-GGUF)
  • Whisper-VAD-EncDec (experimental) — Whisper-base encoder + TransformerDecoder head, 22 MB Q4_K. Trained on Japanese ASMR; may not generalise well to all domains. Pass --vad -vm whisper-vad. Slower than others (~1s vs ~50ms). (cstr/whisper-vad-encdec-asmr-GGUF)

Pass --lid-backend off to skip LID entirely.

Text language identification (post-ASR / standalone)

Audio LID (above) tags what was spoken; text LID tags what was written. Text LID runs on a transcript or any UTF-8 string and is useful for routing post-ASR pipelines (translation, punctuation, sub selection) without re-running an audio model. Three GGUF families, one binary — the dispatcher picks by general.architecture:

BackendLabelsSize (F16)LicenseHF repo
CLD3 (Google compact language detector v3)109 ISO 639-1440 KBApache-2.0cstr/cld3-GGUF
GlotLID-V3 (cis-lmu fastText)2102 ISO 639-3 + script250 MBApache-2.0cstr/glotlid-GGUF
LID-176 (Facebook fastText)176 ISO 639-163 MBCC-BY-NC-4.0¹cstr/fasttext-lid176-GGUF

¹ LID-176 is CC-BY-NC-4.0 — non-commercial use only. CLD3 + GlotLID-V3 are Apache-2.0 with no such constraint. Pick CLD3 for the smallest, fastest path; GlotLID for maximum coverage (low-resource languages); LID-176 only if you need its specific 176-label space and accept its non-commercial terms.

Standalone CLI — auto-routes by GGUF arch, with auto-download:

crispasr-lid -m auto --text "Bonjour le monde"        # → cstr/cld3-GGUF (default, ~440 KB)
crispasr-lid -m auto:glotlid --text "Bonjour le monde" -k 5
crispasr-lid -m auto:lid-fasttext176 --text "Hallo Welt"
# Or pass an explicit path / canonical filename (looked up in the registry):
crispasr-lid -m cld3-f16.gguf --text "你好世界"
# zh	0.997816
echo "Привет мир" | crispasr-lid -m auto --quiet
# ru	0.907322

Post-ASR pipeline--lid-on-transcript runs the same dispatcher on the assembled transcript (also accepts auto[:variant]):

crispasr -m ggml-tiny.bin -f speech.wav --lid-on-transcript auto
# (transcript on stdout)
# lang=de	conf=0.997123	backend=lid-cld3

The dispatcher (src/text_lid_dispatch.{h,cpp}) is a thin C ABI façade — one integer compare per call; per-stage diff harness is green at cos≥0.999 across 8 multilingual smoke samples.


Install & build

git clone --recursive https://github.com/CrispStrobe/CrispASR
cd CrispASR
# already cloned without --recursive? initialize the bundled ggml submodule:
#   git submodule update --init --recursive
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)

The ggml/ submodule is required. If you cloned without --recursive, run git submodule update --init --recursive first — otherwise CMake stops with a message telling you to do exactly that.

Produces build/bin/crispasr (main CLI), build/bin/crispasr-quantize, and build/bin/crispasr-diff. No Python, PyTorch, or pip required at runtime — just a C++17 compiler and CMake 3.14+.

For GPU acceleration, add the matching ggml flag at configure time:

cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON     # NVIDIA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON    # Apple Silicon
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON   # cross-vendor

See docs/install.md for the full guide: all GPU backends (CUDA / Metal / Vulkan / MUSA / SYCL), Windows convenience scripts, ffmpeg ingestion, optional BLAS, glibc notes, and the scripts/dev-build.sh wrapper.


Quick start

ASR examples below; for TTS see the Text-to-Speech section.

Whisper (historical path, byte-identical to upstream whisper.cpp)

# Download a whisper model (same as upstream whisper.cpp)
./models/download-ggml-model.sh base.en

./build/bin/crispasr -m models/ggml-base.en.bin -f samples/jfk.wav
# [00:00:00.000 --> 00:00:07.940]   And so my fellow Americans ask not what your country can do for you
# [00:00:07.940 --> 00:00:10.760]   ask what you can do for your country.

Parakeet (multilingual, free word timestamps, fastest)

# Grab the quantized model (~467 MB)
curl -L -o parakeet.gguf \
    https://huggingface.co/cstr/parakeet-tdt-0.6b-v3-GGUF/resolve/main/parakeet-tdt-0.6b-v3-q4_k.gguf

./build/bin/crispasr -m parakeet.gguf -f samples/jfk.wav
# Auto-detected backend 'parakeet' from GGUF metadata.
# And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.

# Word-level timestamps (one line per word)
./build/bin/crispasr -m parakeet.gguf -f samples/jfk.wav -ml 1

Canary (explicit language, speech translation)

# Transcription (source == target)
./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -sl de -tl de

# Translation (German speech → English text)
./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -sl de -tl en

# ...or use the familiar crispasr flag:
./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -l de --translate

Voxtral (speech-LLM with auto-download)

# First run downloads ~2.5 GB to ~/.cache/crispasr/ via curl, then runs
./build/bin/crispasr --backend voxtral -m auto -f samples/jfk.wav

# Subsequent runs use the cached file
./build/bin/crispasr --backend voxtral -m auto -f samples/jfk.wav -l en

Qwen3-ASR (30 languages + Chinese dialects)

# 0.6B (default, ~500 MB)
./build/bin/crispasr --backend qwen3 -m auto -f audio.zh.wav

# 1.7B (higher quality, ~1.3 GB) — supports both -hf and non-hf source models
./build/bin/crispasr --backend qwen3 -m qwen3-1.7b --auto-download -f audio.wav

# Japanese anime/galgame fine-tune (~1.3 GB)
./build/bin/crispasr --backend qwen3 -m qwen3-ja-anime --auto-download -f anime.wav

Long audio: the default is safe 30 s chunking. --chunk-seconds 0 decodes the whole file in ONE pass (matches the reference model verbatim on multi-minute clips, #218) — but the encoder's full attention is O(N²) in audio length, so keep single-pass clips under ~10 minutes on 16 GB machines. For long-form use prefer the plain -q4_k/-q8_0 GGUFs over the -imatrix variants (see the model card).

GLM-ASR-Nano (Mandarin + dialects + Cantonese + English, 1.5B)

./build/bin/crispasr --backend glm-asr -m auto -f audio.wav

# Long audio in one pass (up to 655 s — 30 s encoder windows, one LLM prompt,
# same layout as the HF/zai reference; matches it verbatim on the #218 clip):
./build/bin/crispasr --backend glm-asr -m auto --chunk-seconds 0 -f long.wav

Note: in single-pass mode the model (like the reference) skips leading non-speech audio; the default 30 s-chunked mode transcribes more of such clips. Custom --ask / non-English --language instructions need a GGUF with baked BPE merges (re-published 2026-07; older GGUFs fall back to the default transcription prompt with a warning).

MiMo-V2.5-ASR (Mandarin + dialects + English, 7.5B Qwen2 LM)

# Download the LM + audio tokenizer (the tokenizer is a separate model)
huggingface-cli download cstr/mimo-asr-GGUF mimo-asr-q4_k.gguf \
    --local-dir ~/.cache/crispasr
huggingface-cli download cstr/mimo-tokenizer-GGUF mimo-tokenizer-q4_k.gguf \
    --local-dir ~/.cache/crispasr

# Transcribe (auto-discovers tokenizer if it sits next to the LM)
./build/bin/crispasr \
    --backend mimo-asr \
    -m ~/.cache/crispasr/mimo-asr-q4_k.gguf \
    --codec-model ~/.cache/crispasr/mimo-tokenizer-q4_k.gguf \
    -f samples/jfk.wav
# Output: And so, my fellow Americans, ask not what your country can do
# for you. Ask what you can do for your country.

The 4.5 GB Q4_K is the recommended quant; F16 (14.9 GB) needs ~16 GB RAM during inference. JFK matches the upstream Python MimoAudio.asr_sft reference verbatim; performance on M1+Metal is ~0.3× realtime (Q4_K dequant per step is the bottleneck — F16 + KV-reuse follow-ups are queued under PLAN #51a/b/c).

Wav2Vec2 (lightweight CTC, any HF Wav2Vec2ForCTC model)

# English (Q4_K quantized, 212 MB — 6x smaller than F16)
curl -L -o wav2vec2-en-q4k.gguf \
    https://huggingface.co/cstr/wav2vec2-large-xlsr-53-english-GGUF/resolve/main/wav2vec2-xlsr-en-q4_k.gguf

./build/bin/crispasr -m wav2vec2-en-q4k.gguf -f samples/jfk.wav
# and so my fellow americans ask not what your country can do for you ask what you can do for your country

# German
curl -L -o wav2vec2-de-q4k.gguf \
    https://huggingface.co/cstr/wav2vec2-large-xlsr-53-german-GGUF/resolve/main/wav2vec2-xlsr-de-q4_k.gguf

./build/bin/crispasr -m wav2vec2-de-q4k.gguf -f audio.de.wav

# Convert any HuggingFace Wav2Vec2ForCTC model:
python models/convert-wav2vec2-to-gguf.py \
    --model-dir jonatasgrosman/wav2vec2-large-xlsr-53-german \
    --output wav2vec2-de.gguf --dtype f32
# Then optionally quantize:
./build/bin/crispasr-quantize wav2vec2-de.gguf wav2vec2-de-q4k.gguf q4_k

Streaming, TTS, and HTTP server

CrispASR has three feature areas that warrant their own docs pages:

  • Streaming & live transcription--stream, --mic, --live, sliding-window chunking, per-token confidence.
  • Text-to-Speech (TTS) — Kokoro (multilingual, smallest), Qwen3-TTS (highest fidelity, voice cloning), VibeVoice (lowest-latency streaming), Orpheus (3 B Llama + SNAC), Chatterbox (flow-matching + HiFT vocoder, German via Kartoffelbox), IndexTTS, VoxCPM2, and CosyVoice3 (9 langs + 18 zh dialects; baked-voice bank + arbitrary-WAV cloning). Voice packs, language routing, and qwen3-tts environment switches. All TTS output is watermarked; post-embed verification warns if confidence is low. Use --detect-watermark file.wav to check any WAV for AI watermarks.
  • Server mode (HTTP API) — persistent model, OpenAI-compatible /v1/audio/transcriptions (ASR) and /v1/audio/speech + /v1/voices (TTS, automatic on any loaded CAP_TTS backend), per-request voice + speed + instructions, CORS, long-form sentence chunking, API keys, Docker Compose, prebuilt CUDA images.
  • Concurrency, parallelism & scaling — one transcription already uses multiple cores; the server accepts requests concurrently but serializes inference on one model by default; --server-workers N runs N model instances so pure-ASR requests run concurrently; and for bulk/throughput workloads, process-level fan-out (xargs -P / GNU parallel) or N replicas behind a load balancer. Also covers what is not supported (batched multi-stream inference, PagedAttention) and why.

Quickest taste of each:

# Streaming from microphone
crispasr --mic -m model.gguf

# TTS via auto-downloaded VibeVoice (~636 MB on first run)
crispasr --backend vibevoice-tts -m auto --tts "Hello world" --tts-output hello.wav

# CosyVoice3 on GPU; companions auto-download beside the LLM
crispasr --backend cosyvoice3-tts -m auto --tts "Hello world" --tts-output cosy.wav

# CosyVoice3 fast mode: 5 flow steps instead of the quality-default 10
COSYVOICE3_FLOW_STEPS=5 crispasr --backend cosyvoice3-tts -m auto \
  --tts "Hello world" --tts-output cosy-fast.wav

# Persistent HTTP server, OpenAI-compatible
crispasr --server -m model.gguf --port 8080
curl -F "file=@audio.wav" http://localhost:8080/v1/audio/transcriptions

# TTS over HTTP — load a TTS backend, hit /v1/audio/speech
crispasr --server --backend qwen3-tts-customvoice -m auto --voice-dir ./voices --port 8080
curl http://localhost:8080/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{"input":"Hello world","voice":"vivian"}' -o out.wav

CosyVoice3 uses batched classifier-free guidance and request-sized KV caching by default. Baked voices load only the LLM, flow, HiFT, and voice bank; the larger S3 tokenizer and CAMPPlus companions load lazily when a .wav cloning voice is first requested.


CLI reference

Common flags:

crispasr -m auto --backend parakeet -f audio.wav --vad -osrt --split-on-punct
FlagMeaning
-m FNAME / --backend NAMEModel path (or auto) and forced backend
-f FNAMEInput audio (repeatable; positional accepted)
--vadSilero VAD chunking — strongly recommended for multi-minute audio
-osrt / -ovtt / -otxt / -oj / -ojfOutput formats (also -ocsv, -olrc)
-am FNAMECTC aligner GGUF for word-level timestamps on LLM backends
--align-onlyStandalone forced alignment: text/.srt + audio → timestamped SRT/JSON (no ASR needed); .srt input keeps its cues and gets re-timed (--align-granularity auto|word|segment)
-tp F / -bs NSampling temperature / beam search width
-n N / --frequency-penalty FGenerated-token cap / opt-in repeated-token penalty for supported autoregressive ASR backends
-l auto / --detect-languageLID pre-step for backends without native lang detect
--hotwords "A,B,C"Contextual biasing — boost named terms during CTC/TDT decode or LLM prompt
-ck NFallback chunk size when VAD is off (default 30 s)
--list-backendsPrint the capability matrix and exit

See docs/cli.md for the full reference: every flag, VAD details, CTC alignment workflow, output JSON layout, the auto-download registry, and supported audio formats. See docs/bindings.md for Python / Rust / Dart / Go / Java / JavaScript / Ruby / mobile.


Architecture, contributing, regression matrix

CrispASR is structured as a stable C-ABI in src/ (every algorithm: VAD, diarize, LID, alignment, cache, registry) consumed by all language wrappers, with thin presentation layers in examples/cli/. Per-model runtimes live in src/{whisper,parakeet,canary,...}.cpp, sharing primitives from src/core/ (mel, ffn, attention, GGUF loader, FastConformer / Conformer / Granite-LLM blocks, etc.).

  • docs/architecture.md — full layered layout, file-by-file tour of src/ and examples/cli/, per-backend internals table, regression discipline.
  • docs/contributing.md — adding a new backend in five files, clang-format-18 setup, the crispasr-diff PyTorch-ground-truth workflow, and the TTS audio-cosine-vs-reference regression target.
  • docs/regression-matrix.mdtools/test-all-backends.py capability tiers, cache modes (keep / ephemeral), --skip-missing for CI.

Shared libraries (cross-repo with CrispEmbed):

  • crisp_audio/ — Whisper-shape audio encoder (Conv-stem + Transformer)
  • crisp_punc/ — punctuation restoration (FireRedPunc + PCS)
  • crisp_lid/ — text-based language identification (fastText + CLD3)
  • crisp_truecase/ — truecasing (statistical + CRF + BiLSTM)

Both are self-contained static libraries with CMakeLists.txt. CrispEmbed links them via add_subdirectory(../CrispASR/crisp_*/); CrispASR uses them directly. If the shared dir is absent, both repos fall back to local copies of the source files.

For benchmarks see PERFORMANCE.md; for the session-by-session port log and the bug-class lessons, see LEARNINGS.md.


Quantize models

build/bin/crispasr-quantize is a single, model-agnostic GGUF re-quantization tool that works across all supported model families (Whisper, Parakeet, Canary, Cohere, Voxtral, Qwen3, Granite, Wav2Vec2, MiMo-ASR, GLM-ASR, Moonshine, VibeVoice, Kokoro, Qwen3-TTS, …):

./build/bin/crispasr-quantize input.gguf output.gguf q4_k

See docs/quantize.md for the full guide: supported quant types, K-quant alignment fallback, recommended quant per backend, and worked examples for each architecture.


GPU backend selection

All backends use ggml_backend_init_best() which automatically picks the highest-priority compiled backend: CUDA > Metal > Vulkan > CPU. To force a specific backend:

# Force Vulkan even when CUDA is available
crispasr --gpu-backend vulkan -m model.gguf -f audio.wav

# Pin a specific GPU (useful on Vulkan systems with iGPU + dGPU)
crispasr --gpu-backend vulkan -dev 1 -m model.gguf -f audio.wav

# Force CPU (useful for benchmarking)
crispasr -ng -m model.gguf -f audio.wav

# CUDA unified memory (swap to RAM when VRAM exhausted)
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 crispasr -m model.gguf -f audio.wav

Build flags: -DGGML_CUDA=ON, -DGGML_METAL=ON, -DGGML_VULKAN=ON.

Notes:

  • --gpu-backend vulkan selects the Vulkan backend, but it does not choose which physical GPU to use. Use -dev N to select the Vulkan device index.
  • On some Windows laptops, Vulkan device 0 is the Intel iGPU and the NVIDIA GPU is 1. If Vulkan looks unexpectedly slow, rerun with -dev 1.
  • The Windows convenience script build-vulkan.bat creates a separate Vulkan-capable binary at build-vulkan\bin\crispasr.exe.

Debugging & profiling

For most backends, -v / --verbose surfaces per-stage timings and device picks. For headless / library use (where the CLI flag isn't plumbed through), set CRISPASR_VERBOSE=1 instead.

# Per-stage timing breakdown (mel / encoder / prefill / decode):
crispasr -v --backend gemma4-e2b -m model.gguf -f audio.wav
# gemma4_e2b: mel 128x1099 (17.2 ms)
# gemma4_e2b: encoder done: 1536x275 (719.0 ms)
# gemma4_e2b: prefill done, first_token=3133 (1464.0 ms)
# gemma4_e2b: decoded 25 tokens (7748.3 ms total)
# crispasr: transcribed 11.0s audio in 7.75s (1.4x realtime)

# Hugging Face access for gated models (Voxtral, Gemma4-E2B, …):
HF_TOKEN=hf_xxx crispasr -m auto --backend gemma4-e2b -f audio.wav

The server has its own auth env: CRISPASR_API_KEYS (see Server mode).

Per-backend debug / bench / dump-dir env vars (developer)

These are useful when porting a new backend or chasing a regression. The *_BENCH=1 toggles emit per-stage timings even without -v; the *_DEBUG=1 toggles emit per-step diagnostic prints; the *_DUMP_DIR= paths write per-stage F32 tensors for diff-testing against a PyTorch reference (see Debug a new backend against PyTorch ground truth).

Env varPurpose
CRISPASR_VERBOSE=1Forces verbose mode for any backend (parallel to the -v flag).
CRISPASR_DUMP_DIR=path/Generic per-stage F32 tensor dump for the crispasr-diff harness.
GEMMA4_E2B_BENCH=1Per-stage timings for the Gemma-4-E2B backend.
COHERE_BENCH=1 / COHERE_DEBUG=1Cohere transcribe per-stage timings / per-step diagnostics.
COHERE_PROF=1Cohere graph-level profiling (per-op timings).
COHERE_THREADS=NOverride thread count for the Cohere backend.
COHERE_DEVICE=cpu|cuda|metal|vulkanForce the Cohere backend onto a specific device.
COHERE_DUMP_ATTN=path/Dump attention activations for Cohere (used by the diff harness).
FIRERED_BENCH=1Per-stage timings for the FireRedASR backend.
FIREREDPUNC_DEBUG=1Per-step diagnostics for the FireRed punctuation post-step.
MOONSHINE_STREAMING_BENCH=1Per-stage timings for moonshine-streaming.
OMNIASR_BENCH=1 / OMNIASR_DEBUG=1 / OMNIASR_DUMP_DIR=OmniASR per-stage timings, diagnostics, and stage dumps.
PARAKEET_DEBUG=1Parakeet TDT per-step diagnostics (joint network, blank-id sanity).
QWEN3_TTS_BENCH=1 / QWEN3_TTS_DEBUG=1 / QWEN3_TTS_DUMP_DIR=Qwen3-TTS per-stage timings, diagnostics, and stage dumps.
VIBEVOICE_BENCH=1 / VIBEVOICE_DEBUG=1 / VIBEVOICE_DUMP_DIR=VibeVoice ASR per-stage timings, diagnostics, and stage dumps.
VIBEVOICE_REF_FEATURES=pathReplace the live encoder with a saved feature tensor (regression harness).
VIBEVOICE_TTS_DUMP=path/VibeVoice TTS per-stage dumps (token IDs, base/TTS hidden, neg condition, frame-0 noise/v_cfg/latent/acoustic_embed) for the diff harness.
VIBEVOICE_TTS_DUMP_PERFRAME=1Per-frame VibeVoice TTS dumps written as perframe_<stage>_f<NNN>.bin. Pair with VIBEVOICE_TTS_DUMP=path/ and VIBEVOICE_TTS_NOISE=path for stage-by-stage AR diff against tools/run_official_vibevoice.py.
VIBEVOICE_TTS_TRACE=1Extra one-line traces (negative-condition prefill rms, scaling/bias factors loaded). Same effect as -vv.
VIBEVOICE_VOICE_AUDIO=path.wavReference voice WAV for 1.5B-base TTS without a .gguf voice cache.
VIBEVOICE_TTS_NOISE=pathOverride the per-frame Gaussian init noise. Flat little-endian float32 [N_frames, vae_dim] — typically the noise.bin written by tools/run_official_vibevoice.py.
VIBEVOICE_VAE_BACKEND=cpu|metal|cuda|vulkanPin the VAE decoder onto a specific backend.
WAV2VEC2_BENCH=1 / WAV2VEC2_VERBOSE=1 / WAV2VEC2_DUMP_DIR=wav2vec2 per-stage timings, verbose graph traces, and stage dumps.
CRISPASR_VOXTRAL4B_STREAM_TIMING=1Per-stage timings for the voxtral4b streaming path (encoder drain / prefill / first-text-token / decode-step p50/p95).
CRISPASR_VOXTRAL4B_STREAM_CHUNK_MS=NOverride the internal encoder chunk size (default 240 ms). Must be a multiple of 80 ms. Larger = faster feed (kernel-launch amortisation), longer live-caption latency floor.
CRISPASR_VOXTRAL4B_STREAM_BATCH_ENCODER=1Regression-debug: ignore the streaming encoder's audio_embeds and re-run the whole batch encoder at flush.
CRISPASR_VOXTRAL4B_STREAM_DEBUG=1 / CRISPASR_VOXTRAL4B_STREAM_DIFF=1Per-step decode prints / side-by-side encoder cosine vs the batch encoder.
CRISPASR_VOXTRAL4B_STREAM_LIVE=1Live-captions decode-during-feed (PLAN #7 phase 3). get_text() polled during feed returns progressive transcript. Default OFF (PTT semantics). Wrappers: Python Session.stream_open(live=True), Rust stream_open_ex(.., live: true).
CRISPASR_VOXTRAL4B_STREAM_DECODER_THREAD=1Decoder worker thread (PLAN #7 phase 4, implies live mode). Lets feed() return between encoder chunks without waiting for the decode loop — useful for mic-driven workloads. On M1 the Metal queue serializes encoder and decoder so total wall-clock is unchanged; faster GPUs with kernel-level parallelism see real overlap.
CRISPASR_VOXTRAL4B_FUSED_QKV=0Opt out of the runtime fused-QKV LLM path (default-on, ~7-8 % decode speedup on M1 Q4_K, ~500 MB extra memory).
CRISPASR_QWEN3_ASR_FUSED_QKV=0Opt out of the runtime fused-QKV LLM path for qwen3-asr (default-on; works on F16/F32/Q4_K/Q8_0/...).
CRISPASR_VOXTRAL_FUSED_QKV=1Opt in to the runtime fused-QKV LLM path for voxtral 3B. Off by default (no measurable speedup on JFK-shape decodes; useful for long-form workloads where decode dominates).
QWEN3_TTS_FUSED_QKV=1Opt in to the runtime fused-QKV talker path.
GRANITE_DISABLE_ENCODER_GRAPH=1Force the granite-speech / -plus / -nar encoder back to the per-layer CPU loop (slower but kept around for debugging). The single ggml-graph encoder with per-layer Shaw RPE is the default and is bit-near-identical to the CPU loop while being ~2× faster end-to-end across all three variants.
CRISPASR_NO_REL_POS=1Ablate the relative-position bias in the Gemma-4 audio encoder (development only).
ECAPA_REF_FBANK=pathReference filterbank tensor for the ECAPA-TDNN LID model (regression harness).
CRISPASR_SHERPA_LID_BIN=pathOverride the auto-detected sherpa-onnx LID binary.
CRISPASR_ARG_DEVICE=NDefault GPU device index when -dev isn't passed.
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1Let CUDA swap to RAM when VRAM is exhausted.
GGML_VK_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICESStandard ggml/CUDA device-visibility filters.

HF_TOKEN and HUGGING_FACE_HUB_TOKEN are both honoured for gated-model downloads (in that order).


Credits

  • whisper.cpp — the original ggml inference engine and Whisper runtime this fork is built on
  • ggml — the tensor library everything runs on
  • NVIDIA NeMo — parakeet-tdt-{0.6b-v2,0.6b-v3,1.1b}, parakeet-tdt_ctc-{110m,1.1b,0.6b-ja}, parakeet-ctc-{0.6b,1.1b}, canary-1b-v2, canary-ctc aligner, and the FastConformer-CTC family (stt_en_fastconformer_ctc_{large,xlarge,xxlarge} plus CTC branches of the stt_*_fastconformer_hybrid_large[_pc] fleet: en-pc, de, es, fr, it, nl, pl, ru, ua, hr, be, ar, fa, ka, hy, uz, kk-ru — all usable both as ASR backends and as compact ~82 MB -am forced aligners)
  • Cohere — cohere-transcribe-03-2026
  • Qwen team (Alibaba) — Qwen3-ASR-0.6B, Qwen3-ASR-1.7B, Qwen3-ForcedAligner-0.6B
  • Mistral AI — Voxtral Mini 3B and 4B Realtime
  • IBM Granite team — Granite Speech 3.2-8b, 3.3-2b, 3.3-8b, 4.0-1b
  • Meta / wav2vec2 — wav2vec2 CTC models (XLSR-53 English, German, multilingual via any Wav2Vec2ForCTC checkpoint)
  • sherpa-onnx — optional diarization via subprocess (ONNX models)
  • Silero — VAD (native GGUF) and language identification (native GGUF, 95 languages)
  • pyannote — speaker diarization segmentation (native GGUF port)
  • miniaudio + stb_vorbis + libopus/opusfile — embedded/linked audio decoders (WAV/MP3/FLAC/AIFF/OGG/Opus, no ffmpeg; AAC/M4A/ALAC via Apple AudioToolbox)
  • glint (MIT) — in-tree clean-room codec suite: MP3/AAC-LC/Ogg-Opus output encoders (TTS .mp3/.aac/.opus) and cross-platform ADTS AAC-LC + Ogg Opus input decoders (no libopus needed)
  • Claude Code (Anthropic) — significant portions of the crispasr integration layer, all model converters, and the FastConformer/attention/mel/FFN/BPE core helpers were co-authored with Claude

License

Same as upstream whisper.cpp: MIT.

Per-model weights are covered by their respective HuggingFace model licenses (see Supported backends). The crispasr binary itself links model runtimes that are mostly permissively licensed (MIT / Apache-2.0 / CC-BY-4.0 for weights).