Whisper large

June 28, 2026 · View on GitHub

OpenAI's openai/whisper-large ported to transcribe.cpp. A 1.55B-parameter encoder-decoder transformer (audio encoder + autoregressive text decoder with cross-attention).

What it's for

Offline multilingual speech-to-text and any-language → English speech translation. The model auto-detects the audio's language (99 languages covered) and emits a transcript in that language; passing language="<code>" and task="translate" to the underlying whisper_full_params produces an English translation instead. transcribe-cli reads a 16 kHz mono WAV and returns the transcript text. Long audio is handled via 30-second chunked decoding.

See the upstream model card for training data, intended use, and the original evaluation methodology.

Licensed Apache-2.0. Ported from upstream commit 4ef9b41, pinned 2026-04-25. Validated against the transformers reference at transcribe.cpp commit 5.6.1 on 2026-04-26.

Download

QuantizationDownloadSizeWER (LibriSpeech test-clean)
F32whisper-large-F32.gguf5.75 GB2.72%
F16whisper-large-F16.gguf2.89 GB2.74%
Q8_0whisper-large-Q8_0.gguf1.55 GB2.74%
Q6_Kwhisper-large-Q6_K.gguf1.21 GB2.62%
Q5_K_Mwhisper-large-Q5_K_M.gguf1.08 GB2.70%
Q4_K_Mwhisper-large-Q4_K_M.gguf950 MB2.67%

WER measured on the full LibriSpeech test-clean split (2620 utterances) with transcribe.cpp's default greedy decode and segment timestamps enabled — the same runs summarized in the Whisper family table. Numbers come from a single Metal-backed run; Metal's non-deterministic parallel reductions add ~0.1pp of run-to-run variance on the noise floor, and quantization is otherwise generally WER-neutral. See the WER methodology for the harness.

Quick Start

cmake -B build
cmake --build build

build/bin/transcribe-cli \
  -m models/whisper-large/whisper-large-Q8_0.gguf \
  samples/jfk.wav

If your audio is not already 16 kHz mono WAV, convert it first:

ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav

Performance

Cells are wall-clock latency (mel + encode + decode, mean over the recorded iterations after warmup), with speedup over realtime in parentheses. Units: ms below 1 s, s above (2 decimal places). Decode latency dominates as model size grows; the encoder is only run once per 30-second window.

Apple M4 Max

BackendSampleQ8_0Q4_K_M
Metaljfk (11.0s)476.5 ms (23.1×)465.1 ms (23.6×)
Metaldots (35.3s)1.33 s (26.6×)1.26 s (28.0×)
CPUjfk (11.0s)9.63 s (1.1×)7.43 s (1.5×)
CPUdots (35.3s)19.88 s (1.8×)15.49 s (2.3×)

macOS 26.4.1, transcribe.cpp e0fa0f6.

Benchmark reproduction:

uv run scripts/bench/run.py \
  --models whisper-large \
  --quants q8_0,q4_k_m \
  --samples jfk,dots \
  --backends metal,cpu \
  --iters 3 --warmup 1 \
  --name whisper-large-publication

AMD Ryzen 7 PRO 4750U

BackendSampleQ8_0Q4_K_M
Vulkanjfk (11.0s)6.26 s (1.8×)6.13 s (1.8×)
Vulkandots (35.3s)14.41 s (2.5×)13.72 s (2.6×)
CPUjfk (11.0s)26.18 s (0.4×)19.83 s (0.6×)
CPUdots (35.3s)55.64 s (0.6×)43.98 s (0.8×)

Fedora 43, transcribe.cpp e0fa0f6. Vulkan device: AMD Radeon Graphics (RADV RENOIR).

Benchmark reproduction:

uv run scripts/bench/run.py \
  --models whisper-large \
  --quants q8_0,q4_k_m \
  --samples jfk,dots \
  --backends cpu,vulkan \
  --iters 3 --warmup 1 \
  --name whisper-large-publication

Numerical Validation

transcribe.cpp is validated tensor-by-tensor against the transformers reference (WhisperForConditionalGeneration, fp32 CPU) on the manifest's case (samples/jfk.wav). All 23 checkpointed tensors fall within per-variant tolerance, and the transcript matches the HF reference verbatim. Tolerance budget lives at tests/tolerances/whisper-large.json. Last validated at commit 1854f57.

FieldValue
Referencetransformers 5.6.1 (WhisperForConditionalGeneration, CPU fp32)
Manifesttests/golden/whisper/whisper-large.manifest.json
Tolerance filetests/tolerances/whisper-large.json
Commanduv run scripts/validate.py all --family whisper --variant whisper-large

Selected tensors (worst observed across cases; see tolerance file for per-tensor budgets):

TensorMax abs diffMean abs diffNotes
enc.mel.in2.229e-053.381e-08fp32 mixed-radix FFT vs torch fp64 frontend
enc.conv1.out2.146e-063.695e-08fp32 conv stem
enc.conv2.out2.098e-053.419e-07stride-2 conv stem (matches enc.embed.out)
enc.block.0.out2.718e-058.861e-07first encoder block
enc.block.31.out3.262e-017.382e-06final encoder block (peak signal grows with depth)
enc.final1.916e-035.219e-06post-LN encoder output
dec.token_emb0.000e+000.000e+00exact zero-drift (ggml_get_rows on the F32 GGUF)
dec.block.0.out1.907e-053.328e-07first decoder block, prompt pass
dec.block.31.out2.441e-046.015e-06final decoder block (accumulated)
dec.out_before_head2.522e-041.035e-05post final LN, pre-vocab projection
dec.logits_raw9.346e-051.508e-05vocab projection (raw logits)
dec.logits8.965e-052.757e-05log-softmax over vocab
dec.logits_raw.gen203.052e-051.148e-05step-20 logits (KV-cached path)

The C++ mel frontend (Slaney filterbank + Hann periodic window + whisper-style log-mel compression) drives enc.mel.in to fp32-vs-fp64 STFT precision drift; downstream tensors stay within budget. KV-cached decoder runs through F16 self/cross caches by default — flip with --kv-type f32 for tighter parity.

Reproduction

Convert

The whisper converter loads from a Hugging Face checkpoint and emits a reference-dtype GGUF.

uv run --project scripts/envs/whisper \
  scripts/convert-whisper.py openai/whisper-large \
  --revision 4ef9b41

Quantize

Run transcribe-quantize once per target quant. Example for Q8_0; repeat for the other shipped presets:

build/bin/transcribe-quantize \
  models/whisper-large/whisper-large-F32.gguf \
  models/whisper-large/whisper-large-Q8_0.gguf \
  --quant Q8_0

Validate

uv run scripts/validate.py all --family whisper --variant whisper-large

Run real-model tests

cmake -B build -DTRANSCRIBE_BUILD_REAL_MODEL_TESTS=ON
cmake --build build

TRANSCRIBE_WHISPER_GGUF=$PWD/models/whisper-large/whisper-large-Q8_0.gguf \
  ctest --test-dir build --output-on-failure -R whisper