Whisper tiny.en

June 28, 2026 · View on GitHub

OpenAI's openai/whisper-tiny.en ported to transcribe.cpp. A 39M-parameter encoder-decoder transformer (audio encoder + autoregressive text decoder with cross-attention).

What it's for

Offline English speech-to-text. The model takes a 16 kHz mono WAV and returns a transcript. English-only checkpoints are typically faster and slightly more accurate than the multilingual model at the same parameter count, but they cannot transcribe other languages and cannot translate. Long audio is handled via 30-second chunked decoding.

See the upstream model card for training data, intended use, and the original evaluation methodology.

Licensed Apache-2.0. Ported from upstream commit 87c7102, pinned 2026-04-25. Validated against the transformers reference at transcribe.cpp commit 5.6.1 on 2026-04-26.

Download

QuantizationDownloadSizeWER (LibriSpeech test-clean)
F32whisper-tiny.en-F32.gguf146 MB5.77%
F16whisper-tiny.en-F16.gguf76 MB5.77%
Q8_0whisper-tiny.en-Q8_0.gguf44 MB5.72%
Q6_Kwhisper-tiny.en-Q6_K.gguf43 MB5.80%
Q5_K_Mwhisper-tiny.en-Q5_K_M.gguf42 MB5.89%
Q4_K_Mwhisper-tiny.en-Q4_K_M.gguf42 MB5.99%

WER measured on the full LibriSpeech test-clean split (2620 utterances) with transcribe.cpp's default greedy decode and segment timestamps enabled — the same runs summarized in the Whisper family table. Numbers come from a single Metal-backed run; Metal's non-deterministic parallel reductions add ~0.1pp of run-to-run variance on the noise floor, and quantization is otherwise generally WER-neutral. See the WER methodology for the harness.

Quick Start

cmake -B build
cmake --build build

build/bin/transcribe-cli \
  -m models/whisper-tiny.en/whisper-tiny.en-Q8_0.gguf \
  samples/jfk.wav

If your audio is not already 16 kHz mono WAV, convert it first:

ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav

Performance

Cells are wall-clock latency (mean over 3 iterations after 1 warmup), with speedup over realtime in parentheses. Units: ms below 1 s, s above (2 decimal places). Decode latency dominates as model size grows; the encoder is only run once per 30-second window.

Apple M4 Max

BackendSampleQ8_0Q4_K_M
Metaljfk (11.0s)39.1 ms (281.2×)34.0 ms (323.8×)
Metaldots (35.3s)127.0 ms (278.3×)125.8 ms (280.9×)
CPUjfk (11.0s)165.0 ms (66.7×)161.4 ms (68.1×)
CPUdots (35.3s)389.4 ms (90.7×)381.8 ms (92.5×)

macOS 26.4.1, transcribe.cpp e0fa0f6.

Benchmark reproduction:

uv run scripts/bench/run.py \
  --models whisper-tiny.en \
  --quants q8_0,q4_k_m \
  --samples jfk,dots \
  --backends metal,cpu \
  --iters 3 --warmup 1 \
  --name whisper-tiny.en-publication

AMD Ryzen 7 PRO 4750U

BackendSampleQ8_0Q4_K_M
Vulkanjfk (11.0s)197 ms (56.0×)193 ms (56.9×)
Vulkandots (35.3s)540 ms (65.4×)541 ms (65.3×)
CPUjfk (11.0s)493 ms (22.3×)436 ms (25.2×)
CPUdots (35.3s)1.19 s (29.8×)1.09 s (32.5×)

Fedora 43, transcribe.cpp e0fa0f6. Vulkan device: AMD Radeon Graphics (RADV RENOIR).

Benchmark reproduction:

uv run scripts/bench/run.py \
  --models whisper-tiny.en \
  --quants q8_0,q4_k_m \
  --samples jfk,dots \
  --backends cpu,vulkan \
  --iters 3 --warmup 1 \
  --name whisper-tiny.en-publication

Numerical Validation

transcribe.cpp is validated tensor-by-tensor against the transformers reference (WhisperForConditionalGeneration, fp32 CPU) on the manifest's case (samples/jfk.wav). All 21 checkpointed tensors fall within per-variant tolerance. Tolerance budget lives at tests/tolerances/whisper-tiny.en.json. Last validated at commit 1854f57.

FieldValue
Referencetransformers 5.6.1 (WhisperForConditionalGeneration, CPU fp32)
Manifesttests/golden/whisper/whisper-tiny.en.manifest.json
Tolerance filetests/tolerances/whisper-tiny.en.json
Commanduv run scripts/validate.py all --family whisper --variant whisper-tiny.en

Selected tensors (worst observed across cases; see tolerance file for per-tensor budgets):

TensorMax abs diffMean abs diffNotes
enc.mel.in2.229e-053.381e-08fp32 mixed-radix FFT vs torch fp64 frontend
enc.conv1.out7.272e-067.293e-08fp32 conv stem
enc.conv2.out1.240e-052.603e-07stride-2 conv stem (matches enc.embed.out)
enc.block.0.out2.623e-058.075e-07first encoder block
enc.block.3.out1.709e-023.026e-06final encoder block (peak signal grows with depth)
enc.final4.005e-052.161e-06post-LN encoder output
dec.token_emb0.000e+000.000e+00exact zero-drift (ggml_get_rows on the F32 GGUF)
dec.block.0.out2.861e-053.417e-07first decoder block, prompt pass
dec.block.3.out4.196e-051.334e-06final decoder block (accumulated)
dec.out_before_head1.450e-042.361e-05post final LN, pre-vocab projection
dec.logits_raw6.807e-053.062e-05vocab projection (raw logits)
dec.logits9.584e-053.682e-05log-softmax over vocab
dec.logits_raw.gen207.248e-054.521e-05step-20 logits (KV-cached path)

The C++ mel frontend (Slaney filterbank + Hann periodic window + whisper-style log-mel compression) drives enc.mel.in to fp32-vs-fp64 STFT precision drift; downstream tensors stay within budget. KV-cached decoder runs through F16 self/cross caches by default — flip with --kv-type f32 for tighter parity.

Reproduction

Convert

The whisper converter loads from a Hugging Face checkpoint and emits a reference-dtype GGUF.

uv run --project scripts/envs/whisper \
  scripts/convert-whisper.py openai/whisper-tiny.en \
  --revision 87c7102

Quantize

Run transcribe-quantize once per target quant. Example for Q8_0; repeat for the other shipped presets:

build/bin/transcribe-quantize \
  models/whisper-tiny.en/whisper-tiny.en-F32.gguf \
  models/whisper-tiny.en/whisper-tiny.en-Q8_0.gguf \
  --quant Q8_0

Validate

uv run scripts/validate.py all --family whisper --variant whisper-tiny.en

Run real-model tests

cmake -B build -DTRANSCRIBE_BUILD_REAL_MODEL_TESTS=ON
cmake --build build

TRANSCRIBE_WHISPER_GGUF=$PWD/models/whisper-tiny.en/whisper-tiny.en-Q8_0.gguf \
  ctest --test-dir build --output-on-failure -R whisper