rav1d-safe [](https://github.com/imazen/rav1d-safe/actions/workflows/ci.yml) [](https://crates.io/crates/rav1d-safe) [](https://lib.rs/crates/rav1d-safe) [](https://docs.rs/rav1d-safe) [](https://github.com/imazen/rav1d-safe#license)

June 16, 2026 · View on GitHub

A safe Rust AV1 decoder. Forked from rav1d, with 160k lines of hand-written x86/ARM assembly replaced by safe Rust SIMD intrinsics.

578 commits since the fork (+90,813 / -6,958 lines across 1,194 files). 68k of those net new lines are safe SIMD in src/safe_simd/.

Quick Start

Add to your Cargo.toml:

[dependencies]
rav1d-safe = "0.5"

Decode an AV1 bitstream:

use rav1d_safe::{Decoder, Planes};

fn decode(obu_data: &[u8]) -> Result<(), Box<dyn std::error::Error>> {
    let mut decoder = Decoder::new()?;

    // Feed raw OBU data (not IVF/WebM containers)
    if let Some(frame) = decoder.decode(obu_data)? {
        println!("{}x{} @ {}bpc", frame.width(), frame.height(), frame.bit_depth());

        match frame.planes() {
            Planes::Depth8(planes) => {
                for row in planes.y().rows() {
                    // row is &[u8] — zero-copy, no allocation
                }
            }
            Planes::Depth16(planes) => {
                let px = planes.y().pixel(0, 0); // 10 or 12-bit value
            }
        }
    }

    // Drain any buffered frames
    for frame in decoder.flush()? {
        // ...
    }
    Ok(())
}

API Overview

The public API lives in src/managed.rs and is re-exported at the crate root, so every public type is reachable directly from rav1d_safe — no src::managed:: path needed. One canonical import covering the whole surface:

use rav1d_safe::{
    Decoder, Settings, CpuLevel, Error, Frame, Result,
    Planes,                       // enum; you match Planes::Depth8(_) / Planes::Depth16(_)
    Planes8, Planes16,            // the inner per-bit-depth plane sets the variants wrap
    PlaneView8, PlaneView16,      // zero-copy 2D plane views
    PixelLayout,                  // I400 / I420 / I422 / I444 (chroma subsampling)
    DecodeFrameType, InloopFilters,
    // HDR / color metadata:
    ColorInfo, ColorPrimaries, TransferCharacteristics, MatrixCoefficients,
    ColorRange, ContentLightLevel, MasteringDisplay,
    enabled_features,
};

Core types:

TypePurpose
DecoderDecodes AV1 OBU data into frames
FrameDecoded frame with metadata (cloneable, Arc-backed)
PlanesEnum with variants Planes::Depth8(Planes8) / Planes::Depth16(Planes16), dispatched by bit depth
Planes8 / Planes16Per-bit-depth plane set: .y() → luma view, .u() / .v()Option<view> (None for I400 monochrome)
PlaneView8 / PlaneView16Zero-copy 2D view: row(y), pixel(x, y), rows(), width(), height(), stride()
SettingsThread count, film grain, frame size limit, inloop filters, CPU level, etc.
CpuLevelSIMD dispatch level (Scalar, X86V2, X86V3, X86V4, Neon, Native)
ErrorEnum: InvalidData, OutOfMemory, NeedMoreData, InvalidSettings(&str), InitFailed, Other(String)

Naming note (reconciles the example above): Depth8 / Depth16 are the variants of the Planes enum — you write Planes::Depth8(..) in a match. Planes8 / Planes16 are the struct types those variants wrap, and are also re-exported at the crate root so you can name them in signatures. Both are correct; they refer to different things.

Metadata types: ColorInfo, ColorPrimaries, TransferCharacteristics, MatrixCoefficients, ColorRange, ContentLightLevel, MasteringDisplay, PixelLayout

Input Format

The decoder expects raw AV1 Open Bitstream Unit (OBU) data. If you have IVF or WebM containers, strip the container framing first and pass the OBU payload. See tests/ivf_parser.rs for an IVF parser example. For AVIF images, use zenavif-parse to extract the OBU data from the ISOBMFF container.

Output Format

decode() / flush() yield a Frame of planar YUV pixels — there is no built-in RGB conversion (do that downstream, e.g. with zenpixels-convert, or via zenavif for AVIF). Read the planes through the bit-depth-dispatched Planes enum:

let layout = frame.pixel_layout();        // PixelLayout: I400 | I420 | I422 | I444
let bpc    = frame.bit_depth();           // 8, 10, or 12

match frame.planes() {
    Planes::Depth8(p) => {                 // 8-bit content: planes are u8
        let y = p.y();                     // PlaneView8 (luma, always present)
        let u = p.u();                     // Option<PlaneView8> — None for I400 monochrome
        let v = p.v();                     // Option<PlaneView8>
        let (w, h, stride) = (y.width(), y.height(), y.stride());
        for row in y.rows() { /* row: &[u8], one luma scanline, zero-copy */ }
        let _ = (u, v, w, h, stride);
    }
    Planes::Depth16(p) => {                // 10/12-bit content: planes are u16
        let _y = p.y();                    // PlaneView16; pixel(x, y) -> u16
    }
}
  • Subsampling is reported by frame.pixel_layout(): I420 (4:2:0, chroma half-width & half-height), I422 (4:2:2, half-width), I444 (4:4:4, full-res chroma), I400 (monochrome — u()/v() return None).
  • Bit depth is frame.bit_depth() (8 / 10 / 12). Planes::Depth8 carries u8 samples; Planes::Depth16 carries u16 (the 10/12-bit value is in the low bits).
  • Each PlaneView{8,16} is a strided 2D view: row(y) borrows one scanline, rows() iterates them, pixel(x, y) reads one sample, and width()/height()/stride() give its geometry. All accessors are zero-copy borrows into the decoder's frame buffer (held alive by the Frame).

Threading

use rav1d_safe::{Decoder, Settings, CpuLevel};

// Single-threaded (default) — synchronous, deterministic
let decoder = Decoder::new()?;

// Multi-threaded — frame threading, better throughput
let decoder = Decoder::with_settings(Settings {
    threads: 0, // auto-detect core count
    ..Default::default()
})?;

// Constrained decoding — limit frame size and CPU features
let decoder = Decoder::with_settings(Settings {
    // Unit is TOTAL PIXELS (width * height), NOT bytes and NOT max-dimension.
    // 3840 * 2160 here means "reject any frame whose width*height exceeds ~8.3 MP" (a 4K cap).
    frame_size_limit: 3840 * 2160,
    cpu_level: CpuLevel::Native,   // use best available SIMD
    ..Default::default()
})?;

frame_size_limit — the pre-decode DoS guard (read this before decoding untrusted AV1)

This is the only resource bound applied before a frame is decoded, so it is the knob that matters for untrusted input. Verified against src/managed.rs and src/obu.rs:

  • Unit: the maximum width * height in pixels (total luma sample count). Not bytes, not the longest dimension. The check is literally width * height > frame_size_limit during OBU header parsing — a frame is rejected before any pixel buffer is allocated or decoded.
  • Default: 120_000_000 (120 megapixels) — chosen to admit ~108 MP phone photos. It is not unlimited by default, but 120 MP is large; set it lower for untrusted servers (e.g. 7680 * 4320 for an 8K cap, or 3840 * 2160 for 4K).
  • Disable: set frame_size_limit: 0 to turn the limit off entirely (no cap).
  • On rejection: the offending OBU fails parsing and decode() returns an Err (surfaced as Error::Other carrying the internal ERANGE range error — not Error::InvalidData). The decoder logs Frame size WxH exceeds limit N.
  • 32-bit hosts: on targets where usize is < 64 bits, the effective cap is additionally clamped to 8192 * 8192 regardless of the value you set (a non-zero larger value is silently reduced; 0/unlimited still means unlimited).

Caveat for servers: frame_size_limit bounds the declared frame dimensions only. It does not bound total decode time, memory of a large-but-legal frame, or the number of OBUs/frames in a stream. There is no per-decode time budget — see Cancellation below.

With threads >= 2 or threads == 0, the decoder uses tile threading to parallelize decode within each frame. decode() may return None for complete frames because processing is asynchronous — call it repeatedly or use flush() to drain.

Tile threading works under forbid(unsafe_code) without the unchecked feature. No special feature flags needed:

rav1d-safe = { version = "0.5", features = ["bitdepth_8", "bitdepth_16"] }

With 2 threads, expect ~2x speedup on photo decode. Frame threading (max_frame_delay > 1) still requires the unchecked feature.

HDR Metadata

if let Some(cll) = frame.content_light() {
    println!("MaxCLL: {} nits", cll.max_content_light_level);
}
if let Some(mdcv) = frame.mastering_display() {
    println!("Peak: {} nits", mdcv.max_luminance_nits());
}
let color = frame.color_info();
// color.primaries, color.transfer_characteristics, color.matrix_coefficients

Error Handling

Fallible operations return the crate's Result<T> alias, which is Result<T, whereat::At<Error>> — the Error is wrapped in [whereat]'s At<…>, recording the source location where the failure surfaced (handy for server logs). Unwrap the inner Error with err.error() (borrow) or err.decompose().0 (owned). Error variants: InvalidData, OutOfMemory, NeedMoreData, InitFailed, InvalidSettings(&str), Other(String). (From<Rav1dError> maps the internal EAGAIN → NeedMoreData, ENOMEM → OutOfMemory, EINVAL → InvalidData, and everything else — including the frame_size_limit ERANGE — to Other.)

Cancellation

There is no in-flight decode cancellation. The decoder exposes no stop/cancel/abort token — once decode() (or flush()) starts processing a frame, it runs to completion on the calling thread (plus any worker threads); you cannot interrupt a slow-but-valid decode from another thread, and dropping the Decoder joins its workers rather than aborting mid-frame.

For untrusted input this means the only pre-decode guard is frame_size_limit (see above), which bounds declared frame dimensions but not wall-clock decode time. If you need a hard time bound on a server, enforce it around the decoder — e.g. run decode() on a worker thread/task with your own timeout and abandon the result (the decode still finishes in the background; budget for that), or pre-screen stream size/frame count before feeding OBUs. A first-class cooperative-cancellation token (e.g. via the enough crate) is not currently implemented; this is tracked upstream.

Safety Model

The default build (forbid(unsafe_code) crate-wide) contains zero unsafe in the main crate. The only unsafe code lives in the rav1d-disjoint-mut workspace sub-crate, a provably sound RefCell-for-ranges abstraction with always-on bounds checking.

The SIMD path uses:

  • archmage for token-based target-feature dispatch (no manual #[target_feature])
  • safe_unaligned_simd for reference-based SIMD load/store (no raw pointers)
  • Value-type SIMD intrinsics, which are safe functions since Rust 1.93
  • Slice-based APIs throughout — no pointer arithmetic in SIMD code

Verify at runtime with rav1d_safe::enabled_features() — returns a comma-delimited list including the active safety level (e.g. "bitdepth_8, bitdepth_16, safety:forbid-unsafe").

What's Been Ported

The default build compiles under forbid(unsafe_code) in the main crate. All SIMD work lives in src/safe_simd/ (59k lines of safe Rust replacing 233k lines of hand-written assembly across x86 and ARM).

Ported: All DSP Kernels (AVX2 + NEON)

Every DSP kernel family has a safe Rust SIMD implementation that compiles under forbid(unsafe_code):

Modulex86 ASM replacedARM ASM replacedSafe Rust
mc (motion compensation)3 files (SSE/AVX2/AVX-512) x2 bitdepths4 files (32+64-bit) x2 bitdepths + SVE/dotprodmc.rs + mc_arm.rs
itx (inverse transforms)3 files x2 bitdepths2 files x2 bitdepthsitx.rs + itx_arm.rs
ipred (intra prediction)3 files x2 bitdepths2 files x2 bitdepthsipred.rs + ipred_arm.rs
cdef (directional enhancement)3 files x2 bitdepths2+tmpl files x2 bitdepthscdef.rs + cdef_arm.rs
loopfilter3 files x2 bitdepths2 files x2 bitdepthsloopfilter.rs + loopfilter_arm.rs
looprestoration (Wiener + SGR)3 files x2 bitdepths2+common+tmpl files x2 bitdepthslooprestoration.rs + looprestoration_arm.rs
filmgrain3+common files x2 bitdepths2 files x2 bitdepthsfilmgrain.rs + filmgrain_arm.rs
pal (palette)1 file(none — ARM uses scalar)pal.rs
refmvs (reference MVs)1 file2 files (32+64-bit)refmvs.rs + refmvs_arm.rs
msac (entropy decoder)1 file (shared)1 file (shared)inline in msac.rs
cpuid1 file (55 lines)replaced by std::arch detection in cpu.rs

msac uses branchless scalar for adapt4/adapt8 and a serial loop with early exit for adapt16. When the unchecked feature is enabled on x86_64, adapt4/adapt8/hi_tok switch to inlined SSE2 intrinsics via the sse2!() macro pattern (no function call overhead). The bool functions (bool_adapt, bool_equi) stay scalar — they have no data parallelism to exploit.

Not Ported (With Rationale)

Scaled MC (put_8tap_scaled, prep_8tap_scaled, put_bilin_scaled, prep_bilin_scaled) — These functions use per-pixel variable step sizes with per-pixel filter selection, making them fundamentally different from fixed-block MC. The ASM versions are heavily register-scheduled for this pattern. Falls back to scalar Rust. ~2% of profile on inter-frame content.

SSE-only paths — 14 files, ~52k lines. The safe SIMD dispatch jumps straight to AVX2 when available. On pre-AVX2 hardware (pre-Haswell, 2013), the decoder falls back to scalar Rust rather than SSE intrinsics. SSE-only x86 hardware is rare enough that maintaining a second intrinsics tier isn't worth the code.

ARM SVE2, dotprod, i8mm extensionsmc_dotprod.S (1,880 lines) and mc16_sve.S (1,649 lines) are optional fast paths for newer ARM cores. The safe SIMD covers baseline NEON; these extension paths fall back to the NEON implementation.

Remaining AVX-512 paths — Some AVX-512 paths have been ported (itx, mc, ipred, looprestoration Wiener), but others remain unported. Falls back to AVX2 where not implemented.

ASM infrastructure filesx86inc.asm (1,983 lines), asm.S, util.S, *_tmpl.S, *_common.S are macro libraries and constants that only exist to support the raw assembly. No independent functionality to port.

Performance

All benchmarks: x86_64 (Zen 4, AVX2), single-threaded, Rust 1.93+, fat LTO. Run with just profile (500 iterations for IVF, 20 iterations for AVIF).

Real photographs (AVIF decode, single image)

Single still images at web-typical quality (YUV420, q60). Source: Google-native 8K photo, downscaled with ImageMagick. These numbers reflect real-world AVIF decode performance where SIMD kernels dominate.

ResolutionASMSafe (checked)Safe (unchecked)Safe vs ASM
4K (3840x2561)120.7 ms187.8 ms179.5 ms1.56x
8K (8192x5464)714.1 ms1103.2 ms1066.9 ms1.54x

(Measured at v0.5.6. The 0.5.6 SIMD work — i16-packed pmaddwd transform row+col passes and YMM-widened loopfilter — brought the 4K checked ratio from 2.0× down to 1.56×.)

Small test vectors (IVF, multi-frame decode)

dav1d-test-data allintra 352x288 (39 frames). Entropy-heavy bitstream where the serial msac decoder dominates, which compresses the ratio compared to pixel-heavy workloads.

Buildms/iterms/framevs ASM
ASM104.32.671.0x
Safe (checked)158.64.071.52x
Safe (unchecked)153.33.931.47x
Partial ASM141.23.621.35x

Where the gap comes from

The safe build is ~1.56x slower on 4K real images and ~1.52x on entropy-heavy vectors. The gap breaks down:

  • Entropy decoder (msac): ~45% of decode time. Serial dependency chain — the core symbol decode loop can't be parallelized; ~95% of decode_coefs is irreducible algorithmic cost (only ~3-5% is Rust-specific bounds-check/index overhead). The partial_asm feature uses hand-tuned ASM for msac and loopfilter, bringing 4K photo decodes to 1.25x vs full ASM.
  • Calling conventions: The ASM kernels use custom register allocation across function boundaries. Rust's ABI reloads registers at each call site.
  • Scaled MC: Falls back to scalar Rust (~2% of inter-frame content). The ASM version uses per-pixel variable-step register scheduling that doesn't map cleanly to safe intrinsics.
  • Bounds checking: The unchecked feature skips DisjointMut borrow tracking. This saves ~5% on photos, confirming the tracking overhead is modest.

Reproduce locally

just generate-bench-avif  # create 4K/8K AVIF test images (requires avifdec + avifenc)
just profile              # all four modes side-by-side (ASM, partial ASM, checked, unchecked)
just profile-quick        # same with fewer iterations

Conformance

Tested against the dav1d-test-data suite. MD5 hashes verified at all CPU dispatch levels (scalar, SSE4.2, AVX2, native).

784 of 803 test vectors pass across all levels.

19 vectors are not exercised by the test harness:

CategoryCountReason
sframe1Requires S-frame support
svc (operating points)6Always decodes at default operating point
argon (vq_suite)12Various decode modes and operating point selection tests

Run conformance tests with cargo test --release --test decode_cpu_levels.

Building

Requires Rust 1.93+ (stable). Install via rustup.rs.

# Default safe-SIMD build (recommended)
cargo build --release

# With original hand-written assembly (for benchmarking)
cargo build --features asm --release

# Run tests
cargo test --release

Feature Flags

FeatureDefaultDescription
bitdepth_8on8-bit pixel support
bitdepth_16on10/12-bit pixel support
uncheckedoffSkip DisjointMut borrow tracking; enables frame threading and SSE2 msac on x86_64
partial_asmoffASM for entropy decoding (msac) and loopfilter only; safe SIMD everything else. Implies unchecked
c-ffioffC API entry points (dav1d_* symbols). Implies unchecked
asmoffFull hand-written assembly. Implies c-ffi

Safety chain: default (forbid(unsafe_code), tile threading) -> unchecked (frame threading) -> c-ffi -> asm. Each level relaxes the safety constraint.

Cross-Compilation

# aarch64
RUSTFLAGS="-C linker=aarch64-linux-gnu-gcc" \
  cargo build --target aarch64-unknown-linux-gnu --release

# Verify aarch64 NEON compiles
cargo check --target aarch64-unknown-linux-gnu

Supported targets: x86_64-unknown-linux-gnu, aarch64-unknown-linux-gnu, i686-unknown-linux-gnu, armv7-unknown-linux-gnueabihf, riscv64gc-unknown-linux-gnu.

Image tech I maintain

State of the art codecs*zenjpeg · zenpng · zenwebp · zengif · zenavif (rav1d-safe · zenrav1e · zenavif-parse · zenavif-serialize) · zenjxl (jxl-encoder · zenjxl-decoder) · zentiff · zenbitmaps · heic · zenraw · zenpdf · ultrahdr · mozjpeg-rs · webpx
Compressionzenflate · zenzop
Processingzenresize · zenfilters · zenquant · zenblend
Metricszensim · fast-ssim2 · butteraugli · resamplescope-rs · codec-eval · codec-corpus
Pixel types & colorzenpixels · zenpixels-convert · linear-srgb · garb
Pipelinezenpipe · zencodec · zencodecs · zenlayout · zennode
ImageResizerImageResizer (C#) — 24M+ NuGet downloads across all packages
ImageflowImage optimization engine (Rust) — .NET · node · go — 9M+ NuGet downloads across all packages
Imageflow ServerThe fast, safe image server (Rust+C#) — 552K+ NuGet downloads, deployed by Fortune 500s and major brands

* as of 2026

General Rust awesomeness

archmage · magetypes · enough · whereat · zenbench · cargo-copter

And other projects · GitHub @imazen · GitHub @lilith · lib.rs/~lilith · NuGet (over 30 million downloads / 87 packages)

License

Dual-licensed: AGPL-3.0 or commercial.

I've maintained and developed open-source image server software — and the 40+ library ecosystem it depends on — full-time since 2011. Fifteen years of continual maintenance, backwards compatibility, support, and the (very rare) security patch. That kind of stability requires sustainable funding, and dual-licensing is how we make it work without venture capital or rug-pulls. Support sustainable and secure software; swap patch tuesday for patch leap-year.

Our open-source products

Your options:

  • Startup license — $1 if your company has under $1M revenue and fewer than 5 employees. Get a key →
  • Commercial subscription — Governed by the Imazen Site-wide Subscription License v1.1 or later. Apache 2.0-like terms, no source-sharing requirement. Sliding scale by company size. Pricing & 60-day free trial →
  • AGPL v3 — Free and open. Share your source if you distribute.

See LICENSE-COMMERCIAL for details.

Upstream code from memorysafety/rav1d is licensed under BSD-2-Clause. Our additions and improvements are dual-licensed (AGPL-3.0 or commercial) as above.

Upstream Contribution

We are willing to release our improvements under the original BSD-2-Clause license if upstream takes over maintenance of those improvements. We'd rather contribute back than maintain a parallel codebase. Open an issue or reach out.

Acknowledgments

Built on the work of the dav1d team (VideoLAN) and the rav1d team (ISRG/Prossimo). The original C and assembly implementations are exceptional — this fork demonstrates that safe Rust SIMD can get within 2x of hand-written assembly while eliminating entire classes of memory safety bugs.