WRBench

July 12, 2026 · View on GitHub

Official toolkit and release artifacts for WRBench, a diagnostic benchmark for testing whether video world models preserve world state across camera motion.

Paper Artifacts Leaderboard

WRBench overview: Natural-25 prompts, controlled camera paths, model-native generation, and D1-D6 evaluation

WRBench provides three composable pieces:

  • Compile: turn a camera instruction into a model-native payload and auditable sidecars. No GPU is required.
  • Generate (optional): execute that payload through a configured model backend.
  • Evaluate: score existing videos with the separable D1–D6 diagnostic contract.

The repository also bundles Natural-25 prompts and first frames, the frozen 23-model paper table, and the versioned paper reproducibility contract.

Install

Python 3.10 or newer is required.

git clone https://github.com/JinPLu/WRBench.git
cd WRBench
pip install -e .

Install optional prompt, first-frame, and development dependencies with pip install -e ".[all]". Model weights and GPU frameworks are intentionally not package dependencies.

1. Inspect and compile a camera request

List the available models and camera presets:

wrbench models
wrbench presets

Compile a go-and-return yaw request with a bundled Natural-25 first frame:

IMAGE="$(python - <<'PY'
from wrbench.datasets import natural25_first_frame_path
print(natural25_first_frame_path("bedroom_cat_bed_jump"))
PY
)"

wrbench generate \
  --model wan22-fun-5b-cam \
  --camera preset:yaw_LR \
  --image "$IMAGE" \
  --prompt "A house cat waits on the bedroom floor, facing the bed." \
  --out out.mp4

wrbench generate is compile-only by default: it writes the target trajectory, native payload, and metadata without loading a model. Use wrbench actions --camera "yaw:left:37@49" --fps 16 to inspect a custom camera script. See the camera-control reference and more compile examples.

2. Evaluate an existing video set

First inspect the metric contract; this requires no model configuration:

wrbench eval contract

For D1–D6 scoring, copy wrbench.runtime.example.json to wrbench.runtime.json, fill in the scorer paths, and run:

wrbench eval --runtime-config wrbench.runtime.json run \
  --manifest videos.json \
  --out-dir eval_out \
  --scorer-profile wrbench_default \
  --sidecar-profile-gate main

Scorer requirements and granular commands are documented in the evaluation guide. D1–D6 remain separate diagnostics; WRBench does not collapse them into one quality score.

3. Generate videos (optional)

Real generation is an explicit opt-in because each upstream model has its own environment and weights. Copy wrbench.runtime.example.json, configure a supported backend, then add --runtime-config /path/to/wrbench.runtime.json --no-dry-run to the compile command above. See the backend guide.

4. Add a model

A generation-capable model needs one registry record, one adapter module, and one adapter import. Follow Adding a model, then validate it with:

wrbench doctor --model <model-key>

Evaluation dimensions

DimensionWhat it measuresApplies when
D1-CamPrecRequested-camera trajectory precisionAn explicit target trajectory exists
D1-CamAlignPrompt-camera intent alignmentCamera control is prompt/API based
D2Visual integrityAll videos
D3Visible spatial consistencyAll videos
D4Visible state consistencyAll videos
D5Re-observation spatial consistencyThe relevant content returns into view
D6Re-observation event-state consistencyThe relevant content returns into view

Rows without re-observation support are excluded from D5/D6 rather than counted as failures. The exact contract lives in src/wrbench/eval/contract/.

Published artifacts

ArtifactSource
PaperarXiv:2606.20545
All releasesHugging Face collection
Natural-25dataset · bundled data
Frozen paper prompts, first frames, camera scopes, and source mappingspaper_main_20260608
Resultsrolling dataset · bundled frozen 23-model tables
Videos and per-video scoresdataset
Human annotationsdataset
LeaderboardSpace

The 9,600-row paper table and its provenance are frozen. Corrections or newly verified assets belong to a separately versioned rolling release; they do not silently rewrite the paper aggregates. Exact prompt, first-frame, TV2V-source, and HyDRA policies are owned by the linked release contract rather than repeated here.

Documentation

Citation

@article{wrbench2026,
  title   = {Current World Models Lack a Persistent State Core},
  author  = {Jinpeng Lu and Dexu Zhu and Haoyuan Shi and Linghan Cai and Guo Tang and Yinda Chen and Jie Cao and Duyu Tang and Yi Zhang and Yong Dai and Xiaozhu Ju},
  journal = {arXiv preprint arXiv:2606.20545},
  year    = {2026},
  url     = {https://arxiv.org/abs/2606.20545},
}

License

Apache 2.0 — see LICENSE.