Inference Performance
September 11, 2026 · View on GitHub
This report provides practical runtime and GPU-memory references for running 4DAnyone on common GPUs.
Turbo vs. Base
Generation Quality
The 4-step Turbo model achieves quality comparable to a 50-step Base reference:
| Model | Denoising steps | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|
| 4DAnyone-Turbo | 4 | 25.48 | 0.895 | 0.072 |
| 4DAnyone-Base | 50 | 25.69 | 0.899 | 0.073 |
Denoising Speedup
For the 6-view example, Turbo uses 4 denoising steps and Base uses 24. The values show Turbo's speedup over Base on the same GPU and attention backend.
| GPU | SDPA | FlashAttention-3 | SageAttention |
|---|---|---|---|
| H200 | 5.49× | 5.35× | 5.12× |
| H20-3E | 5.83× | 5.78× | 5.73× |
| RTX 5090 D v2 | 5.77× | — | 5.70× |
| RTX 4090 | 5.85× | — | 5.78× |
| RTX 3090 | 5.85× | — | 5.80× |
| RTX 5880 Ada | 5.79× | — | 5.80× |
| RTX A6000 | 5.77× | — | 5.75× |
4DAnyone-Turbo Performance
The performance tables below use the following terms:
SDPA,FlashAttention-3, andSageAttention: denoising time for each attention backend.Peak CUDA allocated: maximum live tensor memory during generation.GPU memory: total memory capacity per GPU.—: backend or workload unavailable on that GPU.
Note
SageAttention results use the official sageattn entry point, which automatically dispatches an architecture-specific kernel for each GPU.
6-View Full Orbit
The 6-view full-orbit benchmark uses the following command:
python inference.py \
--video_path "data/source/pexels/2785536-uhd_2160_3840_25fps.mp4" \
--views_per_layer 6
| GPU | SDPA (min) | FlashAttention-3 (min) | SageAttention (min) | Peak CUDA allocated (GB) | GPU memory (GB) |
|---|---|---|---|---|---|
| H200 | 0.97 | 0.76 | 0.80 | 24.06 | 139.81 |
| H20-3E | 3.21 | 2.54 | 2.09 | 24.06 | 139.81 |
| RTX 5090 D v2 | 2.06 | — | 1.61 | 22.11 | 23.40 |
| RTX 4090 | 2.71 | — | 2.13 | 22.18 | 23.52 |
| RTX 3090 | 5.48 | — | 4.58 | 21.55 | 23.56 |
| RTX 5880 Ada | 3.47 | — | 2.99 | 24.06 | 47.37 |
| RTX A6000 | 3.78 | — | 3.69 | 24.06 | 47.53 |
24-View Full Orbit
The 24-view full-orbit benchmark runs 4DAnyone-Turbo on a single GPU:
python inference.py \
--video_path "data/source/pexels/2785536-uhd_2160_3840_25fps.mp4" \
--views_per_layer 24
For 24-view generation, the pipeline first denoises six RCP views to use as references and then denoises the 24 target views. The table reports the total time for both stages.
| GPU | SDPA (min) | FlashAttention-3 (min) | SageAttention (min) | Peak CUDA allocated (GB) | GPU memory (GB) |
|---|---|---|---|---|---|
| H200 | 5.11 | 3.90 | 3.90 | 24.06 | 139.81 |
| H20-3E | 17.66 | 13.86 | 11.24 | 24.06 | 139.81 |
| RTX 5090 D v2 | 11.22 | — | — | 21.94 | 23.40 |
| RTX 4090 | 16.20 | — | — | 22.18 | 23.52 |
| RTX 3090 | 31.98 | — | — | 21.94 | 23.56 |
| RTX 5880 Ada | 18.84 | — | 16.14 | 24.06 | 47.37 |
| RTX A6000 | 20.55 | — | 19.98 | 24.06 | 47.53 |
4DAnyone-Base Performance
The 6-view 4DAnyone-Base benchmark uses 24 denoising steps:
python inference.py \
--video_path "data/source/pexels/2785536-uhd_2160_3840_25fps.mp4" \
--views_per_layer 6 \
--enable_turbo=False
| GPU | SDPA (min) | FlashAttention-3 (min) | SageAttention (min) | Peak CUDA allocated (GB) | GPU memory (GB) |
|---|---|---|---|---|---|
| H200 | 5.33 | 4.05 | 4.09 | 24.06 | 139.81 |
| H20-3E | 18.68 | 14.70 | 11.98 | 24.06 | 139.81 |
| RTX 5090 D v2 | 11.87 | — | 9.17 | 22.11 | 23.40 |
| RTX 4090 | 15.86 | — | 12.29 | 22.18 | 23.52 |
| RTX 3090 | 32.03 | — | 26.52 | 21.55 | 23.56 |
| RTX 5880 Ada | 20.10 | — | 17.36 | 24.06 | 47.37 |
| RTX A6000 | 21.82 | — | 21.20 | 24.06 | 47.53 |
Preprocessing Performance
GVHMR and skeleton extraction run before denoising and are shared by both models. These 6-view timings can vary with CPU and storage performance.
GVHMR: tracking, ViTPose, image-feature extraction, and motion prediction.Skeleton extraction: target-view skeleton-condition rendering.
| GPU | GVHMR (min) | Skeleton extraction (min) |
|---|---|---|
| H200 | 1.20 | 1.25 |
| H20-3E | 1.08 | 1.27 |
| RTX 5090 D v2 | 0.72 | 0.82 |
| RTX 4090 | 0.72 | 0.73 |
| RTX 3090 | 1.66 | 1.72 |
| RTX 5880 Ada | 0.98 | 1.00 |
| RTX A6000 | 1.57 | 2.12 |