Speed:

June 8, 2026 · View on GitHub

Measured on a 3090 at 1024x1024, 26 steps with Flux2 Klein Base 9B.

FormatSpeed (s/it) ↓Relative Speedup
bf162.071.00×
bf16 compile2.240.92×
fp82.061.00×
int81.641.26×
int8 compile ★1.041.99×
gguf8_0 compile2.031.02×

3090, Qwen Image 2512.

FormatSpeed (s/it) ↓
Nunchaku INT4 Best Quality1.21
Nunchaku INT4 with R128 Lora1.36
INT8 ConvRot compile1.26
INT8 Row compile ★1.18
INT8 R128 LoraNo slowdown, except if dynamic.

I would also like to point out that we beat Nunchaku INT4 on every quality measurement in the Quality Metrics

Additionally, the quality of loras applied with this nunchaku lora node appears to be degraded.

Klein 9B, Measured on an 8gb 5060, same settings as the 3090 run:

FormatSpeed (s/it) ↓Relative Speedup
fp83.041.00×
fp8 fast3.001.00×
fp8 compilecouldn't get to work??×
int82.531.20×
int8 compile ★2.251.35×

8gb RTX 5060, Anima, Comfy version from 2026-05-02, Pytorch 2.11+CU13.0, latest kitchen triton and everything else

FormatSpeed (it/s) ↑
bf160.78
INT8 ConvRot1.12
INT8 Row1.24
INT8 ConvRot Compile1.47
MXFP80.89
MXFP8 --fast0.93
MXFP8 + CompileStill failing.

Finally have gotten compile with --fast to work with mxfp8, PyTorch 2.13.0.dev20260511+cu132, RTX5060 same as before.

Quality results for this run, can be found here: Anima Results

FormatSpeed (it/s) ↑
MXFP8 --fast + Compile1.37it
INT8 ConvRot + Compile1.47it