Stable Diffusion 1.5

July 7, 2025 · View on GitHub

Train Benchmark

Fig

Figure_1

ModelStagePaddle training speed(ips)ContrastPytorch training speed(ips)Paddle GPU memory uage(G)
LLaVA1.6 7BPretrain82+26%6519/22
SFT52+6%4933/49
LoRA56+14%4916/17
LLaVA1.6 13BPretrain52+18%4433/36
SFT24+4%2350/68
LoRA36+5%3429/30
Qwen2VL 2BSFT33+43%23-
Qwen2VL 7BSFT13+18%11-
Stable Diffusion 1.5Pretrain560-12%63828/34
LoRA200+6%18730/34
Stable Diffusion 3SFT (Dreambooth)34034  -
LoRA66-0.01%67-

Notes:

  • All models were tested on the H800 (8 * 80G) platform
  • For GPU menory usage, the table shows max_memory_allocated/max_memory_reserved
  • Please see below for the testing configuration details.
See
SoftwareVersion
CUDA12.3
CUDNN9.0
PaddlePaddle3.0beta2
PaddleNLP3.0beta3
Pytorch2.5