SimpleTuner 💹

August 5, 2026 · View on GitHub

ℹ️ No data is sent to any third parties except through opt-in flag report_to, push_to_hub, or webhooks which must be manually configured.

SimpleTuner is geared towards simplicity, with a focus on making the code easily understood. This codebase serves as a shared academic exercise, and contributions are welcome.

If you'd like to join our community, we can be found on Discord via Terminus Research Group. If you have any questions, please feel free to reach out to us there.

image

Table of Contents

Design Philosophy

  • Simplicity: Aiming to have good default settings for most use cases, so less tinkering is required.
  • Versatility: Designed to handle a wide range of image quantities - from small datasets to extensive collections.
  • Cutting-Edge Features: Only incorporates features that have proven efficacy, avoiding the addition of untested options.

Tutorial

Please fully explore this README before embarking on the new web UI tutorial or the class command-line tutorial, as this document contains vital information that you might need to know first.

For a manually configured quick start without reading the full documentation or using any web interfaces, you can use the Quick Start guide.

For memory-constrained systems, see the DeepSpeed document which explains how to use 🤗Accelerate to configure Microsoft's DeepSpeed for optimiser state offload. For DTensor-based sharding and context parallelism, read the FSDP2 guide which covers the new FullyShardedDataParallel v2 workflow inside SimpleTuner.

For multi-node distributed training, this guide will help tweak the configurations from the INSTALL and Quickstart guides to be suitable for multi-node training, and optimising for image datasets numbering in the billions of samples.


Features

SimpleTuner provides comprehensive training support across multiple diffusion model architectures with consistent feature availability:

Core Training Features

  • User-friendly web UI - Manage your entire training lifecycle through a sleek dashboard
  • Multi-modal training - Unified pipeline for Image, Video, and Audio generative models
  • Multi-GPU training - Distributed training across multiple GPUs with automatic optimization
  • Advanced caching - Image, video, audio, and caption embeddings cached to disk for faster training
  • CaptionFlow integration - Generate dataset captions from local GPUs through the Web UI job queue using bghira/CaptionFlow; see the CaptionFlow integration guide
  • Aspect bucketing - Support for varied image/video sizes and aspect ratios
  • Concept sliders - Slider-friendly targeting for LoRA/LyCORIS/full (via LyCORIS full) with positive/negative/neutral sampling and per-prompt strength; see Slider LoRA guide
  • Memory optimization - Most models trainable on 24G GPU, many on 16G with optimizations
  • DeepSpeed & FSDP2 integration - Train large models on smaller GPUs with optim/grad/parameter sharding, context parallel attention, gradient checkpointing, and optimizer state offload
  • S3 training - Train directly from cloud storage (Cloudflare R2, Wasabi S3)
  • EMA support - Exponential moving average weights for improved stability and quality
  • Custom experiment trackers - Drop an accelerate.GeneralTracker into simpletuner/custom-trackers and use --report_to=custom-tracker --custom_tracker=<name>

Multi-User & Enterprise Features

SimpleTuner includes a complete multi-user training platform with enterprise-grade features—free and open source, forever.

  • Worker Orchestration - Register distributed GPU workers that auto-connect to a central panel and receive job dispatch via SSE; supports ephemeral (cloud-launched) and persistent (always-on) workers; see Worker Orchestration Guide
  • SSO Integration - Authenticate with LDAP/Active Directory or OIDC providers (Okta, Azure AD, Keycloak, Google); see External Auth Guide
  • Role-Based Access Control - Four default roles (Viewer, Researcher, Lead, Admin) with 17+ granular permissions; define resource rules with glob patterns to restrict configs, hardware, or providers per team
  • Organizations & Teams - Hierarchical multi-tenant structure with ceiling-based quotas; org limits enforce absolute maximums, team limits operate within org bounds
  • Quotas & Spending Limits - Enforce cost ceilings (daily/monthly), job concurrency limits, and submission rate limits at org, team, or user scope; actions include block, warn, or require approval
  • Job Queue with Priorities - Five priority levels (Low → Critical) with fair-share scheduling across teams, starvation prevention for long-waiting jobs, and admin priority overrides
  • Approval Workflows - Configurable rules trigger approval for jobs exceeding cost thresholds, first-time users, or specific hardware requests; approve via UI, API, or email reply
  • Email Notifications - SMTP/IMAP integration for job status, approval requests, quota warnings, and completion alerts
  • API Keys & Scoped Permissions - Generate API keys with expiration and limited scope for CI/CD pipelines
  • Audit Logging - Track all user actions with chain verification for compliance; see Audit Guide

For deployment details, see the Enterprise Guide.

Model Architecture Support

SimpleTuner supports the following model families. Detailed training feature support lives in the Quickstart Guide.

ModelParametersLicenseCommercial use
ACE-Step3.5BApache-2.0Yes
AnimaNot specifiedCircleStone Labs Non-Commercial License v1.2No (model); outputs allowed
Auraflow6BApache-2.0Yes
Boogu-ImageNot specifiedApache-2.0Yes
Chroma 18.9BApache-2.0Yes
Cosmos22B-14BNVIDIA Open Model LicenseYes
Cosmos316B-65BOpenMDW-1.1Yes
DeepFloyd IF0.4B-4.3B stagesDeepFloyd IF LicenseAbandonware
ERNIE-ImageNot specifiedApache-2.0Yes
Flux.18B-12BApache-2.0 (schnell); FLUX.1 [dev] Non-Commercial License (dev/Kontext)Mixed by checkpoint
Flux.24B-32BApache-2.0 (klein 4B); FLUX Non-Commercial License (dev/klein 9B)Mixed by checkpoint
HeartMuLa3BNot specified in SimpleTunerSee upstream terms
HiDream17B (8.5B MoE)MITYes
Hunyuan Video8.3BAGPL-3.0Yes (copyleft)
Ideogram 49BIdeogram 4 Non-CommercialNo
Kandinsky 5.0 Image6B (lite)MITYes
Kandinsky 5.0 Video2B lite, 19B proMITYes
Kwai Kolors2.7BApache-2.0Abandonware
Krea2Not specifiedKrea 2 Community LicenseYes (under $1M revenue; safeguards required)
LongCat Image6BApache-2.0Yes
LongCat Video13.6BMITYes
LTX Video~2.5BApache-2.0Yes
LTX Video 219BApache-2.0Yes
Lumina22BApache-2.0Yes
Mage-Flow4BMITYes
MiniMax H333BMiniMax H3 Community LicenseConditions apply (territory exclusions; authorization required in US/EU/UK/KR)
OmniGen3.8BMITYes
PixArt Sigma0.6B-0.9BOpenRAIL++Yes (restricted)
Qwen Image20BApache-2.0Yes
Sana0.6B-4.8BApache-2.0Yes
Sana Video2BApache-2.0Yes
SD 1.x/2.x (Legacy)0.9BOpenRAIL++Yes (restricted)
Stable Diffusion 32B-8BStability AI Community LicenseYes (under $1M revenue)
Stable Diffusion XL3.5BCreativeML OpenRAIL-MYes (restricted)
Stable Cascade (Stage C)1B, 3.6B priorNot specified in SimpleTunerAbandonware
Wan Video1.3B-14BApache-2.0Yes
Wan S2V14BApache-2.0Yes
Z-Image6BApache-2.0Yes
Z-Image Omni6BApache-2.0Yes
ZLab I13BMITYes

License values are taken from SimpleTuner model helpers when available and upstream model cards/licenses for entries that were previously unspecified. Not specified in SimpleTuner means the helper does not name a license and no upstream term is summarized here; check the upstream model card before use.

Advanced Training Techniques

  • TREAD - Token-wise dropout for transformer models, including Kontext training
  • Masked loss training - Superior convergence with segmentation/depth guidance
  • Prior regularization - Enhanced training stability for character consistency
  • Gradient checkpointing - Configurable intervals for memory/speed optimization
  • Loss functions - L2, Huber, Smooth L1 with scheduling support
  • SNR weighting - Min-SNR gamma weighting for improved training dynamics
  • Group offloading - Diffusers v0.33+ module-group CPU/disk staging with optional CUDA streams
  • Validation adapter sweeps - Temporarily attach LoRA adapters (single or JSON presets) during validation to measure adapter-only or comparison renders without touching the training loop
  • External validation hooks - Swap the built-in validation pipeline or post-upload steps for your own scripts, so you can run checks on another GPU or forward artifacts to any cloud provider of your choice (details)
  • AnyFlow distillation - FlowMap interval conditioning for flow-matching models with online teacher targets (guide)
  • CREPA regularization - Cross-frame representation alignment for video DiTs (guide)
  • LoRA I/O formats - Load/save PEFT LoRAs in standard Diffusers layout or ComfyUI-style diffusion_model.* keys (Flux/Flux2/Lumina2/Z-Image auto-detect ComfyUI inputs)

Model-Specific Features

  • Flux Kontext - Edit conditioning and image-to-image training for Flux models
  • Reference-input training - Existing paired reference/edit/I2V paths for Flux Kontext, Flux.2, LTX Video 2, Qwen Edit, LongCat edit/I2V, Boogu edit, Hunyuan I2V, and Kandinsky I2I/I2V
  • PixArt two-stage - eDiff training pipeline support for PixArt Sigma
  • Flow matching models - Advanced scheduling with beta/uniform distributions
  • HiDream MoE - Mixture of Experts gate loss augmentation
  • T5 masked training - Enhanced fine details for Flux and compatible models
  • QKV fusion - Memory and speed optimizations (Flux, Lumina2)
  • TREAD integration - Selective token routing for most models
  • Wan 2.x I2V - High/low stage presets plus a 2.1 time-embedding fallback (see Wan quickstart)
  • Classifier-free guidance - Optional CFG reintroduction for distilled models

Quickstart Guides

Detailed quickstart guides are available for all supported models:


Hardware Requirements

General Requirements

  • NVIDIA: RTX 3080+ recommended (tested up to H200)
  • AMD: 7900 XTX 24GB and MI300X verified (higher memory usage vs NVIDIA)
  • Apple: M3 Max+ with 24GB+ unified memory for LoRA training

Memory Guidelines by Model Size

  • Large models (12B+): A100-80G for full-rank, 24G+ for LoRA/Lycoris
  • Medium models (2B-8B): 16G+ for LoRA, 40G+ for full-rank training
  • Small models (<2B): 12G+ sufficient for most training types

Note: Quantization (int8/fp8/nf4) significantly reduces memory requirements. See individual quickstart guides for model-specific requirements.

Setup

SimpleTuner can be installed via pip for most users:

# Base installation (CPU-only PyTorch)
pip install simpletuner

# CUDA users (NVIDIA GPUs)
pip install 'simpletuner[cuda]'

# CUDA 13 / Blackwell users (NVIDIA B-series GPUs)
pip install 'simpletuner[cuda13]' --extra-index-url https://download.pytorch.org/whl/cu130

# CUDA 13 with TransformerEngine FP8 support
pip install 'simpletuner[cuda13-transformerengine]' --extra-index-url https://download.pytorch.org/whl/cu130

# ROCm users (AMD GPUs)
pip install 'simpletuner[rocm]' --extra-index-url https://download.pytorch.org/whl/rocm7.1

# Apple Silicon users (M1/M2/M3/M4 Macs)
pip install 'simpletuner[apple]'

For manual installation or development setup, see the installation documentation.

Troubleshooting

Enable debug logs for a more detailed insight by adding export SIMPLETUNER_LOG_LEVEL=DEBUG to your environment (config/config.env) file.

For performance analysis of the training loop, setting SIMPLETUNER_TRAINING_LOOP_LOG_LEVEL=DEBUG will have timestamps that highlight any issues in your configuration.

For a comprehensive list of options available, consult this documentation.