README.md

July 2, 2026 ยท View on GitHub

UniEdit-Flow
Unleashing Inversion and Editing in the Era of Flow Models

Guanlong Jiao1,3, Biqing Huang1, Kuan-Chieh Wang2, Renjie Liao3

1Tsinghua University, 2Snap Inc., 3The University of British Columbia

arXiv

TL;DR: A highly accurate and efficient, model-agnostic, training and tuning-free sampling strategy for inversion and editing tasks. Support text-driven image ๐ŸŽจ (FLUX, Stable Diffusion 3, Stable Diffusion XL, etc.) and video ๐ŸŽฅ (Wan, flow-based video generation model) editing.

๐Ÿ’œ Overview

In this work, we introduce a predictor-corrector-based framework for inversion and editing in flow models. First, we propose Uni-Inv, an effective inversion method designed for accurate reconstruction. Building on this, we extend the concept of delayed injection to flow models and introduce Uni-Edit, a region-aware, robust image editing approach. Our methodology is tuning-free, model-agnostic, efficient, and effective, enabling diverse edits while ensuring strong preservation of edit-irrelevant regions.

โœจ Feature: Text-driven Image / Video Editing

More results can be found in our project page.

๐ŸŽจ Image Editing

Editing Prompt Source Image FLUX Stable Diffusion 3 Stable Diffusion XL
A long short haired cat with blue eyes looking up at something.
Two origami birds sitting on a branch.
A clown in pixel art style with colorful hair.

๐ŸŽฅ Video Editing

Editing Prompt Source Video Wan + Uni-Edit
A young rider wearing full protective gear, including a black helmet and motocross-style outfit, is navigating a BMX bike motorcycle over a series of sandy dirt bumps on a track enclosed by a fence...
A koala cat with thick gray fur is captured mid-motion as it reaches out with its front paws to climb or move between tree branches, surrounded by lush green leaves and dappled sunlight in a forested area.

๐Ÿ‘จโ€๐Ÿ’ป Implementation

Here we provide two implementation options:

  • Implementation by diffusers: Support FLUX (e.g., black-forest-labs/FLUX.1-dev), Stable Diffusion 3 (e.g., stabilityai/stable-diffusion-3-medium), Stable Diffusion XL (e.g., SG161222/RealVisXL_V4.0), etc., for text-driven image editing tasks. As well, support Wan (e.g., Wan-AI/Wan2.1-T2V-1.3B-Diffusers) for text-driven video editing tasks.
  • Implementation on official FLUX repository: Implementation based on original FLUX. The performance is slightly better than the diffusers-based FLUX pipeline.

๐Ÿ”ฎ Acknowledgements

We sincerely thank FireFlow, RF-Solver, and FLUX for their awesome work! Additionally, we would also like to thank PnpInversion for providing comprehensive baseline survey and implementations, as well as their great benchmark.

๐Ÿ“‘ Cite Us

If you like our work, you can cite our paper through the bibtex below. Thank for your attention!

@inproceedings{jiao2026unieditflow,
	title={UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models},
	author={Guanlong Jiao and Biqing Huang and Kuan-Chieh Jackson Wang and Renjie Liao},
	booktitle={The Fourteenth International Conference on Learning Representations},
	year={2026},
	url={https://openreview.net/forum?id=ArU2CeB7Tm}
}