< Back to Main README
250+ papers · 46 arena-ranked models · 6 sub-domains
The progressive evolution of text-conditioned visual synthesis — from GANs through diffusion models to autoregressive transformers
| Sub-domain | Focus | Key Models |
|---|
| Models & Face Synthesis | Foundational T2I models, face generation | FLUX.2, Seedream 3.0, Imagen 3, PixArt-α, Stable Diffusion |
| Control & Composition | Controllable generation, layout control | ControlNet, GLIGEN, Uni-ControlNet, DreamBooth |
| Editing & Personalization | Text-guided editing, personalization | InstructPix2Pix, DreamBooth, Blended Diffusion |
| Safety & Applications | Safety, bias, evaluation, applications | SafeGen, OpenBias, DIAGNOSIS |
| Cross-Modal Extensions | Video, 3D, motion from T2I | AnimateDiff, DreamFusion, Text2LIVE |
| Arena & Benchmarks | Leaderboards, benchmarks, metrics | LM Arena (46 models), HEIM, GenEval |
| Model | Title | Venue | Year |
|---|
| FLUX.2 | Next-generation flow matching T2I model | Black Forest Labs | 2025 |
| Seedream 3.0 | Scaled-up diffusion transformer | ByteDance | 2025 |
| Imagen 3 | Photorealistic T2I from Google DeepMind | Google | 2024 |
| PixArt-α | Fast Training of Diffusion Transformer | ICLR | 2024 |
| PixArt-Σ | Weak-to-Strong Training for 4K T2I | arXiv | 2024 |
| SDXL-Lightning | Progressive Adversarial Diffusion Distillation | arXiv | 2024 |
| Kolors | Diffusion Model for Photorealistic T2I | Kuaishou | 2024 |
| RealCompo | Dynamic Equilibrium between Realism and Compositionality | arXiv | 2024 |
| Playground v2.5 | Enhancing Aesthetic Quality in T2I | arXiv | 2024 |
| Model | Title | Venue | Year |
|---|
| Stable Diffusion (LDM) | High-Resolution Image Synthesis with Latent Diffusion Models | CVPR | 2022 |
| Imagen | Photorealistic T2I with Deep Language Understanding | NeurIPS | 2022 |
| DALL·E 2 | Hierarchical Text-Conditional Image Generation with CLIP Latents | arXiv | 2022 |
| Parti | Scaling Autoregressive Models for Content-Rich T2I | TMLR | 2022 |
| ControlNet | Adding Conditional Control to T2I Diffusion Models | ICCV | 2023 |
| GLIGEN | Open-Set Grounded Text-to-Image Generation | CVPR | 2023 |
| Muse | Text-To-Image Generation via Masked Generative Transformers | arXiv | 2023 |
| Model | Title | Venue | Year |
|---|
| ControlNet | Adding Conditional Control to T2I Diffusion Models | ICCV | 2023 |
| Uni-ControlNet | All-in-One Control to T2I Diffusion Models | arXiv | 2023 |
| GLIGEN | Open-Set Grounded T2I Generation | CVPR | 2023 |
| Multi-Concept Customization | Multi-Concept Customization of T2I Diffusion | CVPR | 2023 |
| DreamBooth | Fine Tuning T2I Diffusion for Subject-Driven Generation | CVPR | 2023 |
| Blended Latent Diffusion | Blended Latent Diffusion | SIGGRAPH | 2023 |
| Model | Title | Venue | Year |
|---|
| PreciseControl | Enhancing T2I with Fine-Grained Attribute Control | ECCV | 2024 |
| CosmicMan | A T2I Foundation Model for Humans | CVPR | 2024 |
| DreamFace | Progressive Generation of Animatable 3D Faces | SIGGRAPH | 2023 |
| AnyFace | Free-style Text-to-Face Synthesis and Manipulation | CVPR | 2022 |
| TediGAN | Text-Guided Diverse Image Generation and Manipulation | CVPR | 2021 |
| Benchmark | Focus | Venue | Year |
|---|
| GenExam | Multidisciplinary T2I examination | arXiv | 2025 |
| KRIS-Bench | Next-level intelligent image editing | NeurIPS | 2025 |
| DreamBench++ | Human-aligned personalized T2I | ICLR | 2025 |
| T2I-CompBench++ | Enhanced compositional T2I evaluation | TPAMI | 2025 |
| GenAI-Bench | Compositional text-to-visual generation | CVPR | 2024 |
| GenEval | Object-focused T2I alignment | NeurIPS | 2023 |
| TIFA | T2I faithfulness via QA | ICCV | 2023 |
| HEIM | Holistic evaluation of T2I models | NeurIPS | 2023 |