Text-to-Image Generation

February 15, 2026 · View on GitHub

< Back to Main README

Text-to-Image Generation

250+ papers · 46 arena-ranked models · 6 sub-domains

The progressive evolution of text-conditioned visual synthesis — from GANs through diffusion models to autoregressive transformers


Sub-domains

Sub-domainFocusKey Models
Models & Face SynthesisFoundational T2I models, face generationFLUX.2, Seedream 3.0, Imagen 3, PixArt-α, Stable Diffusion
Control & CompositionControllable generation, layout controlControlNet, GLIGEN, Uni-ControlNet, DreamBooth
Editing & PersonalizationText-guided editing, personalizationInstructPix2Pix, DreamBooth, Blended Diffusion
Safety & ApplicationsSafety, bias, evaluation, applicationsSafeGen, OpenBias, DIAGNOSIS
Cross-Modal ExtensionsVideo, 3D, motion from T2IAnimateDiff, DreamFusion, Text2LIVE
Arena & BenchmarksLeaderboards, benchmarks, metricsLM Arena (46 models), HEIM, GenEval

Foundational Models — Diffusion & Transformer Era (2024–2025)

ModelTitleVenueYear
FLUX.2Next-generation flow matching T2I modelBlack Forest Labs2025
Seedream 3.0Scaled-up diffusion transformerByteDance2025
Imagen 3Photorealistic T2I from Google DeepMindGoogle2024
PixArt-αFast Training of Diffusion TransformerICLR2024
PixArt-ΣWeak-to-Strong Training for 4K T2IarXiv2024
SDXL-LightningProgressive Adversarial Diffusion DistillationarXiv2024
KolorsDiffusion Model for Photorealistic T2IKuaishou2024
RealCompoDynamic Equilibrium between Realism and CompositionalityarXiv2024
Playground v2.5Enhancing Aesthetic Quality in T2IarXiv2024

Foundational Models — Latent Diffusion & GAN Era (2022–2023)

ModelTitleVenueYear
Stable Diffusion (LDM)High-Resolution Image Synthesis with Latent Diffusion ModelsCVPR2022
ImagenPhotorealistic T2I with Deep Language UnderstandingNeurIPS2022
DALL·E 2Hierarchical Text-Conditional Image Generation with CLIP LatentsarXiv2022
PartiScaling Autoregressive Models for Content-Rich T2ITMLR2022
ControlNetAdding Conditional Control to T2I Diffusion ModelsICCV2023
GLIGENOpen-Set Grounded Text-to-Image GenerationCVPR2023
MuseText-To-Image Generation via Masked Generative TransformersarXiv2023

Controllable Generation

ModelTitleVenueYear
ControlNetAdding Conditional Control to T2I Diffusion ModelsICCV2023
Uni-ControlNetAll-in-One Control to T2I Diffusion ModelsarXiv2023
GLIGENOpen-Set Grounded T2I GenerationCVPR2023
Multi-Concept CustomizationMulti-Concept Customization of T2I DiffusionCVPR2023
DreamBoothFine Tuning T2I Diffusion for Subject-Driven GenerationCVPR2023
Blended Latent DiffusionBlended Latent DiffusionSIGGRAPH2023

Text-to-Face Synthesis

ModelTitleVenueYear
PreciseControlEnhancing T2I with Fine-Grained Attribute ControlECCV2024
CosmicManA T2I Foundation Model for HumansCVPR2024
DreamFaceProgressive Generation of Animatable 3D FacesSIGGRAPH2023
AnyFaceFree-style Text-to-Face Synthesis and ManipulationCVPR2022
TediGANText-Guided Diverse Image Generation and ManipulationCVPR2021

Image Generation Benchmarks

BenchmarkFocusVenueYear
GenExamMultidisciplinary T2I examinationarXiv2025
KRIS-BenchNext-level intelligent image editingNeurIPS2025
DreamBench++Human-aligned personalized T2IICLR2025
T2I-CompBench++Enhanced compositional T2I evaluationTPAMI2025
GenAI-BenchCompositional text-to-visual generationCVPR2024
GenEvalObject-focused T2I alignmentNeurIPS2023
TIFAT2I faithfulness via QAICCV2023
HEIMHolistic evaluation of T2I modelsNeurIPS2023

< Back to Main README · View on Website