Arena Leaderboard & Rankings
February 15, 2026 · View on GitHub
Arena Leaderboard & Rankings
Human-preference-based Elo rankings from millions of blind pairwise comparisons
LM Arena Text-to-Image Leaderboard
Source: lmarena.ai/leaderboard/text-to-image · 3,918,094 votes · 46 models · Updated Feb 2026
| Rank | Model | Organization | Elo Score | CI | License |
|---|---|---|---|---|---|
| 1 | GPT-Image-1.5 High Fidelity | OpenAI | 1248 | ±5 | Proprietary |
| 2 | Gemini-3-Pro Image Preview 2K | 1237 | ±5 | Proprietary | |
| 3 | Gemini-3-Pro Image Preview | 1233 | ±5 | Proprietary | |
| 4 | Grok Imagine Image | xAI | 1174 | ±6 | Proprietary |
| 5 | FLUX.2 Max | Black Forest Labs | 1169 | ±4 | Proprietary |
| 6 | Grok Imagine Image Pro | xAI | 1166 | ±6 | Proprietary |
| 7 | FLUX.2 Flex | Black Forest Labs | 1158 | ±4 | Proprietary |
| 8 | Gemini 2.5 Flash Image | 1157 | ±3 | Proprietary | |
| 9 | FLUX.2 Pro | Black Forest Labs | 1156 | ±4 | Proprietary |
| 10 | HunyuanImage 3.0 | Tencent | 1151 | ±3 | Community |
| 11 | FLUX.2 Dev | Black Forest Labs | 1150 | ±5 | Proprietary |
| 12 | Imagen Ultra 4.0 | 1149 | ±4 | Proprietary | |
| 13 | Seedream 4 2K | ByteDance | 1141 | ±6 | Proprietary |
| 14 | Seedream 4.5 | ByteDance | 1141 | ±4 | Proprietary |
| 15 | Qwen-Image 2512 | Alibaba | 1139 | ±5 | Apache 2.0 |
| 16 | Imagen 4.0 | 1135 | ±3 | Proprietary | |
| 17 | Wan2.6 T2I | Alibaba | 1126 | ±6 | Proprietary |
| 18 | Seedream 4 (Fal) | ByteDance | 1119 | ±6 | Proprietary |
| 19 | Wan2.5 T2I Preview | Alibaba | 1117 | ±4 | Proprietary |
| 20 | GPT-Image-1 | OpenAI | 1115 | ±3 | Proprietary |
| 21 | Seedream 4 High-Res | ByteDance | 1114 | ±4 | Proprietary |
| 22 | GPT-Image-1 Mini | OpenAI | 1100 | ±4 | Proprietary |
| 23 | MAI-Image-1 | Microsoft AI | 1094 | ±4 | Proprietary |
| 24 | Seedream 3 | ByteDance | 1084 | ±5 | Proprietary |
| 25 | Z-Image Turbo | Alibaba | 1083 | ±7 | Apache 2.0 |
| 26 | FLUX.1 Kontext Max | Black Forest Labs | 1076 | ±3 | Proprietary |
| 27 | FLUX.2 Klein 9B | Black Forest Labs | 1065 | ±4 | Non-Commercial |
| 28 | Qwen-Image (Prompt Extend) | Alibaba | 1061 | ±3 | Apache 2.0 |
| 29 | FLUX.1 Kontext Pro | Black Forest Labs | 1060 | ±3 | Proprietary |
| 30 | Imagen 3.0 (002) | 1059 | ±3 | Proprietary | |
| 31 | Qwen-Image | Alibaba | 1058 | ±2 | Apache 2.0 |
| 32 | P-Image | Pruna | 1053 | ±5 | Proprietary |
| 33 | Ideogram v3 Quality | Ideogram | 1050 | ±4 | Proprietary |
| 34 | Photon | Luma AI | 1037 | ±4 | Proprietary |
| 35 | FLUX.2 Klein 4B | Black Forest Labs | 1021 | ±4 | Apache 2.0 |
| 36 | Recraft v3 | Recraft | 1021 | ±3 | Proprietary |
| 37 | FLUX 1.1 Pro | Black Forest Labs | 1017 | ±3 | Proprietary |
| 38 | Lucid Origin | Leonardo AI | 1015 | ±3 | Proprietary |
| 39 | Ideogram v2 | Ideogram | 1015 | ±3 | Proprietary |
| 40 | GLM-Image | Z.ai | 1013 | ±9 | MIT |
| 41 | Gemini 2.0 Flash Image | 976 | ±3 | Proprietary | |
| 42 | FLUX.1 Dev FP8 | Black Forest Labs | 970 | ±4 | Open |
| 43 | DALL·E 3 | OpenAI | 969 | ±4 | Proprietary |
| 44 | FLUX.1 Kontext Dev | Black Forest Labs | 942 | ±4 | Non-Commercial |
| 45 | Stable Diffusion v3.5 Large | Stability AI | 939 | ±4 | Open |
| 46 | BAGEL | ByteDance | 900 | ±6 | Apache 2.0 |
Artificial Analysis Text-to-Image Arena
Source: artificialanalysis.ai/text-to-image/arena · 15 specialized visual categories
| Category | Description |
|---|---|
| General & Photorealistic | Overall image quality and photorealism |
| Anime | Japanese animation-style generation |
| Text & Typography | Accuracy of rendered text in images |
| People: Portraits | Single-subject human portrait quality |
| People: Groups & Activities | Multi-person scenes and interactions |
| Nature & Landscapes | Environmental and natural scenes |
| Traditional Art | Painting, watercolor, and fine art styles |
| Cartoon & Illustration | Stylized cartoon and illustration |
| Vintage & Retro | Period-accurate retro aesthetic |
| Futuristic & Sci-Fi | Science fiction and futuristic imagery |
| Graphic Design & Digital Rendering | Design-oriented visual outputs |
| Fantasy & Mythical | Fantasy creatures and mythology |
| UI/UX Design | User interface mockups |
| Commercial | Product photography and marketing imagery |
| Physical Spaces | Architectural and interior rendering |
Recently added models (Feb 2026): Grok Imagine Image Pro, Z-Image Base, Qwen Image Plus 2601, HunyuanImage 3.0 Instruct, Wan2.6 T2I, FLUX.2 Klein (9B & 4B), GLM-Image.
Key Observations
| Insight | Detail |
|---|---|
| Proprietary Dominance | Top 5 positions held exclusively by OpenAI, Google, and xAI; highest open-weight model (Qwen-Image, Apache 2.0) at rank 15 |
| Rapid Capability Growth | GPT-Image-1.5 (Elo 1248) vs. GPT-Image-1 (Elo 1115) — 133 Elo improvement within a single model family |
| Chinese Model Competitiveness | Alibaba (Qwen-Image, Wan2.6, Z-Image), ByteDance (Seedream 4.5), and Tencent (HunyuanImage 3.0) in top-20 |
| Open-Weight Progress | Qwen-Image (Apache 2.0, Elo 1139) and FLUX.2 Klein 4B (Apache 2.0, Elo 1021) demonstrate open models closing the gap |
| FLUX Ecosystem | Black Forest Labs holds 9 of 46 positions spanning Max, Pro, Flex, Dev, Klein, and Kontext variants |
Quantitative Metrics Reference
| Metric | Full Name | Description |
|---|---|---|
| FID | Frechet Inception Distance | Distributional similarity between generated and real images |
| IS | Inception Score | Quality and diversity of generated images |
| CLIP Score | CLIP-based Alignment Score | Text-image semantic alignment via CLIP embeddings |
| VQAScore | VQA-based Score | Compositional evaluation via visual question answering |
| HPSv2 | Human Preference Score v2 | Human preference prediction for T2I outputs |
| ImageReward | Image Reward | Learned human preference model |
| LPIPS | Learned Perceptual Image Patch Similarity | Perceptual similarity using deep features |
| TIFA | Text-to-Image Faithfulness Assessment | Faithfulness evaluation via QA |
| DSG | Davidsonian Scene Graph | Fine-grained compositional evaluation |
| DreamSim | Dream Similarity | Human visual similarity judgment |
| SSIM | Structural Similarity Index | Structural image similarity |