Arena Leaderboard & Rankings

February 15, 2026 · View on GitHub

< Back to Main README

Arena Leaderboard & Rankings

Human-preference-based Elo rankings from millions of blind pairwise comparisons


LM Arena Text-to-Image Leaderboard

Source: lmarena.ai/leaderboard/text-to-image · 3,918,094 votes · 46 models · Updated Feb 2026

RankModelOrganizationElo ScoreCILicense
1GPT-Image-1.5 High FidelityOpenAI1248±5Proprietary
2Gemini-3-Pro Image Preview 2KGoogle1237±5Proprietary
3Gemini-3-Pro Image PreviewGoogle1233±5Proprietary
4Grok Imagine ImagexAI1174±6Proprietary
5FLUX.2 MaxBlack Forest Labs1169±4Proprietary
6Grok Imagine Image ProxAI1166±6Proprietary
7FLUX.2 FlexBlack Forest Labs1158±4Proprietary
8Gemini 2.5 Flash ImageGoogle1157±3Proprietary
9FLUX.2 ProBlack Forest Labs1156±4Proprietary
10HunyuanImage 3.0Tencent1151±3Community
11FLUX.2 DevBlack Forest Labs1150±5Proprietary
12Imagen Ultra 4.0Google1149±4Proprietary
13Seedream 4 2KByteDance1141±6Proprietary
14Seedream 4.5ByteDance1141±4Proprietary
15Qwen-Image 2512Alibaba1139±5Apache 2.0
16Imagen 4.0Google1135±3Proprietary
17Wan2.6 T2IAlibaba1126±6Proprietary
18Seedream 4 (Fal)ByteDance1119±6Proprietary
19Wan2.5 T2I PreviewAlibaba1117±4Proprietary
20GPT-Image-1OpenAI1115±3Proprietary
21Seedream 4 High-ResByteDance1114±4Proprietary
22GPT-Image-1 MiniOpenAI1100±4Proprietary
23MAI-Image-1Microsoft AI1094±4Proprietary
24Seedream 3ByteDance1084±5Proprietary
25Z-Image TurboAlibaba1083±7Apache 2.0
26FLUX.1 Kontext MaxBlack Forest Labs1076±3Proprietary
27FLUX.2 Klein 9BBlack Forest Labs1065±4Non-Commercial
28Qwen-Image (Prompt Extend)Alibaba1061±3Apache 2.0
29FLUX.1 Kontext ProBlack Forest Labs1060±3Proprietary
30Imagen 3.0 (002)Google1059±3Proprietary
31Qwen-ImageAlibaba1058±2Apache 2.0
32P-ImagePruna1053±5Proprietary
33Ideogram v3 QualityIdeogram1050±4Proprietary
34PhotonLuma AI1037±4Proprietary
35FLUX.2 Klein 4BBlack Forest Labs1021±4Apache 2.0
36Recraft v3Recraft1021±3Proprietary
37FLUX 1.1 ProBlack Forest Labs1017±3Proprietary
38Lucid OriginLeonardo AI1015±3Proprietary
39Ideogram v2Ideogram1015±3Proprietary
40GLM-ImageZ.ai1013±9MIT
41Gemini 2.0 Flash ImageGoogle976±3Proprietary
42FLUX.1 Dev FP8Black Forest Labs970±4Open
43DALL·E 3OpenAI969±4Proprietary
44FLUX.1 Kontext DevBlack Forest Labs942±4Non-Commercial
45Stable Diffusion v3.5 LargeStability AI939±4Open
46BAGELByteDance900±6Apache 2.0

Artificial Analysis Text-to-Image Arena

Source: artificialanalysis.ai/text-to-image/arena · 15 specialized visual categories

CategoryDescription
General & PhotorealisticOverall image quality and photorealism
AnimeJapanese animation-style generation
Text & TypographyAccuracy of rendered text in images
People: PortraitsSingle-subject human portrait quality
People: Groups & ActivitiesMulti-person scenes and interactions
Nature & LandscapesEnvironmental and natural scenes
Traditional ArtPainting, watercolor, and fine art styles
Cartoon & IllustrationStylized cartoon and illustration
Vintage & RetroPeriod-accurate retro aesthetic
Futuristic & Sci-FiScience fiction and futuristic imagery
Graphic Design & Digital RenderingDesign-oriented visual outputs
Fantasy & MythicalFantasy creatures and mythology
UI/UX DesignUser interface mockups
CommercialProduct photography and marketing imagery
Physical SpacesArchitectural and interior rendering

Recently added models (Feb 2026): Grok Imagine Image Pro, Z-Image Base, Qwen Image Plus 2601, HunyuanImage 3.0 Instruct, Wan2.6 T2I, FLUX.2 Klein (9B & 4B), GLM-Image.


Key Observations

InsightDetail
Proprietary DominanceTop 5 positions held exclusively by OpenAI, Google, and xAI; highest open-weight model (Qwen-Image, Apache 2.0) at rank 15
Rapid Capability GrowthGPT-Image-1.5 (Elo 1248) vs. GPT-Image-1 (Elo 1115) — 133 Elo improvement within a single model family
Chinese Model CompetitivenessAlibaba (Qwen-Image, Wan2.6, Z-Image), ByteDance (Seedream 4.5), and Tencent (HunyuanImage 3.0) in top-20
Open-Weight ProgressQwen-Image (Apache 2.0, Elo 1139) and FLUX.2 Klein 4B (Apache 2.0, Elo 1021) demonstrate open models closing the gap
FLUX EcosystemBlack Forest Labs holds 9 of 46 positions spanning Max, Pro, Flex, Dev, Klein, and Kontext variants

Quantitative Metrics Reference

MetricFull NameDescription
FIDFrechet Inception DistanceDistributional similarity between generated and real images
ISInception ScoreQuality and diversity of generated images
CLIP ScoreCLIP-based Alignment ScoreText-image semantic alignment via CLIP embeddings
VQAScoreVQA-based ScoreCompositional evaluation via visual question answering
HPSv2Human Preference Score v2Human preference prediction for T2I outputs
ImageRewardImage RewardLearned human preference model
LPIPSLearned Perceptual Image Patch SimilarityPerceptual similarity using deep features
TIFAText-to-Image Faithfulness AssessmentFaithfulness evaluation via QA
DSGDavidsonian Scene GraphFine-grained compositional evaluation
DreamSimDream SimilarityHuman visual similarity judgment
SSIMStructural Similarity IndexStructural image similarity

< Back to Main README · View on Website