OpenMM-Arena

February 16, 2026 · View on GitHub

OpenMM-Arena Logo

OpenMM-Arena

A Comprehensive Compendium for Multimodal Artificial Intelligence Research

Awesome License: MIT PRs Welcome Website GitHub stars GitHub forks GitHub watchers

Papers Models Benchmarks Datasets Research Pillars Last Updated

A systematically curated and taxonomically organized knowledge base that unifies seven converging research frontiers of multimodal AI — encompassing 2000+ papers, 46 arena-ranked models, 100+ benchmarks, and 90+ datasets — providing researchers with a single, authoritative cross-disciplinary reference.

Explore the Compendium · Arena Leaderboard · Technical Blogs · Star on GitHub


If you find this resource valuable for your research, please consider giving us a :star: to increase its visibility.


Table of Contents


Overview

The field of multimodal artificial intelligence has undergone an unprecedented convergence of historically disparate research traditions. Text-to-image synthesis and text-to-video generation have evolved from nascent subfields into foundational pillars of generative AI, while image-to-video generation bridges static visual content with temporal dynamics through learned motion priors. 3D vision — underpinned by Neural Radiance Fields, 3D Gaussian Splatting, LLM-driven scene understanding, and visual SLAM — extends generative modeling into the spatial domain. 4D spatial intelligence introduces the temporal axis through depth estimation, dense 3D/4D tracking, and physics-grounded simulation. These areas intersect profoundly with world models that seek to learn predictive environment dynamics, and unified multimodal architectures that dissolve the boundary between visual perception and generation.

OpenMM-Arena serves as a definitive, centralized knowledge base that systematically organizes, cross-references, and catalogues this expansive and rapidly growing literature across seven research pillars:

                              ┌──────────────────────────────────────┐
                              │         OpenMM-Arena                 │
                              │   Multimodal AI Research Compendium  │
                              └──────────────┬───────────────────────┘

              ┌──────────┬──────────┬────────┼────────┬──────────┬──────────┐
              ▼          ▼          ▼        ▼        ▼          ▼          ▼
         ┌─────────┐┌─────────┐┌────────┐┌──────┐┌──────┐┌──────────┐┌────────┐
         │  T2I    ││  T2V    ││  I2V   ││  3D  ││  4D  ││ Unified  ││ World  │
         │ 250+    ││ 200+    ││ 170+   ││ 500+ ││ 500+ ││ 120+     ││ 200+   │
         │ papers  ││ papers  ││ papers ││papers││papers││ models   ││ papers │
         └────┬────┘└────┬────┘└───┬────┘└──┬───┘└──┬───┘└────┬─────┘└───┬────┘
              │          │         │        │       │         │          │
              ▼          ▼         ▼        ▼       ▼         ▼          ▼
          Models     Foundation  Animation  3DGS   Depth   Diffusion   Theory
          Control    Controllable Editing   NeRF   Tracking  AR        Games
          Editing    Benchmarks  Portraits  SLAM   Recon    Hybrid     Driving
          Safety     Datasets    Transfer   LLM-3D Dynamic  Any2Any   Embodied
          Arena                             Robot  Human    Eval       Sim
          Benchmarks                        Nav    Physics  Datasets   MBRL

2000+ Papers · 7 Research Pillars · 46 Arena-Ranked Models · 100+ Benchmarks · 90+ Datasets · 73 Curated Blog Posts · 11 Source Repositories


Seven Research Pillars

Text-to-Image Generation

250+ papers · 46 arena-ranked models · 6 sub-domains

The progressive evolution of text-conditioned image synthesis — from generative adversarial networks through denoising diffusion models to autoregressive transformers. Encompasses foundational architectures (FLUX.2, Seedream 3.0, GPT-Image, Imagen 3), controllable generation via spatial and semantic conditioning (ControlNet, GLIGEN), text-guided editing and subject-driven personalization (InstructPix2Pix, DreamBooth), safety and bias mitigation, and comprehensive arena leaderboards derived from 3.9M+ human preference votes.

Sub-domains: Models & Face Synthesis · Control & Composition · Editing & Personalization · Safety & Applications · Cross-Modal Extensions · Arena & Benchmarks


Text-to-Video Generation

200+ papers · 25+ datasets · 3 sub-domains

Foundation video synthesis models spanning GANs to diffusion transformers — covering Sora, Wan 2.1, Veo 2, CogVideoX, and contemporary controllable and efficient video generation approaches. Addresses long-form video synthesis, temporal coherence benchmarks, and large-scale video-text corpora.

Sub-domains: Foundation T2V Models · Controllable & Efficient Synthesis · Benchmarks & Datasets


Image-to-Video Generation

170+ papers · 20+ applications · 2 sub-domains

The synthesis of dynamic video sequences from static images via learned motion priors — encompassing image animation, character-driven video synthesis, talking-head generation, temporally consistent video editing, motion transfer, and audio-driven synthesis. (Documented within the T2V reference.)

Sub-domains: Animation & Portraits · Video Editing & Enhancement


3D Vision

500+ papers · 6 sub-domains

A comprehensive survey spanning 3D Gaussian Splatting, Neural Radiance Fields (NeRF), text/image-to-3D generation, LLM-driven 3D understanding, NeRF-SLAM, GS-SLAM, visual and LiDAR SLAM, robotic manipulation, autonomous navigation, and spatial localization. Extends generative AI into the spatial domain, forging connections to real-time robotics and embodied agents.

Sub-domains: 3D Gaussian Splatting · NeRF & Generation · NeRF-SLAM · GS-SLAM · LLM-3D Understanding · Robotics & Navigation


4D Spatial Intelligence

500+ papers · 6 sub-domains

The temporal dimension of 3D understanding — monocular and multi-view depth estimation, camera pose recovery, dense 3D/4D point tracking, scene reconstruction, 4D dynamic scenes via deformable NeRFs and 4D Gaussian Splatting, human-centric motion capture, and physics-grounded simulation.

Sub-domains: Geometry & Depth · 3D/4D Tracking · Reconstruction · Dynamic Scenes · Human-Centric · Physics-Based Simulation


Unified Multimodal Models

120+ models · 30+ benchmarks · 3 sub-domains

Architectures that jointly perform visual understanding and generation within a single parametric framework — diffusion-based unified models, autoregressive multimodal LLMs with pixel and semantic encoding, hybrid AR-diffusion architectures, and any-to-any paradigms. Complemented by comprehensive evaluation benchmarks and curated training corpora.

Sub-domains: Models & Architectures · Datasets & Training Corpora · Evaluation & Benchmarks


World Models

200+ papers · 6 theory themes · 6 sub-domains

Learned environment dynamics for game simulation, autonomous driving, embodied manipulation, model-based reinforcement learning, video generation as world simulation, and theoretical underpinnings. Traces the lineage from Ha & Schmidhuber's seminal formulation through contemporary systems such as Sora, NVIDIA Cosmos, and Genie.

Sub-domains: Theory & Surveys · Game Simulation · Video Generation · LiDAR Generation · Occupancy Generation · Embodied AI


Arena Leaderboard Spotlight

Source: LM Arena · 3.9M+ human preference votes · 46 models · Updated February 2026

RankModelOrganizationElo ScoreLicense
1GPT-Image-1.5OpenAI1248Proprietary
2Gemini-3-Pro Image 2KGoogle1237Proprietary
3Seedream 3.0ByteDance1233Proprietary
4Grok Imagine ImagexAI1174Proprietary
5FLUX.2 MaxBlack Forest Labs1169Proprietary
6Grok Imagine Image ProxAI1166Proprietary
7FLUX.2 FlexBlack Forest Labs1158Proprietary
8Gemini 2.5 Flash ImageGoogle1157Proprietary
9FLUX.2 ProBlack Forest Labs1156Proprietary
10HunyuanImage 3.0Tencent1151Community

View Complete Leaderboard (46 Models) →

Key Observations

FindingAnalysis
Proprietary-model dominanceThe top five positions are held by OpenAI, Google, ByteDance, and xAI, indicating that proprietary training infrastructure and data remain decisive advantages
Rapid intra-family progressGPT-Image-1.5 (Elo 1248) surpasses GPT-Image-1 (Elo 1115) by 133 points — a substantial within-family improvement
Narrowing open-weight gapQwen-Image (Apache 2.0, Elo 1139, rank 15) and FLUX.2 Klein 4B (Apache 2.0, Elo 1021) demonstrate that open-weight models are increasingly competitive
Growing Chinese-lab presenceModels from Alibaba, ByteDance, and Tencent now occupy multiple top-20 positions, reflecting significant investment in multimodal generation research

Benchmarks & Evaluation Metrics

Evaluation Benchmarks

CategoryRepresentative Benchmarks
Multimodal UnderstandingGeneral-Bench, MMMU, MM-Vet v2, MMBench, SEED-Bench-2, GQA
Image Generation QualityGenExam, KRIS-Bench, WISE, DreamBench++, T2I-CompBench++, GenAI-Bench, TIFA, HEIM
World Model EvaluationWorldScore, WorldSimBench, PhyWorld, Newton, WorldGym, EWMBench
Interleaved GenerationUniBench, OpenING, ISG, MMIE
Human-Preference RankingsLM Arena T2I (3.9M votes), Artificial Analysis T2I (15 style categories)

Quantitative Metrics

MetricFull NamePrimary Domain
FIDFréchet Inception DistanceImage Generation Quality
CLIP ScoreCLIP-based Text–Image Alignment ScoreText-to-Image Faithfulness
VQAScoreVQA-based Compositional ScoreT2I Semantic Faithfulness
HPSv2Human Preference Score v2Human Preference Alignment
ImageRewardImage Reward ModelHuman Preference Alignment
LPIPSLearned Perceptual Image Patch SimilarityPerceptual Image Quality
TIFAText-to-Image Faithfulness AssessmentAttribute Faithfulness
DSGDavidsonian Scene GraphCompositional Correctness
SSIMStructural Similarity Index MeasureStructural Image Quality

Datasets

90+ datasets catalogued across all modalities

CategoryRepresentative DatasetsApproximate Scale
Multimodal UnderstandingLAION-5B, DataComp, Infinity-MM, Cambrian-10M5.9B – 10M samples
Text-to-ImageLAION-Aesthetics, PixelProse, PD12M, CC-12M, SAM120M – 11M samples
Image EditingByteMorph-6M, UltraEdit, AnyEdit, ImgEdit6M – 1.2M samples
Interleaved MultimodalOmniCorpus (8B), OBELICS (141M), Multimodal C4 (101M)8B – 101M samples
Video-TextWebVid-10M, InternVid, HD-VILA-100M, Panda-70M100M – 10M samples

Documentation Structure

OpenMM-Arena/
├── README.md                    ← This document (project overview)
├── docs/
│   ├── T2I.md                   ← Text-to-Image: in-depth reference
│   ├── T2V.md                   ← Text-to-Video & Image-to-Video: in-depth reference
│   ├── 3D_VISION.md             ← 3D Vision: in-depth reference
│   ├── 4D_VISION.md             ← 4D Spatial Intelligence: in-depth reference
│   ├── UNIFIED.md               ← Unified Multimodal Models: in-depth reference
│   ├── WORLD_MODELS.md          ← World Models: in-depth reference
│   ├── ARENA.md                 ← Arena Leaderboard: full rankings & analysis
│   │
│   ├── index.html               ← Website: landing page
│   ├── blog.html                ← Website: curated technical blog posts (73 posts)
│   ├── arena.html               ← Website: interactive arena leaderboards
│   ├── style.css                ← Website: design system
│   ├── script.js                ← Website: interactive features & search
│   ├── sitemap.xml              ← Website: SEO sitemap
│   ├── robots.txt               ← Website: crawler directives
│   ├── 404.html                 ← Website: custom error page
│   ├── t2i.html / t2v.html ... ← Website: pillar hub pages
│   ├── t2i/ t2v/ i2v/ ...      ← Website: sub-domain detail pages
│   └── blog/                    ← Website: all blog posts page
DocumentContent ScopeCoverage
T2I.mdFoundational architectures, face synthesis, controllable generation, editing, safety, cross-modal extensions, arena rankings250+ papers
T2V.mdFoundation T2V models, controllable synthesis, I2V animation, video editing, benchmarks, datasets370+ papers
3D_VISION.md3D Gaussian Splatting, NeRF, SLAM, LLM-3D understanding, robotics, navigation500+ papers
4D_VISION.mdDepth estimation, 3D/4D tracking, reconstruction, dynamic scenes, human motion, physics simulation500+ papers
UNIFIED.mdDiffusion, autoregressive, hybrid, any-to-any unified models, evaluation benchmarks120+ models
WORLD_MODELS.mdTheory, game simulation, autonomous driving, video generation, embodied AI, MBRL200+ papers
ARENA.mdLM Arena & Artificial Analysis leaderboards, 46 ranked models, Elo methodology46 models

Source Repositories

OpenMM-Arena systematically consolidates and cross-references knowledge from 11 foundational open-source repositories:

RepositoryPrimary DomainMaintainer
Awesome-Text-to-ImageText-to-Image GenerationYutong Zhou et al.
Awesome-T2V-GenerationText-to-Video Generationsoraw-ai
Awesome-Video-DiffusionVideo Diffusion ModelsShowLab
Awesome-3DGS3D Gaussian SplattingMrNeRF
Awesome-3DReconstruction3D ReconstructionOpenMVG
Awesome-LLM-3DLLM-3D UnderstandingActiveVisionLab
Awesome-3D-Vision3D Vision (General)Hardy-Uint
Awesome-NeRF-3DGS-SLAMNeRF & 3DGS-based SLAM3D-Vision-World
Awesome-4D-SI4D Spatial IntelligenceYukang Cao
Awesome-World-ModelsWorld ModelsSiqiao Huang et al.
Awesome-Unified-MMUnified Multimodal ModelsAIDC-AI / Zhang et al.

Key Features

FeatureDescription
2000+ Papers CataloguedAmong the most comprehensive multimodal AI paper collections available — meticulously organized with full bibliographic citations, arXiv links, and venue information
Taxonomic OrganizationNavigate a carefully structured hierarchy — from high-level research pillars through thematic sub-domains to individual papers, with principled categorization by chronological era and methodological paradigm
Arena LeaderboardsReal-time rankings derived from 3.9M+ human preference votes — enabling head-to-head model comparison via Elo scores, win rates, and comprehensive performance statistics
73 Curated Blog PostsAn anthology of authoritative technical blog posts from leading AI laboratories (BFL, Google DeepMind, OpenAI, Meta AI, NVIDIA, Stability AI, ByteDance, and others)
Smart Search & FilteringReal-time search, sortable tables, year-based filtering, and keyboard shortcuts for efficient navigation
Dark ModeFull dark-mode support with a carefully designed color scheme preserving readability and contrast
Responsive DesignOptimized for desktop, tablet, and mobile viewing with collapsible sidebar navigation
Continuously UpdatedMaintained in pace with the latest developments — new publications from CVPR, ICLR, NeurIPS, ECCV, and arXiv are integrated regularly
Open SourceReleased under the MIT License — community-driven, with contributions welcome from the global research community

Quick Start

Browse online (recommended):

Visit openenvision-lab.github.io/OpenMM-Arena

Run locally:

git clone https://github.com/OpenEnvision-Lab/OpenMM-Arena.git
cd OpenMM-Arena/docs
python -m http.server 8000
# Navigate to http://localhost:8000 in your browser

Keyboard shortcuts (on the website):

KeyAction
Cmd/Ctrl + KOpen global search
DToggle dark mode
[Toggle sidebar
?Display all shortcuts

Contributing

We welcome contributions from the research community. There are several ways to participate:

  • Adding papers — Submit a pull request incorporating new publications into any of the seven research pillars
  • Updating leaderboards — Help maintain current arena rankings as new models are evaluated
  • Correcting errors — Report or resolve citation errors, broken links, or taxonomic misclassifications
  • Proposing improvements — Open an issue with suggestions for better organization or new content areas
  • Disseminating the resource — Star the repository, share on academic and social platforms, and cite in your publications

Increasing Visibility

If you find OpenMM-Arena valuable for your research, please consider:

  1. Starring this repository to increase its visibility within the GitHub community
  2. Sharing on academic platforms — Twitter/X, Reddit r/MachineLearning, Hacker News, Papers With Code
  3. Citing in your publications — see the Citation section below
  4. Cross-referencing — if you maintain a related awesome list or survey, consider linking to OpenMM-Arena

For repository maintainers: please add these topics in the GitHub repository settings to maximize discoverability.

multimodal-ai text-to-image text-to-video image-to-video 3d-vision 4d-vision world-models diffusion-models generative-ai computer-vision deep-learning awesome-list neural-radiance-fields gaussian-splatting research-papers benchmark leaderboard arena survey paper-collection


Citation

If you find OpenMM-Arena useful for your research, please consider citing:

@misc{openmmarena2025,
    title     = {OpenMM-Arena: A Comprehensive Compendium for Multimodal AI Research},
    author    = {OpenMM-Arena Contributors},
    year      = {2025},
    journal   = {GitHub repository},
    url       = {https://github.com/OpenEnvision-Lab/OpenMM-Arena}
}
Cite foundational works
@inproceedings{zhou2023vision,
  title     = {Vision + Language Applications: A Survey},
  author    = {Zhou, Yutong and Shimada, Nobutaka},
  booktitle = {CVPRW},
  year      = {2023}
}

@misc{huang2025awesomeworldmodels,
  title  = {Awesome-World-Models},
  author = {Siqiao Huang},
  year   = {2025},
  url    = {https://github.com/knightnemo/Awesome-World-Models}
}

@article{zhang2025unified,
  title   = {Unified Multimodal Understanding and Generation Models},
  author  = {Zhang, Xinjie and others},
  journal = {arXiv preprint arXiv:2505.02567},
  year    = {2025}
}

OpenMM-Arena belongs to OpenEnvision Lab

Curated with scholarly rigor for the multimodal AI research community

Last updated: February 2026

Website · GitHub · Issues