< Back to Main README
120+ models · 30+ benchmarks · 3 sub-domains
Architectures that jointly understand and generate across modalities within a shared representational space
| Model | Title | Date |
|---|
| UniModel | A Visual-Only Framework for Unified Multimodal Understanding and Generation | 2025/11 |
| MMaDA | Multimodal Large Diffusion Language Models | 2025/05 |
| Muddit | Liberating Generation Beyond T2I with Unified Discrete Diffusion | 2025/05 |
| UniDisc | Unified Multimodal Discrete Diffusion | 2025/03 |
| Dual Diffusion | Dual Diffusion for Unified Image Generation and Understanding | 2024/12 |
| Model | Encoding | Title | Date |
|---|
| Emu3.5 | Pixel | Native Multimodal Models are World Learners | 2025/10 |
| Qwen-Image | Semantic | Qwen-Image Technical Report | 2025/08 |
| OmniGen2 | Semantic | Exploration to Advanced Multimodal Generation | 2025/06 |
| Selftok | Pixel | Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning | 2025/05 |
| Harmon | Pixel | Harmonizing Visual Representations for Unified Understanding and Generation | 2025/03 |
| Liquid | Semantic | Language Models are Scalable and Unified Multi-modal Generators | 2024/12 |
| Emu3 | Pixel | Next-Token Prediction is All You Need | 2024/09 |
| Chameleon | Pixel | Mixed-Modal Early-Fusion Foundation Models | 2024/05 |
| Model | Architecture | Title | Date |
|---|
| BAGEL | AR + Diffusion | Emerging Properties in Unified Multimodal Pretraining | 2025/05 |
| Show-o2 | AR + Diffusion | Improved Native Unified Multimodal Models | 2025/06 |
| Janus-Pro | Hybrid Enc. | Unified Understanding and Generation with Data and Model Scaling | 2025/01 |
| JanusFlow | AR + Flow | Harmonizing Autoregression and Rectified Flow | 2024/11 |
| Show-o | AR + Diffusion | One Transformer to Unify Understanding and Generation | 2024/08 |
| Transfusion | AR + Diffusion | Predict Next Token and Diffuse Images with One Model | 2024/08 |
| Model | Title | Date |
|---|
| Qwen3-Omni | Qwen3-Omni Technical Report | 2025/09 |
| Ming-Omni | Unified Multimodal Model for Perception and Generation | 2025/06 |
| OmniFlow | Any-to-Any Generation with Multi-Modal Rectified Flows | 2024/12 |
| NExT-GPT | Any-to-Any Multimodal LLM | 2023/09 |
| CoDi | Any-to-Any Generation via Composable Diffusion | 2023 |
| ImageBind | One Embedding Space To Bind Them All | 2023 |
| Benchmark | Focus | Venue | Year |
|---|
| General-Bench | Multimodal generalist evaluation | ICML | 2025 |
| MMMU | Multi-discipline multimodal understanding | CVPR | 2023 |
| MM-Vet v2 | Integrated capabilities of LMMs | arXiv | 2024 |
| MMBench | All-around multi-modal evaluation | ECCV | 2023 |
| UniBench | Unified evaluation for unified models | — | 2024 |
| OpenING | Open-ended interleaved generation | — | 2024 |