< Back to Main README
200+ papers · 6 sub-domains · 6 theory themes
Learning internal representations of how the world works — from game simulation to autonomous driving and embodied AI
| Sub-domain | Focus | Key Models |
|---|
| Theory & Surveys | Theoretical foundations, surveys, scaling laws | 20+ surveys |
| Game Simulation | Interactive game engines, virtual worlds | GameNGen, DIAMOND, Oasis |
| Video Generation | Video generation as world simulation | Sora, Cosmos, Genie |
| LiDAR Generation | LiDAR point cloud world models | LiDAR-based driving models |
| Occupancy Generation | 3D occupancy prediction and generation | OccWorld, occupancy models |
| Embodied AI | Manipulation, navigation, locomotion, MBRL | TesserAct, DreamGen, Dreamer v4 |
| Model | Title | Representation | Year |
|---|
| Genie 3 | A new frontier for world models | Interactive Video | 2025 |
| V-JEPA 2 | Self-Supervised Video Models Enable Understanding, Prediction and Planning | Latent (JEPA) | 2025 |
| Cosmos 2.5 | Evolving the World Foundation Models for Physical AI | Video Generation | 2025 |
| PAN | A World Model for General, Interactable, Long-Horizon Simulation | Interactive | 2025 |
| Sora | Video generation models as world simulators | Video Generation | 2024 |
| Genie 2 | A Large-Scale Foundation World Model | Interactive Video | 2024 |
| Cosmos | Cosmos World Foundation Model Platform for Physical AI | Video Generation | 2024 |
| UniSim | Learning Interactive Real-World Simulators | Interactive | 2023 |
| Pandora | General World Model with Natural Language Actions and Video States | Interactive | 2024 |
| Model | Title | Space |
|---|
| GameNGen | Diffusion Models Are Real-Time Game Engines | Pixel |
| MineWorld | Real-Time Interactive World Model on Minecraft | Pixel |
| DIAMOND | Diffusion for World Modeling: Visual Details Matter in Atari | Pixel |
| Matrix-Game 2.0 | Open-Source Real-Time Streaming Interactive World Model | Pixel |
| Oasis | A Universe in a Transformer | Pixel |
| HunyuanWorld 1.0 | Generating Immersive Explorable Interactive 3D Worlds | 3D Mesh |
| Matrix-3D | Omnidirectional Explorable 3D World Generation | 3D Mesh |
| Model | Title |
|---|
| Cosmos-Drive-Dreams | Scalable Synthetic Driving Data with World Foundation Models |
| GAIA-2 | A Controllable Multi-View Generative World Model for Driving |
| Vista | Generalizable Driving World Model with High Fidelity |
| DrivingWorld | Constructing World Model for Driving via Video GPT |
| OccWorld | Learning a 3D Occupancy World Model for Driving |
| Drive-WM | Multiview Visual Forecasting and Planning with World Model |
| Model | Title | Domain |
|---|
| TesserAct | Learning 4D Embodied World Models | Foundation |
| DreamGen | Unlocking Generalization via Video World Models | Foundation |
| iVideoGPT | Interactive VideoGPTs are Scalable World Models | Foundation |
| AgiBot-World | Large-scale Manipulation Platform for Embodied Systems | Manipulation |
| NWM | Navigation World Models | Navigation |
| DWL | Advancing Humanoid Locomotion with World Model Learning | Locomotion |
| Model | Title | Focus |
|---|
| CoT-VLA | Visual Chain-of-Thought for Vision-Language-Action Models | VLA |
| WorldVLA | Towards Autoregressive Action World Model | VLA |
| Dreamer v4 | Training Agents Inside of Scalable World Models | MBRL |
| Dreamer v3 | Mastering Diverse Domains through World Models | MBRL |
| TD-MPC2 | Scalable, Robust World Models for Continuous Control | MBRL |
| DINO-WM | World Models on Pre-trained Visual Features for Zero-shot Planning | Latent |
| Title | Key Contribution |
|---|
| General agents contain world models | Theoretical framework for world representation emergence |
| When Do Neural Networks Learn World Models? | Conditions under which networks develop world models |
| Transformers Use Causal World Models in Maze-Solving | Evidence of causal reasoning in next-token prediction |
| Video as the New Language for Real-World Decision Making | Video as universal reasoning substrate |
| Scaling Laws for Pre-training Agents and World Models | Compute-optimal training strategies |
| Title | Domain | Year |
|---|
| Is Sora a World Simulator? | General | 2024 |
| A Comprehensive Survey on World Models for Embodied AI | Embodied | 2024 |
| A Survey of World Models for Autonomous Driving | Driving | 2024 |
| 3D and 4D World Modeling: A Survey | 3D/4D | 2024 |
| World Models: The Safety Perspective | Safety | 2024 |