Awesome-World-Models

July 11, 2026 · View on GitHub

arXiv PDF Maintenance Last Commit Contribution Welcome

📢 Updates

  • 2026.07: Add recent ICML 2026 world model papers
  • 2026.05: Add recent world model papers
  • 2026.01: We release our paper. "Learning to Model the World: A Survey of World Models in Artificial Intelligence"

👀 Introduction

Welcome to the repository for our survey paper, "Learning to Model the World: A Survey of World Models in Artificial Intelligence". This repository provides resources and updates related to our research. For a detailed introduction, please refer to our survey paper.

World models (WMs) provide a unified approach for modeling how environments evolve over time by learning predictive representations of states and observations. Recent advances in large-scale generative modeling and multimodal foundation models have substantially broadened their applicability across a wide range of interactive and multimodal domains; however, existing research remains fragmented across modeling paradigms, application domains, and evaluation protocols. This survey provides a systematic and in-depth review of WMs in artificial intelligence. Based on the world modeling paradigms of existing methods, we first categorize WMs into four major branches with formal mathematical formulations: reinforcement learning-based, observation-level generative, latent space, and object-centric world models. We further review a broad range of WM applications spanning robotics, autonomous driving, scientific discovery, virtual game simulation, GUI-based agents, as well as interpretability and trustworthiness, and summarize benchmark datasets, evaluation metrics, simulation platforms, and comparative results across WMs. Finally, we discuss key challenges, including long-horizon consistency, controllability, robustness, evaluation limitations, and generalization, and outline promising directions for future research. This survey aims to offer a unified reference for understanding, comparing, and advancing WMs.

image

The recent timeline of world models, covering core methods and the release of open-source and closed-source reproduction projects.

📒 Table of Contents

Part 0: Survey Papers

  • Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses [Paper]

  • From Specialist to Generalist: A Comprehensive Survey on World Models [Paper]

  • Understanding World or Predicting Future? A Comprehensive Survey of World Models [Paper]

  • World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child [Paper]

  • Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond [Paper]

  • A Survey of World Models for Autonomous Driving [Paper]

  • 3D and 4D World Modeling: A Survey [Paper]

  • A Comprehensive Survey on World Models for Embodied AI [Paper]

Part 1: Reinforcement Learning-Based World Models

  • WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control [Paper]
  • HAUWM: Learning to Be Uncertain: Pre-training World Models with Horizon-Calibrated Uncertainty [Paper]
  • Newt: Learning Massively Multitask World Models for Continuous Control [Paper]
  • R2-Dreamer: Redundancy-Reduced World Models without Decoders or Augmentation [Paper]
  • MoW: Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent Dynamics [Paper]
  • ScaleZero: One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning [Paper]
  • EAWM: From Observations to Events: Event-Aware World Models for Reinforcement Learning [Paper]
  • Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models [Paper]
  • Sparse Imagination for Efficient Visual World Model Planning [Paper]
  • Compositional Planning with Jumpy World Models [Paper]
  • Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination [Paper]
  • Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling [Paper]
  • Learning Reward-Cost Balance in Safe RL via Score-Based World Models [Paper]
  • Parallel Stochastic Gradient-Based Planning for World Models [Paper]
  • Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning [Paper]
  • Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning [Paper]
  • Boosting World Models Learning via Latent-Space Value Alignment [Paper]
  • SOLAR for Offline MARL: Plateau-Triggered Potential Shaping under World-Model Uncertainty [Paper]
  • WorldCompass: Reinforcement Learning for Long-Horizon World Models [Paper]
  • Learning Disentangled Multi-Agent World Model for Decentralized Control [Paper]
  • Synthesizing world models for bilevel planning [Paper]
  • Deep SPI: Safe Policy Improvement via World Models [Paper]
  • Context and Diversity Matter: The Emergence of In-Context Learning in World Models [Paper]
  • World Models as Reference Trajectories for Rapid Motor Adaptation [Paper]
  • Adversarial Diffusion for Robust Reinforcement Learning [Paper]
  • World Modeling with Probabilistic Structure Integration [Paper]
  • DreamerV3: Mastering diverse control tasks through world models [Paper]
  • Dreamer: Dream to control: Learning behaviors by latent imagination [Paper]
  • DreamSmooth: Improving model-based reinforcement learning via reward smoothing [Paper]
  • PlaNet: Learning latent dynamics for planning from pixels [Paper]
  • DreamerV2: Mastering atari with discrete world models [Paper]
  • PIGDreamer: Privileged information guided world models for safe partially [Paper]
  • HarmonyDream: Task Harmonization Inside World Models [Paper]
  • DyMoDreamer: World modeling with dynamic modulation [Paper]
  • TD-MPC2: Scalable, robust world models for continuous control [Paper]
  • Hieros: Hierarchical imagination on structured state space sequence world models [Paper]
  • THICK: Learning hierarchical world models with adaptive temporal abstractions from discrete latent dynamics [Paper]
  • MoSim: Neural motion simulator pushing the limit of world models in reinforcement learning [Paper]
  • R2I: Mastering memory tasks with world models [Paper]
  • LEQ: Model-based offline reinforcement learning with lower expectile q-learning [Paper]
  • DIMA: Revisiting multi-agent world modeling from a diffusion-inspired perspective [Paper]
  • PCM: Policy-conditioned environment models are more generalizable [Paper]
  • CoWorld: Making offline RL online: Collaborative world models for offline visual reinforcement learning [Paper]
  • IQ-MPC: Reward-free world models for online imitation learning [Paper]
  • WAKER: Reward-free curricula for training robust world models [Paper]
  • REM: Improving token-based world models with parallel observation prediction [Paper]
  • cRSSM: Dreaming of many worlds: Learning contextual world models aids zero-shot generalization [Paper]
  • Adaptive world models: Learning behaviors by latent imagination under non-stationarity $$$ [Paper]
  • PWM: Policy learning with multi-task world models [Paper]

Part 2: Observation-Level Generative World Models

Language Observations

*ByteSized32-State-Prediction: Can language models serve as text-based world simulators? $$$[Paper]

  • WorldLLM: Improving LLMs' World Modeling Using Curiosity-Driven Theory-Making [Paper]
  • GPT-4: Gpt-4 technical report [Paper]
  • Llama 3: The llama 3 herd of models [Paper]
  • LLMCWM: Language agents meet causality – bridging LLMs and causal world models [Paper]
  • Surge: On the potential of large language mode [Paper]
  • RAP: Reasoning with language model is planning with world model [Paper]
  • Making large language models into world models with precondition and effect knowledge [Paper]
  • LWM: World model on million-length video and language with blockwise ringattention [Paper]
  • Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech [Paper]

Visual Observations

  • Astra: General Interactive World Model with Autoregressive Denoising [Paper]
  • Composition of Memory Experts for Diffusion World Models [Paper]
  • Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention [Paper]
  • Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Video Generation [Paper]
  • Olaf-World: Orienting Latent Actions for Video World Modeling [Paper]
  • EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression Recognition [Paper]
  • World Guidance: World Modeling in Condition Space for Action Generation [Paper]
  • Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model [Paper]
  • LIVE: Long-horizon Interactive Video World Modeling [Paper]
  • Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics [Paper]
  • VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model [Paper]
  • StableWorld: Towards Stable and Consistent Long Interactive Video Generation [Paper]
  • Learning World Models for Interactive Video Generation [Paper]
  • How Far is Video Generation from World Model: A Physical Law Perspective [Paper]
  • Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions [Paper]
  • Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion [Paper]
  • Sora: Video generation models as world simulators [Paper]
  • TeleWorld: Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation [Paper]
  • Gen-3 [Blog]
  • T2v-turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback [Paper]
  • Emu3: Next-token prediction is all you need [Paper]
  • SPMEM: Video world models with long-term spatial memory [Paper]
  • Wan: Open and Advanced Large-Scale Video Generative Models [Paper]
  • LLaVA: Visual instruction tuning [Paper]
  • Vid2world: Crafting video diffusion models to interactive world models [Paper]
  • VideoCrafter2: Overcoming data limitations for high-quality video diffusion models [Paper]
  • DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors [Paper]
  • IRASim: A Fine-Grained World Model for Robot Manipulation [Paper]
  • WISA: World simulator assistant for physics-aware text-to-video generation [Paper]
  • Co-Evolving Latent Action World Models [Paper]

3D and 4D Observations

  • WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving [Paper]
  • Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling [Paper]
  • Unified 3D Scene Understanding Through Physical World Modeling [Paper]
  • WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-Chains [Paper]
  • Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion [Paper]
  • MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation [Paper]
  • Towards Practical World Model for 4D Occupancy Forecasting in Autonomous Driving [Paper]
  • WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting [Paper]
  • VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs [Paper]
  • PERSIST: Beyond Pixel Histories: World Models with Persistent 3D State [Paper]
  • PointWorld: Scaling 3D World Models for In-the-wild Robotic Manipulation [Paper]
  • MagicWorld: Interactive Geometry-driven Video World Exploration [Paper]
  • LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation [Paper]
  • 4D-fy: Text-to-4d generation using hybrid score distillation sampling [Paper]
  • WonderJourney: Going from anywhere to everywhere [Paper]
  • SceneScape: Text-driven consistent scene generation [Paper]
  • LiDARCrafter: Dynamic 4d world modeling from lidar sequences [Paper]
  • Text2room: Extracting textured 3d meshes from 2d text-to-image models [Paper]
  • WonderWorld: Interactive 3d scene generation from a single image [Paper]
  • Invisible Stitch: Generating smooth 3d scenes with depth inpainting [Paper]

Part 3: Latent Space World Models

  • I-JEPA: Self-supervised learning from images with a joint-embedding predictive architecture [Paper]
  • Causal-JEPA: Learning World Models through Object-Level Latent Interventions [Paper]
  • VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models [Paper]
  • Identifiable Token Correspondence for World Models [Paper]
  • Self-supervised Hierarchical Visual Reasoning with World Model [Paper]
  • V-JEPA: Revisiting Feature Prediction for Learning Visual Representations from Video [Paper]
  • V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning [Paper]
  • seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models [Paper]
  • MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features [Paper]
  • DINO-WM: World Models on Pre-trained Visual Features Enable Zero-shot Planning [Paper]
  • DINO-world: Back to the Features: DINO as a Foundation for Video World Models [Paper]
  • DINO-Foresight: Looking into the Future with DINO [Paper]
  • World Models Group Latents: Learning Abstract World Models with a Group-Structured Latent Space [Paper]
  • Structure Abstraction and Generalization in a Hippocampus-Entorhinal Inspired World Model [Paper]

Part 4: Object-centric World Models

  • Object-centric Learning with Slot Attention [Paper]
  • SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models [Paper]
  • FIOC-WM: Learning Interactive World Model for Object-Centric Reinforcement Learning [Paper]
  • LPWM: Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling [Paper]
  • G-SWM: Improving Generative Imagination in Object-Centric World Models [Paper]
  • When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks [Paper]
  • Objects matter: object-centric world models improve reinforcement learning in visually complex environments [Paper]
  • LSlotFormer: Object-Centric World Model for Language-Guided Manipulation [Paper]
  • Compositional OCL: Unifying Causal and Object-centric Representation Learning allows Causal Composition [Paper]
  • Dyn-O: Building Structured World Models with Object-Centric Representations [Paper]
  • MEAD: Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction [Paper]
  • Object-Centric Latent Action Learning [Paper]
  • Object-Centric Representations Generalize Better Compositionally with Less Compute [Paper]
  • CarFormer: Self-driving with Learned Object-Centric Representations [Paper]
  • FOCUS: object-centric world models for robotic manipulation [Paper]

Part 5: World Models for Robotics

Manipulation

Visual Futrue Prediction

  • Structured 4D Latent World Model for Robot Planning [Paper]
  • Visuo-Tactile World Models [Paper]
  • DDP-WM: Disentangled Dynamics Prediction for Efficient World Models [Paper]
  • DiLA: Disentangled Latent Action World Models [Paper]
  • RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation [Paper]
  • DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos [Paper]
  • Flow Equivariant World Models: Structured Memory for Dynamic Environments [Paper]
  • WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching [Paper]
  • RoboDreamer: Learning Compositional World Models for Robot Imagination [Paper]
  • Grounding Video Models to Actions through Goal Conditioned Exploration [Paper]
  • ViPRA: Video Prediction for Robot Actions [Paper]
  • FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation [Paper]
  • TesserAct: Learning 4D Embodied World Models [Paper]
  • ORV: 4D Occupancy-centric Robot Video Generation [Paper]
  • WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation [Paper]
  • Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control [Paper]
  • Vidar: Embodied Video Diffusion Model for Generalist Bimanual Manipulation [Paper]
  • LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation [Paper]
  • KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models [Paper]

Latent Action State Imagination

  • Cross-Embodiment Robot Foundation World Models with Latent Actions [Paper]
  • Factored Latent Action World Models [Paper]
  • Learning Latent Action World Models In The Wild [Paper]
  • Multi-view Consistent Latent Action Learning for World Modeling and Control [Paper]
  • FLARE: Robot Learning with Implicit World Modeling [Paper]
  • AdaWorld: Learning Adaptable World Models with Latent Actions [Paper]
  • DyWA: Dynamic World Adaptation for Generalizable World Models [Paper]
  • OSVI-WM: One-Shot Visual Imitation via World Models [Paper]
  • DEMO3^3: Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning [Paper]
  • Strengthening Generative Robot Policies through Predictive World Modeling [Paper]
  • In-Context Policy Improvement for Contact-Rich Manipulation with Pretrained Generative Models [Paper]

Control-Oriented Planning

  • Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (PETS) [Paper]
  • TD-MPC: Temporal Difference Learning for Model Predictive Control [Paper]
  • Time-Aware World Model for Adaptive Prediction and Control (TAWM) [Paper]
  • Planning with Diffusion for Flexible Behavior Synthesis [Paper]
  • Decision Diffuser: Is Conditional Generative Modeling All You Need for Decision-Making? [Paper]
  • Potential-Based Diffusion Motion Planning [Paper]
  • HiP: Compositional Foundation Models for Hierarchical Planning [Paper]
  • PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation [Paper]
  • Hdflow: Hierarchical diffusion-flow planning for long-horizon robotic assembl [Paper]
  • ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning [Paper]
  • TaskLoom: Knowledge Graph-Driven Commonsense World Models for Robotic Task Planning [Paper]
  • ManipDreamer: Learning Manipulation World Models with Action-Tree Supervisions [Paper]
  • RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation [Paper]
  • Mobile Manipulation with Active Inference [Paper]

World Action Modeling

  • Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation [Paper]
  • WoW: Towards a World Omniscient World Model through Embodied Interaction [Paper]
  • PAR: Physical Autoregressive Model for Robotic Manipulation without Action Pretraining [Paper]
  • iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation [Paper]
  • PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models [Paper]
  • WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems [Paper]
  • World4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic Manipulation [Paper]
  • NWM: Navigation World Models [Paper] 会议徽章
  • Kinodynamic: Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models [Paper] 会议徽章
  • Neuro-Symbolic: Perspective-Shifted Neuro-Symbolic World Models for Socially-Aware Robot Navigation [Paper] 会议徽章
  • Abs-Sim2Rea: Abstract Sim2Real through Approximate Information States [Paper] 会议徽章
  • ESWM: Building Spatial World Models from Sparse Transitional Episodic Memories [Paper]
  • TMoW: Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments [Paper]
  • Unified WM: Unified world models: Memory-augmented planning and foresight for visual navigation [Paper] 会议徽章
  • Scene Graph World: Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation [Paper] 会议徽章
  • Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension [Paper] 会议徽章
  • X-Mobility: End-To-End Generalizable Navigation via World Modeling [Paper] 会议徽章
  • SC2^{2}-WM: A Self-Correcting World Model with Closed-Loop Feedback for Embodied Navigation [Paper]
  • WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation [Paper] 会议徽章
  • RECON: Rapid Exploration for Open-World Navigation with Latent Goal Models [Paper] 会议徽章
  • NaVi-WM: Deductive Chain-of-Thought Augmented Socially-aware Robot Navigation World Model [Paper] 会议徽章
  • MindJourney: Test-Time Scaling with World Models for Spatial Reasoning [Paper] 会议徽章
  • NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments [Paper] 会议徽章
  • Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation [Paper] 会议徽章
  • World Model Implanting for Test-time Adaptation of Embodied Agents [Paper] 会议徽章
  • After Persistent Embodied WM: Learning 3D Persistent Embodied World Models [Paper] 会议徽章

Policy Learning

  • WMPO: World Model-based Policy Optimization for Vision-Language-Action Models [Paper]
  • World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation [Paper]
  • DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions [Paper]
  • Latent Action World Models for Control with Unlabeled Trajectories [Paper]
  • Prelar: World model pre-training with learnable action representation [Paper]
  • World models can leverage human videos for dexterous manipulation [Paper]
  • TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos [Paper]
  • 3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model [Paper]
  • Ctrl-World: A Controllable Generative World Model for Robot Manipulation [Paper]
  • Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics [Paper]
  • WorldGym: World Model as An Environment for Policy Evaluation [Paper]
  • Worldeval: World model as realworld robot policies evaluator [Paper]

Locomotion

  • WMP: World Model-based Perception for Visual Legged Locomotion [Paper] 会议徽章
  • WMR: Learning Humanoid Locomotion with World Model Reconstruction [Paper] 会议徽章
  • World Model Predictive Control for Robust Locomotion [Paper] 会议徽章
  • Ego: Ego-Vision World Model for Humanoid Contact Planning [Paper] 会议徽章

Part 6: World Models for Autonomous Driving

Predictive Modeling

  • DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving [Paper]
  • CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving [Paper]
  • DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving [Paper]
  • Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks [Paper]
  • ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask [Paper]
  • ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving [Paper]
  • UniWorld: Autonomous Driving Pre-training via World Models [Paper]
  • HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving [Paper]
  • WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation [Paper]
  • NeMo: Neural Volumetric World Models for Autonomous Driving [Paper]
  • UnO: Unsupervised Occupancy Fields for Perception and Forecasting [Paper]
  • Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion [Paper]
  • Renderworld: World Model with Self-Supervised 3D Label [Paper]
  • DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation [Paper]

Action-Conditioned Imagination

  • GAIA-1: A Generative World Model for Autonomous Driving [Paper]
  • WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens [Paper]
  • OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving [Paper]
  • DriveDreamer: Towards Real-World-Drive World Models for Autonomous Driving [Paper]
  • Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability [Paper]
  • Drive-WM: Driving Into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving [Paper]
  • DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving [Paper]
  • InfinityDrive: Breaking Time Limits in Driving World Models [Paper]
  • DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation [Paper]

Decision-Centric Integration

  • Think2Drive: Efficient Reinforcement Learning by Thinking with Latent World Model for Autonomous Driving (in CARLA-V2) [Paper]
  • Doe-1: Closed-Loop Autonomous Driving with Large World Model [Paper]
  • AdaWM: Adaptive World Model based Planning for Autonomous Driving [Paper]
  • DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving [Paper]

Part 7: World Models for Science

Social Science and Socioeconomic Systems

  • Building Social World Model with Large Language Models [Paper]
  • SWM: Social World Models [Paper]
  • SWM-AP: Social World Model-Augmented Mechanism Design Policy Learning [Paper]
  • SocioVerse: A World Model for Social Simulation Powered by LLM Agents and a Pool of 10 Million Real-World Users [Paper]
  • TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets [Paper]
  • A Virtual Reality-Integrated System for Behavioral Analysis in Neurological Decline [Paper]
  • Bidding for Influence: Auction-Driven Diffusion Image Generation [Paper]
  • Typhon: Effectively Designing 2-Dimensional Sequence Models for Multivariate Time Series [Paper]

Physical and Natural Sciences

  • VCWorld: A Biological World Model for Virtual Cell Simulation [Paper]
  • LithoDreamer: A Process-Level Lithography World Model [Paper]
  • Surgical Vision World Model [Paper]
  • CellFlux: Simulating Cellular Morphology Changes via Flow Matching [Paper]
  • ODesign: A World Model for Biomolecular Interaction Design [Paper]
  • ORBIT: A Prognostic World Model for Ocular Reasoning Based on Imagined Breakdown Trajectories [Paper]
  • Medical world model: Generative simulation of tumor evolution for treatment planning [Paper]
  • CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning [Paper]
  • Xray2Xray: World Model from Chest X-rays with Volumetric Context [Paper]
  • EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance [Paper]
  • PINT: Physics-Informed Neural Time Series Models with Applications to Long-term Inference on WeatherBench 2m-Temperature Data [Paper]
  • Reconstructing Dynamics from Steady Spatial Patterns with Partial Observations [Paper]
  • HEP-JEPA: A foundation model for collider physics [Paper]

Part 8: World Models for Virtual Game Simulation

Pixel-Level Observation Prediction

  • Diffusion For World Modeling: Visual Details Matter In Atari [Paper]
  • Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory [Paper]
  • WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling [Paper]
  • GameNGen: Diffusion Models Are Real-Time Game Engines [Paper]
  • ActionParty: Multi-Subject Action Binding in Generative Video Games [Paper]
  • Code World Models for General Game Playing [Paper]
  • One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration [Paper]
  • GameGen-X: Interactive Open-world Game Video Generation [Paper]
  • iVideoGPT: Interactive VideoGPTs are Scalable World Models [Paper]
  • Oasis: A Universe In A Transformer [Paper]
  • AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction [Paper]
  • Matrix-Game 2.0: An Open-Source, Real-Time, And Streaming Interactive World Model [Paper]
  • Matrix-Game: Interactive World Foundation Model [Paper]
  • RealPlay: From Virtual Games To Real-World Play [Paper]
  • GameFactory: Creating New Games With Generative Interactive Videos [Paper]

3D Mesh-Level Observation Prediction

  • HunyuanWorld 1.0: Generating Immersive, Explorable, And Interactive 3D Worlds From Words Or Pixels [Paper]
  • Matrix-3D: Omnidirectional Explorable 3D World Generation [Paper]

Part 9: World Models for GUI-Based Agents

  • Generative Visual Code Mobile World Models [Paper]
  • Agent World Model: Playable Agentically-Grounded World Generation [Paper]
  • Code2Worlds: Web Interactive World Generation as Code Generation [Paper]
  • WebWorld: A Large-Scale World Model for Web Agent Training [Paper]
  • PathWise: Planning through World Model for Automated Heuristic Design with LLMs [Paper]
  • NeuralOS: Towards Simulating Operating Systems via Neural Generative Models [Paper]
  • ViMo: A Generative Visual GUI World Model for App Agents [Paper]
  • Unlocking Smarter Device Control: Foresighted Planning With A World Model-Driven Code Execution Approach [Paper]
  • R-WoM: Retrieval-augmented World Model For Computer-use Agents [Paper]
  • WKM: Agent Planning With World Knowledge Model [Paper]
  • WebDreamer: Is Your LLM Secretly A World Model Of The Internet? Model-Based Planning For Web Agents [Paper]
  • WMA: Web Agents With World Models: Learning And Leveraging Environment Dynamics In Web Navigation [Paper]
  • WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis [Paper]
  • WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model [Paper]
  • WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents [Paper]
  • SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents [Paper]
  • Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents [Paper]

Part 10: Interpretable and Trustworthy World Models

  • General agents need world models [Paper] 会议徽章
  • Position: Express Your Doubts: Probabilistic World Modeling Should not be Based on Token logprobs [Paper]
  • Position: We Need A Unified Definition of Hallucination in NLP - It's The World Model, Stupid! [Paper]
  • Position: World Models as an Intermediary between Agents and the Real World [Paper]
  • Cold-Start Personalization via Training-Free Priors from Structured World Models [Paper]
  • From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers [Paper]
  • Interpreting Physics in Video World Models [Paper]
  • MetaOthello: Toward Studying the Emergence of World Models in Language Models through a Deep Redundancy-Free Framework [Paper]
  • World-Model Inspired Emotion-aware Image Captioning Using Multimodal Large Language Model [Paper]
  • GPT:A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment [Paper] 会议徽章
  • Stan&Terry:Transformers Use Causal World Models in Maze-Solving Tasks [Paper] 会议徽章
  • MLP:When Do Neural Networks Learn World Models? [Paper] 会议徽章
  • LLMs:Linear spatial world models emerge in large language models. [Paper] 会议徽章
  • Mistral:Revisiting the othello world model hypothesis. [Paper] 会议徽章
  • GPT-2-style transformer:Scaling laws for pre-training agents and world models. [Paper] 会议徽章
  • VAE:How hard is it to confuse a world model? [Paper] 会议徽章
  • DeepSeek:Utilizing world models for adaptively covariate acquisition under limited budget for causal decision making. [Paper] 会议徽章

Part 11: Benchmark of World Models

Benchmark Datasets & Evaluation Metrics

  • World-In-World: World Models in a Closed-Loop World [Paper]
  • OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling [Paper]
  • DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving [Paper]
  • iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework [Paper]
  • WorldMark: A Unified Benchmark Suite for Interactive Video World Models [Paper]
  • Benchmarking World-Model Learning with Environment-Level Queries [Paper]
  • Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World Model [Paper]
  • Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models [Paper]
  • Frozen in time: A joint video and image encoder for end-to-end retrieval. [Paper] 会议徽章
  • Panda-70m: Captioning 70m videos with multiple cross-modality teachers. [Paper] 会议徽章
  • Ego4d: Around the world in 3,000 hours of egocentric video. [Paper] 会议徽章
  • Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. [Paper] 会议徽章
  • Worldscore:A unified evaluation benchmark for world generation. [Paper] 会议徽章
  • Open x-embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0. [Paper] 会议徽章
  • Room-Across-Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. [Paper] 会议徽章
  • Ewmbench: Evaluating scene, motion, and semantic quality in embodied world models. [Paper] 会议徽章
  • Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. [Paper] 会议徽章
  • nuscenes: A multimodal dataset for autonomous driving. [Paper] 会议徽章
  • ACT-Bench: Towards Action Controllable World Models for Autonomous Driving. [Paper] 会议徽章
  • Jump cell painting dataset: morphological impact of 136,000 chemical and genetic perturbations. [Paper] 会议徽章
  • Random Forest Classifier:A machine learning modelto predict hepatocellular carcinoma response to transcatheter arterial chemoembolization. [Paper] 会议徽章
  • The arcade learning environment: An evaluation platform for general agents. [Paper] 会议徽章
  • Minerl: A large-scale dataset of minecraft demonstrations. [Paper] 会议徽章
  • OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. [Paper] 会议徽章
  • Windows agent arena: Evaluating multi-modal OS agents at scale. [Paper] 会议徽章

Physics Engines & Simulation Platforms

image

  • Gazebo: Design and use paradigms for Gazebo, an open-source multi-robot simulator [Paper]
  • WebotsTM: Professional Mobile Robot Simulation [Paper]
  • Bullet Physics SDK [github] [Blog]
  • PyBullet Quickstart Guide [SDK] [Blog]
  • MuJoCo [github] [Blog]
  • Simbody: multibody dynamics for biomedical research [Paper]
  • NVIDIA PhysX SDK [SDK]
  • Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning [Paper] [Blog]
  • Omniverse [Notion] [SDK]
  • Genesis: A Generative and Universal Physics Engine for Robotics and Beyond [github] [SDK]

Part 12: Performance Comparison

image

Citation

If you find this work useful, welcome to cite us.

@article{WM_Survey,
  author={Jiahua Dong and Qi Lyu and Baichen Liu and Xudong Wang and Wenqi Liang and Duzhen Zhang and Jiahang Tu and Hongliu Li and Hanbin Zhao and Henghui Ding and Yulun Zhang and Zhi Han and Nicu Sebe and Fahad Shahbaz Khan and Salman Khan and Mubarak Shah and Philip Torr and Ming-Hsuan Yang and Dacheng Tao},
  journal={TechRxiv}, 
  title={Learning to Model the World: A Survey of World Models in Artificial Intelligence}, 
  year={2026},
}

⭐ Star History

Star History Chart