Awesome-World-Model [](https://github.com/sindresorhus/awesome)

November 18, 2025 Β· View on GitHub

A curated list of awesome resources on World Models, based on the comprehensive survey "Understanding World or Predicting Future? A Comprehensive Survey of World Models".

Loading roadmap

NewsπŸ”₯

  • [2024/11/21] Initial release of our survey is available on arXiv.
  • [2025/06/13] Our survey paper "Understanding World or Predicting Future? A Comprehensive Survey of World Models" has been accepted by ACM Computing Surveys.
  • [2025/06/25] Second version of our survey is available on arXiv.
  • [2025/07/18] Initial release of the Awesome-World-Model GitHub repository.
  • [2025/11/18] Third version of our survey is available on arXiv.

Contact

If you have any suggestions or find our work helpful, feel free to contact us
Email: dingjt15@tsinghua.org.cn

If this list helps your research, please ⭐ and cite:

@article{ding2025understanding,
  title={Understanding World or Predicting Future? A Comprehensive Survey of World Models},
  author={Ding, Jingtao and Zhang, Yunke and Shang, Yu and Zhang, Yuheng and Zong, Zefang and Feng, Jie and Yuan, Yuan and Su, Hongyuan and Li, Nian and Sukiennik, Nicholas and others},
  journal={ACM Computing Surveys},
  volume={58},
  number={3},
  pages={1--38},
  year={2025},
  publisher={ACM New York, NY}
}

Table of Contents πŸƒ

Roadmap of world models in deep learning era

Loading roadmap

Model-based RL

TitlePub. & DateCode/Project URL
Recurrent world models facilitate policy evolution (RWM)NeurIPS 2018Website
Learning Latent Dynamics for Planning from Pixels (PlaNet)ICML 2019Star
Dream to control: Learning behaviors by latent imagination (Dreamer V1)ICLR 2020Star
Mastering atari with discrete world models (Dreamer V2)ICLR 2021Star
Temporal Difference Learning for Model Predictive Control (TD-MPC1)ICML 2023Star
Mastering Diverse Domains through World Models (Dreamer V3)2023Star
TD-MPC2: Scalable, Robust World Models for Continuous Control (TD-MPC2)ICLR 2024Star
PWM: Policy Learning with Multi-Task World Models (PWM)ICLR 2025Star

Self-supervised learning

TitlePub.&DateCode/Project URL
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27 (JEPA)2024β€”
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (DINO-WM)2024Star
Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA)2024Star
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (V-JEPA2)2025Star

LLM/MLLM

TitlePub.&DateCode/Project URL
Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning (LLM-DM)NeurIPS 2023Star
WorldGPT: Empowering LLM as Multimodal World Model (WorldGPT)ACM MM 2024Star
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation (Text2World)ACL 2025Star

Video generation

TitlePub.&DateCode/Project URL
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers (CogVideo)ICLR 2023Star
Structure and Content-Guided Video Synthesis with Diffusion Models (Gen‑1)ICCV 2023Website
UniSim: Learning Interactive Real-World Simulators (Unisim)ICLR 2024Website
Sora: Creating video from text (Sora)OpenAI 2024β€”
World model on million-length video and language with ring-attention (LWM)ICLR 2025Star
Genie: Generative Interactive Environmentsn (Genie)ICML 2024Website
iVideoGPT: Interactive VideoGPTs are Scalable World Models (iVideoGPT)NeurIPS 2024Star
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer (CogVideoX)ICLR 2025Star
Wan: Open and Advanced Large-Scale Video Generative Models (Wan)2025Star
Cosmos World Foundation Model Platform for Physical AI (Cosmos)2025Star

Interactive 3D environment

TitlePub.&DateCode/Project URL
Interactive 3D Scene Generation from a Single Image (WonderWorld)CVPR 2025Star
Matrix-3D: Omnidirectional Explorable 3D World Generation (Matrix-3D)2025Star
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels (Hunyuan World)2025Star

Application

TitlePub.&DateCode/Project URL
DayDreamer: World Models for Physical Robot Learning (DayDreamer)2023Star
Generative Agents: Interactive Simulacra of Human Behavior (Generative Agents)UIST 2023Star
GAIA-1: A generative world model for autonomous driving (GAIA-1)2023Website
OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving (OccWorld)ECCV 2024Star
Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (GR1)ICLR 2024Star
DriveDreamer: Towards real-world-driven world models (DriveDreamer)ECCV 2024Star
Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving (Drive-WM)CVPR 2024Star
Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2) (Think2Drive)ECCV 2024Website
Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving (Drive-WM)CVPR 2024Star
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving (GaussianWorld)CVPR 2025Star
Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving (Drive-WM)2025Website
World and Human Action Models towards Gameplay Ideation (WHAM)Nature 2025β€”
Mineworld: a Real-time and Open-source Interactive World Model on Minecraft (MineWorld)2025Star
GameFactory: Creating New Games with Generative Interactive Videos (Gamefactory)ICCV 2025Star
AgentSociety: Large-scale simulation of LLM-driven generative agents (AgentSociety)ACL 2025, COLM 2025Star
EnerVerse-AC: Envisioning Embodied Environments with Action Condition (EnerVerse)2025Star
GR-3 Technical Report (GR3)2025Website
Aether: Geometric-Aware Unified World Modeling (Aether)2025Star
GWM: Towards Scalable Gaussian World Models for Robotic Manipulation (GWM)ICCV 2025Website
AirScape: An Aerial Generative World Model with Motion Controllability (AirScape)ACM MM 2025Website
RoboScape: Physics-informed Embodied World Model (RoboScape)2025Star
DreamGen: Unlocking Generalization in Robot Learning through Video World Models (DreamGen)2025Star
Matrix-Game: Interactive World Foundation Model (Matrix-Game)2025Star
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation (Genie Envisioner)2025Star

3 Implicit Representation of the External World

3.1 World Model in Decision Making

TitlePub. & DateCode/Project URL
Deep reinforcement learning in a handful of trials using probabilistic dynamics modelsNeurIPS 2018Star
PWM: Policy Learning with Multi-Task World ModelsICLR 2025Star
Recurrent world models facilitate policy evolutionNeurIPS 2018Website
Dream to control: Learning behaviors by latent imaginationICLR 2020Star
Leveraging pre-trained large language models to construct and utilize world models for model-based task planningNeurIPS 2023Star
Mastering atari with discrete world modelsICLR 2021Star
Mastering diverse control tasks through world modelsNature 2024Star
TD-MPC2: Scalable, Robust World Models for Continuous ControlICLR 2024Star
When to trust your model: Model-based policy optimizationNeurIPS 2019Star
Offline reinforcement learning as one big sequence modeling problemNeurIPS 2021Star
Model predictive controlSpringerβ€”
Algorithmic framework for model-based deep reinforcement learning with theoretical guaranteesICLR 2019Star
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuningICRA 2018Star
A game theoretic framework for model based reinforcement learningICML 2021Star
General agents need world modelsICML 2025β€”
Mastering memory tasks with world modelsICLR 2024Star
A generalist dynamics model for controlarXiv 2023β€”
Exploring model-based planning with policy networksICLR 2020Star
A0c: Alpha zero in continuous action spacearXiv 2018Star
Probabilistic adaptation of text-to-video modelsICLR 2024Website
RoboDreamer: Learning Compositional World Models for Robot ImaginationICML 2024Star
Discuss before moving: Visual language navigation via multi-expert discussionsICRA 2024Star
OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and Structured RepresentationCVPR 2024Star
RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual NavigationCVPR 2024Website
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language ModelsarXiv 2025β€”
Position: LLMs can't plan, but can help planning in LLM-modulo frameworksICML 2024β€”
Language models meet world models: Embodied experiences enhance language modelsNeurIPS 2023Star
Virtualhome: Simulating household activities via programsCVPR 2018Star
Learning to Model the World with LanguageICML 2024Star
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample EfficiencyICML 2024Star
Alfworld: Aligning text and embodied environments for interactive learningICLR 2021Star
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web AgentsEMNLP 2024Star
Agent Planning with World Knowledge ModelNeurIPS 2024Star
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentNeurIPS 2024Star
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web NavigationICLR 2025Star

3.2 World Knowledge Learned by Models

TitlePub. & DateCode / Project URL
Does the chimpanzee have a theory of mind?Behav. & Brain Sci. 1978β€”
GPT4GEO: How a Language Model Sees the World’s GeographyNeurIPS 2023Star
LLMs achieve adult human performance on higher-order theory of mind tasksarXiv 2024β€”
COKE: A cognitive knowledge graph for machine theory of mindACL 2024Star
Think Twice: Perspective-Taking Improves LLM Theory-of-MindACL 2024Star
Language Models Represent Space and TimeICLR 2024Star
GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsICLR 2024Star
Large language models are geographically biasedICML 2024Star
Emergent Representations of Program Semantics in Language Models Trained on ProgramsICML 2024Star
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and LanguagesNeurIPS 2024Star
SafeWorld: Geo-Diverse Safety AlignmentNeurIPS 2024Star
EAI: Emotional Decision-Making of LLMs in Strategic Games and Ethical DilemmasNeurIPS 2024Star
Testing theory of mind in large language models and humansNature Human Behaviour 2024Website
Automated construction of cognitive maps with visual predictive codingNature Machine Intelligence 2024Star
Evaluating Large Language Models in Theory of Mind TasksPNAS 2024Website
Elements of World Knowledge (EWOK)Transactions of the ACL 2025Website
The Geometry of Concepts: Sparse Autoencoder Feature StructureEntropy 2025Star
AgentMove: A large language model based agentic framework for zero-shot next location predictionNAACL 2025Star
CityGPT: Empowering Urban Spatial Cognition of Large Language ModelsKDD 2025Star
CityBench: Evaluating the Capabilities of Large Language Model as World ModelKDD 2025Star
LocalGPT: Benchmarking and Advancing Large Language Models for Local Life ServicesKDD 2025Star
UrbanLLaVA: A Multi-modal Large Language Model for Urban IntelligenceICCV 2025Star
Open-Set Living Need Prediction with Large Language ModelsACL 2025 FindingsStar
Do Vision-Language Models Have Internal World Models? Towards an Atomic EvaluationACL 2025 FindingsWebsite
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality AligningEMNLP 2025 FindingsStar
GPS as a Control Signal for Image GenerationCVPR 2025Star
All Languages Matter: Evaluating LMMs on Culturally Diverse 100 LanguagesCVPR 2025Website
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal ModelsCVPR 2025Star
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall SpacesCVPR 2025Website
A Survey of Large Language Model-Powered Spatial Intelligence Across ScalesarXiv 2025β€”
AI's Blind Spots: Geographic Knowledge and Diversity Deficit in Generated Urban ScenarioarXiv 2025β€”
Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsarXiv 2025β€”

4 Future Prediction of the Physical World

4.1 World Model as Video Generation

TitlePub. & DateCode / Project URL
Video generation models as world simulatorsOpenAI Blog 2024β€”
Sora: Creating video from textOpenAI 2024β€”
Is Sora a world simulator? A comprehensive survey on general world models and beyondarXiv 2024Star
Sora as an AGI world model? A complete survey on text-to-video generationarXiv 2024β€”
How Far is Video Generation from World Model: A Physical Law PerspectiveICML 2025Star
Do generative video models learn physical principles from watching videos?arXiv 2025Star
Genesis: A Generative and Universal Physics Engine for Robotics and BeyondICML 2024Star
PhysGen: Rigid-body physics-grounded image-to-video generationECCV 2024Star
NUWA-XL: Diffusion over Diffusion for Extremely Long Video GenerationACL 2023Website
OccWorld: Learning a 3D Occupancy World Model for Autonomous DrivingECCV 2024Star
OccSora: 4D Occupancy Generation Models as World SimulatorsICLR 2025Star
World model on million-length video and language with ring-attentionICLR 2025Star
GAIA-1: A generative world model for autonomous drivingarXiv 2023Website
DriveDreamer: Towards real-world-driven world modelsECCV 2024Star
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationAAAI 2025Star
Driving into the Future: Multiview Visual Forecasting and Planning with World ModelCVPR 2024Star
Vista: A Generalizable Driving World Model with High FidelityNeurIPS 2024Star
WorldDreamer: Towards general world models for video generationarXiv 2024Star
WorldGPT: a Sora-inspired video AI agentarXiv 2024β€”

4.2 World Model as Embodied Environment

TitlePub. & DateCode / Project URL
Holodeck: Language guided generation of 3d embodied ai environmentsCVPR 2024Star
GRUtopia: Dream General Robots in a City at ScalearXiv 2024Star
Anyhome: Open-vocabulary generation of structured and textured 3d homesECCV 2024Star
LEGENT: Open Platform for Embodied AgentsarXiv 2024Star
UrbanWorld: An Urban World Model for 3D City GenerationarXiv 2024Star
MetaUrban: An Embodied AI Simulation Platform for Urban MicromobilityICLR 2025Star
Minedojo: Building open-ended embodied agents with internet-scale knowledgeNeurIPS 2022Star
UniSim: Learning Interactive Real-World SimulatorsICLR 2024Website
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City EnvironmentarXiv 2024Star
Empowering World Models with Reflection for Embodied Video PredictionICML 2025β€”
Streetscapes: Large-scale consistent street view generation using autoregressive video diffusionSIGGRAPH 2024β€”
AVID: Adapting Video Diffusion Models to World ModelsarXiv 2024Star
Pandora: Towards General World Model with Natural Language Actions and Video StatesarXiv 2024Star
RoboScape: Physics-informed Embodied World ModelarXiv 2025Star
TesserAct: Learning 4D Embodied World ModelsarXiv 2025Star

5 Applications of World Models

5.1 Game Intelligence

TitlePub. & DateCode / Project URL
World and Human Action Models towards Gameplay IdeationNature 2025β€”
GameFactory: Creating New Games with Generative Interactive VideosICCV 2025Star
Unbounded: A Generative Infinite Game of Character Life SimulationCVPR 2025Website
GameGen-𝕏: Interactive Open-world Game Video GenerationICLR 2025Star
Diffusion Models Are Real-Time Game EnginesICLR 2025Website
Exploration-Driven Generative Interactive EnvironmentsICLR 2025Star
Matrix-Game: Interactive World Foundation ModelarXiv 2025Star
Mineworld: a Real-time and Open-source Interactive World Model on MinecraftarXiv 2025Star
Model as a Game: On Numerical and Spatial Consistency for Generative GamesarXiv 2025β€”

5.2 Embodied Intelligence

TitlePub. & DateCode / Project URL
OpenEQA: Embodied Question Answering in the Era of Foundation ModelsCVPR 2024Star
iVideoGPT: Interactive VideoGPTs are Scalable World ModelsNeurIPS 2024Star
IRASim: A Fine-Grained World Model for Robot ManipulationICCV 2025Star
RoboScape: Physics-informed Embodied World ModelarXiv 2025Star
TesserAct: Learning 4D Embodied World ModelsarXiv 2025Star
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningarXiv 2025Star
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsICML 2025Star
DreamGen: Unlocking Generalization in Robot Learning through Video World ModelsarXiv 2025Star
EnerVerse: Envisioning Embodied Future Space for Robotics ManipulationarXiv 2025Website
EnerVerse-AC: Envisioning Embodied Environments with Action ConditionarXiv 2025Star
Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationarXiv 2025Star
Vidar: Embodied Video Diffusion Model for Generalist Bimanual ManipulationarXiv 2025β€”
WorldVLA: Towards Autoregressive Action World ModelarXiv 2025Star
ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World ModelarXiv 2025Star
ORV: 4D Occupancy-centric Robot Video GenerationarXiv 2025Star
GWM: Towards Scalable Gaussian World Models for Robotic ManipulationICCV 2025Website
WorldEval: World Model as Real-World Robot Policies EvaluatorarXiv 2025Star

5.3 Urban Intelligence

Autonomous Driving

TitlePub. & DateCode / Project URL
Video generation models as world simulatorsOpenAI Research (2024)Website
GPT-4 technical reportarXiv 2023β€”
Visual Instruction TuningNeurIPS 2023Star
World models for autonomous driving: An initial surveyIEEE T-IV 2024β€”
Waymax: An accelerated, data-driven simulator for large-scale autonomous driving researcharXiv 2023Star
Planning-oriented autonomous drivingCVPR 2023Star
A survey on trajectory-prediction methods for autonomous drivingIEEE T-IV 2022β€”
BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformersECCV 2022Star
Transfusion: Robust lidar-camera fusion for 3D object detection with transformersCVPR 2022Star
YOLOP: You only look once for panoptic driving perceptionMIR 2022Star
Wayformer: Motion forecasting via simple & efficient attention networksICRA 2023β€”
Motion Transformer with Global Intention Localization and Local Movement RefinementNeurIPS 2022Star
Query-Centric Trajectory PredictionCVPR 2023Star
HPTR: Real-time motion prediction via heterogeneous polyline transformer with relative pose encodingNeurIPS 2023Star
MotionDiffuser: Controllable multi-agent motion prediction using diffusionCVPR 2023β€”
Tokenize the world into object-level knowledge to address long-tail events in autonomous drivingarXiv 2024β€”
GAIA-1: A generative world model for autonomous drivingarXiv 2023Website
DriveDreamer: Towards real-world-driven world models for autonomous drivingECCV 2024Star
Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingCVPR 2024Star
OccWorld: Learning a 3D occupancy world model for autonomous drivingECCV 2024Star
OccSora: 4D occupancy generation models as world simulators for autonomous drivingarXiv 2024Star
Vista: A generalizable driving world model with high fidelity and versatile controllabilityNeurIPS 2024Star
Copilot4D: Learning unsupervised world models for autonomous driving via discrete diffusionICLR 2024β€”
MUVO: A multimodal generative world model for autonomous driving with geometric representationsIEEE T-IV 2025Star
UniWorld: Autonomous driving pre-training via world modelsarXiv 2023Star
MetaUrban: A simulation platform for embodied AI in urban spacesICLR 2025Star
UrbanWorld: An urban world model for 3D city generationarXiv 2024Star
Streetscapes: Large-scale consistent street view generation using autoregressive video diffusionSIGGRAPH 2024Website

Autonomous Logistics & Urban Analytics

TitlePub. & DateCode / Project URL
Navigation World ModelsCVPR 2025Star
Towards Autonomous Micromobility through Scalable Urban SimulationCVPR 2025Star
Vid2Sim: Realistic and Interactive Simulation from Video for Urban NavigationCVPR 2025Star
CityWalker: Learning Embodied Urban Navigation from Web-Scale VideosCVPR 2025Star
AirScape: An Aerial Generative World Model with Motion Controllability (AirScape)ACM MM 2025Website
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global MemoryACL 2025Star
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City SpaceEMNLP 2025Star
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban SpacesACL 2025Star
GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsICLR 2024Star
CityGPT: Empowering Urban Spatial Cognition of Large Language ModelsKDD 2025Star
UrbanLLaVA: A Multi-modal Large Language Model for Urban IntelligenceICCV 2025Star
GPS as a Control Signal for Image GenerationCVPR 2025Star
AI's Blind Spots: Geographic Knowledge and Diversity Deficit in Generated Urban ScenarioarXiv 2025β€”
AgentMove: A Large Language Model based Agentic Framework for Zero-shot Next Location PredictionNAACL 2025Star
CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility SimulationarXiv 2025Star
Open-Set Living Need Prediction with Large Language ModelsACL 2025Star

5.4 Societal Intelligence

TitlePub. & DateCode / Project URL
AgentSociety: Large-scale simulation of LLM-driven generative agentsACL 2025, COLM 2025Star
GenSim: A General Social Simulation Platform with Large Language Model based AgentsNAACL 2025Star
Simulating Human-like Daily Activities with Desire-driven AutonomyICLR 2025Star
EconAgent: Large language model-empowered agents for simulating macroeconomic activitiesACL 2024Star
Agent-Pro: Learning to evolve via policy-level reflection and optimizationACL 2024Star
Exploring collaboration mechanisms for LLM agents: A social psychology viewACL 2024Star
Cooperate or Collapse: Emergence of sustainability behaviors in a society of LLM agentsNeurIPS 2024Star
SocioDojo: Building Lifelong Analytical Agents with Real-world Text and Time SeriesICLR 2024Star
SRAP-Agent: Simulating and optimizing scarce resource allocation policy with LLM-based agentEMNLP 2024Star
Generative agents: Interactive simulacra of human behaviorUIST 2023Star
SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World UsersarXiv 2025Star
YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language ModelsarXiv 2025Star
OASIS: Open Agent Social Interaction Simulations with One Million AgentsarXiv 2024Star
Project Sid: Many-agent simulations toward AI civilizationarXiv 2024Star
Network Formation and Dynamics Among Multi-LLMsarXiv 2024Star
S3: Social-network Simulation System with Large Language Model-Empowered AgentsarXiv 2023β€”
Exploring large language models for communication games: An empirical study on werewolfarXiv 2023Star