| τ0-WM: A Unified Video-Action World Model for Robotic Manipulation | arXiv 2026 | C4 · C6 · C8 |  |
| 3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation | AAAI 2026 | C1 · C3 |  |
| ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment | arXiv 2026 | C2 · C4 |  |
| ActionParty: Multi-Subject Action Binding in Generative Video Games | arXiv 2026 | C3 · C4 |  |
| ActWorld: From Explorable to Interactive World Model via Action-Aware Memory | arXiv 2026 | C2 · C4 · C5 |  |
| Advancing Open-source World Models | arXiv 2026 | C3 · C5 · C7 |  |
| Being-H0.7: A Latent World-Action Model from Egocentric Videos | arXiv 2026 | C4 · C6 · C7 |  |
| BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks | arXiv 2026 | C1 · C3 · C4 |  |
| Comp4D: Compositional 4D Scene Generation | WACV 2026 | C1 · C7 |  |
| Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning | arXiv 2026 | C3 · C4 · C6 |  |
| CP4D: Compositional Physics-aware 4D Scene Generation | arXiv 2026 | C1 · C2 · C5 |  |
| DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory | arXiv 2026 | C1 · C5 |  |
| DiLA: Disentangled Latent Action World Models | arXiv 2026 | C4 · C5 · C7 |  |
| DisCo: World Models with Discrete Camera Motion Control | arXiv 2026 | C4 · C5 | — |
| DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos | arXiv 2026 | C3 · C4 · C7 |  |
| DreamX-World 1.0: A General-Purpose Interactive World Model | arXiv 2026 | C3 · C4 · C5 |  |
| DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving | arXiv 2026 | C4 · C7 · C8 |  |
| Factored Latent Action World Models | arXiv 2026 | C3 · C4 |  |
| Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention | arXiv 2026 | C3 · C5 | — |
| Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World Model | AAAI 2026 | C4 · C5 |  |
| FlowDreamer: A RGB-D World Model With Flow-Based Motion Representations for Robot Manipulation | ICRA 2026 | C4 · C6 |  |
| Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion | arXiv 2026 | C1 · C5 · C7 |  |
| GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents | arXiv 2026 | C3 · C8 |  |
| GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation | arXiv 2026 | C1 · C2 · C5 |  |
| Grounding World Simulation Models in a Real-World Metropolis | arXiv 2026 | C4 · C5 | — |
| Hierarchical Latent Action Model | arXiv 2026 | C4 | — |
| iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework | arXiv 2026 | C4 · C5 · C8 |  |
| Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation | arXiv 2026 | C2 · C4 · C6 |  |
| Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models | arXiv 2026 | C4 · C5 |  |
| Learning Latent Action World Models In The Wild | arXiv 2026 | C3 · C4 · C7 | — |
| LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models | arXiv 2026 | C5 · C8 |  |
| Lyra 2.0: Explorable Generative 3D Worlds | arXiv 2026 | C1 · C3 · C5 |  |
| MAGICITY4D: Controllable and Editable 4D City Scene Generation Using MLLM-Enhanced Procedural Content Generation | ICASSP 2026 | C1 · C4 |  |
| MANIPDREAMER: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance | ICASSP 2026 | C3 · C4 |  |
| Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory | arXiv 2026 | C3 · C4 · C5 |  |
| minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models | arXiv 2026 | C3 · C4 |  |
| MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold | arXiv 2026 | C1 · C3 · C5 |  |
| MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation | arXiv 2026 | C4 · C6 |  |
| Nano World Models: A Minimalist Implementation of Future Video Prediction | arXiv 2026 | C4 · C7 · C8 |  |
| NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos | arXiv 2026 | C1 · C4 · C7 |  |
| NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation | arXiv 2026 | C3 · C4 · C7 |  |
| OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving | TIP 2026 | C4 · C5 · C6 |  |
| Olaf-World: Orienting Latent Actions for Video World Modeling | arXiv 2026 | C3 · C4 · C7 |  |
| OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation | arXiv 2026 | C4 · C5 |  |
| Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models | arXiv 2026 | C3 · C8 |  |
| OpenGame: Open Agentic Coding for Games | arXiv 2026 | C8 |  |
| OrbiSim: World Models as Differentiable Physics Engines for Embodied Intelligence | arXiv 2026 | C2 · C4 · C6 |  |
| Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space | AAAI 2026 | C3 · C4 | — |
| Physics Consistent World Models via Schrödinger-Bridge Optimal Transport for Computational Imaging and 3D-Consistent Video Generations | AAAI 2026 | C2 · C5 | — |
| PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation | arXiv 2026 | C2 · C3 · C6 |  |
| Pre-Trained Video Generative Models as World Simulators | AAAI 2026 | C3 · C4 · C5 | — |
| Prisma-World: Camera-Controllable Multi-Agent Video World Model | arXiv 2026 | C1 · C4 · C5 |  |
| ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving | arXiv 2026 | C4 · C5 · C8 |  |
| RealWonder: Real-Time Physical Action-Conditioned Video Generation | arXiv 2026 | C2 · C3 |  |
| RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation | arXiv 2026 | C2 · C4 · C8 |  |
| Self-evolving World Model | Autonomous Embodied AI 2026 | C3 · C4 | — |
| Simulating the Real World: A Unified Survey of Multimodal Generative Models | TPAMI 2026 | C7 · C8 | — |
| Solaris: Building a Multiplayer Video World Model in Minecraft | arXiv 2026 | C3 · C5 |  |
| SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks | arXiv 2026 | C4 · C7 · C8 |  |
| STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation | arXiv 2026 | C4 · C6 | — |
| Towards Dynamic World Model Generation with Monocular Video | ICASSP 2026 | C1 · C4 | — |
| Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends | arXiv 2026 | C3 · C4 · C8 |  |
| UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving | arXiv 2026 | C3 · C4 |  |
| UWM-JEPA: Predictive World Models That Imagine in Belief Space | arXiv 2026 | C4 · C5 |  |
| VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs | arXiv 2026 | C3 · C5 · C6 |  |
| Vega: Learning to Drive with Natural Language Instructions | arXiv 2026 | C4 |  |
| VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control | arXiv 2026 | C1 · C2 · C4 |  |
| Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms | arXiv 2026 | C2 · C5 | — |
| VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models | arXiv 2026 | C4 · C5 |  |
| WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation | arXiv 2026 | C3 · C8 |  |
| What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? | TMLR 2026 | C3 · C4 |  |
| What-If World: A Causal Benchmark for General World Models in Embodied Scenarios | arXiv 2026 | C2 · C5 · C8 | — |
| What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction | arXiv 2026 | C4 · C7 | — |
| WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform | arXiv 2026 | C6 · C7 · C8 |  |
| WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models | arXiv 2026 | C2 · C8 | — |
| WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models | arXiv 2026 | C3 · C4 · C5 |  |
| World model-based long-tail and scenario-specific generation for autonomous driving | JICV 2026 | C3 · C7 | — |
| World Model for Robot Learning: A Comprehensive Survey | arXiv 2026 | C8 |  |
| World Models for Robotic Manipulation: A Survey | arXiv 2026 | C4 · C8 | — |
| Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving | arXiv 2026 | C1 · C4 · C6 |  |
| A Comprehensive Survey on World Models for Embodied AI | arXiv 2025 | C2 · C5 · C8 |  |
| AdaWorld: Learning Adaptable World Models with Latent Actions | ICML 2025 | C3 · C4 |  |
| AETHER: Geometric-Aware Unified World Modeling | ICCV 2025 | C1 · C4 · C7 |  |
| A Survey: Learning Embodied Intelligence from Physical Simulators and World Models | arXiv 2025 | C2 · C3 · C8 |  |
| A Survey of World Models for Autonomous Driving | arXiv 2025 | C8 | — |
| ChronoDreamer: Action-Conditioned World Model as an Online Simulator for Robotic Planning | arXiv 2025 | C2 · C3 · C6 |  |
| Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models | arXiv 2025 | C4 · C7 |  |
| Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control | arXiv 2025 | C4 · C7 |  |
| Cosmos World Foundation Model Platform for Physical AI | arXiv 2025 | C1 · C7 |  |
| Ctrl-World: A Controllable Generative World Model for Robot Manipulation | arXiv 2025 | C4 · C5 · C8 |  |
| Diffusion Models Are Real-Time Game Engines | ICLR 2025 | C3 · C4 · C5 |  |
| DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation | ICCV 2025 | C4 · C6 |  |
| DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation | AAAI 2025 | C4 · C7 · C8 |  |
| DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive Transformers | ICCV 2025 | C4 · C5 · C8 |  |
| Embodied World Models Emerge from Navigational Task in Open-Ended Environments | arXiv 2025 | C3 · C6 | — |
| ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction | arXiv 2025 | C3 · C6 · C8 |  |
| Free4D: Tuning-Free 4D Scene Generation with Spatial-Temporal Consistency | ICCV 2025 | C1 · C4 |  |
| From 2D to 3D Cognition: A Brief Survey of General World Models | arXiv 2025 | C1 · C2 · C3 | — |
| GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation | arXiv 2025 | C1 · C4 · C6 |  |
| GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving | arXiv 2025 | C4 · C5 · C7 |  |
| GameFactory: Creating New Games with Generative Interactive Videos | ICCV 2025 | C3 · C4 · C7 |  |
| GameGen-X: Interactive Open-world Game Video Generation | ICLR 2025 | C1 · C4 · C7 |  |
| GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation | arXiv 2025 | C4 · C6 |  |
| GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control | CVPR 2025 | C4 · C6 · C8 |  |
| Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report | arXiv 2025 | C4 · C6 · C8 |  |
| Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation | arXiv 2025 | C2 · C4 · C8 |  |
| GigaWorld-0: World Models as Data Engine to Empower Embodied AI | arXiv 2025 | C1 · C2 · C7 |  |
| GWM: Towards Scalable Gaussian World Models for Robotic Manipulation | ICCV 2025 | C1 · C2 · C4 |  |
| HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation | ACM MM 2025 | C1 · C7 |  |
| Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model | arXiv 2025 | C3 · C4 · C8 |  |
| Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation | arXiv 2025 | C3 · C5 · C8 |  |
| IRASim: A Fine-Grained World Model for Robot Manipulation | ICCV 2025 | C2 · C3 · C4 |  |
| Latent Action World Models for Control with Unlabeled Trajectories | arXiv 2025 | C4 · C7 | — |
| Learning 3D Persistent Embodied World Models | NeurIPS 2025 | C5 · C6 |  |
| Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression | arXiv 2025 | C3 · C4 · C8 |  |
| Mastering diverse control tasks through world models | Nature 2025 | C5 · C7 |  |
| Matrix-game 2.0: An open-source real-time and streaming interactive world model | arXiv 2025 | C3 · C4 · C5 |  |
| Matrix-Game: Interactive World Foundation Model | arXiv 2025 | C3 · C4 · C8 |  |
| Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation | arXiv 2025 | C5 · C8 | — |
| MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving | arXiv 2025 | C4 · C5 |  |
| MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft | arXiv 2025 | C3 · C8 |  |
| Motus: A Unified Latent Action World Model | arXiv 2025 | C4 · C7 |  |
| MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations | IEEE IV 2025 | C4 · C6 |  |
| Navigation World Models | CVPR 2025 | C4 · C5 · C7 | — |
| PAN: A World Model for General, Interactable, and Long-Horizon World Simulation | arXiv 2025 | C3 · C4 · C5 | — |
| PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation | RSS 2025 | C1 · C2 · C7 |  |
| PlayerOne: Egocentric World Simulator | NeurIPS 2025 | C1 · C4 · C5 |  |
| Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins | RSS 2025 | C2 · C6 |  |
| ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos | arXiv 2025 | C4 · C5 · C6 |  |
| ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction | arXiv 2025 | C2 · C3 · C8 |  |
| RELIC: Interactive Video World Model with Long-Horizon Memory | arXiv 2025 | C3 · C4 · C5 |  |
| RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies | arXiv 2025 | C3 · C8 |  |
| RoboScape: Physics-informed Embodied World Model | arXiv 2025 | C1 · C2 · C6 |  |
| RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation | arXiv 2025 | C2 · C4 · C6 |  |
| Simulating the Visual World with Artificial Intelligence: A Roadmap | arXiv 2025 | C2 · C3 · C8 |  |
| SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model | IROS 2025 | C6 · C8 |  |
| STORM: Search-Guided Generative World Models for Robotic Manipulation | arXiv 2025 | C4 · C6 · C8 | — |
| TesserAct: Learning 4D Embodied World Models | ICCV 2025 | C1 · C6 |  |
| The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control | NeurIPS 2025 | C3 · C4 · C5 |  |
| The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey | arXiv 2025 | C8 |  |
| Understanding World or Predicting Future? A Comprehensive Survey of World Models | ACM Computing Surveys 2025 | C8 |  |
| Unified Video Action Model | RSS 2025 | C3 · C6 |  |
| UniScene: Unified Occupancy-centric Driving Scene Generation | CVPR 2025 | C1 · C7 · C8 |  |
| Video World Models with Long-term Spatial Memory | arXiv 2025 | C1 · C5 |  |
| V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning | arXiv 2025 | C4 · C5 · C7 |  |
| VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory | ICCV 2025 | C4 · C5 |  |
| Whole-Body Conditioned Egocentric Video Prediction | NeurIPS 2025 | C2 · C4 · C8 | — |
| WonderWorld: Interactive 3D Scene Generation from a Single Image | CVPR 2025 | C1 · C3 · C5 |  |
| WorldGen: From Text to Traversable and Interactive 3D Worlds | arXiv 2025 | C1 · C3 · C4 |  |
| World-in-World: World Models in a Closed-Loop World | arXiv 2025 | C3 · C4 · C8 |  |
| World model-based end-to-end scene generation for accident anticipation in autonomous driving | Communications Engineering 2025 | C4 · C7 | — |
| World Model Enhanced Embodied Intelligence for Deformable Object Manipulation of Dynamic Targets | CINTI 2025 | C2 · C5 | — |
| WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling | arXiv 2025 | C4 · C5 · C8 | — |
| WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling | arXiv 2025 | C3 · C4 · C5 |  |
| WorldScore: A Unified Evaluation Benchmark for World Generation | ICCV 2025 | C4 · C8 |  |
| WorldVLA: Towards Autoregressive Action World Model | arXiv 2025 | C2 · C4 · C5 |  |
| WoW: Towards a World omniscient World model Through Embodied Interaction | arXiv 2025 | C2 · C5 · C8 |  |
| Yan: Foundational Interactive Video Generation | arXiv 2025 | C3 · C4 · C5 |  |
| Yume-1.5: A Text-Controlled Interactive World Generation Model | arXiv 2025 | C3 · C4 · C5 |  |
| Yume: An Interactive World Generation Model | arXiv 2025 | C3 · C4 · C5 |  |
| A Unified Approach for Text-and Image-Guided 4D Scene Generation | CVPR 2024 | C1 · C7 | — |
| AVID: Adapting Video Diffusion Models to World Models | arXiv 2024 | C3 · C4 | — |
| Diffusion for World Modeling: Visual Details Matter in Atari | NeurIPS 2024 | C3 · C5 · C6 |  |
| Doe-1: Closed-Loop Autonomous Driving with Large World Model | arXiv 2024 | C3 · C4 · C6 |  |
| DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving | ECCV 2024 | C1 · C4 |  |
| DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model | arXiv 2024 | C3 · C7 · C8 |  |
| Driving Into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving | CVPR 2024 | C4 · C5 · C7 |  |
| DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT | arXiv 2024 | C4 · C5 |  |
| Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey | arXiv 2024 | C8 | — |
| GenAD: Generalized Predictive Model for Autonomous Driving | arXiv 2024 | C3 · C4 · C7 |  |
| Generative World Explorer | arXiv 2024 | C3 · C5 · C6 |  |
| Genie: Generative Interactive Environments | ICML 2024 | C1 · C3 · C4 |  |
| Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond | arXiv 2024 | C2 · C5 |  |
| iVideoGPT: Interactive VideoGPTs are Scalable World Models | NeurIPS 2024 | C3 · C4 · C6 |  |
| Learning Interactive Real-World Simulators | ICLR 2024 | C3 · C4 · C7 |  |
| Learning to Act from Actionless Videos through Dense Correspondences | ICLR 2024 | C4 · C6 · C7 |  |
| MagicDrive: Street View Generation with Diverse 3D Geometry Control | ICLR 2024 | C1 · C4 · C5 |  |
| MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation | ICRA 2024 | C3 · C6 |  |
| OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving | arXiv 2024 | C3 · C4 · C6 | — |
| Pandora: Towards General World Model with Natural Language Actions and Video States | arXiv 2024 | C3 · C4 · C7 |  |
| Peekaboo: Interactive Video Generation via Masked-Diffusion | CVPR 2024 | C3 · C4 |  |
| Playable Game Generation | arXiv 2024 | C3 · C5 · C8 |  |
| RoboDreamer: Learning Compositional World Models for Robot Imagination | ICML 2024 | C4 · C5 · C7 |  |
| ADriver-I: A General World Model for Autonomous Driving | arXiv 2023 | C3 · C4 · C5 | — |
| GAIA-1: A Generative World Model for Autonomous Driving | arXiv 2023 | C4 · C5 · C7 |  |
| Learning Universal Policies via Text-Guided Video Generation | NeurIPS 2023 | C4 · C6 · C7 |  |
| Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions | ICLR 2023 | C4 · C5 | — |
| TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction | ICRA 2023 | C3 · C4 |  |
| Transformers are Sample-Efficient World Models | ICLR 2023 | C5 · C6 |  |
| DayDreamer: World Models for Physical Robot Learning | CoRL 2022 | C2 · C3 |  |
| Planning with Diffusion for Flexible Behavior Synthesis | ICML 2022 | C4 · C5 |  |
| Playable Environments: Video Manipulation in Space and Time | CVPR 2022 | C1 · C4 · C5 |  |
| Mastering Atari with Discrete World Models | ICLR 2021 | C5 · C6 | — |
| Playable Video Generation | CVPR 2021 | C3 · C4 |  |
| VideoGPT: Video Generation using VQ-VAE and Transformers | arXiv 2021 | C1 · C7 |  |
| Dream to Control: Learning Behaviors by Latent Imagination | ICLR 2020 | C5 · C6 | — |
| Mastering Atari, Go, chess and shogi by planning with a learned model | Nature 2020 | C5 · C6 | — |
| Model Based Reinforcement Learning for Atari | ICLR 2020 | C5 · C6 | — |
| Learning Latent Dynamics for Planning from Pixels | ICML 2019 | C5 · C6 | — |
| World Models | arXiv 2018 | C3 · C5 |  |