Awesome Agentic World Modeling

August 22, 2026 · View on GitHub

Awesome PR's Welcome Visitors arXiv Hugging Face Papers

Awesome Agentic World Modeling

This repository accompanies the position paper "Quo Vadis, World Modeling? Towards Interactive World Proxies for Continually Improving Agents," and curates the works it reviews. The survey argues for a shift from a World Model that passively predicts the next physical state toward an Agent-Centric World Proxy: an environment-grounded interface that returns the information transition an agent needs (a future state, rendered view, execution result, retrieved memory or skill, or verdict on a plan) to plan, learn, and continually improve.

The taxonomy has two orthogonal axes:

  • Six proxy functions (what feedback the proxy returns): Dynamics, Spatial, Execution, Memory / Experience, Skill, and Reward / Verification.
  • Three empowerment levels (how the proxy improves the agent): L1 inference-time guidance, L2 training-time optimization, and L3 Agent-Proxy co-evolution.

Call for Contributions

Agentic world modeling moves fast, and the boundary between a passive world model and an agent-facing world proxy is still being drawn. This list aims to be a living, community-maintained map of that space, and contributions of new work or corrections are very welcome.

To contribute, open an issue or a pull request:

  • Where: add each entry to the section matching its proxy function (or the closest one), and keep every table in chronological order.
  • Format: follow the existing row layout: a short model name in Model (or - if none), the title with an arXiv badge in Paper, and the abbreviated venue with year in Venue.
  • Verify: check the title, authors, arXiv ID, and venue against the primary source before adding, and prefer the formally published version whenever one exists.
  • Fix: corrections to venues, broken links, or mis-categorized entries are equally welcome; flag them in an issue or fix them directly in a pull request.

Table of Contents

Background

Classical World Models & Model-Based RL

Foundations: state-transition world models and model-based RL that predict the next physical state given a state and an action.

ModelPaperVenue
-Dyna, an integrated architecture for learning, planning, and reactingSIGART 1991
-arXiv
Recurrent World Models Facilitate Policy Evolution
NeurIPS 2018
-arXiv
When to Trust Your Model: Model-Based Policy Optimization
NeurIPS 2019
-arXiv
Dream to Control: Learning Behaviors by Latent Imagination
ICLR 2020
-arXiv
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Nature 2020
-A Path Towards Autonomous Machine IntelligenceOpenReview preprint 2022
-arXiv
Model-Based Reinforcement Learning: A Survey
FTML 2023
-arXiv
Mastering diverse control tasks through world models
Nature 2025

Continual Improvement & Interaction Cost

Why real-environment interaction alone cannot carry continual improvement: cost, safety, latency, and the need for scalable feedback.

ModelPaperVenue
MuJoCoMuJoCo: A Physics Engine for Model-Based ControlIROS 2012
-arXiv
Isaac Gym: High Performance GPU-Based Physics Simulation for Robot Learning
NeurIPS 2021
DayDreamerarXiv
DayDreamer: World Models for Physical Robot Learning
CoRL 2023
CodeItarXiv
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
ICML 2024
VoyagerarXiv
Voyager: An Open-Ended Embodied Agent with Large Language Models
TMLR 2024
OSWorldarXiv
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
NeurIPS 2024
WebArenaarXiv
WebArena: A Realistic Web Environment for Building Autonomous Agents
ICLR 2024
CASCADEarXiv
CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution
arXiv 2025
-arXiv
Continual Learning and Catastrophic Forgetting
Learning and Memory: A Comprehensive Reference 2025
-arXiv
Agent Learning via Early Experience
ICML 2026

The Six Proxy Functions

Each proxy function answers a different question for the agent, but all return agent-usable feedback.

Dynamics Proxy

How would the environment change if the agent executed a given action? Video, driving, game, and model-based dynamics predictors.

ModelPaperVenue
-Learning Latent Dynamics for Planning from PixelsICML 2019
FitVidarXiv
FitVid: Overfitting in Pixel-Level Video Prediction
arXiv 2021
VideoGPTarXiv
VideoGPT: Video Generation using VQ-VAE and Transformers
arXiv 2021
-arXiv
Video Diffusion Models
NeurIPS 2022
-arXiv
Diffusion Models for Video Prediction and Infilling
TMLR 2022
MCVDarXiv
MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation
NeurIPS 2022
GAIA-1arXiv
GAIA-1: A Generative World Model for Autonomous Driving
arXiv 2023
-arXiv
Transformers Are Sample-Efficient World Models
ICLR 2023
STORMSTORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningNeurIPS 2023
-arXiv
Diffusion for World Modeling: Visual Details Matter in Atari
NeurIPS 2024
-Video generation models as world simulators2024
GeniearXiv
Genie: Generative Interactive Environments
ICML 2024
VistaarXiv
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
NeurIPS 2024
TD-MPC2arXiv
TD-MPC2: Scalable, Robust World Models for Continuous Control
ICLR 2024
RoboCasaRoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist RobotsRSS 2024
DriveDreamerarXiv
DriveDreamer: Towards Real-World-Driven World Models for Autonomous Driving
ECCV 2024
iVideoGPTiVideoGPT: Interactive VideoGPTs are Scalable World ModelsNeurIPS 2024
-arXiv
Learning Interactive Real-World Simulators
ICLR 2024
DynamicCityarXiv
DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes
ICLR 2025
GameGen-XarXiv
GameGen-X: Interactive Open-world Game Video Generation
ICLR 2025
-Genie 3: A New Frontier for World Models2025
-arXiv
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
arXiv 2025
-arXiv
World Modelling Improves Language Model Agents
arXiv 2025
MineWorldarXiv
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
arXiv 2025
MobileWorldBencharXiv
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
arXiv 2025
-World Model on Million-Length Video And Language With Blockwise RingAttentionICLR 2025
VisionarXiv
Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models
arXiv 2025
-arXiv
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
arXiv 2025
-Diffusion Models Are Real-Time Game EnginesICLR 2025
MarbleMarble: A Multimodal World Model2025
-arXiv
Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving
AAAI 2025
Matrix-GamearXiv
Matrix-Game: Interactive World Foundation Model
arXiv 2025
LanguagearXiv
Language-Conditioned World Modeling for Visual Navigation
arXiv 2026
-arXiv
Advancing Open-source World Models
arXiv 2026
NavThinkerarXiv
NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation
arXiv 2026
LiDARCrafterarXiv
LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
AAAI 2026
SelfarXiv
Self-Execution Simulation Improves Coding Models
arXiv 2026
-arXiv
Is Your Driving World Model an All-Around Player?
CVPR 2026
U4DarXiv
U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
CVPR 2026
AD-R1arXiv
AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
CVPR 2026
SPIRALarXiv
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
arXiv 2026
FlowWAMarXiv
FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
arXiv 2026
LeapBotarXiv
LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
arXiv 2026
EnfoldarXiv
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
arXiv 2026
-arXiv
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
arXiv 2026
AutoarXiv
Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving
arXiv 2026
FasterarXiv
Faster-WAM: Do World Action Models Need Deep Action Modules?
arXiv 2026
DF$^3$arXiv
DF3^3: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation
arXiv 2026
SGarXiv
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
arXiv 2026
LAWMarXiv
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
arXiv 2026
XEWorldarXiv
XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
arXiv 2026
AdaptivearXiv
Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features
arXiv 2026
GWMarXiv
GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning
arXiv 2026
4DarXiv
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
arXiv 2026
WAarXiv
WA-SpecDec: World-Aware Speculative Decoding for Vision-Language-Action Models
arXiv 2026
JEPAarXiv
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
arXiv 2026
-arXiv
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
arXiv 2026
SLIMarXiv
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
arXiv 2026
EnergyarXiv
Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning
arXiv 2026
DreamerarXiv
Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving
arXiv 2026
BrainWAMarXiv
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
arXiv 2026
DreamXarXiv
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
arXiv 2026
ForgeWMarXiv
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
arXiv 2026
-arXiv
Q-Learning With World Models
arXiv 2026
HydraarXiv
Hydra-0: Action Flow for Generalist World Modeling and Control
arXiv 2026
GigaBrainarXiv
GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
arXiv 2026
-arXiv
Reinforced Planning with Latent World Models
arXiv 2026
DAarXiv
DA-WAM: Decision-Aligned Future Latents for Driving World Models
arXiv 2026
-arXiv
Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
arXiv 2026
-arXiv
Overcoming Statistical Bias in Action-Controllable World Models
arXiv 2026
-arXiv
Mitigating Compounding Error via Video Representation Regularization
arXiv 2026
ODEWorldarXiv
ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
arXiv 2026
GeniWorldarXiv
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
arXiv 2026
ContactFlowarXiv
ContactFlow: A video action conditioning that transfers across embodiments
arXiv 2026
EmbodiedVAEarXiv
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
arXiv 2026
-arXiv
τ0τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
arXiv 2026
DeptharXiv
Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data
arXiv 2026
-arXiv
Is Forward Prediction Enough? Physical State Grounding for JEPA World Models
arXiv 2026
XiaomiarXiv
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
arXiv 2026
MiniWorldarXiv
MiniWorld: Democratizing the Training of Video World Models from Scratch
arXiv 2026
BWMarXiv
BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
arXiv 2026
S2arXiv
S2-HWM: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation
arXiv 2026
-arXiv
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
arXiv 2026
PopulationarXiv
Population-Scalable Multi-Agent World Modeling
arXiv 2026
DreamQASarXiv
DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
arXiv 2026
FedWorldarXiv
FedWorld: Scope-Aware Federation of Agent World Models
arXiv 2026
-arXiv
Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control
arXiv 2026
DreamTrajectoryarXiv
DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation
arXiv 2026
EndoWAMarXiv
EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation
arXiv 2026
-arXiv
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
arXiv 2026

Spatial Proxy

What would the agent observe from another viewpoint or position? NeRF/Gaussian rendering, 3D/4D reconstruction and generation, and visual imagination for navigation.

ModelPaperVenue
NeRFarXiv
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
ECCV 2020
Mip-NeRFarXiv
Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
ICCV 2021
PathDreamerarXiv
PathDreamer: A World Model for Indoor Navigation
ICCV 2021
MiparXiv
Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
CVPR 2022
SceneDreamerarXiv
SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections
TPAMI 2023
SceneScapearXiv
SceneScape: Text-Driven Consistent Scene Generation
NeurIPS 2023
Text2RoomarXiv
Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
ICCV 2023
VoxPoserarXiv
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
CoRL 2023
-arXiv
3D Gaussian Splatting for Real-Time Radiance Field Rendering
TOG 2023
NerfstudioarXiv
Nerfstudio: A Modular Framework for Neural Radiance Field Development
SIGGRAPH Asia 2023
-arXiv
A Survey on 3D Gaussian Splatting
CSUR 2024
-arXiv
3D Gaussian as a New Era: A Survey
TVCG 2024
-arXiv
Grounding Image Matching in 3D with MASt3R
ECCV 2024
DreamScenearXiv
DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
ECCV 2024
Scaffold-GSarXiv
Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
CVPR 2024
DreamforgearXiv
Dreamforge: Motion-aware autoregressive video generation for multi-view driving scenes
arXiv 2024
DriveWorldDriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous DrivingCVPR 2024
DimensionXarXiv
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
arXiv 2024
ManiSkill3arXiv
ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
arXiv 2024
DUSt3RarXiv
DUSt3R: Geometric 3D Vision Made Easy
CVPR 2024
-arXiv
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
CVPR 2024
CityDreamerarXiv
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
CVPR 2024
Mip-SplattingarXiv
Mip-Splatting: Alias-Free 3D Gaussian Splatting
CVPR 2024
Text2NeRFarXiv
Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
TVCG 2024
Copilot4DCopilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionICLR 2024
OccWorldarXiv
OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
ECCV 2024
ReCamMasterarXiv
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
arXiv 2025
-arXiv
Navigation world models
CVPR 2025
PhysX-AnythingarXiv
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
arXiv 2025
GAFarXiv
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
arXiv 2025
ParticleFormerarXiv
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
arXiv 2025
-arXiv
3D and 4D World Modeling: A Survey
arXiv 2025
GWMGWM: Towards Scalable Gaussian World Models for Robotic ManipulationICCV 2025
MindCubearXiv
MindCube: Spatial Mental Modeling from Limited Views
arXiv 2025
GAIA-2arXiv
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
arXiv 2025
-arXiv
3D Reconstruction with Spatial Memory
3DV 2025
-arXiv
Continuous 3D Perception Model with Persistent State
arXiv 2025
Fast3RarXiv
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
arXiv 2025
WonderWorldarXiv
WonderWorld: Interactive 3D Scene Generation from a Single Image
CVPR 2025
-arXiv
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
IJCAI 2025
-arXiv
Occupancy World Model for Robots
arXiv 2025
TesserActarXiv
TesserAct: Learning 4D Embodied World Models
arXiv 2025
AetherarXiv
Aether: Geometric-Aware Unified World Modeling
ICCV 2025
OccuBencharXiv
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
arXiv 2026
PointWorldarXiv
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
arXiv 2026
CounterScenearXiv
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
arXiv 2026
MVISTA-4DarXiv
MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
arXiv 2026
X-scenearXiv
X-scene: Large-scale driving scene generation with high fidelity and flexible controllability
NeurIPS 2026
Code2WorldarXiv
Code2World: A GUI World Model via Renderable Code Generation
arXiv 2026
UniNavarXiv
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
arXiv 2026
-arXiv
SC2^{2}-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments
arXiv 2026
-arXiv
Latent World Models with Monotone Planning Costs for Image-Goal Navigation
arXiv 2026
OrbitarXiv
Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents
arXiv 2026
WorldarXiv
World-Model-Grounded LLM Planning for AUV and ASV Navigation Near Offshore Wind Farms
arXiv 2026

Execution Proxy

What would the digital environment return after an operation? Web, GUI, code, shell, and API/tool execution predictors.

ModelPaperVenue
RT-2arXiv
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
CoRL 2023
-arXiv
Code as Policies: Language Model Programs for Embodied Control
ICRA 2023
ToolformerarXiv
Toolformer: Language Models Can Teach Themselves to Use Tools
NeurIPS 2023
-arXiv
Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search
NeurIPS 2024
AgentBencharXiv
AgentBench: Evaluating LLMs as Agents
ICLR 2024
-arXiv
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
NeurIPS 2024
Math-ShepherdarXiv
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations
ACL 2024
-arXiv
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
ICLR 2025
CWMarXiv
CWM: An Open-Weights LLM for Research on Code Generation with World Models
arXiv 2025
-arXiv
Web World Models
arXiv 2025
WebSynthesisarXiv
WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
arXiv 2025
-arXiv
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
TMLR 2025
MegaSaMarXiv
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
CVPR 2025
ToolSandboxarXiv
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
NAACL 2025
ViMoarXiv
ViMo: A Generative Visual GUI World Model for App Agents
arXiv 2025
-arXiv
Cosmos World Foundation Model Platform for Physical AI
arXiv 2025
GTMarXiv
GTM: Simulating the World of Tools for AI Agents
arXiv 2025
-arXiv
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
arXiv 2025
GameFactoryGameFactory: Creating New Games with Generative Interactive VideosICCV 2025
ComActarXiv
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
arXiv 2026
MobileDreamerarXiv
MobileDreamer: Generative Sketch World Model for GUI Agent
arXiv 2026
ComputerarXiv
Computer-Using World Model
arXiv 2026
MCP-CosmosarXiv
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
arXiv 2026
-arXiv
Generative Visual Code Mobile World Models
arXiv 2026
IterCADarXiv
IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing
arXiv 2026
-arXiv
Debugging Code World Models
arXiv 2026
NeuralOSarXiv
NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
ICLR 2026
World-ModelarXiv
World-Model-Augmented Web Agents with Action Correction
arXiv 2026
WebWorldarXiv
WebWorld: A Large-Scale World Model for Web Agent Training
arXiv 2026
AppDeltaWorldarXiv
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
arXiv 2026
WebRiderarXiv
WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance
arXiv 2026
TwinarXiv
Twin: Playing an Unknown Game with a Test-Time Digital Twin
arXiv 2026

Memory, Experience & Skill Proxy

What prior experience, constraints, or reusable skills does the agent need now? Store-retrieve memory systems and reusable skill/behavior libraries.

ModelPaperVenue
-arXiv
Reasoning with Language Model is Planning with World Model
EMNLP 2023
-arXiv
Generative Agents: Interactive Simulacra of Human Behavior
UIST 2023
-arXiv
Cognitive Architectures for Language Agents
TMLR 2024
VismemarXiv
Vismem: Latent vision memory unlocks potential of vision-language models
arXiv 2025
MemHarnessarXiv
MemHarness: Memory Is Reconstructed, Not Replayed
arXiv 2026
MemWMarXiv
MemWM: Memory-Augmented Text-Based World Model
arXiv 2026
-arXiv
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
arXiv 2026
-arXiv
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
arXiv 2026
PBDarXiv
PBD-AG: Persistent Baseline-Delta Active Graphs with Uncertainty-Aware Inspection for Long-Horizon Service Robots
arXiv 2026
LUCIDarXiv
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
arXiv 2026
-arXiv
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
arXiv 2026
-arXiv
Progressive Experience Fusion for Multi-Task World Model Control in Endovascular Navigation
arXiv 2026

Reward / Verification Proxy

Is the agent's behavior correct, safe, and feasible, and how should it improve? Reward models, verifiers, critics, preference models, and LLM-as-judge.

ModelPaperVenue
-arXiv
Deep Reinforcement Learning from Human Preferences
NeurIPS 2017
FinearXiv
Fine-Tuning Language Models from Human Preferences
arXiv 2019
-arXiv
Training Verifiers to Solve Math Word Problems
arXiv 2021
-arXiv
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
NeurIPS 2023
-arXiv
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
NeurIPS 2023
-arXiv
A General Theoretical Paradigm to Understand Learning from Human Preferences
AISTATS 2024
-arXiv
A Survey on LLM-as-a-Judge
arXiv 2024
ORPOarXiv
ORPO: Monolithic Preference Optimization without Reference Model
EMNLP 2024
-arXiv
Let's Verify Step by Step
ICLR 2024
-arXiv
LLM Critics Help Catch LLM Bugs
arXiv 2024
SimPOarXiv
SimPO: Simple Preference Optimization with a Reference-Free Reward
NeurIPS 2024
WorldSimBencharXiv
WorldSimBench: Towards Video Generation Models as World Simulators
arXiv 2024
WorldModelBencharXiv
WorldModelBench: Judging Video Generation Models As World Models
arXiv 2025
ComputerarXiv
Computer-Use Agents as Judges for Generative User Interface
arXiv 2025
Critic-VarXiv
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
CVPR 2025
WCMarXiv
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
arXiv 2026
CheckVLAarXiv
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
arXiv 2026
ContactGuardarXiv
ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models
arXiv 2026
-arXiv
Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling
arXiv 2026
-arXiv
Calibrated Predictive Safety for Heterogeneous Robots: An Action-Conditioned JEPA Framework with Model-Based Safety Shields
arXiv 2026
OntologyarXiv
Ontology-Grounded World Models for Failure Diagnosis and Closed-Loop Repair in Physical AI Systems
arXiv 2026
StructRewardarXiv
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning
arXiv 2026
-arXiv
Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction
arXiv 2026
-arXiv
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
arXiv 2026
ARCarXiv
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
arXiv 2026

Empowerment Levels (L1 / L2 / L3)

How proxies empower agents: L1 inference-time guidance (advise a frozen agent), L2 training-time optimization (generate experience/signal to train the agent), and L3 Agent-Proxy co-evolution (both improve each other in a loop).

ModelPaperVenue
-arXiv
Scaling Agent Learning via Experience Synthesis
arXiv 2025
RAGENarXiv
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
arXiv 2025
VAGENarXiv
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
NeurIPS 2025
WebEvolverarXiv
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
EMNLP 2025
-arXiv
Aligning Agentic World Models via Knowledgeable Experience Learning
arXiv 2026
-arXiv
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
ICML 2026
SearchGymarXiv
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
arXiv 2026
hint$^2$arXiv
hint2^2: Hierarchical World Models for Inference-Time Temporal Logic Guidance
arXiv 2026

Evaluation, Fidelity & Safety

Benchmarks and metrics for world models, fidelity and calibration limits, and the safety/attack surface of proxy-driven learning.

ModelPaperVenue
G-EvalarXiv
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
EMNLP 2023
LengtharXiv
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
COLM 2024
PrometheusarXiv
Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models
ICLR 2024
BEHAVIOR-1KarXiv
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
arXiv 2024
-arXiv
Evaluating the World Model Implicit in a Generative Model
NeurIPS 2024
-arXiv
World Models: The Safety Perspective
arXiv 2024
-arXiv
Understanding World or Predicting Future? A Comprehensive Survey of World Models
CSUR 2025
-arXiv
How Far is Video Generation from World Model: A Physical Law Perspective
ICML 2025
-arXiv
When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
arXiv 2026
-arXiv
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
arXiv 2026
JailWAMarXiv
JailWAM: Jailbreaking World Action Models in Robot Control
arXiv 2026
-arXiv
Current Agents Fail to Leverage World Model as Tool for Foresight
arXiv 2026
CtrlAttackarXiv
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
arXiv 2026
SafeDreamarXiv
SafeDream: Safety World Model for Proactive Early Jailbreak Detection
arXiv 2026
KineBencharXiv
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
arXiv 2026
ApplearXiv
Apple-ππ: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
arXiv 2026
WorldExamarXiv
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
arXiv 2026
GAUGEarXiv
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
arXiv 2026
H2RarXiv
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
arXiv 2026
WorldSimProbearXiv
WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
arXiv 2026
VIScorearXiv
VIScore: Diagnosing Planning-Relevant Quality in Latent World Models
arXiv 2026
HarnessEvalarXiv
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
arXiv 2026
CGarXiv
CG-World: A Large-Scale World-State Dataset and Protocol for World Models
arXiv 2026
-arXiv
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
arXiv 2026
-arXiv
Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes
arXiv 2026

Surveys & Related Work

Related surveys and the paper's core position statements.

ModelPaperVenue
-arXiv
Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
arXiv 2024
-arXiv
A Survey of World Models for Autonomous Driving
arXiv 2025
-arXiv
A Comprehensive Survey on World Models for Embodied AI
arXiv 2025
-arXiv
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
arXiv 2026
-arXiv
Reinforcement Learning: From Algorithms To Foundation Models
arXiv 2026
PlanningarXiv
Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms
arXiv 2026
-arXiv
Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent
arXiv 2026

Additional References

Further supporting works cited in the survey.

ModelPaperVenue
-arXiv
Proximal Policy Optimization Algorithms
arXiv 2017
HabitatHabitat: A Platform for Embodied AI ResearchICCV 2019
-Model-Based Reinforcement Learning for AtariICLR 2020
-arXiv
Learning to Summarize from Human Feedback
NeurIPS 2020
-arXiv
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
CoRL 2020
CLIPortarXiv
CLIPort: What and Where Pathways for Robotic Manipulation
CoRL 2021
-arXiv
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
CoRL 2022
-arXiv
Constitutional AI: Harmlessness from AI Feedback
arXiv 2022
-arXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
arXiv 2022
TensoRFarXiv
TensoRF: Tensorial Radiance Fields
ECCV 2022
-arXiv
Inner Monologue: Embodied Reasoning through Planning with Language Models
CoRL 2022
-arXiv
Large Language Models are Zero-Shot Reasoners
NeurIPS 2022
-arXiv
Instant Neural Graphics Primitives with a Multiresolution Hash Encoding
TOG 2022
-arXiv
Training Language Models to Follow Instructions with Human Feedback
NeurIPS 2022
Perceiver-ActorarXiv
Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
CoRL 2022
-arXiv
Solving Math Word Problems with Process- and Outcome-Based Feedback
arXiv 2022
Chain-ofarXiv
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
NeurIPS 2022
PlenoxelsarXiv
Plenoxels: Radiance Fields without Neural Networks
CVPR 2022
SelfarXiv
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
CVPR 2023
RT-1arXiv
RT-1: Robotics Transformer for Real-World Control at Scale
RSS 2023
Self-RefinearXiv
Self-Refine: Iterative Refinement with Self-Feedback
NeurIPS 2023
MemGPTarXiv
MemGPT: Towards LLMs as Operating Systems
arXiv 2023
ReflexionarXiv
Reflexion: Language Agents with Verbal Reinforcement Learning
NeurIPS 2023
ReActarXiv
ReAct: Synergizing Reasoning and Acting in Language Models
ICLR 2023
-arXiv
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
NeurIPS 2023
RRHFarXiv
RRHF: Rank Responses to Align Language Models with Human Feedback without Tears
NeurIPS 2023
KTOarXiv
KTO: Model Alignment as Prospect Theoretic Optimization
ICML 2024
SWE-bencharXiv
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
ICLR 2024
OasisOasis: A Universe in a Transformer2024
-arXiv
Agent Planning with World Knowledge Model
NeurIPS 2024
DeepSeekMatharXiv
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
arXiv 2024
VGGTarXiv
VGGT: Visual Geometry Grounded Transformer
arXiv 2025
MonST3RarXiv
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
ICLR 2025
-arXiv
The latent space: Foundation, evolution, mechanism, ability, and outlook
arXiv 2026

Acknowledgements

This list is maintained alongside the survey. Thanks to all contributors and to the authors of the works listed here.

Citation

If you find this work useful, please consider citing:

@article{yang2026worldproxy,
  title   = {Quo Vadis, World Modeling?},
  author  = {Yu Yang and Xuemeng Yang and Licheng Wen and Lingdong Kong and Xiaobin Hu and Dongyue Lu and Wei Chow and Xiyan Huang and Yuxiang Feng and Yue Liao and Jianbiao Mei and Daocheng Fu and Rong Wu and Pinlong Cai and Ran Yi and Ying Tai and Jiangning Zhang and Botian Shi and Yong Liu and Shuicheng Yan},
  journal = {arXiv preprint arXiv:2608.02713},
  year    = {2026}
}