World Action Models: The Next Frontier in Embodied AI

July 21, 2026 · View on GitHub

World Action Models: The Next Frontier in Embodied AI

arXiv Hugging Face Project Page Leaderboard Feishu/Lark

Siyin Wang1,2,*,‡, Junhao Shi1,2,*, Zhaoyang Fu1,*, Xinzhe He1,*, Feihong Liu1,*,
Chenchen Yang1,2, Yikang Zhou2, Zhaoye Fei1, Jingjing Gong2, Jinlan Fu1
, Mike Zheng Shou3 · Xuanjing Huang1,2 · Xipeng Qiu1,2 · Yu-Gang Jiang1,†

1Fudan University, 2Shanghai Innovation Institute, 3National University of Singapore
*Equal Contribution, Project Lead, Corresponding Author

This repository accompanies our survey on World Action Models (WAMs) — the emerging paradigm that unifies predictive world modeling with action generation for embodied AI. We will keep this repo continuously updated as the field evolves.

  • 📄 The first systematic WAM survey, covering architecture taxonomy (Cascaded & Joint WAMs), training data, evaluation protocols, and world models for VLA learning.
  • 📖 Reading blogs included: we also provide a concise, structured summary blog for each paper to help you quickly grasp the key ideas, architecture, and contributions. The summarization skill used to generate them is also open-sourced in this repository (Check Paper2Blog).
  • 🤝 Community-driven: found a missing paper or have a suggestion? Feel free to open an issue or submit a pull request!

Temporal evolution and taxonomy of representative works on World Action Models (WAMs).


Welcome to join our Feishu discussion group for in-depth academic exchange on World-Action Models, robot learning, and embodied intelligence. Our bot will also periodically share the latest WAM-related papers and updates, helping everyone stay informed about recent progress and discuss new ideas together.

Feishu/Lark QR Code

🔔 News

  • [2026-07-01] We are excited to share that the repository has surpassed 1,000 stars. Thank you all for your support!

  • [2026-06-24] Added the community discussion group QR code, and everyone is welcome to join and exchange ideas. (Use the QR code above or check Feishu/Lark)

  • [2026-05-25] Added the benchmark leaderboard page and benchmark-level performance trend visualization.

  • [2026-05-21] Updated the link to the paper-reading skill.

  • [2026-05-13] Initial release of the survey paper and repository.

Contents

Tag Legend

World Action Model tags

  • Cascaded WAM

  • Pixel-space Representations

    • Learned Action Extraction
    • Geometric Extraction
  • Implicit Planning via Latent Representations

  • Joint WAM

  • Autoregressive Generation

    • Explicit Decoupled Representation
    • Unified Discrete Representations
    • Predictive Latent Representation
  • Diffusion-based Generation

    • Unified Stream
      • Explicit Future Generation
      • Implicit Future Prediction
    • Multi-Stream
      • Cross-Attention Coupling
      • Hidden-State Coupling
      • Shared Representation
  • C3^3ache: "C3^3ache: Accelerating World Action Models with Cross Inference Chunk Cache", arXiv 2026. [📄 Paper]

World Action Model

Cascaded World-Action-Model

Joint World-Action-Model

Autoregressive Generation

Diffusion-based Generation

World Model for VLA

World Model for Imitation Learning

  • DREMA: "Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination", ICLR 2025. [📄 Paper] [🌍 Webpage] [💻 Code]

  • RoboScape: "RoboScape: Physics-informed Embodied World Model", arXiv 2025. [📄 Paper]

  • Ctrl-World: "Ctrl-World: A Controllable Generative World Model for Robot", ICLR 2026. [📄 Paper] [🌍 Webpage] [💻 Code]

World Model for Reinforcement Learning

World Model for Evaluation

Training Data


🤖 Robot-Centric

PaperReleasedLinks
QT-Opt - QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018-06📄 Paper · 🌍 Web · 📦 Dataset
MIME - Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation2018-10📄 Paper · 🌍 Web · 📦 Dataset
RoboNet - RoboNet: Large-Scale Multi-Robot Learning2019-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
RoboTurk - Scaling Robot Supervision to Hundreds of Hours with RoboTurk: Robotic Manipulation Dataset through Human Reasoning and Dexterity2019-11📄 Paper · 🌍 Web · 📦 Dataset
Bridge - Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets2021-09📄 Paper · 🌍 Web · 📦 Dataset
MT-Opt - MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale2021-04📄 Paper · 🌍 Web · 📦 Dataset
BC-Z - BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning2022-02📄 Paper · 🌍 Web · 📦 Dataset
Language-Table - Interactive Language: Talking to Robots in Real Time2022-10📄 Paper · 🌍 Web · 📦 Dataset
RT-1 - RT-1: Robotics Transformer for Real-World Control at Scale2022-12📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
BridgeData V2 - BridgeData V2: A Dataset for Robot Learning at Scale2023-08📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
Jaco-Play - CLVR Jaco Play Dataset2023-04💻 Code· 📦 Dataset
Cable-Routing-Dataset - Multi-Stage Cable Routing through Hierarchical Imitation Learning2023-07📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
RH20T - RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot2023-07📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
OXE - Open X-Embodiment: Robotic Learning Datasets and RT-X Models2023-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
DROID - DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset2024-03📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
RH20T-P - RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents2024-03📄 Paper · 🌍 Web · 📦 Dataset
RoboMIND - RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation2024-12📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
ARIO - All Robots in One: A New Standard and Unified Dataset for Versatile, General-Purpose Embodied Agents2024-08📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
RoboData - RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation2024-12📄 Paper · 📦 Dataset · 💻 Code
DexCap - DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation2024-03📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
FuSe - Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding2025-01📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
AgiBot World Colosseo - AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems2025-03📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
REASSEMBLE - REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly2025-02📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
OmniAction - RoboOmni: Proactive Robot Manipulation in Omni-modal Context2025-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code ·
UnifoLM-WBT - UnifoLM-WBT-Dataset2026-03📦 Dataset

🖐️ UMI

PaperReleasedLinks
UMI - Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots2024-02📄 Paper · 🌍 Web · 💻 Code· 📦 Dataset
FastUMI - FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset2024-09📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
FastUMI-100K - FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset2025-10📄 Paper · 💻 Code
RealOmin - 10Kh-RealOmin-OpenData: A Large-Scale Real-World Manipulation Dataset2026-01🌍 Web · 📦 Dataset · 💻 Code
Hoi! - Hoi! -- A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation2025-12📄 Paper · 🌍 Web · 📦 Dataset
DexUMI - DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation2025-05📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code

🖥️ Simulation

PaperReleasedLinks
MimicGen - MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations2023-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
ManiSkill2 - ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills2023-02📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
RoboCasa - RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots2024-06📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
RoboTwin - RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins2025-04📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
DexMimicGen - DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning2024-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
QUARD-Auto - GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot2024-03📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
TesserAct - TesserAct: Learning 4D Embodied World Models2025-04📄 Paper · 🌍 Web · 💻 Code
RoboCerebra - RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation2025-06📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
SynGrasp-1B - GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data2025-05📄 Paper · 🌍 Web · 💻 Code
RoboTwin 2.0 - RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation2025-06📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
TLA - TLA: Tactile-Language-Action Model for Contact-Rich Manipulation2025-03📄 Paper · 🌍 Web
InternVLA-M1 - InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy2025-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
InternData-A1 - InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation2026-01📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code

👁️ Human / Egocentric

PaperReleasedLinks
SSv2 - The "something something" video database for learning and evaluating visual common sense2017-06📄 Paper · 📦 Dataset
EPIC-KITCHENS - Scaling Egocentric Vision: The EPIC-KITCHENS Dataset2018-04📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
HowTo100M - HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips2019-06📄 Paper · 🌍 Web · 📦 Dataset
Kinetics-700 - A Short Note on the Kinetics-700 Human Action Dataset2019-07📄 Paper · 📦 Dataset · 💻 Code
EGTEA Gaze+ - In the Eye of the Beholder: Gaze and Actions in First Person Video2020-06📄 Paper · 🌍 Web · 📦 Dataset
Ego4D - Ego4D: Around the World in 3,000 Hours of Egocentric Video2021-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
H2O - H2O: Two Hands Manipulating Objects for First Person Interaction Recognition2021-04📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
HOI4D - HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction2022-03📄 Paper · 🌍 Web · 📦 Dataset
Assembly101 - Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities2022-03📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
EgoPAT3D - Egocentric Prediction of Action Target in 3D2022-03📄 Paper · 📦 Dataset · 💻 Code
ARCTIC - ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation2022-04📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
HoloAssist - HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World2023-09📄 Paper · 🌍 Web · 📦 Dataset
Ego-Exo4D - Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives2023-11📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
TACO - TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding2024-01📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
Aria Everyday Activities - Aria Everyday Activities Dataset2024-02📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
OAKINK2 - OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion2024-03📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
Nymeria - Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild2024-06📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
COM Kitchens - COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark2024-08📄 Paper · 📦 Dataset · 💻 Code
EgoVid-5M - EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation2024-11📄 Paper · 🌍 Web · 📦 Dataset
EgoMimic - EgoMimic: Scaling Imitation Learning via Egocentric Video2024-10📄 Paper · 🌍 Web · 📦 Dataset
HOT3D - HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos2024-11📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
Egocentric-10K - Egocentric-10K2025-11🌍 Web · 📦 Dataset
DreamDojo - DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos2026-02📄 Paper · 🌍 Web · 💻 Code
PH²D - Humanoid Policy ~ Human Policy2025-03📄 Paper · 🌍 Web · 📦 Dataset
Humanoid Everyday - Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation2025-10📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
IndEgo - IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants2025-11📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
PLAICraft - PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI2025-05📄 Paper · 🌍 Web · 📦 Dataset
HD-EPIC - HD-EPIC: A Highly-Detailed Egocentric Video Dataset2025-02📄 Paper · 🌍 Web · 📦 Dataset
UniHand - Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos2025-07📄 Paper · 🌍 Web · 📦 Dataset · 💻 Code
Ego-Centric Human Manipulation Dataset - EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos2025-07📄 Paper · 🌍 Web · 📦 Dataset
Kaiwu - Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction2025-03📄 Paper · 📦 Dataset
EgoDex - EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video2025-05📄 Paper · 📦 Dataset · 💻 Code

Evaluation


🌐 World Model — Visual Fidelity

PaperReleasedLinks
PSNR / SSIM - Image Quality Assessment: From Error Visibility to Structural Similarity2004📄 Paper · 🌍 Web · 💻 Code
LPIPS - The Unreasonable Effectiveness of Deep Features as a Perceptual Metric2018-01📄 Paper · 🌍 Web · 💻 Code
DreamSim - DreamSim: Learning New Dimensions of Human Visual Similarity Using Synthetic Data2023-06📄 Paper · 🌍 Web · 💻 Code
DINOv2 - DINOv2: Learning Robust Visual Features without Supervision2023-04📄 Paper · 🌍 Web · 💻 Code
FVD - Towards Accurate Generative Models of Video: A New Metric & Challenges2018-12📄 Paper · 💻 Code

🌐 World Model — Physical Commonsense

PaperReleasedLinks
VideoPhy - VideoPhy: Evaluating Physical Commonsense for Video Generation2024-06📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
PhyGenBench - Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation2024-10📄 Paper · 🌍 Web · 💻 Code
VBench-2.0 - VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness2025-03📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
WorldModelBench - WorldModelBench: Judging Video Generation Models as World Models2025-02📄 Paper · 🌍 Web · 💻 Code
Physics-IQ - Do Generative Video Models Understand Physical Principles?2025-01📄 Paper · 🌍 Web · 💻 Code
WorldScore - WorldScore: A Unified Evaluation Benchmark for World Generation2025-04📄 Paper · 🌍 Web · 💻 Code
EWMBench - EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models2025-05📄 Paper · 💻 Code

🌐 World Model — Action Plausibility

PaperReleasedLinks
WorldSimBench - WorldSimBench: Towards Video Generation Models as World Simulators2024-10📄 Paper · 🌍 Web · 💻 Code
Wow, wo, val! - A Comprehensive Embodied World Model Evaluation Turing Test2026-01📄 Paper

🤖 Action Policy — General

PaperReleasedLinks
MetaWorld - MetaWorld: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning2019-10📄 Paper · 🌍 Web · 💻 Code
RLBench - RLBench: The Robot Learning Benchmark & Learning Environment2019-09📄 Paper · 🌍 Web · 💻 Code
Robomimic - What Matters in Learning from Offline Human Demonstrations for Robot Manipulation2021-08📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
Franka Kitchen - Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning2019-10📄 Paper · 🌍 Web · 💻 Code
ManiSkill - ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations2021-07📄 Paper · 🌍 Web · 💻 Code
ManiSkill2 - ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills2023-02📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
ManiSkill3 - ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI2024-10📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
RoboCasa - RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots2024-06📄 Paper · 🌍 Web · 💻 Code
CALVIN - CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks2021-12📄 Paper · 🌍 Web · 💻 Code
VIMAbench - VIMA: General Robot Manipulation with Multimodal Prompts2022-10📄 Paper · 🌍 Web · 💻 Code
VLMbench - VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation2022-06📄 Paper · 💻 Code
LIBERO - LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning2023-06📄 Paper · 🌍 Web · 💻 Code
LIBERO-Plus - LIBERO-Plus: In-Depth Robustness Analysis of Vision-Language-Action Models2025-10📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
LIBERO-PRO - LIBERO-PRO: Towards Robust and Fair Evaluation of VLA Models beyond Memorization2025-10📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
LIBERO-X - LIBERO-X: Robustness Litmus for Vision-Language-Action Models2026-02📄 Paper· 🌍 Web · 💻 Code
COLOSSEUM - THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation2024-02📄 Paper · 🌍 Web · 💻 Code
AGNOSTOS - Exploring the Limits of Vision-Language-Action Manipulations in Cross-Task Generalization2025-05📄 Paper · 🌍 Web · 💻 Code
RoboEval - RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation2025-07📄 Paper · 🌍 Web· 💻 Code
RoboVerse - RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning2025-04📄 Paper · 🌍 Web · 💻 Code
PolaRiS - PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies2025-12📄 Paper · 🌍 Web · 💻 Code · 📦 Dataset
RoboMME - RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies2026-03📄 Paper · 🌍 Web · 💻 Code
GenManip - GenManip: LLM-Driven Simulation for Generalizable Instruction-Following Manipulation2025-06📄 Paper · 🌍 Web
VLABench - VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks2024-12📄 Paper · 💻 Code
RoboSuite - Robosuite: A Modular Simulation Framework and Benchmark for Robot Learning2020-09📄 Paper · 🌍 Web · 💻 Code
RoboLab - RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies2026-04📄 Paper
SimplerEnv - Evaluating Real-World Robot Manipulation Policies in Simulation2024-05📄 Paper · 🌍 Web · 💻 Code
ARNOLD - ARNOLD: A Benchmark for Language-Grounded Task Learning with Continuous States in Realistic 3D Scenes2023-04📄 Paper · 🌍 Web · 💻 Code
GemBench - Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-Guided 3D Policy2024-10📄 Paper · 🌍 Web · 💻 Code

🤖 Action Policy — Bimanual and Humanoid Form

PaperReleasedLinks
RoboTwin - RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins2025-04📄 Paper · 🌍 Web · 💻 Code
BiGym - BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark2024-07📄 Paper · 🌍 Web · 💻 Code
HumanoidBench - HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation2024-03📄 Paper · 🌍 Web · 💻 Code
HumanoidGen - HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM Reasoning2025-07📄 Paper · 🌍 Web · 💻 Code

🤖 Action Policy — Mobile Manipulation

PaperReleasedLinks
ManipulaTHOR - ManipulaTHOR: A Framework for Visual Object Manipulation2021-04📄 Paper · 🌍 Web · 💻 Code
HomeRobot - HomeRobot: Open-Vocabulary Mobile Manipulation2023-06📄 Paper · 🌍 Web · 💻 Code
BEHAVIOR-1K - BEHAVIOR-1K: A Benchmark for Embodied AI with 1,000 Everyday Activities and Realistic Simulation2024-03📄 Paper · 🌍 Web · 💻 Code

🤖 Action Policy — Contact and Deformation Manipulation

PaperReleasedLinks
SoftGym - SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation2020-11📄 Paper· 🌍 Web · 💻 Code
PlasticineLab - PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable Physics2021-04📄 Paper · 🌍 Web · 💻 Code
DaXBench - DaXBench: Benchmarking Deformable Object Manipulation with Differentiable Physics2022-10📄 Paper · 🌍 Web · 💻 Code
TacSL - TacSL: A Library for Visuotactile Sensor Simulation and Learning2024-08📄 Paper · 🌍 Web · 💻 Code
ManiFeel - ManiFeel: Benchmarking and Understanding Visuotactile Manipulation Policy Learning2025-05📄 Paper · 🌍 Web· 💻 Code

🤖 Action Policy — Real-Device

PaperReleasedLinks
RoboArena - RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies2025-06📄 Paper · 🌍 Web · 💻 Code
RoboChallenge - RoboChallenge: Large-Scale Real-Robot Evaluation of Embodied Policies2025-10📄 Paper · 🌍 Web· 💻 Code
ManipArena - ManipArena: Comprehensive Real-World Evaluation of Reasoning-Oriented Generalist Robot Manipulation2026-03📄 Paper · 🌍 Web · 💻 Code

🌟 Star History

Star History Chart

🙏 Acknowledgements

We thank the AllenAI VLA Evaluation Harness team for curating and maintaining the VLA leaderboard data that supports our leaderboard.

👋 Citation

If you find this survey or repository helpful for your research, please consider citing our paper:

@article{wang2026world,
  title={World Action Models for Generalist Robotics: From Next Token Prediction to Next State Synthesis},
  author={Wang, Siyin and Shi, Junhao and Fu, Zhaoyang and He, Xinzhe and Liu, Feihong and Yang, Chenchen and Zhou, Yikang and Fei, Zhaoye and Gong, Jingjing and Fu, Jinlan and Shou, Mike Zheng and Huang, Xuanjing and Qiu, Xipeng and Jiang, Yu-Gang},
  year={2026}
}