Awesome-VLA-Papers

July 3, 2025 · View on GitHub

This repository contains the list of representative VLA works in the survey “A Survey on Vision-Language-Action Models: An Action Tokenization Perspective”, along with relevant reference materials.

Foundation Models

Language Foundation Models

Vision Foundation Models

Vision Language Models

Language Description as Action Tokens

Language Plan

  • Language Planner, Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, 2022.01, ICML 2022. [📄 Paper] [🌍 Website] [💻 Code]
  • Socratic Models, Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language, 2022.04, ICLR 2023. [📄 Paper] [🌍 Website] [💻 Code]
  • SayCan, Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, 2022.04. [📄 Paper] [🌍 Website] [💻 Code]
  • Inner Monologue, Inner Monologue: Embodied Reasoning through Planning with Language Models, 2022.07, CoRL 2022. [📄 Paper] [🌍 Website]
  • PaLM-E, PaLM-E: An Embodied Multimodal Language Model, 2023.03, ICML 2023. [📄 Paper] [🌍 Website]
  • EmbodiedGPT, EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, 2023.05, NeurIPS 2023. [📄 Paper] [🌍 Website] [💻 Code] [📊 Dataset]
  • DoReMi, DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment, 2023.07, IROS 2024. [📄 Paper] [🌍 Website]
  • ViLa, Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning, 2023.11, Workshop on Vision-Language Models for Navigation and Manipulation, ICRA 2024. [📄 Paper] [🌍 Website]
  • 3D-VLA, 3D-VLA: A 3D Vision-Language-Action Generative World Model, 2024.03, ICML 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • Bi-VLA, Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous Manipulations, 2024.05, SMC 2024. [📄 Paper]
  • RoboMamba, RoboMamba: Multimodal State Space Model for Efficient Robot Reasoning and Manipulation, 2024.06, NeurIPS 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • ReplanVLM, ReplanVLM: Replanning Robotic Tasks with Visual Language Models, 2024.07. [📄 Paper]
  • BUMBLE, BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation, 2024.10, ICRA 2025. [📄 Paper] [🌍 Website] [💻 Code]
  • ReflectVLM, Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation, 2025.02. [📄 Paper] [🌍 Website] [💻 Code] [📊 Dataset] [🤗 Model]
  • Hi Robot, Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models, 2025.02. [📄 Paper] [🌍 Website]
  • RoboBrain, RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete, 2025.02, CVPR 2025. [📄 Paper] [🌍 Website] [💻 Code] [📊 Dataset]
  • π0.5\pi_{0.5}, π0.5\pi_{0.5}: a Vision-Language-Action Model with Open-World Generalization, 2025.04. [📄 Paper] [🌍 Website]

Language Motion

Code as Action Tokens

  • Code as Policies, Code as Policies: Language Model Programs for Embodied Control, 2022.09, ICRA 2023. [📄 Paper] [🌍 Website] [💻 Code]
  • ProgPrompt, ProgPrompt: Generating Situated Robot Task Plans using Large Language Models, 2022.09, ICRA 2023. [📄 Paper] [🌍 Website] [💻 Code]
  • ChatGPT for Robotics, ChatGPT for Robotics: Design Principles and Model Abilities, 2023.02, IEEE Access 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • Text2Motion, Text2Motion: From Natural Language Instructions to Feasible Plans, 2023.03, ICRL 2023. [📄 Paper] [🌍 Website]
  • Instruct2Act, Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model, 2023.05. [📄 Paper] [💻 Code]
  • RoboScript, RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation, 2024.02. [📄 Paper]
  • RoboCodeX, RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis, 2024.02, ICML 2024. [📄 Paper] [🌍 Website] [💻 Code]

Affordance as Action Tokens

Keypoint

  • KITE, KITE: Keypoint-Conditioned Policies for Semantic Manipulation, 2023.6, CoRL 2023. [📄 Paper] [🌍 Website] [💻 Code]
  • CoPa, CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models, 2024.3, IROS 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • RoboPoint, RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics, 2024.6, CoRL 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • RAM, RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation, 2024.7, CoRL 2024 Oral. [📄 Paper] [🌍 Website] [💻 Code]
  • ReKep, ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation, 2024.9, CoRL 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • OmniManip, OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints, 2025.1, CVPR 2025 Highlight. [📄 Paper] [🌍 Website]
  • Magma, Magma: A Foundation Model for Multimodal AI Agents, 2025.2, CVPR 2025. [📄 Paper] [🌍 Website] [💻 Code]
  • KUDA, KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation, 2025.3, ICRA 2025. [📄 Paper] [🌍 Website] [💻 Code]

Bounding Box

Segmentation Mask

  • MOO, Open-World Object Manipulation using Pre-Trained Vision-Language Models, 2023.3, CoRL 2023. [📄 Paper] [🌍 Website]
  • ROCKET-1, ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting, 2024.11, CVPR 2025. [📄 Paper] [🌍 Website]
  • SoFar, SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation, 2025.2. [📄 Paper] [🌍 Website] [💻 Code]
  • RoboDexVLM, RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation, 2025.3. [📄 Paper] [🌍 Website]

Affordance Map

  • CLIPort, CLIPort: What and Where Pathways for Robotic Manipulation, 2021.9, CoRL 2021. [📄 Paper] [🌍 Website] [💻 Code]
  • VoxPoser, VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models, 2023.7, CoRL 2023 Oral. [📄 Paper] [🌍 Website] [💻 Code]
  • ManipLLM, ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation, 2023.12, CVPR 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • ManiFoundation, ManiFoundation Model for General-Purpose Robotic Manipulation of Contact Synthesis with Arbitrary Objects and Robots, 2024.5, IROS 2024 Oral. [📄 Paper] [🌍 Website] [💻 Code]
  • MOKA, MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting, 2024.5, RSS 2024. [📄 Paper] [🌍 Website] [💻 Code]

Trajectory as Action Tokens

Robotic Manipulation

Autonomous Driving

  • DriveVLM, DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models, 2024.02, CoRL 2024. [📄 Paper] [🌍 Website]
  • CoVLA, CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving, 2024.08, WACV 2025 Oral. [📄 Paper] [🌍 Website] [📊 Dataset]
  • EMMA, EMMA: End-to-End Multimodal Model for Autonomous Driving, 2024.10. [📄 Paper]
  • VLM-E2E, VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion, 2025.02. [📄 Paper]

Goal State as Action Tokens

Single-Frame Image / Point Cloud

  • SuSIE, Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models, 2023.10, ICLR 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • 3D-VLA, 3D-VLA: A 3D Vision-Language-Action Generative World Model, 2024.03, ICML 2024. [📄 Paper] [🌍 Website]
  • CoTDiffusion, Generate Subgoal Images before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal Prompts, 2024.06, CVPR 2024. [📄 Paper] [🌍 Website]
  • CoT-VLA, CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, 2025.03, CVPR 2025. [📄 Paper] [🌍 Website]

Multi-Frame Video

  • UniPi, Learning Universal Policies via Text-Guided Video Generation, 2023.02, NeurIPS 2023 spotlight. [📄 Paper] [🌍 Website]
  • AVDC, Learning to Act from Actionless Videos through Dense Correspondences, 2023.10, ICLR 2024 spotlight. [📄 Paper] [🌍 Website] [💻 Code]
  • VLP, Video Language Planning, 2023.10. [📄 Paper] [🌍 Website] [💻 Code]
  • Gen2Act, Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation, 2024.09, CoRL-X-Embodiment-WS 2024. [📄 Paper] [🌍 Website]
  • Video Prediction Policy, Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations, 2024.12, ICML 2025 Spotlight. [📄 Paper] [🌍 Website] [💻 Code]
  • FLIP, FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model, 2024.12, ICLR 2025 [📄 Paper] [🌍 Website] [💻 Code]
  • GEVRM, GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation, 2025.02, ICLR 2025. [📄 Paper]

Latent Representation as Action Tokens

Raw Action as Action Tokens

Reasoning as Action Tokens

  • Inner Monologue, Inner Monologue: Embodied Reasoning through Planning with Language Models, 2022.07, CoRL 2022. [📄 Paper] [🌍 Website]
  • DriveVLM, DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models, 2024.02, CoRL 2024. [📄 Paper] [🌍 Website]
  • ECoT, Robotic Control via Embodied Chain-of-Thought Reasoning, 2024.07, CoRL 2024. [📄 Paper] [🌍 Website]
  • RAD, Action-Free Reasoning for Policy Generalization, 2025.02. [📄 Paper] [🌍 Website]
  • AlphaDrive, AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning, 2025.03. [📄 Paper] [🌍 Website]
  • Cosmos-Reason1, Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning, 2025.03. [📄 Paper] [🌍 Website] [💻 Code]

Scalable Data Sources

Bottom Layer: Web Data and Human Video

  • Something-Something V2, The" something something" video database for learning and evaluating visual common sense, 2017.06. [📄 Paper] [🌍 Website]
  • EPIC-KITCHENS-100, Scaling Egocentric Vision: The EPIC-KITCHENS Dataset, 2018.04. [📄 Paper] [🌍 Website]
  • Ego4D, Ego4D: Around the World in 3,000 Hours of Egocentric Video, 2021.10. [📄 Paper] [🌍 Website]
  • Ego-Exo4D, Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives, 2023.11. [📄 Paper] [🌍 Website]

Middle Layer: Synthetic and Simulation Data

  • MimicGen, MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations, 2023.10, CoRL 2023. [📄 Paper]
  • RoboCase, RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots, 2024.06, RSS 2024. [📄 Paper] [🌍 Website] [💻 Code]
  • DexMimicGen, DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning, 2024.10, ICRA 2025. [📄 Paper] [🌍 Website]
  • AgiBot DigitalWorld, AgiBot DigitalWorld, 2025.02. [🌍 Website]

Top Layer: Real-world Robot Data

  • nuScenes, nuScenes: A multimodal dataset for autonomous driving, 2019.03, CVPR 2020. [📄 Paper] [🌍 Website] [💻 Code]
  • WOMD, Large Scale Interactive Motion Forecasting for Autonomous Driving: The Waymo Open Motion Dataset, 2021.04, ICCV 2021. [📄 Paper] [🌍 Website] [💻 Code]
  • RT-1, RT-1: Robotics Transformer for Real-World Control at Scale, 2022.12. [📄 Paper] [🌍 Website]
  • RH20T, RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot, 2023.06, ICRA 2024. [📄 Paper] [🌍 Website]
  • BridgeData V2, BridgeData V2: A Dataset for Robot Learning at Scale, 2023.08, CoRL 2o23. [📄 Paper] [🌍 Website]
  • OXE, Open X-Embodiment: Robotic Learning Datasets and RT-X Models, 2023.10, ICRA 2024. [📄 Paper] [🌍 Website]
  • HoNY, On Bringing Robots Home, 2023.11. [📄 Paper] [🌍 Website]
  • DROID, DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset, 2024.03, RSS 2024. [📄 Paper] [🌍 Website]
  • CoVLA, CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving, 2024.08, WACV 2025. [📄 Paper] [🌍 Website]
  • RoboMIND, RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation, 2024.12, RSS 2025. [📄 Paper] [🌍 Website]
  • AgiBot World, AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems, 2025.03. [📄 Paper] [🌍 Website] [💻 Code]

Related Surveys

  • Robot Learning in the Era of Foundation Models: A Survey, 2023.11, Neurocomputing Volume 638. [📄 Paper]
  • A Survey on Robotics with Foundation Models: toward Embodied AI, 2024.02. [📄 Paper]
  • A Survey on Integration of Large Language Models with Intelligent Robots, 2024.04, Intelligent Service Robotics 2024. [📄 Paper]
  • What Foundation Models can Bring for Robot Learning in Manipulation: A Survey, 2024.04. [📄 Paper]
  • A Survey on Vision-Language-Action Models for Embodied AI, 2024.05. [📄 Paper]
  • Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI, 2024.07. [📄 Paper]
  • Exploring Embodied Multimodal Large Models: Development, Datasets, and Future Directions, 2025.02, Information Fusion Volume 122. [📄 Paper]
  • Generative Artificial Intelligence in Robotic Manipulation: A Survey, 2025.03. [📄 Paper]
  • OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation, 2025.05. [📄 Paper]
  • Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, 2025.05. [📄 Paper]