Awesome-Dual-System-VLA

November 21, 2025 · View on GitHub

To address the challenges in vision-language-action (VLA) models, such as the difficulty of achieving efficient real-time performance, the high cost of pre-training, and the complexity of end-to-end fine-tuning on embodied data due to domain shift and catastrophic forgetting, Dual-System VLA models were introduced.

The development and architectural details of Dual-System VLA models are discussed in the paper.

This repository will be continuously updated, and we warmly welcome contributions from the community. If you have papers, projects, or resources that are not yet included, please feel free to submit them via a pull request or open an issue for discussion.

Current Results

CALVIN ABC→D

Method12345Avg. Len.
Single-System
OpenVLA91.377.862.052.143.53.27
UniVLA95.585.875.466.956.53.80
Seer94.487.279.972.264.33.98
Dual-System
LCB73.650.228.516.09.91.78
RationalVLA74.358.342.330.020.72.26
Robodual94.482.772.162.454.43.66
OpenHelix97.191.482.872.664.14.08

LIBERO

MethodLIBERO-SpatialLIBERO-ObjectLIBERO-GoalLIBERO-LongAvg.
Single-System
OpenVLA84.788.479.253.776.5
π096.898.895.885.294.2
OpenVLA-OFT97.698.497.994.597.1
GR00T N194.497.690.693.993.9
UniVLA96.596.895.692.095.2
Seer---87.7-
Dual-System
DexVLA97.299.195.6--
Hume98.699.899.498.698.6

✅ Dual-System VLA

Robot Manipulation

TitleVenueDateCode
Galaxea Open-World Dataset and G0 Dual-System VLA ModelarXiv2025-08-30Star Github
TriVLA: A Unified Triple-System-Based Unified Vision-Language-Action Model for General Robot ControlarXiv2025-07-01Project
RationalVLA: A Rational Vision-Language-Action Model with Dual SystemarXiv2025-06-12Project
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow ReasoningarXiv2025-06-02Star Github
Hume: Introducing System-2 Thinking in Visual-Language-Action ModelarXiv2025-05-27Star Github
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic ManipulationarXiv2025-05-06Star Github
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot ControlarXiv2025-02-09Star Github
RoboDual: Towards Synergistic, Generalized, and Efficient Dual-System for Robotic ManipulationarXiv2024-10-10Star Github
DP-VLA: A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLMCoRL 20242024-10-21-
HiRT: Enhancing Robotic Control with Hierarchical Robot TransformersCoRL 20242024-09-12-
LCB: From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot ControlIROS 20242024-05-08Project

Humanoid Robot

TitleVenueDateCode
Helix: A Vision-Language-Action Model for Generalist Humanoid Control--Porject

❌ Not a Dual-System VLA

Robot Manipulation

TitleVenueDateCode
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive ReasoningarXiv2025-05-17Star Github
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied TasksarXiv2025-04-28Star Github
π0.5: a Vision-Language-Action Model with Open-World GeneralizationarXiv2025-04-22Project
π0: A Vision-Language-Action Flow Model for General Robot ControlarXiv2024-10-31Project
PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic ManipulationNeurIPS 20242024-10-14Star Github
MResT: Multi-Resolution Sensing for Real-Time Control with Vision-Language ModelsCoRL 20232024-01-25Star Github

Humanoid Robot

TitleVenueDateCode
GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsarXiv2025-03-18Porject