Awesome-Robot-Learning-from-Human-Videos
June 22, 2026 ยท View on GitHub

๐ Overview
This repository provides a curated reading list for robot learning from human videos (LfHV), with an emphasis on human-robot skill transfer techniques in task, observation, and action levels. It also covers representative works on human-object interaction analysis and human video sources that are widely used in the literature.
Papers with publicly released code are marked with a star ๐. Papers with real-world robot experiments are marked with a robot ๐ค.
Contributions are welcome. If you find missing papers or inaccurate classifications, feel free to open a pull request or contact me via email.
The related survey paper can be found at this link. In this survey, you can find answers to the following interesting questions:
๐ How much more efficient is human video collection compared to robot teleoperation?
๐ What types of information can be transferred from human videos to robot manipulation?
๐ Which is more widely adopted in LfHV, egocentric or exocentric video data?
๐ How do imitation learning and reinforcement learning incorporate information from human videos?
๐ How can you design your own LfHV framework for specific application needs?
๐ How have open-source human video datasets evolved over time?
๐ What are the most promising future research directions in LfHV?
๐ ...
We hope that the answers presented in this survey can inspire future research and contribute to the development of generalist robotic systems. A more concise version of this survey is coming soon.
We sincerely appreciate these blogs for helping distill and promote our work:
- ๆ ้ๆผๅINFINITY: ๆบๅจไบบไปไบบ็ฑปไธ็ๅญฆไน ๏ผRobot Learning from Human Videos
- human five: ๆบๅจไบบๅฆไฝไปไบบ็ฑป่ง้ขไธญๅญฆไผๆไฝ?
- ๅ ท่บซๆบ่ฝไนๅฟ: ๅฆไฝไปไบบ็ฑป่ง้ขไธญๅญฆไน ๆบๅจไบบๆไฝ๏ผ่ฟ400็ฏๅทฅไฝไธ่งๆฐๆฎ้ ็ฝฎใๅญฆไน ๆนๅผใๆ ธๅฟๆๆ๏ผ
- ๆบ้ฉพๅ ท่บซๆฐๆฎๆๆ: ไธไบค็่ดบๅๅข้LfHV Survey๏ผไปไบบ็ฑป่ง้ขๅญฆไน ๆบๅจไบบๆไฝๅ จๆฏ็ปผ่ฟฐ๏ผ
- ๅคง่ฏญ่จๆจกๅๅๅ ท่บซๆบไฝๅ่ชๅจ้ฉพ้ฉถ: ๆบๅจไบบไปไบบ็ฑป่ง้ขไธญๅญฆไน ๏ผ็ปผ่ฟฐ
๐ Citation
If you find this work helpful for your research, please kindly consider citing our paper:
@article{ma2026robot,
title={Robot Learning from Human Videos: A Survey},
author={Ma, Junyi and Zhang, Erhang and Yang, Haoran and Li, Ditao and Xu, Chenyang and Wang, Guangming and Wang, Hesheng},
journal={arXiv preprint arXiv:2604.27621},
year={2026}
}
๐๏ธ Table of Contents

๐ Human-Robot Skill Transfer
Task-Oriented Transfer
This section covers methods that infer task structures and intents from human videos to guide robot decision-making at the task level.
2026
- ๐ค [arXiv 2026.02] Act, Sense, Act: Learning Non-Markovian Active Perception Strategies from Large-Scale Egocentric Human Data [Code] [Page] [
Task intents]
2025
- ๐ค [IJRR] FMimic: Foundation Models Are Fine-Grained Action Learners from Human Videos [Code] [Page] [
Task structures] - ๐ค [ICRA] Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models [Code] [Page] [
Task structures] - ๐ค [CoRL] Action-Free Reasoning for Policy Generalization [Code] [Page] [
Task structures] - [ICASSP] Interactive Robot Action Replanning Using Multimodal LLM Trained from Human Demonstration Videos [Code] [[Page]] [
Task structures] - ๐ค [arXiv 2025.12] PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence [Code] [Page] [
Task structures] - [arXiv 2025.11] Robot Confirmation Generation and Action Planning Using Long-Context Q-Former Integrated with Multimodal LLM [Code] [[Page]] [
Task structures] - ๐๐ค [arXiv 2025.09] From Watch to Imagine: Steering Long-Horizon Manipulation via Human Demonstration and Future Envisionment [Code] [Page] [
Task structures] - ๐๐ค [arXiv 2025.08] EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos [Code] [[Page]] [
Task structures]
2024
- ๐๐ค [RA-L] GPT-4V(ision) for Robotics: Multimodal Task Planning From Human Demonstration [Code] [Page] [
Task structures] - ๐ค [NeurIPS] VLMimic: Vision Language Models Are Visual Imitation Learner for Fine-Grained Actions [Code] [Page] [
Task structures] - ๐ค [RSS] Vid2Robot: End-to-End Video-Conditioned Policy Learning with Cross-Attention Transformers [Code] [Page] [
Task intents] - ๐๐ค [IROS] Knowledge-Based Programming by Demonstration Using Semantic Action Models for Industrial Assembly [Code] [Page] [
Task structures] - ๐๐ค [arXiv 2024.10] VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model [Code] [Page] [
Task structures]
2023
- ๐ค [T-SMC] Watch and Act: Learning Robotic Manipulation From Visual Demonstration [Code] [Page] [
Task structures] - ๐๐ค [CoRL] XSkill: Cross Embodiment Skill Discovery [Code] [Page] [
Task intents] - ๐ค [arXiv 2023.12] Learning Multi-Step Manipulation Tasks from A Single Human Demonstration [Code] [[Page]] [
Task structures]
2022
- ๐ค [T-MECH] Explicit-to-Implicit Robot Imitation Learning by Exploring Visual Content Change [Code] [Page] [
Task structures] - ๐ค [ICRA] Learning Periodic Tasks from Human Demonstrations [Code] [Page] [
Task structures] - ๐ค [CoRL] BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning [Code] [Page] [
Task intents] - ๐ [CoRL] Cross-Domain Transfer via Semantic Skill Imitation [Code] [Page] [
Task structures]
2019
- ๐๐ค [NeurIPS] Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller [Code] [Page] [
Task intents] - ๐ค [IROS] Learning Actions from Human Demonstration Video for Robotic Manipulation [Code] [[Page]] [
Task structures]
2018
- ๐ค [RSS] One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning [Code] [[Page]] [
Task intents] - ๐๐ค [ICRA] Translating Videos to Commands for Robotic Manipulation with Deep Recurrent Neural Networks [Code] [[Page]] [
Task structures] - ๐ค [arXiv 2018.10] One-Shot Hierarchical Imitation Learning of Compound Visuomotor Tasks [Code] [Page] [
Task intents]
2017
- ๐๐ค [RSS] Unsupervised Perceptual Rewards for Imitation Learning [Code] [Page] [
Task intents] - [RA-L] Unsupervised Linking of Visual Features to Textual Descriptions in Long Manipulation Activities [Code] [[Page]] [
Task structures]
2015
- [AAAI] Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide Web [Code] [[Page]] [
Task structures]
Observation-Oriented Transfer
This section focuses on bridging the observation gap between humans and robots via transformed videos and visual embeddings.
2026
- ๐ค [Science Robotics] Visual-Tactile Pretraining and Online Multitask Learning for Humanlike Manipulation Dexterity [Code] [Page] [
Visual embeddings] - ๐๐ค [ICRA] Masquerade: Learning from In-the-Wild Human Videos Using Data-Editing [Code] [Page] [
Transformed videos] - ๐๐ค [ICRA] MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos [Code] [Page] [
Visual embeddings] - ๐ค [AAAI] Human2Robot: Learning Robot Actions from Paired Human-Robot Videos [Code] [Page] [
Transformed videos] - ๐ค [arXiv 2026.04] WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations [Code] [Page] [
Transformed videos] - ๐๐ค [arXiv 2026.02] FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment[Code] [Page] [
Visual embeddings,VLA]
2025
- ๐ค [NeurIPS] EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data [Code] [Page] [
Visual embeddings] - ๐๐ค [CVPR] Mitigating the Human-Robot Domain Discrepancy in Visual Pre-Training for Robotic Manipulation [Code] [Page] [
Visual embeddings] - ๐๐ค [RA-L] GR-MG: Leveraging Partially-Annotated Data via Multi-Modal Goal-Conditioned Policy [Code] [Page] [
Visual embeddings] - ๐๐ค [ICRA] One-Shot Imitation Under Mismatched Execution [Code] [Page] [
Visual embeddings] - ๐ [IROS] Ag2x2: Robust Agent-Agnostic Visual Representations for Zero-Shot Bimanual Manipulation [Code] [Page] [
Transformed videos] - ๐ค [IROS] VTAO-BiManip: Masked Visual-Tactile-Action Pre-Training with Object Understanding for Bimanual Dexterous Manipulation [Code] [Page] [
Visual embeddings] - ๐ค [CoRL] Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation [Code1] [Code2] [Page] [
Visual embeddings] - ๐๐ค [CoRL] ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation [Code] [Page] [
Visual embeddings] - ๐๐ค [CoRL] Phantom: Training Robots Without Robots Using Only Human Videos [Code] [Page] [
Transformed videos] - ๐ค [CoRL] Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt [Code] [Page] [
Visual embeddings] - ๐ [arXiv 2025.12] H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos [Code] [Page] [
Transformed videos] - ๐ [arXiv 2025.12] Mitty: Diffusion-Based Human-to-Robot Video Generation [Code] [Page] [
Transformed videos] - ๐ [arXiv 2025.12] X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale [Code] [Page] [
Transformed videos] - ๐๐ค [arXiv 2025.12] mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs [Code] [Page] [
Visual embeddings] - ๐๐ค [arXiv 2025.12] World Models Can Leverage Human Videos for Dexterous Manipulation [Code] [Page] [
Visual embeddings] - ๐ค [arXiv 2025.10] Trajectory-Conditioned Cross-Embodiment Skill Transfer [Code] [Page] [
Transformed videos] - ๐๐ค [arXiv 2025.09] MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training [Code] [Page] [
Transformed videos] - ๐๐ค [arXiv 2025.09] RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation [Code] [Page] [
Visual embeddings] - ๐ค [CoRL] FLARE: Robot Learning with Implicit World Modeling [Code] [Page] [
Visual embeddings,VLA] - ๐ค [arXiv 2025.05] H2R: A Human-to-Robot Data Augmentation for Robot Pre-Training from Videos [Code] [Page] [
Transformed videos]
2024
- ๐ค [T-ASE] Contrast, Imitate, Adapt: Learning Robotic Skills from Raw Human Videos [Code] [Page] [
Visual embeddings] - ๐๐ค [ICLR] Unleashing Large-Scale Video Generative Pre-Training for Visual Robot Manipulation [Code] [Page] [
Visual embeddings] - ๐ค [RSS] Vid2Robot: End-to-End Video-Conditioned Policy Learning with Cross-Attention Transformers [Code] [Page] [
Visual embeddings] - ๐๐ค [RSS] Learning Manipulation by Predicting Interaction [Code] [Page] [
Visual embeddings] - [ICRA] Masked Visual-Tactile Pre-Training for Robot Manipulation [Code] [Page] [
Visual embeddings] - ๐๐ค [ICRA] Robotic Offline RL from Internet Videos via Value-Function Pre-Training [Code] [Page] [
Visual embeddings] - ๐๐ค [IROS] Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations [Code] [Page] [
Transformed videos] - ๐๐ค [CoRL] What Makes Pre-Trained Visual Representations Successful for Robust Manipulation? [Code] [Page] [
Visual embeddings] - ๐ค [arXiv 2024.10] GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation [Code] [Page] [
Visual embeddings]
2023
- ๐๐ค [ICML] LIV: Language-Image Representations and Rewards for Robotic Control [Code] [Page] [
Visual embeddings] - ๐๐ค [ICLR] VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training [Code] [Page] [
Visual embeddings] - ๐๐ค [NeurIPS] Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos [Code] [Page] [
Transformed videos] - ๐๐ค [NeurIPS] Where Are We in the Search for an Artificial Visual Cortex for Embodied Intelligence? [Code] [Page] [
Visual embeddings] - ๐ค [RSS] Structured World Models from Human Videos [Code] [Page] [
Visual embeddings] - ๐๐ค [RSS] Language-Driven Representation Learning for Robotics [Code] [Page] [
Visual embeddings] - [ICRA] Learning Video-Conditioned Policies for Unseen Manipulation Tasks [Code] [Page] [
Visual embeddings] - ๐๐ค [CoRL] AR2-D2: Training a Robot Without a Robot [Code] [Page] [
Transformed videos] - ๐๐ค [CoRL] An Unbiased Look at Datasets for Visuo-Motor Pre-Training [Code] [Page] [
Visual embeddings] - ๐ค [arXiv 2023.04] Efficient Robot Skill Learning with Imitation from a Single Video for Contact-Rich Fabric Manipulation [Code] [Page] [
Visual embeddings]
2022
- ๐๐ค [CoRL] R3M: A Universal Visual Representation for Robot Manipulation [Code] [Page] [
Visual embeddings] - ๐ค [RSS] Human-to-Robot Imitation in the Wild [Code] [Page] [
Transformed videos] - ๐๐ค [CoRL] Real-World Robot Learning with Masked Visual Pre-Training [Code] [Page] [
Visual embeddings] - ๐ค [CoRL] RoboTube: Learning Household Manipulation from Human Videos with Simulated Twin Environments [Code] [Page] [
Visual embeddings] - [Machines] Learning by Watching via Keypoint Extraction and Imitation Learning [Code] [Page] [
Transformed videos] - ๐ [arXiv 2022.03] Masked Visual Pre-Training for Motor Control [Code] [Page] [
Visual embeddings]
2021
- ๐๐ค [RSS] Learning Generalizable Robotic Reward Functions from "In-the-Wild" Human Videos [Code] [Page] [
Visual embeddings] - [IROS] Learning by Watching: Physical Imitation of Manipulation Skills from Human Videos [Code] [Page] [
Transformed videos] - ๐ [CoRL] XIRL: Cross-Embodiment Inverse Reinforcement Learning [Code] [Page] [
Visual embeddings] - ๐ค [arXiv 2021.03] Manipulator-Independent Representations for Visual Imitation [Code] [Page] [
Transformed videos]
2020
- ๐ค [RSS] AVID: Learning Multi-Stage Tasks Via Pixel-Level Translation of Human Videos [Code] [Page] [
Transformed videos] - ๐๐ค [CoRL] Reinforcement Learning with Videos: Combining Offline Observations with Interaction [Code] [Page] [
Visual embeddings]
2018
- ๐๐ค [RA-L] Deep Episodic Memory: Encoding, Recalling, and Predicting Episodic Experiences for Robot Action Execution [Code] [Page] [
Visual embeddings] - ๐๐ค [ICRA] Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation [Code] [Page] [
Visual embeddings]
2017
- ๐๐ค [CVPRW] Time-Contrastive Networks: Self-Supervised Learning from Multi-view Observation [Code] [Page] [
Visual embeddings]
Action-Oriented Transfer
This section includes methods that transfer actionable motion knowledge from human videos to robot control, including affordances (explicit interaction info) and latent actions.
2026
- ๐๐ค [TRO&IROS] AgiBot World Colosseo: A Large-Scale Manipulation Platform for Scalable and Intelligent Embodied Systems [Code] [Page] [
Latent actions] - ๐๐ค [ICLR&NeurIPS Workshop] ViPRA: Video Prediction for Robot Actions [Code] [Page] [
Latent actions] - ๐๐ค [CVPR] Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild [Code] [Page] [
Latent actions] - ๐๐ค [CVPR] UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos [Code] [Page][
Affordances] - ๐๐ค [CVPR] Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos [Code] [Page] [
Affordances] - ๐๐ค [AAAI] H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation [Code] [Page] [
Affordances] - ๐ [AAAI] Towards Affordance-Aware Robotic Dexterous Grasping with Human-Like Priors [Code] [Page] [
Affordances] - [ICASSP] CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control [Code] [Page] [
Latent actions] - ๐ค [RA-L] EMMA: Scaling Mobile Manipulation via Egocentric Human Data [Code] [Page] [
Affordances] - ๐๐ค [ICRA] RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation [Code] [Page] [
Affordances] - ๐๐ค [ICRA] DemoDiffusion: One-Shot Human Imitation Using Pre-Trained Diffusion Policy [Code] [Page] [
Affordances] - ๐๐ค [ICRA] NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos [Code] [Page] [
Affordances] - ๐๐ค [ICRA] Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2026.06] Unifying Egocentric Human and Robotic Data for VLA Pretraining [Code] [Page] [
VLA] - ๐ค [arXiv 2026.06] Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data [Code] [Page] [
VLA] - ๐๐ค [arXiv 2026.05] HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos [Code] [Page] [
Affordances] - ๐ค [arXiv 2026.04] ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration [Code] [Page] [
Affordances] - ๐ค [arXiv 2026.04] Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2026.04] Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2026.03] PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos [Code] [Page] [
Affordances] - ๐ค [arXiv 2026.02] EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data [Code] [Page] [
Affordances] - [arXiv 2026.02] MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2026.02] DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos [Code] [Page] [
Latent actions] - ๐ [CVPR] VideoWorld 2: Learning Transferable Knowledge from Real-world Videos [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2026.02] UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2026.02] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model [Code] [Page] [
Latent actions] - ๐ค [arXiv 2026.02] Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2026.02] ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2026.02] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2026.01] Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization [Code] [Page][
Affordances] - [arXiv 2026.01] Learning Latent Action World Models in the Wild [Code] [Page] [
Latent actions] - [arXiv 2026.01] ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos [Code] [Page] [
Affordances] - ๐ค [arXiv 2026.01] CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos [Code] [Page] [
Latent actions]
2025
- ๐๐ค [CVPR] ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning [Code] [Page] [
Affordances] - ๐๐ค [CVPR] Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning [Code] [Page] [
Affordances] - ๐ค [CVPR] GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning [Code] [Page] [
Affordances] - ๐๐ค [CVPR] VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation [Code] [Page] [
Affordances] - ๐๐ค [ICCV] AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning [Code] [Page] [
Affordances] - ๐๐ค [ICCV] DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot Handover [Code] [Page] [
Affordances] - ๐๐ค [ICCV] Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos [Code] [Page] [
Latent actions] - ๐๐ค [ICCV&CoRL Workshop] 2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos [Code] [Page] [
Affordances] - ๐๐ค [RSS] UniVLA: Learning to Act Anywhere with Task-Centric Latent Actions [Code] [Page] [
Latent actions] - ๐ค [RSS] You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations [Code] [Page] [
Affordances] - ๐ค [RSS Workshop] ViSA-Flow: Accelerating Robot Skill Learning via Large-Scale Video Semantic Action Flow [Code] [Page] [
Affordances] - [RSS Workshop] Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos [Code] [Page] [
Affordances] - ๐๐ค [ICRA] ViViDex: Learning Vision-Based Dexterous Manipulation from Human Videos [Code] [Page] [
Affordances] - ๐๐ค [ICRA] R+X: Retrieval and Execution from Everyday Human Videos [Code] [Page] [
Affordances] - ๐๐ค [ICRA] ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos [Code] [Page] [
Affordances] - ๐๐ค [ICRA] SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation [Code] [Page] [
Affordances] - ๐๐ค [ICRA] Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards [Code] [Page] [
Affordances] - ๐๐ค [ICRA] EgoMimic: Scaling Imitation Learning via Egocentric Video [Code] [Page] [
Affordances] - ๐๐ค [ICRA] Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning [Code] [Page] [
Affordances] - ๐๐ค [ICRA] Hand-Object Interaction Pretraining from Videos [Code] [Page] [
Affordances] - ๐๐ค [ICRA] Hand-Object Interaction Pretraining from Videos [Code] [Page] [
Latent actions] - ๐๐ค [CoRL] Generalist Robot Manipulation beyond Action Labeled Data [Code] [Page] [
Affordances] - ๐๐ค [CoRL] In-N-On: Scaling Egocentric Manipulation with In-the-Wild and On-Task Data [Code] [Page] [
Affordances] - ๐๐ค [CoRL] Humanoid Policy ~ Human Policy [Code] [Page] [
Affordances] - ๐๐ค [CoRL] X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real [Code] [Page] [
Affordances] - ๐๐ค [CoRL] Crossing the Human-Robot Embodiment Gap with Sim-to-Real RL Using One Human Demonstration [Code] [Page] [
Affordances] - ๐๐ค [CoRL] FUNCTO: Function-Centric One-Shot Imitation Learning for Tool Manipulation [Code] [Page] [
Affordances] - ๐๐ค [CoRL] MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence [Code] [Page] [
Affordances] - ๐ [CoRL] Articulated Object Estimation in the Wild [Code] [Page] [
Affordances] - ๐๐ค [CoRL] Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation [Code] [Page] [
Affordances] - ๐๐ค [CoRL] UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations [Code] [Page] [
Latent actions] - ๐๐ค [CoRL] General Flow as Foundation Affordance for Scalable Robot Learning [Code] [Page] [
Affordances] - ๐ค [CoRL Workshop] HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation [Code] [Page] [
Affordances] - ๐๐ค [Autonomous Robots] View: Visual Imitation Learning with Waypoints [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.12] GR-Dexter Technical Report [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.12] Emergence of Human to Robot Transfer in Vision-Language-Action Models [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.12] Motus: A Unified Latent Action World Model [Code] [Page] [
Latent actions] - ๐ค [arXiv 2025.12] SurgWorld: Learning Surgical Robot Policies from Videos via World Modeling [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.11] Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.11] Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.11] LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models [Code] [Page] [
Latent actions] - ๐ค [arXiv 2025.10] Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos [Code] [Page] [
Affordances] - ๐ [arXiv 2025.10] DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.10] From Human Hands to Robot Arms: Manipulation Skills Transfer via Trajectory Alignment [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.10] Traj2Action: A Co-Denoising Framework for Trajectory-Guided Human-to-Robot Skill Transfer [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.09] Developing Vision-Language-Action Model from Egocentric Videos [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.09] MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies [Code] [Page] [
Affordances] - ๐ [arXiv 2025.08] Deep Sensorimotor Control by Imitating Predictive Models of Human Motion [Code] [Page] [
Affordances] - ๐ [arXiv 2025.07] EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.07] Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.07] GR-3 Technical Report [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.07] villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models [Code] [Page] [
Latent actions] - ๐๐ค [arXiv 2025.06] Object-Centric 3D Motion Field for Robot Learning from Human Videos [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.05] EgoZero: Robot Learning from Smart Glasses [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.05] Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.04] Slot-Level Robotic Placement via Visual Imitation from Single Human Video [Code] [Page] [
Affordances] - ๐๐ค [arXiv 2025.03] GR00T N1: An Open Foundation Model for Generalist Humanoid Robots [Code] [Page] [
Affordances] - ๐ค [arXiv 2025.02] Video2Policy: Scaling Up Manipulation Tasks in Simulation Through Internet Videos [Code] [Page] [
Affordances]
2024
- ๐๐ค [ICLR] Learning to Act from Actionless Videos through Dense Correspondences [Code] [Page] [
Affordances] - ๐๐ค [ICLR] Latent Action Pretraining from Videos [Code] [Page] [
Latent actions] - ๐ค [ICLR] RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches [Code] [Page] [
Affordances] - ๐๐ค [ECCV] Track2Act: Predicting Point Tracks from Internet Videos Enables Generalizable Robot Manipulation [Code] [Page] [
Affordances] - ๐๐ค [ECCV] Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation [Code] [Page] [
Affordances] - ๐๐ค [RSS] HRP: Human Affordances for Robotic Pre-Training [Code] [Page] [
Affordances] - ๐๐ค [RSS] ScrewMimic: Bimanual Imitation from Human Videos with Screw Space Projection [Code] [Page] [
Affordances] - ๐ค [ICRA] Towards Generalizable Zero-Shot Manipulation via Translating Human Interaction Plan [Code] [Page] [
Affordances] - ๐๐ค [IROS] Ditto: Demonstration Imitation by Trajectory Transformation [Code] [Page] [
Affordances] - ๐๐ค [CoRL] RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation [Code] [Page] [
Affordances] - ๐๐ค [CoRL] Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction [Code] [Page] [
Affordances] - ๐๐ค [CoRL] OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation [Code] [Page] [
Affordances] - ๐๐ค [CoRL] Im2Flow2Act: Flow as the Cross-Domain Manipulation Interface [Code] [Page] [
Affordances] - [arXiv 2024.11] IGOR: Image-Goal Representations Are the Atomic Control Units for Foundation Models in Embodied AI [Code] [Page] [
Latent actions] - ๐ค [arXiv 2024.05] Vision-Based Manipulation from Single Human Video with Open-World Object Graphs [Code] [Page] [
Affordances]
2023
- ๐๐ค [TRO] K-VIL: Keypoints-Based Visual Imitation Learning [Code] [Page] [
Affordances] - ๐๐ค [ICRA] Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous Manipulation [Code] [Page] [
Affordances] - ๐๐ค [CVPR] Affordances from Human Videos as a Versatile Representation for Robotics [Code] [Page] [
Affordances] - ๐๐ค [CoRL] MimicPlay: Long-Horizon Imitation Learning by Watching Human Play [Code] [Page] [
Affordances] - ๐๐ค [CoRL] DEFT: Dexterous Fine-Tuning for Hand Policies [Code] [Page] [
Affordances] - ๐๐ค [RSS] Any-Point Trajectory Modeling for Policy Learning [Code] [Page] [
Affordances] - ๐๐ค [RA-L&IROS] Learning Continuous Grasping Function with a Dexterous Hand from Human Demonstrations [Code] [Page] [
Affordances] - ๐ค [ICRA Workshop] Zero-Shot Robot Manipulation from Passive Human Videos [Code] [Page] [
Affordances] - ๐ค [Sensors] Robot Programming from a Single Demonstration for High-Precision Industrial Insertion [Code] [Page] [
Affordances]
2022
- [RSS] You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration [Code] [Page] [
Affordances] - ๐ค [CoRL] VideoDex: Learning Dexterity from Internet Videos [Code] [Page] [
Affordances] - ๐๐ค [CoRL] Graph Inverse Reinforcement Learning from Diverse Videos [Code] [
Affordances] - ๐ค [RSS] Robotic Telekinesis: Learning a Robotic Hand Imitator by Watching Humans on YouTube [Code] [Page] [
Affordances] - ๐๐ค [RA-L&IROS] From One Hand to Multiple Hands: Imitation Learning for Dexterous Manipulation from Single-Camera Teleoperation [Code] [Page] [
Affordances] - ๐ค [RSS] Human-to-Robot Imitation in the Wild [Code] [Page] [
Affordances] - ๐ [ECCV] DexMV: Imitation Learning for Dexterous Manipulation from Human Videos [Code] [Page] [
Affordances] - [arXiv 2022.11] Learning to Imitate Object Interactions from Internet Videos [Code] [Page] [
Affordances]
2021
- [CoRL] DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video [Code] [Page] [
Affordances]
2020
- ๐ค [CoRL] Model-Based Inverse Reinforcement Learning from Visual Demonstrations [Code] [Page] [
Affordances]
2019
- ๐ค [CoRL] Graph-Structured Visual Imitation [Code] [Page] [
Affordances]
2017
- ๐ค [CVPR] Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression [Code] [Page] [
Affordances]
๐งฉ Data Foundations
Open-Source Datasets
This section collects open-source human video datasets.
| Dataset | Year | Venue | Website | Viewpoint | Organization (first author) |
|---|---|---|---|---|---|
| HumanNet | 2026 | arXiv 2026.05 | [Code] [Page] | Ego+Exo | Peking University |
| EgoLive | 2026 | arXiv 2026.04 | [Code] [Page] | Ego | JD |
| DreamDojo-HV | 2026 | arXiv 2026.02 | [Code] [Page] | Ego | NVIDIA |
| UniHand-Mix | 2026 | arXiv 2026.02 | [Code] [Page] | Ego | BeingBeyond |
| UniHand-2.0 | 2026 | arXiv 2026.01 | [Code] [Page] | Ego | BeingBeyond |
| Action100M | 2026 | arXiv 2026.01 | [Code] [Page] | Ego+Exo | Meta |
| EgoVid-5M | 2025 | NeurIPS | [Code] [Page] | Ego | Alibaba Group |
| HO-Cap | 2025 | NeurIPS | [Code] [Page] | Ego+Exo | University of Texas at Dallas |
| IndEgo | 2025 | NeurIPS | [Code] [Page] | Ego+Exo | Fraunhofer IPK |
| HD-EPIC | 2025 | CVPR | [Code] [Page] | Ego | University of Bristol |
| HOT3D | 2025 | CVPR | [Code] [Page] | Ego | Meta |
| TASTE-Rob | 2025 | CVPR | [Code] [Page] | Ego | CUHK-Shenzhen |
| LVP-1M | 2025 | arXiv 2025.12 | [Code] [Page] | Ego+Exo | MIT |
| UniHand-1.0 | 2025 | arXiv 2025.07 | [Code] [Page] | Ego | BeingBeyond |
| EgoDex | 2025 | arXiv 2025.05 | [Code] [Page] | Ego | Apple |
| PHยฒD | 2025 | arXiv 2025.03 | [Code] [Page] | Ego | UC San Diego |
| Egocentric-100k | 2025 | --- | [Code] [Page] | Ego | Build AI |
| Egocentric-10k | 2025 | --- | [Code] [Page] | Ego | Build AI |
| OakInk2 | 2024 | NeurIPS | [Code] [Page] | Ego+Exo | Shanghai Jiao Tong University |
| CaptainCook4D | 2024 | NeurIPS | [Code] [Page] | Ego | University of Texas at Dallas |
| Panda-70M | 2024 | CVPR | [Code] [Page] | Ego+Exo | Snap Inc. |
| TACO | 2024 | CVPR | [Code] [Page] | Ego+Exo | Tsinghua University |
| Ego-Exo4D | 2024 | CVPR | [Code] [Page] | Ego+Exo | Meta |
| Nymeria | 2024 | ECCV | [Code] [Page] | Ego+Exo | Meta |
| ARCTIC | 2023 | CVPR | [Code] [Page] | Ego+Exo | ETH Zรผrich |
| HoloAssist | 2023 | ICCV | [Code] [Page] | Ego | Microsoft |
| RH20T-Human | 2023 | ICRA | [Code] [Page] | Ego+Exo | Shanghai Jiao Tong University |
| EPIC-KITCHENS-100 | 2022 | IJCV | [Code] [Page] | Ego | University of Bristol |
| Assembly101 | 2022 | CVPR | [Code] [Page] | Ego+Exo | Meta |
| Ego4D | 2022 | CVPR | [Code] [Page] | Ego | Meta |
| OakInk | 2022 | CVPR | [Code] [Page] | Exo | Shanghai Jiao Tong University |
| HOI4D | 2022 | CVPR | [Code] [Page] | Ego | Tsinghua University |
| EgoPAT3D | 2022 | CVPR | [Code] [Page] | Ego | New York University |
| AGD20K | 2022 | CVPR | [Code] [Page] | Ego+Exo | University of Science and Technology of China |
| EgoHOS | 2022 | ECCV | [Code] [Page] | Ego | University of Pennsylvania |
| DexYCB | 2021 | CVPR | [Code] [Page] | Exo | NVIDIA |
| H2O | 2021 | ICCV | [Code] [Page] | Ego | ETH Zurich |
| MOW | 2021 | ICCV | [Code] [Page] | Exo | UC Berkeley |
| 100DOH | 2020 | CVPR | [Code] [Page] | Ego+Exo | University of Michigan |
| Kinetics-700 | 2020 | arXiv 2020.10 | [Code] [Page] | Exo | |
| HowTo100M | 2019 | ICCV | [Code] [Page] | Ego+Exo | ENS |
| FreiHAND | 2019 | ICCV | [Code] [Page] | Exo | University of Freiburg |
| FPHA | 2018 | CVPR | [Code] [Page] | Ego | Imperial College London |
| VLOG | 2018 | CVPR | [Code] [Page] | Exo | University of Michigan |
| EPIC-KITCHENS | 2018 | ECCV | [Code] [Page] | Ego | University of Bristol |
| EGTEA Gaze+ | 2018 | ECCV | [Code] [Page] | Ego | University of Wisconsin-Madison |
| YouCook2 | 2018 | AAAI | [Code] [Page] | Exo | University of Michigan |
| Something-Something | 2017 | ICCV | [Code] [Page] | Exo | TwentyBN |
| Charades | 2016 | ECCV | [Code] [Page] | Exo | Carnegie Mellon University |
| ActivityNet | 2015 | CVPR | [Code] [Page] | Exo | Universidad del Norte |
| EgoHands | 2015 | ICCV | [Code] [Page] | Ego | Indiana University |
| Breakfast | 2014 | CVPR | [Code] [Page] | Exo | Fraunhofer FKIE |
Human Video Generation
This section summarizes works that use video generation techniques to synthesize human videos for robot learning.
2026
- ๐๐ค [ICLR] Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations [Code] [Page]
- ๐ [CVPR] Dexterous World Models [Code] [Page]
- ๐๐ค [ICRA] NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos [Code] [Page]
2025
- ๐๐ค [CoRL] Dreamitate: Real-World Visuomotor Policy Learning via Video Generation [Code] [Page]
- ๐๐ค [arXiv 2025.12] Large Video Planner Enables Generalizable Robot Control [Code] [Page]
- ๐๐ค [arXiv 2025.12] Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow [Code] [Page]
2024
- ๐ค [arXiv 2024.09] Gen2Act: Human Video Generation in Novel Scenarios Enables Generalizable Robot Manipulation [Code] [Page]
2023
2020
HOI Analysis Techniques
This section gathers HOI analysis techniques, including hand/object detection and reconstruction, pose estimation, and point tracking tools widely used in LfHV research.
2026
- ๐ [arXiv 2026.02] SAM 3D Body: Robust Full-Body Human Mesh Recovery [Code] [Page]
2025
- ๐ [CVPR] WiLoR: End-to-End 3D Hand Localization and Reconstruction In-the-Wild [Code] [Page]
- ๐ [CVPR] HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos [Code] [Page]
- ๐ [CVPR] Any6D: Model-Free 6D Pose Estimation of Novel Objects [Code] [Page]
- ๐ [CVPR] Structured 3D Latents for Scalable and Versatile 3D Generation [Code] [Page]
- ๐ [ICCV] CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos [Code] [Page]
- ๐ [ICCV] SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera Motion [Code] [Page]
- [Meshy AI] Meshy AI: The #1 AI 3D Model Generator for Creators [Page]
2024
- ๐ [CVPR] Reconstructing Hands in 3D with Transformers [Code] [Page]
- ๐ [CVPR] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects [Code] [Page]
- ๐ [CVPR] SpatialTracker: Tracking Any 2D Pixels in 3D Space [Code] [Page]
- ๐ [ECCV] CoTracker: It Is Better to Track Together [Code] [Page]
- ๐ [ECCV] Local All-Pair Correspondence for Point Tracking [Code] [Page]
- ๐ [ACCV] Bootstap: Bootstrapped Training for Tracking-Any-Point [Code] [Page]
- ๐ [arXiv 2024.04] InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-View Large Reconstruction Models [Code] [Page]
2023
- ๐ [NeurIPS] Towards a Richer 2D Understanding of Hands at Scale [Code] [Page]
- ๐ [CVPR] BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects [Code] [Page]
- ๐ [ICCV] TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal Refinement [Code] [Page]
2022
- ๐ [CVPR] Human Hands as Probes for Interactive Object Understanding [Code] [Page]
- ๐ [CoRL] MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare [Code] Page]
- ๐ [ICCVW] FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration [Code] [Page]
2020
- ๐ [CVPR] Understanding Human Hands in Contact at Internet Scale [Code] [Page]
- ๐ [CVPRW] MediaPipe Hands: On-Device Real-Time Hand Tracking [Code] [Page]
Other vision foundation models like Grounding DINO, YOLO series, SAM series are also widely used in the LfHV literature.
๐ LICENSE
This project is licensed under the MIT License - see the LICENSE file for details.
๐ฌ Contact
If you have any questions or suggestions, please feel free to contact Junyi Ma.