H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos

December 12, 2025 Β· View on GitHub

Hai Ci, Xiaokang Liu, Pei Yang, Yiren Song, Mike Zheng Shou*
Show Lab, National University of Singapore
*Corresponding author

πŸ“„ Paper (arXiv): https://arxiv.org/abs/2512.09406
🌐 Project Page: https://showlab.github.io/H2R-Grounder/


⚑ TL;DR

H2R-Grounder converts third-person human interaction videos into frame-aligned robot manipulation videos β€” using no paired human–robot data for training.


πŸ“· Method Overview

H2R-Grounder Pipeline

Figure: H2R-Grounder pipeline. We extract pose and background to form H2Rep, then use a diffusion-based in-context model to generate physically grounded robot videos aligned with human actions.


πŸŽ₯ Qualitative Results

Visit our project page for full videos, comparisons, ablations, and failure case analysis:

πŸ‘‰ https://showlab.github.io/H2R-Grounder/


πŸ“¦ Code & Models

Code and models will be released soon.


✏️ Citation

@article{ci2025h2rgrounder,
  title={H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos},
  author={Ci, Hai and Liu, Xiaokang and Yang, Pei and Song, Yiren and Shou, Mike Zheng},
  journal={arXiv preprint arXiv:2512.09406},
  year={2025}
}