MotionHub

July 6, 2026 · View on GitHub

MotionHub is a curated multi-domain human-motion dataset collection released for training and evaluating generalist motion models. The released version contains motion, language, music, speech, and two-person interaction supervision in a unified MotionHub annotation format.

This public release is used by VersatileMotion (ECCV 2026). Every subset listed below has been visually inspected, converted to the repository SMPL-H convention, re-split where needed, and uploaded after data-quality review.

Release links: Hugging Face dataset · GitHub repository and preview assets

Highlights

ScaleLanguage SupervisionAudio, Music, and Interaction
20 released subsets
1.11M clips
1,528 h motion
164.98M frames
3.31M text-to-motion prompts
3.31M motion-to-text references
macro / meso / micro caption levels
5.6K music-to-dance pairs
44.1K speech/audio-to-gesture pairs
44.1K script-to-gesture scripts
23.6K interaction text-to-motion pairs

Modality Previews

The previews below are rendered with Three.js from the released SMPL-H motion files and grouped by modality rather than task direction. Music and speech examples include the paired source audio where available.

Text and Motion
Music and Dance
Speech and Gesture
Two-Person Interaction

Quick Start

Download the complete release:

huggingface-cli download ZeyuLing/MotionHub \
  --repo-type dataset \
  --local-dir MotionHub

Download one subset only:

huggingface-cli download ZeyuLing/MotionHub \
  --repo-type dataset \
  --include "aist/**" "aist/*.json" \
  --local-dir MotionHub

Python access:

from huggingface_hub import snapshot_download

root = snapshot_download(
    repo_id="ZeyuLing/MotionHub",
    repo_type="dataset",
    local_dir="MotionHub",
)

Data Format

The release keeps motion assets, task annotations, and statistics in separate top-level locations:

annotations/
  all/
    train.json
    test.json
  text_motion/
    train.json
    test.json
  music_dance/
    train.json
    test.json
  speech_gesture/
    train.json
    test.json
  two_person_interaction/
    train.json
    test.json
    placement_radii.json    # 2P augmentation placement metadata
  subsets/
    <subset>/               # original per-subset train/test split files
  humanml3d/                # official HumanML3D split helpers

statistics/
  smplh_universal_stats.json
  smplh_universal_stats_aug.json

<subset>/
  smplh_52/                  # SMPL-H motion NPZ files
  hierarchical_caption/       # macro / meso / micro captions
  ...                         # optional music, audio, speech, or pair metadata

The motion files are normalized to the MotionHub SMPL-H convention used in this repository. In particular, trans / transl stores the body-model translation parameter, and the data should not be re-canonicalized in a viewer before quality inspection.

Normalization uses the shared SMPL-H statistics in statistics/; per-subset statistics are intentionally not part of the release surface.

MotionGV is train-only by design and is included in annotations/text_motion/train.json; it does not expose a separate task split.

Released Subsets

FamilyDatasetPathClipsHoursSupervisionCite
Single-person text-motionAMASS_SUPamass_sup7,67324.95T2M, M2TAMASS
Single-person text-motionCombatMotionCombatMotion_seperate25,98723.26T2M, M2TCombatMotion
Single-person text-motionEgoBodyEgoBody9804.06T2M, M2TEgoBody
Single-person text-motionFit3Dfit3d9443.14T2M, M2TFit3D
Single-person text-motionGRABGRAB1,3353.76T2M, M2TGRAB
Single-person text-motionHuman3.6Mhuman36m9252.98T2M, M2THuman3.6M
Single-person text-motionHumanML3D-AMASSHumanML3D_AMASS29,12054.14T2M, M2THumanML3D + AMASS
Single-person text-motionHumanML3D-HumanAct12HumanML3D_HumanACT122,3822.05T2M, M2THumanML3D + HumanAct12
Single-person text-motionHumanSC3Dhumansc3d6881.12T2M, M2THumanSC3D
Single-person text-motionMotionGVMotionGV833,1211,114T2M, M2TMotionMillion / Go to Zero
Single-person text-motionNTU RGB+D 120nturgbd120106,86471.77T2M, M2TNTU RGB+D 120
Single-person text-motionPerMopermo6,6108.56T2M, M2TPersonaBooth / PerMo
Single-person text-motionTRUMANStrumans3,6236.89T2M, M2TTRUMANS
Dance and musicAIST++aist1,4085.20T2M, M2T, music-to-danceAIST++
Dance and musicFineDancefinedance4,19413.98T2M, M2T, music-to-danceFineDance
Speech and gestureBEAT v2.0.0beat_v2.0.021,60357.09T2M, M2T, audio-to-gesture, script-to-gestureBEAT
Speech and gestureTED-DBted_db22,54872.64T2M, M2T, audio-to-gesture, script-to-gestureTED Gesture
Two-person interactionChi3Dchi3d9121.48T2M, M2T, interaction T2MCHI3D
Two-person interactionHi4Dhi4d3000.33T2M, M2T, interaction T2MHi4D
Two-person interactionInterXinterx34,16156.25T2M, M2T, interaction T2MInterX

Task Coverage

TaskSupervision sourceCount
Text-to-motionhierarchical captions to motion3,307,116 prompts
Motion-to-textmotion to macro / meso / micro captions3,307,116 references
Music-to-dancesynchronized music and dance motion5,602 pairs
Speech/audio-to-gesturespeech audio and gesture motion44,134 pairs
Script-to-gesturespeech transcript and gesture motion44,142 scripts
Interaction text-to-motiontwo-person captions and paired motions23,582 pairs

Detailed Inventory

Open the full subset statistics table
DatasetSplitsClipsMotion refsFramesHoursMusic refsT2M promptsM2T motions / refsMusic pairsInvalid skipped
CombatMotion_seperatetrain:25,887; test:10025,98725,9872,512,09323.26077,96125,987 / 77,96100
EgoBodytrain:931; test:49980980438,9564.0602,940980 / 2,94000
GRABtrain:1,268; test:671,3351,335406,2643.7604,0051,335 / 4,00500
HumanML3D_AMASStrain:25,160; test:3,96029,12029,1205,846,94054.14087,36029,120 / 87,36000
HumanML3D_HumanACT12train:2,040; test:3422,3822,382221,2322.0507,1462,382 / 7,14600
MotionGVtrain:833,121833,121833,121120,306,8591,11402,499,351833,117 / 2,499,35100
aisttrain:1,388; test:201,4081,408562,0915.201,4084,2241,408 / 4,2241,4080
amass_suptrain:7,373; test:3007,6737,6732,694,69124.95023,0197,673 / 23,01900
beat_v2.0.0train:21,234; test:36921,60321,6036,165,86157.09055,80318,601 / 55,80300
chi3dtrain:819; test:939121,216159,5641.4802,736912 / 2,73600
finedancetrain:4,097; test:974,1944,1941,509,84013.984,19412,5824,194 / 12,5824,19443
fit3dtrain:934; test:10944944338,9043.1402,832944 / 2,83200
hi4dtrain:231; test:6930040035,8350.330900300 / 90000
human36mtrain:915; test:10925925322,1722.9802,775925 / 2,77500
humansc3dtrain:653; test:35688688120,9781.1202,064688 / 2,06400
interxtrain:29,037; test:5,12434,16145,5486,074,61956.250102,48334,161 / 102,48300
nturgbd120train:105,904; test:960106,864106,8647,751,04971.770320,592106,864 / 320,59200
permotrain:6,543; test:676,6106,610924,7268.56019,8306,610 / 19,83000
ted_dbtrain:22,179; test:36922,54822,5487,845,11272.64067,64422,548 / 67,64400
trumanstrain:3,586; test:373,6233,623743,5846.89010,8693,623 / 10,86900

Quality and Scope Notes

  • This release includes only subsets that have passed visual inspection and format review.
  • Low-quality or ambiguous subsets from earlier internal processing passes are intentionally excluded from the public release.
  • Source datasets retain their own licenses and usage restrictions. Please check and follow the upstream license for every subset you use.
  • The statistics above are counted from MotionHub annotations. Frame counts are read from num_frames when available, otherwise from duration and FPS.

Citation

Please cite MotionHub through the VersatileMotion ECCV 2026 paper and also cite every original subset used in your experiment.

MotionHub / VersatileMotion

@inproceedings{ling2026versatilemotion,
  title={VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension},
  author={Ling, Zeyu and Han, Bo and Li, Shiyang and Cheng, Jikang and Shen, Hongdeng and Zou, Changqing},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}

Source Datasets

AMASS / AMASS_SUP

@conference{AMASS:ICCV:2019,
  title={{AMASS}: Archive of Motion Capture as Surface Shapes},
  author={Mahmood, Naureen and Ghorbani, Nima and Troje, Nikolaus F. and Pons-Moll, Gerard and Black, Michael J.},
  booktitle={International Conference on Computer Vision},
  pages={5442--5451},
  year={2019}
}

AIST++

@inproceedings{li2021aistpp,
  title={AI Choreographer: Music Conditioned 3D Dance Generation with AIST++},
  author={Li, Ruilong and Yang, Shan and Ross, David A. and Kanazawa, Angjoo},
  booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
  year={2021}
}

BEAT

@inproceedings{liu2022beat,
  title={BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis},
  author={Liu, Haiyang and Zhu, Zihao and Iwamoto, Naoya and Peng, Yichen and Li, Zhengqing and Zhou, You and Bozkurt, Elif and Zheng, Bo},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2022}
}

Chi3D

@inproceedings{fieraru2020chi3d,
  title={Three-Dimensional Reconstruction of Human Interactions},
  author={Fieraru, Mihai and Zanfir, Mihai and Oneata, Elisabeta and Popa, Alin-Ionut and Olaru, Vlad and Sminchisescu, Cristian},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2020}
}

CombatMotion

@misc{liao2024animationgpt,
  title={AnimationGPT: An AIGC Tool for Generating Game Combat Motion Assets},
  author={Liao, Yihao and Fu, Yiyu and Cheng, Ziming and Wang, Jiangfeiyang},
  year={2024},
  howpublished={\url{https://github.com/fyyakaxyy/AnimationGPT}}
}

EgoBody

@inproceedings{zhang2022egobody,
  title={EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices},
  author={Zhang, Siwei and Ma, Qianli and Zhang, Yan and Qian, Zhiyin and Kwon, Taein and Pollefeys, Marc and Bogo, Federica and Tang, Siyu},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2022}
}

FineDance

@inproceedings{li2023finedance,
  title={FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation},
  author={Li, Ronghui and Zhao, Junfan and Zhang, Yachao and Su, Mingyang and Ren, Zeping and Zhang, Han and Tang, Yansong and Li, Xiu},
  booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
  year={2023}
}

Fit3D

@inproceedings{fieraru2021aifit,
  title={AIFit: Automatic 3D Human-Interpretable Feedback Models for Fitness Training},
  author={Fieraru, Mihai and Zanfir, Mihai and Pirlea, Silviu-Cristian and Olaru, Vlad and Sminchisescu, Cristian},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2021}
}

GRAB

@inproceedings{taheri2020grab,
  title={{GRAB}: A Dataset of Whole-Body Human Grasping of Objects},
  author={Taheri, Omid and Ghorbani, Nima and Black, Michael J. and Tzionas, Dimitrios},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2020}
}

Hi4D

@inproceedings{yin2023hi4d,
  title={Hi4D: 4D Instance Segmentation of Close Human Interaction},
  author={Yin, Yifei and Guo, Chen and Kaufmann, Manuel and Zarate, Juan Jose and Song, Jie and Hilliges, Otmar},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2023}
}

Human3.6M

@article{h36m_pami,
  title={Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments},
  author={Ionescu, Catalin and Papava, Dragos and Olaru, Vlad and Sminchisescu, Cristian},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  volume={36},
  number={7},
  pages={1325--1339},
  year={2014}
}

HumanAct12

@inproceedings{guo2020action2motion,
  title={Action2Motion: Conditioned Generation of 3D Human Motions},
  author={Guo, Chuan and Zuo, Xinxin and Wang, Sen and Zou, Shihao and Sun, Qingyao and Deng, Annan and Gong, Minglun and Cheng, Li},
  booktitle={ACM International Conference on Multimedia},
  pages={2021--2029},
  year={2020}
}

HumanML3D

@inproceedings{guo2022generating,
  title={Generating Diverse and Natural 3D Human Motions from Text},
  author={Guo, Chuan and Zuo, Xinxin and Wang, Sen and Zou, Shihao and Sun, Qingyao and Deng, Annan and Gong, Minglun and Cheng, Li},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2022}
}

HumanSC3D

@inproceedings{fieraru2021learning,
  title={Learning Complex 3D Human Self-Contact},
  author={Fieraru, Mihai and Zanfir, Mihai and Oneata, Elisabeta and Popa, Alin-Ionut and Olaru, Vlad and Sminchisescu, Cristian},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  volume={35},
  number={2},
  pages={1343--1351},
  year={2021}
}

InterX

@inproceedings{xu2024interx,
  title={Inter-X: Towards Versatile Human-Human Interaction Analysis},
  author={Xu, Liang and Lv, Xintao and Yan, Yichao and Jin, Xin and Wu, Shuwen and Xu, Congsheng and Liu, Yifan and Zhou, Yizhou and Rao, Fengyun and Sheng, Xingdong and Liu, Yunhui and Zeng, Wenjun and Yang, Xiaokang},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2024}
}

MotionGV / MotionMillion

@article{fan2025go,
  title={Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data},
  author={Fan, Ke and Lu, Shunlin and Dai, Minyue and Yu, Runyi and Xiao, Lixing and Dou, Zhiyang and Dong, Junting and Ma, Lizhuang and Wang, Jingbo},
  journal={arXiv preprint arXiv:2507.07095},
  year={2025}
}

NTU RGB+D 120

@article{liu2020ntu,
  title={NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding},
  author={Liu, Jun and Shahroudy, Amir and Perez, Mauricio and Wang, Gang and Duan, Ling-Yu and Kot, Alex C},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  volume={42},
  number={10},
  pages={2684--2701},
  year={2020}
}

PerMo / PersonaBooth

@inproceedings{kim2025personabooth,
  title={PersonaBooth: Personalized Text-to-Motion Generation},
  author={Kim, Boeun and Jeong, Hea In and Sung, JungHoon and Cheng, Yihua and Lee, Jeongmin and Chang, Ju Yong and Choi, Sang-Il and Choi, Younggeun and Shin, Saim and Kim, Jungho and Chang, Hyung Jin},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2025}
}

TED Gesture

@inproceedings{yoon2019ted_gesture,
  title={Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots},
  author={Yoon, Youngwoo and Ko, Woo-Ri and Jang, Minsu and Lee, Jaeyeon and Kim, Jaehong and Lee, Geehyuk},
  booktitle={IEEE International Conference on Robotics and Automation (ICRA)},
  year={2019}
}

TRUMANS

@inproceedings{jiang2024trumans,
  title={Scaling Up Dynamic Human-Scene Interaction Modeling},
  author={Jiang, Nan and Zhang, Zhiyuan and Li, Hongjie and Ma, Xiaoxuan and Wang, Zan and Chen, Yixin and Liu, Tengyu and Zhu, Yixin and Huang, Siyuan},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2024}
}