motion_24-25.md

July 24, 2026 ยท View on GitHub

๐ŸŽ‰ 2025 Motion Accepted Papers

YearTitleVenuePaperCodeProject Page
2025MixerMDM: Learnable Composition of Human Motion Diffusion ModelsCVPR 2025LinkLinkLink
2025Dynamic Motion Blending for Versatile Motion EditingCVPR 2025LinkLinkLink
2025MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial TrainingCVPR 2025 HuMoGen WorkshopLinkLinkLink
2025MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent SpaceICCV 2025LinkLinkLink
2025Go to Zero: Towards Zero-shot Motion Generation with Million-scale DataICCV 2025LinkLinkLink
2025KinMo: Kinematic-aware Human Motion Understanding and GenerationICCV 2025Link--Link
2025MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmICCV 2025LinkLinkLink
2025GENMO: A GENeralist Model for Human MOtionICCV 2025 (Highlight)LinkLinkLink
2025ControlMM: Controllable Masked Motion GenerationICCV 2025 (Oral)LinkLinkLink
2025ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided AlignmentAAAI 2026LinkLinkLink
2025SnapMoGen: Human Motion Generation from Expressive TextsNeurIPS 2025LinkLinkLink
2025Motion Anything: Any to Motion GenerationNeurIPS 2025LinkLinkLink
2025MotionBind: Multi-Modal Human Motion Alignment for Retrieval, Recognition, and GenerationNeurIPS 2025LinkLinkLink
2025MOSPA: Human Motion Generation Driven by Spatial AudioNeurIPS 2025LinkLinkLink
2025Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion ModelsICLR 2026Link--Link
2025EgoTwin: Dreaming Body and View in First PersonICLR 2026Link--Link
2025X-MoGen: Unified Motion Generation across Humans and AnimalsAAAI 2026Link----
2025FlowMotion: Target-Predictive Conditional Flow Matching for Jitter-Reduced Text-Driven Human Motion GenerationComputers & Graphics 2025LinkLinkLink
Accepted Papers References
%accepted papers

@article{ruiz2025mixermdm,
  title={MixerMDM: Learnable Composition of Human Motion Diffusion Models},
  author={Ruiz-Ponce, Pablo and Barquero, German and Palmero, Cristina and Escalera, Sergio and Garc{\'\i}a-Rodr{\'\i}guez, Jos{\'e}},
  journal={arXiv preprint arXiv:2504.01019},
  year={2025}
}

@article{jiang2025dynamic,
  title={Dynamic Motion Blending for Versatile Motion Editing},
  author={Jiang, Nan and Li, Hongjie and Yuan, Ziye and He, Zimo and Chen, Yixin and Liu, Tengyu and Zhu, Yixin and Huang, Siyuan},
  journal={arXiv preprint arXiv:2503.20724},
  year={2025}
}

@inproceedings{uchida2025mola,
  title={Mola: Motion generation and editing with latent diffusion enhanced by adversarial training},
  author={Uchida, Kengo and Shibuya, Takashi and Takida, Yuhta and Murata, Naoki and Tanke, Julian and Takahashi, Shusuke and Mitsufuji, Yuki},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={2910--2919},
  year={2025}
}

@article{xiao2025motionstreamer,
      title={MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space},
      author={Xiao, Lixing and Lu, Shunlin and Pi, Huaijin and Fan, Ke and Pan, Liang and Zhou, Yueer and Feng, Ziyong and Zhou, Xiaowei and Peng, Sida and Wang, Jingbo},
      journal={arXiv preprint arXiv:2503.15451},
      year={2025}
}

@misc{fan2025zerozeroshotmotiongeneration,
      title={Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data}, 
      author={Ke Fan and Shunlin Lu and Minyue Dai and Runyi Yu and Lixing Xiao and Zhiyang Dou and Junting Dong and Lizhuang Ma and Jingbo Wang},
      year={2025},
      eprint={2507.07095},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2507.07095}, 
}

@inproceedings{kinmo2025kinematicawarehumanmotion,
      title={{KinMo: Kinematic-aware Human Motion Understanding and Generation}},
      author={Pengfei Zhang and Pinxin Liu and Pablo Garrido and Hyeongwoo Kim and Bindita Chaudhuri},
      booktitle={IEEE/CVF International Conference on Computer Vision},
      year={2025},
}

@article{guo2025motionlab,
  title={Motionlab: Unified human motion generation and editing via the motion-condition-motion paradigm},
  author={Guo, Ziyan and Hu, Zeyu and Soh, De Wen and Zhao, Na},
  journal={arXiv preprint arXiv:2502.02358},
  year={2025}
}

@article{li2025genmo,
  title={GENMO: A GENeralist Model for Human MOtion},
  author={Li, Jiefeng and Cao, Jinkun and Zhang, Haotian and Rempe, Davis and Kautz, Jan and Iqbal, Umar and Yuan, Ye},
  journal={arXiv preprint arXiv:2505.01425},
  year={2025}
}

@misc{pinyoanuntapong2025maskcontrolspatiotemporalcontrolmasked,
      title={MaskControl: Spatio-Temporal Control for Masked Motion Synthesis}, 
      author={Ekkasit Pinyoanuntapong and Muhammad Usama Saleem and Korrawe Karunratanakul and Pu Wang and Hongfei Xue and Chen Chen and Chuan Guo and Junli Cao and Jian Ren and Sergey Tulyakov},
      year={2025},
      eprint={2410.10780},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2410.10780}, 
}

@article{weng2025realign,
  title={ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment},
  author={Weng, Wanjiang and Tan, Xiaofeng and Wang, Hongsong and Zhou, Pan},
  journal={arXiv preprint arXiv:2505.04974},
  year={2025}
}

@article{guo2025snapmogen,
  title={SnapMoGen: Human Motion Generation from Expressive Texts},
  author={Guo, Chuan and Hwang, Inwoo and Wang, Jian and Zhou, Bing},
  journal={arXiv preprint arXiv:2507.09122},
  year={2025}
}


@article{zhang2025motion,
  title={Motion anything: Any to motion generation},
  author={Zhang, Zeyu and Wang, Yiran and Mao, Wei and Li, Danning and Zhao, Rui and Wu, Biao and Song, Zirui and Zhuang, Bohan and Reid, Ian and Hartley, Richard},
  journal={arXiv preprint arXiv:2503.06955},
  year={2025}
}

@inproceedings{kinfumotionbind,
  title={MotionBind: Multi-Modal Human Motion Alignment for Retrieval, Recognition, and Generation},
  author={Kinfu, Kaleab A and Vidal, Rene},
  booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems}
}

@article{xu2025mospa,
  title={Mospa: Human motion generation driven by spatial audio},
  author={Xu, Shuyang and Dou, Zhiyang and Shi, Mingyi and Pan, Liang and Ho, Leo and Wang, Jingbo and Liu, Yuan and Lin, Cheng and Ma, Yuexin and Wang, Wenping and others},
  journal={arXiv preprint arXiv:2507.11949},
  year={2025}
}

@article{millan2025animating,
  title={Animating the uncaptured: Humanoid mesh animation with video diffusion models},
  author={Mill{\'a}n, Marc Bened{\'\i} San and Dai, Angela and Nie{\ss}ner, Matthias},
  journal={arXiv preprint arXiv:2503.15996},
  year={2025}
}

@article{xiu2025egotwin,
  title={Egotwin: Dreaming body and view in first person},
  author={Xiu, Jingqiao and Hong, Fangzhou and Li, Yicong and Li, Mengze and Wang, Wentao and Han, Sirui and Pan, Liang and Liu, Ziwei},
  journal={arXiv preprint arXiv:2508.13013},
  year={2025}
}

@inproceedings{wang2026x,
  title={X-MoGen: Unified Motion Generation Across Humans and Animals},
  author={Wang, Xuan and Ruan, Kai and Qian, Liyang and Zhi, Guo Zhi and Su, Chang and Wang, Gaoang},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  volume={40},
  number={12},
  pages={10234--10242},
  year={2026}
}

@article{cuba2025flowmotion,
  title={FlowMotion: Target-predictive conditional flow matching for Jitter-Reduced text-driven human motion generation},
  author={Cuba, Manolo Canales and do Carmo Mel{\'\i}cio, Vin{\'\i}cius and Gois, Jo{\~a}o Paulo},
  journal={Computers \& Graphics},
  pages={104374},
  year={2025},
  publisher={Elsevier}
}


๐Ÿ’ก 2025 Motion ArXiv Papers

1. UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes

Zichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Mian (University of Western Australia, Data61 CSIRO Australia)

Abstract Human motion synthesis in complex scenes presents a fundamental challenge, extending beyond conventional Text-to-Motion tasks by requiring the integration of diverse modalities such as static environments, movable objects, natural language prompts, and spatial waypoints. Existing language-conditioned motion models often struggle with scene-aware motion generation due to limitations in motion tokenization, which leads to information loss and fails to capture the continuous, context-dependent nature of 3D human movement. To address these issues, we propose UniHM, a unified motion language model that leverages diffusion-based generation for synthesizing scene-aware human motion. UniHM is the first framework to support both Text-to-Motion and Text-to-Human-Object Interaction (HOI) in complex 3D scenes. Our approach introduces three key contributions: (1) a mixed-motion representation that fuses continuous 6DoF motion with discrete local motion tokens to improve motion realism; (2) a novel Look-Up-Free Quantization VAE (LFQ-VAE) that surpasses traditional VQ-VAEs in both reconstruction accuracy and generative performance; and (3) an enriched version of the Lingo dataset augmented with HumanML3D annotations, providing stronger supervision for scene-specific motion learning. Experimental results demonstrate that UniHM achieves comparative performance on the OMOMO benchmark for text-to-HOI synthesis and yields competitive results on HumanML3D for general text-conditioned motion generation.

2. ReMoMask: Retrieval-Augmented Masked Motion Generation

Zhengdao Li, Siheng Wang, Zeyu Zhang, Hao Tang (Peking University, Jiangsu University)

Abstract Text-to-Motion (T2M) generation aims to synthesize realistic and semantically aligned human motion sequences from natural language descriptions. However, current approaches face dual challenges: Generative models (e.g., diffusion models) suffer from limited diversity, error accumulation, and physical implausibility, while Retrieval-Augmented Generation (RAG) methods exhibit diffusion inertia, partial-mode collapse, and asynchronous artifacts. To address these limitations, we propose ReMoMask, a unified framework integrating three key innovations: 1) A Bidirectional Momentum Text-Motion Model decouples negative sample scale from batch size via momentum queues, substantially improving cross-modal retrieval precision; 2) A Semantic Spatio-temporal Attention mechanism enforces biomechanical constraints during part-level fusion to eliminate asynchronous artifacts; 3) RAG-Classier-Free Guidance incorporates minor unconditional generation to enhance generalization. Built upon MoMask's RVQ-VAE, ReMoMask efficiently generates temporally coherent motions in minimal steps. Extensive experiments on standard benchmarks demonstrate the state-of-the-art performance of ReMoMask, achieving a 3.88% and 10.97% improvement in FID scores on HumanML3D and KIT-ML, respectively, compared to the previous SOTA method RAG-T2M.

3. Pulp Motion: Framing-aware multimodal camera and human motion generation

Robin Courant, Xi Wang, David Loiseaux, Marc Christie, Vicky Kalogeiton

(LIX, Ecole Polytechnique, IP Paris; Inria Saclay; Inria, IRISA, CNRS, Univ. Rennes)

Abstract Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the first to cast this task as a text-conditioned joint generation, aiming to maintain consistent on-screen framing while producing two heterogeneous, yet intrinsically linked, modalities: human motion and camera trajectories. We propose a simple, model-agnostic framework that enforces multimodal coherence via an auxiliary modality: the on-screen framing induced by projecting human joints onto the camera. This on-screen framing provides a natural and effective bridge between modalities, promoting consistency and leading to more precise joint distribution. We first design a joint autoencoder that learns a shared latent space, together with a lightweight linear transform from the human and camera latents to a framing latent. We then introduce auxiliary sampling, which exploits this linear transform to steer generation toward a coherent framing modality. To support this task, we also introduce the PulpMotion dataset, a human-motion and camera-trajectory dataset with rich captions, and high-quality human motions. Extensive experiments across DiT- and MAR-based architectures show the generality and effectiveness of our method in generating on-frame coherent human-camera motions, while also achieving gains on textual alignment for both modalities. Our qualitative results yield more cinematographically meaningful framings setting the new state of the art for this task.

4. Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation

Qingxuan Wu, Zhiyang Dou, Chuan Guo, Yiming Huang, Qiao Feng, Bing Zhou, Jian Wang, Lingjie Liu

(University of Pennsylvania, The University of Hong Kong, Snap Inc.)

Abstract Modeling human-human interactions from text remains challenging because it requires not only realistic individual dynamics but also precise, text-consistent spatiotemporal coupling between agents. Currently, progress is hindered by 1) limited two-person training data, inadequate to capture the diverse intricacies of two-person interactions; and 2) insufficiently fine-grained text-to-interaction modeling, where language conditioning collapses rich, structured prompts into a single sentence embedding. To address these limitations, we propose our Text2Interact framework, designed to generate realistic, text-aligned human-human interactions through a scalable high-fidelity interaction data synthesizer and an effective spatiotemporal coordination pipeline. First, we present InterCompose, a scalable synthesis-by-composition pipeline that aligns LLM-generated interaction descriptions with strong single-person motion priors. Given a prompt and a motion for an agent, InterCompose retrieves candidate single-person motions, trains a conditional reaction generator for another agent, and uses a neural motion evaluator to filter weak or misaligned samples-expanding interaction coverage without extra capture. Second, we propose InterActor, a text-to-interaction model with word-level conditioning that preserves token-level cues (initiation, response, contact ordering) and an adaptive interaction loss that emphasizes contextually relevant inter-person joint pairs, improving coupling and physical plausibility for fine-grained interaction modeling. Extensive experiments show consistent gains in motion diversity, fidelity, and generalization, including out-of-distribution scenarios and user studies. We will release code and models to facilitate reproducibility.
YearTitleArXiv TimePaperCodeProject Page
2025UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes19 May 2025Link----
2025ReMoMask: Retrieval-Augmented Masked Motion Generation4 Aug 2025LinkLinkLink
2025Pulp Motion: Framing-aware multimodal camera and human motion generation6 Oct 2025LinkLinkLink
2025Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation7 Oct 2025Link----
ArXiv Papers References
%axiv papers

@misc{geng2025unihmuniversalhumanmotion,
      title={UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes}, 
      author={Zichen Geng and Zeeshan Hayder and Wei Liu and Ajmal Mian},
      year={2025},
      eprint={2505.12774},
      archivePrefix={arXiv},
      primaryClass={cs.GR},
      url={https://arxiv.org/abs/2505.12774}, 
}

@article{li2025remomask,
  title={ReMoMask: Retrieval-Augmented Masked Motion Generation},
  author={Li, Zhengdao and Wang, Siheng and Zhang, Zeyu and Tang, Hao},
  journal={arXiv preprint arXiv:2508.02605},
  year={2025}
}

@misc{courant2025pulpmotionframingawaremultimodal,
      title={Pulp Motion: Framing-aware multimodal camera and human motion generation}, 
      author={Robin Courant and Xi Wang and David Loiseaux and Marc Christie and Vicky Kalogeiton},
      year={2025},
      eprint={2510.05097},
      archivePrefix={arXiv},
      primaryClass={cs.GR},
      url={https://arxiv.org/abs/2510.05097}, 
}

@misc{wu2025text2interacthighfidelitydiversetexttotwoperson,
      title={Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation}, 
      author={Qingxuan Wu and Zhiyang Dou and Chuan Guo and Yiming Huang and Qiao Feng and Bing Zhou and Jian Wang and Lingjie Liu},
      year={2025},
      eprint={2510.06504},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2510.06504}, 
}

๐Ÿ“š Dataset Works

1. HUMOTO: A 4D Dataset of Mocap Human Object Interactions

Jiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang, Yi Zhou

(University of Texas at Austin, Adobe Research)

Abstract We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 736 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated parts. Our innovations include a scene-driven LLM scripting pipeline creating complete, purposeful tasks with natural progression, and a mocap-and-camera recording setup to effectively handle occlusions. Spanning diverse activities from cooking to outdoor picnics, HUMOTO preserves both physical accuracy and logical task flow. Professional artists rigorously clean and verify each sequence, minimizing foot sliding and object penetrations. We also provide benchmarks compared to other datasets. HUMOTO's comprehensive full-body motion and simultaneous multi-object interactions address key data-capturing challenges and provide opportunities to advance realistic human-object interaction modeling across research domains with practical applications in animation, robotics, and embodied AI systems.
YearTitleArXiv TimePaperDataset PageProject Page
2025HUMOTO: A 4D Dataset of Mocap Human Object InteractionsICCV 2025LinkLinkLink
References
%axiv papers

@article{lu2025humoto,
  title={HUMOTO: A 4D Dataset of Mocap Human Object Interactions},
  author={Lu, Jiaxin and Huang, Chun-Hao Paul and Bhattacharya, Uttaran and Huang, Qixing and Zhou, Yi},
  journal={arXiv preprint arXiv:2504.10414},
  year={2025}
}