README.md

September 10, 2026 ยท View on GitHub

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Models

Xiaomi Robotics

Paper Project Page Hugging Face ModelScope License

Xiaomi-Robotics-U0 model architecture with autoregressive generation and FlashAR acceleration.
HighlightSummary
๐Ÿง World Foundation ModelA 34B autoregressive model for text, images, and embodied observations, initialized from EMU3.5.
๐ŸงฉUnified Token SpaceUses a shared discrete visual tokenizer and a single next-token objective across multimodal sequences.
๐Ÿค–Embodied SynthesisBridges foundation image generation with robot-centric scene, transfer, and video generation.
โšกXiaomi-Robotics-U0-FlashAR AccelerationDecodes visual tokens in anti-diagonal groups and supports vLLM batching for high-resolution inference.
๐Ÿ“ฆOpen Inference RepoProvides inference code, composable configs, Gradio entry points, and AR / FlashAR vLLM patch sets.
๐Ÿ“ˆ1024x1024 T2I SpeedOn one H20, FlashAR vLLM reaches 5.44 s/img, 82.86x faster than AR eager and 3.04x faster than FlashAR eager.
Xiaomi-Robotics-U0 task examples across image generation, embodied scene generation, transfer, and video generation.

Xiaomi-Robotics-U0 exposes six public task types through one autoregressive framework:

TaskInput โ†’ Output
๐ŸŽจT2IText prompt โ†’ image.
๐Ÿ–ผ๏ธX2IReference image plus instruction โ†’ generated or edited image.
๐ŸงญScene GenScene and task description โ†’ multi-view embodied observations.
๐Ÿ”TransferConditioned embodied observation โ†’ target RGB multi-view scene.
๐Ÿฆพinterleave_subtaskInitial observations and task instruction โ†’ interleaved subtask text and observations.
๐ŸŽฌinterleave_videoInitial observation and task context โ†’ embodied video rollout.

News

  • [September 2026] ๐Ÿ”ฅ Released Xiaomi-Robotics-U0-4B, Xiaomi-Robotics-U0-Sequence, and Xiaomi-Robotics-U0-4B-Sequence weights.
  • [September 2026] ๐Ÿ’ป Open-sourced the FSDP training code.
  • [July 2026] ๐ŸŽ‰ Released the Technical Report.
  • [July 2026] ๐Ÿ”ฅ Released Xiaomi-Robotics-U0 and Xiaomi-Robotics-U0-FlashAR weights.
  • [July 2026] ๐Ÿ’ป Inference code and scripts are now live!

Table of Contents

  1. Model & Weights
  2. Inference
  3. Training
  4. Citation

Model & Weights

Xiaomi-Robotics-U0, Xiaomi-Robotics-U0-4B, and Xiaomi-Robotics-U0-FlashAR support Scene Gen, Transfer, T2I, and X2I. The Sequence checkpoints support interleave_subtask and interleave_video with the eager backend.

Model nameHugging Face WeightModelScope Weight
Xiaomi-Robotics-U0Hugging FaceModelScope
Xiaomi-Robotics-U0-FlashARHugging FaceModelScope
Xiaomi-Robotics-U0-4BHugging FaceModelScope
Xiaomi-Robotics-U0-SequenceHugging FaceModelScope
Xiaomi-Robotics-U0-4B-SequenceHugging FaceModelScope
VisionTokenizerHugging FaceModelScope

Inference

The complete inference implementation, environment setup, configuration reference, command-line examples, and distributed inference instructions are available in inference/README.md.

The repository supports both eager execution and vLLM backends for AR and FlashAR inference. A Gradio demo is also provided for interactive T2I, X2I, Scene Gen, and Transfer workflows.

Training

The PyTorch FSDP training framework for Xiaomi-Robotics-U0 is available in training/README.md. It provides 4B and 34B distributed training presets with long-context packed data and Ulysses sequence parallelism.

Citation

If you find this work useful, please cite:

@misc{li2026xiaomiroboticsu0,
  title         = {{Xiaomi-Robotics-U0}: Unified Embodied Synthesis with World Foundation Model},
  author        = {Xinghang Li and Jun Guo and Qiwei Li and Long Qian and Hang Lai and Yueze Wang and Hongyu Yan and Jiahang Cao and Xi Chen and Jingen Qu and Jiaxi Song and Nan Sun and Hanye Zhao and Futeng Liu and Wanli Peng and Heyun Wang and Yunhong Wang and Caoyu Xia and Jack Zhao and Diyun Xiang and Hangjun Ye and Heng Qu and Huaping Liu and Jason Li},
  year          = {2026},
  eprint        = {2607.11643},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2607.11643}
}