README.md
September 10, 2026 ยท View on GitHub
| Highlight | Summary | |
|---|---|---|
| ๐ง | World Foundation Model | A 34B autoregressive model for text, images, and embodied observations, initialized from EMU3.5. |
| ๐งฉ | Unified Token Space | Uses a shared discrete visual tokenizer and a single next-token objective across multimodal sequences. |
| ๐ค | Embodied Synthesis | Bridges foundation image generation with robot-centric scene, transfer, and video generation. |
| โก | Xiaomi-Robotics-U0-FlashAR Acceleration | Decodes visual tokens in anti-diagonal groups and supports vLLM batching for high-resolution inference. |
| ๐ฆ | Open Inference Repo | Provides inference code, composable configs, Gradio entry points, and AR / FlashAR vLLM patch sets. |
| ๐ | 1024x1024 T2I Speed | On one H20, FlashAR vLLM reaches 5.44 s/img, 82.86x faster than AR eager and 3.04x faster than FlashAR eager. |
Xiaomi-Robotics-U0 exposes six public task types through one autoregressive framework:
| Task | Input โ Output | |
|---|---|---|
| ๐จ | T2I | Text prompt โ image. |
| ๐ผ๏ธ | X2I | Reference image plus instruction โ generated or edited image. |
| ๐งญ | Scene Gen | Scene and task description โ multi-view embodied observations. |
| ๐ | Transfer | Conditioned embodied observation โ target RGB multi-view scene. |
| ๐ฆพ | interleave_subtask | Initial observations and task instruction โ interleaved subtask text and observations. |
| ๐ฌ | interleave_video | Initial observation and task context โ embodied video rollout. |
News
- [September 2026] ๐ฅ Released Xiaomi-Robotics-U0-4B, Xiaomi-Robotics-U0-Sequence, and Xiaomi-Robotics-U0-4B-Sequence weights.
- [September 2026] ๐ป Open-sourced the FSDP training code.
- [July 2026] ๐ Released the Technical Report.
- [July 2026] ๐ฅ Released Xiaomi-Robotics-U0 and Xiaomi-Robotics-U0-FlashAR weights.
- [July 2026] ๐ป Inference code and scripts are now live!
Table of Contents
Model & Weights
Xiaomi-Robotics-U0, Xiaomi-Robotics-U0-4B, and Xiaomi-Robotics-U0-FlashAR support Scene Gen, Transfer, T2I, and X2I. The Sequence checkpoints support interleave_subtask and interleave_video with the eager backend.
| Model name | Hugging Face Weight | ModelScope Weight |
|---|---|---|
| Xiaomi-Robotics-U0 | ||
| Xiaomi-Robotics-U0-FlashAR | ||
| Xiaomi-Robotics-U0-4B | ||
| Xiaomi-Robotics-U0-Sequence | ||
| Xiaomi-Robotics-U0-4B-Sequence | ||
| VisionTokenizer |
Inference
The complete inference implementation, environment setup, configuration reference, command-line examples, and distributed inference instructions are available in inference/README.md.
The repository supports both eager execution and vLLM backends for AR and FlashAR inference. A Gradio demo is also provided for interactive T2I, X2I, Scene Gen, and Transfer workflows.
Training
The PyTorch FSDP training framework for Xiaomi-Robotics-U0 is available in
training/README.md. It provides 4B and 34B distributed
training presets with long-context packed data and Ulysses sequence parallelism.
Citation
If you find this work useful, please cite:
@misc{li2026xiaomiroboticsu0,
title = {{Xiaomi-Robotics-U0}: Unified Embodied Synthesis with World Foundation Model},
author = {Xinghang Li and Jun Guo and Qiwei Li and Long Qian and Hang Lai and Yueze Wang and Hongyu Yan and Jiahang Cao and Xi Chen and Jingen Qu and Jiaxi Song and Nan Sun and Hanye Zhao and Futeng Liu and Wanli Peng and Heyun Wang and Yunhong Wang and Caoyu Xia and Jack Zhao and Diyun Xiang and Hangjun Ye and Heng Qu and Huaping Liu and Jason Li},
year = {2026},
eprint = {2607.11643},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.11643}
}