README.md

July 26, 2026 ยท View on GitHub

ABot-World: Infinite Interactive World Rollout on a Single Desktop GPU

Studio Playground Project Paper Code Model Space Paper Model

ABot-World Team


TL;DR: ABot-World turns a single NVIDIA RTX 5090 desktop GPU into a real-time interactive world simulator, enabling infinite action-conditioned world rollout at 720P, 16 FPS, 1.2s latency, and 19GB GPU memory.

๐Ÿš€ Key Highlights

  • ๐ŸŽฎ Action-Driven World Control: Responds to user actions in real time, enabling continuous exploration instead of passive video playback.
  • โšก Real-Time Desktop Inference: Runs at 720p and 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s latency and 19GB GPU memory.
  • โ™พ๏ธ Infinite World Rollout: Supports open-ended interactive world generation beyond fixed video-length limits.
  • ๐Ÿง  Open-Ended World Imagination: Expands the world with new scenes and dynamics during rollout, avoiding scene lock-in, without prompt switching, by our LongForcing training.

๐Ÿ“ข News

  • 2026-07-22: Release ABot-World-0 technical report.
  • 2026-07-13: ABot-World is now on Reactor!
  • 2026-07-10: We have decided to open-source our 500-hour video training dataset with accurate action annotations. Stay tunedโ€”we plan to release it very soon.
  • 2026-07-09: We release the causal student model ABot-World-0-5B-LF, inference code, our local gradio demo and online playground ABot World Studio.

๐Ÿ› ๏ธ Setup

This installation was tested on: Ubuntu 22.04, CUDA 12.8, Python 3.12, NVIDIA RTX 5090. For common hardware and environment questions, see FAQ.md.

  1. Clone the repository:
git clone https://github.com/amap-cvlab/ABot-World.git
cd ABot-World
  1. Create a conda environment:
conda create -n aworld python=3.12 -y
conda activate aworld
  1. Install PyTorch (CUDA 12.8):
pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0 \
  --index-url https://download.pytorch.org/whl/cu128
  1. Install FlashAttention:
wget https://github.com/Dao-AILab/flash-attention/releases/download/v2.8.1/flash_attn-2.8.1+cu12torch2.8cxx11abiFALSE-cp312-cp312-linux_x86_64.whl
pip install flash_attn-2.8.1+cu12torch2.8cxx11abiFALSE-cp312-cp312-linux_x86_64.whl

If you encounter glibc version issues, please refer to flash-attention#1708.

  1. Install SageAttention from source:

Follow the official instructions at SageAttention:

git clone https://github.com/thu-ml/SageAttention.git
cd SageAttention
export EXT_PARALLEL=4 NVCC_APPEND_FLAGS="--threads 8" MAX_JOBS=32  # Optional
python setup.py install
cd ..
  1. Install Python dependencies:
pip install -r requirements.txt
  1. Install lightx2v_kernel:
git clone https://github.com/NVIDIA/cutlass.git
git clone https://github.com/ModelTC/LightX2V.git
cd LightX2V/lightx2v_kernel

# Set CUTLASS_PATH to the absolute path of the cutlass repository cloned above.
MAX_JOBS=$(nproc) && CMAKE_BUILD_PARALLEL_LEVEL=$(nproc) \
uv build --wheel \
    -Cbuild-dir=build . \
    -Ccmake.define.CUTLASS_PATH=/path/to/cutlass \
    --verbose \
    --color=always \
    --no-build-isolation

pip install dist/*whl --force-reinstall --no-deps
cd ../..
  1. Download checkpoints:

Download models using HuggingFace:

pip install -U "huggingface_hub"
hf download acvlab/ABot-World-0-5B-LF --local-dir ./checkpoints/ABot-World-0-5B-LF

Download models using ModelScope:

pip install -U "modelscope"
modelscope download "amap_cvlab/ABot-World-0-5B-LF" --local_dir ./checkpoints/ABot-World-0-5B-LF

After downloading, the project should have the following checkpoint structure:

checkpoints/
โ””โ”€โ”€ ABot-World-0-5B-LF/
    โ”œโ”€โ”€ Wan2.2_VAE.pth
    โ”œโ”€โ”€ taew2_2.pth
    โ”œโ”€โ”€ models_t5_umt5-xxl-enc-bf16.pth
    โ”œโ”€โ”€ diffusion_pytorch_model.safetensors
    โ””โ”€โ”€ google/umt5-xxl/

The checkpoint paths are configured in configs/long_forcing_dmd.yaml and configs/default_config.yaml. The distilled generator weights are already merged into ABot-World-0-5B-LF/diffusion_pytorch_model.safetensors.

๐Ÿค— Gradio Demo

bash web_client/run.sh

Select a GPU with:

CUDA_ID=0 bash web_client/run.sh

โ“ FAQ

For common questions regarding hardware, environment setup, and runtime compatibility, please refer to FAQ.md.

License

This project is released under the Apache License 2.0. See LICENSE, NOTICE, and THIRD_PARTY_NOTICES.md for copyright and third-party attribution details.

๐Ÿค Acknowledgement

This project builds on and is inspired by the following open-source projects: Causal Forcing, AngelSlim, LightX2V, taehv, Wan2.2, Helios, from which the optimized Triton RoPE and normalization kernels in wan/modules/helios_kernels are derived.

๐Ÿ—“๏ธ Roadmap

  • Interactive Web Playground (ABot World Studio)
  • Inference Code Release
  • Local Gradio Demo Release
  • Causal Student Model Release
  • Technical Report (Arxiv)
  • Bidirectional Teacher Model Release
  • 500-Hour Video Training Dataset with Accurate Action Annotations

๐Ÿ“ Citation

If you find our work helpful, please cite our paper:

@misc{jiang2026abotworld0,
      title={{ABot-World-0}: Infinite Interactive World Rollout on a Single Desktop GPU}, 
      author={Fan Jiang and Zhaoxu Sun and Mengchao Wang and Ziyu Zhu and Chiyu Wang and Yunpeng Zhang and Wenlin Liu and Yun Wang and Xue Zheng and Rui Sun and Junfeng Ni and Hongyu Pan and Zhongxu Sun and Fei Yu and Zengye Ge and Mengmeng Du and Nianfei Fan and Mingchao Sun and Yu Liu and Yongchang and Yanqing Zhu and Jiahang Wang and Ning Ying and Yuze Xuan and Di Yang and Zhicheng Liu and Zhe Gao and Tingbing Xu and Jiacheng Sui and Wenjin Yang and Junnan Lai and Shufeng Liu and Yuan Liu and Zheng Zhou and Yingliang Peng and Dawei Cao and Kaifeng Sheng and Yuxiang Cai and Fei Lu and Mu Xu and Ning Guo},
      year={2026},
      eprint={2607.19191},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.19191}, 
}

๐Ÿ›ฐ Contact Us via WeChat Group

Feel free to contact us!