Setup Guide

September 3, 2025 ยท View on GitHub

System Requirements

  • NVIDIA GPUs with Ampere architecture (RTX 30 Series, A100) or newer
  • NVIDIA driver compatible with CUDA 12.6
  • Linux x86-64
  • glibc>=2.31 (e.g Ubuntu >=22.04)
  • Python 3.10

Installation

Clone the repository

git clone git@github.com:nvidia-cosmos/cosmos-predict2.git
cd cosmos-predict2

Option 1: Virtual environment

Install system dependencies:

uv

curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env

Install the package into a new environment:

uv sync --extra cu126
source .venv/bin/activate

Or, install the package into the active environment (e.g. conda):

uv sync --extra cu126 --active --inexact

Option 2: Docker container

Please make sure you have access to Docker on your machine and the NVIDIA Container Toolkit is installed.

Build and run the container:

docker run --gpus all --rm -v .:/workspace -v /workspace/.venv -it $(docker build -q .)

Downloading Checkpoints

  1. Get a Hugging Face Access Token with Read permission
  2. Install Hugging Face CLI: uv tool install -U "huggingface_hub[cli]"
  3. Login: hf auth login
  4. Accept the Llama-Guard-3-8B terms.

To download a specific model:

./scripts/download_checkpoints.py --model_types <model_type> --model_sizes <model_size>
ModelsLinkDownload ArgumentsNotes
Cosmos-Predict2-0.6B-Text2Image๐Ÿค— Huggingface--model_types text2image --model_sizes 0.6BN/A
Cosmos-Predict2-2B-Text2Image๐Ÿค— Huggingface--model_types text2image --model_sizes 2BN/A
Cosmos-Predict2-14B-Text2Image๐Ÿค— Huggingface--model_types text2image --model_sizes 14BN/A
Cosmos-Predict2-2B-Video2World๐Ÿค— Huggingface--model_types video2world --model_sizes 2BDownload 720P, 16FPS by default. Supports 480P and 720P resolution. Supports 10FPS and 16FPS
Cosmos-Predict2-14B-Video2World๐Ÿค— Huggingface--model_types video2world --model_sizes 14BDownload 720P, 16FPS by default. Supports 480P and 720P resolution. Supports 10FPS and 16FPS
Cosmos-Predict2-2B-Sample-Action-Conditioned๐Ÿค— Huggingface--model_types sample_action_conditionedSupports 480P and 4FPS.
Cosmos-Predict2-14B-Sample-GR00T-Dreams-GR1๐Ÿค— Huggingface--model_types sample_gr00t_dreams_gr1Supports 480P and 16FPS.
Cosmos-Predict2-14B-Sample-GR00T-Dreams-DROID๐Ÿค— Huggingface--model_types sample_gr00t_dreams_droidSupports 480P and 16FPS.

To download all the checkpoints (requires ~250GB of disk space), run:

./scripts/download_checkpoints.py

To see the full list of options, run:

./scripts/download_checkpoints.py --help

Troubleshooting

CUDA/GPU Issues

  • CUDA driver version insufficient: Update NVIDIA drivers to latest version compatible with CUDA 12.6+

  • Out of Memory (OOM) errors: Use 2B models instead of 14B, or reduce batch size/resolution

  • Missing CUDA libraries: Set paths with export CUDA_HOME=$CONDA_PREFIX

Installation Issues

  • Conda environment conflicts: Create fresh environment with conda create -n cosmos-predict2-clean python=3.10 -y

  • Flash-attention build failures: Install build tools with apt-get install build-essential

  • Transformer engine linking errors: Reinstall with pip install --force-reinstall transformer-engine==1.12.0

For other issues, check GitHub Issues.