README.md

August 28, 2026 ยท View on GitHub

JoyAI-Echo generated video gallery

JoyAI-Echo 1.5

๐ŸŽฌ Long-form audio-video generation with reference-driven multi-shot memory

๐Ÿ“„ Paper 1.5 | ๐ŸŒ Project Page | ๐Ÿš€ Quickstart | ๐Ÿค— Model Weights | ๐ŸŽฌ Director Agent

Python 3.11 PyTorch 2.8 CUDA 12.8 Inference only

๐Ÿ“ข Whats New

  • ๐ŸŽ‰ JoyAI-Echo 1.5 is now available! Code, model weights, and R2V inference are released.
  • ๐ŸŽฌ We also introduce Echo Director Agent This Time !, an agentic workflow for planning and creating multi-shot videos. See Director_Agent/.
  • ๐ŸŽฎ JoyAI-Echo 1.5 brings high-quality video generation to consumer GPUs. Fire up your RTX GPU and start creating!
  • ๐Ÿ“ฆ JoyAI-Echo 1.0 remains available on the echo1.0 archive branch.

JoyAI-Echo 1.5

JoyAI-Echo 1.5 is an inference-only release for coherent, long-form audio-visual generation. It combines few-step generation with a paired cross-modal memory bank, carrying character appearance, voice, and scene continuity across independently editable shots.

Reference-to-video generation

Echo 1.5 supports reference-to-video (R2V) generation. Each request may include a text prompt, an optional first-frame condition, and up to seven ordered memory slots containing reference images and audio. This makes visual identity, voice, and story context directly reusable from shot to shot.

A complete request schema and portable example are available in schemas/r2v_request.schema.json and examples/the_last_visa/. ๐Ÿ•ถ๏ธ๐Ÿฟ

Echo Director Agent

We also introduce Echo Director Agent, a local-first agent for planning, generating, reviewing and assembling multi-shot videos. It turns a story idea into an editable production workflow and submits generation jobs to an Echo 1.5 service.

See Director_Agent/ for installation and usage.

Quickstart

1. Clone

git clone https://github.com/jd-opensource/JoyAI-Echo.git
cd JoyAI-Echo/echo_longvideo

All commands below are run from echo_longvideo/.

2. Install

The reference environment is Python 3.11, PyTorch 2.8 and CUDA 12.8. We recommend using uv:

uv venv --python 3.11 .venv
source .venv/bin/activate
uv pip install --extra-index-url https://download.pytorch.org/whl/cu128 \
  -r requirements.txt
python scripts/setup_msst.py

ffmpeg must also be available on PATH. For FP4 inference, install the optional NVIDIA ModelOpt dependency:

uv pip install -r requirements-fp4.txt

Conda users can instead run:

conda env create -f environment.yml
conda activate joyai-echo15
python scripts/setup_msst.py

3. Download model weights

Download the release weights from Hugging Face and arrange them as follows:

CheckpointPrecisionDownload
echo15_full_dmdBF16Hugging Face
echo15_fp8FP8Hugging Face
echo15_fp4FP4Hugging Face
gemma-3-12bText encoderHugging Face
checkpoints/
โ”œโ”€โ”€ echo15_full_dmd/          # BF16 reference checkpoint
โ”œโ”€โ”€ echo15_fp8/               # FP8 scaled-matmul checkpoint
โ”œโ”€โ”€ echo15_fp4/               # packed ModelOpt FP4 checkpoint
โ”œโ”€โ”€ gemma-3-12b/              # Gemma text encoder
โ””โ”€โ”€ msst/                     # installed by scripts/setup_msst.py

Each Echo checkpoint directory includes its own checkpoint.json manifest. See checkpoints/README.md for the exact layout.

4. Run batch inference

The default command loads the model once and processes all R2V JSON requests under examples/the_last_visa/requests/:

python inference.py --config configs/inference.bf16.yaml

Use FP8 or FP4 by selecting the corresponding configuration:

python inference.py --config configs/inference.fp8.yaml
python inference.py --config configs/inference.fp4.yaml

Outputs are written to inference_result/<work-id>/<shot-id>/.

Consumer GPU support

Echo 1.5 includes low-memory profiles for consumer GPUs. They combine layer-wise DiT weight streaming with tiled Video VAE decoding, trading some latency for substantially lower peak VRAM usage.

# BF16 (requires substantial system RAM)
python inference.py --config configs/inference.consumer.bf16.yaml

# FP8
python inference.py --config configs/inference.consumer.fp8.yaml

# FP4 standalone (recommended)
python inference.py --config configs/inference.consumer.fp4.yaml

For a 24 GiB VRAM target, precompute conditioning before online generation. Actual headroom depends on the GPU, driver and request shape.

Local inference server

The repository includes a small local server for Echo Director Agent and other R2V clients. The root server.py is its command-line entry point; the service implementation lives in the server/ package. It provides an in-memory queue and dynamically keeps model weights on the GPU when memory permits.

uv pip install -r requirements-server.txt
uv run python server.py --config configs/server.consumer.yaml

On Linux or macOS, after setup, start it with one command:

./scripts/start_server.sh

On Windows, start it from Command Prompt with:

scripts\start_server.cmd

Extra server arguments may be appended, for example ./scripts/start_server.sh --port 8222 or scripts\start_server.cmd --port 8222.

The server YAML owns deployment settings and references a separate inference YAML, which owns the checkpoint and pipeline settings.

See docs/LOCAL_SERVER.md for deployment options.

Acknowledgements

We gratefully acknowledge the open-source projects that make this release possible, especially LTX-2.3, Gemma and MSST-WebUI. See THIRD_PARTY_NOTICES.md for details.

Citation

If JoyAI-Echo helps your research, please cite:

@article{duan2026joyaiecho15,
  title         = {Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds},
  author        = {Duan, Nan and Huang, Haoyang and Jin, Weiyang and Li, Haoran and Li, Yaowei and Li, Yuming and Liu, Yijun and Lu, Xin and Ma, Xiaoxiao and Ma, Yanwen and Su, Yaofeng and Sun, Yilang and Wang, Haoyu and Xue, Zeyue and Zhang, Songchun and Zhuang, Junhao},
  journal       = {arXiv preprint arXiv:2608.23383},
  year          = {2026},
  eprint        = {2608.23383},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2608.23383}
}

For academic research and non-commercial use only.

License

This project is based on LTX-2 by Lightricks Ltd.

Portions of the original LTX-2 codebase have been modified by JD.com for academic and research purposes only. This project is not intended for commercial use. For commercial use of LTX-2 or its derivatives, please contact Lightricks Ltd.

All original copyright, license, patent, trademark, and attribution notices from LTX-2 are retained. This project remains subject to the LTX-2 Community License Agreement.