Implementation of Semantic and Sequential Alignment for Referring Video Object Segmentation

June 28, 2026 ยท View on GitHub

[paper]

Installation

Please read INSTALL.md for details.

Data Preparation

This project uses MeViS, Ref-YouTube-VOS and Ref-DAVIS17 datasets.

Please read DATA.md for details.

Inference

MeViS

python train_net_ssa.py \
    --config-file configs/mevis/ssa_clip_bs8.yaml \
    --num-gpus 8 --dist-url auto --eval-only \
    MODEL.WEIGHTS [path_to_weights] \
    OUTPUT_DIR [path_to_outputs]

We provide the model checkpoints here.

model weightsoutput filesJ&F
fine-tuningmodeloutputs48.56
joint trainingmodeloutputs49.37

The val split is evaluated through the online evaluation server, whereas the valu split can be evaluated locally with tools/eval_mevis.py

Ref-Youtube-VOS

python train_net_ssa.py \
    --config-file configs/mevis/ssa_clip_bs8.yaml \
    --num-gpus 8 --dist-url auto --eval-only \
    MODEL.WEIGHTS [path_to_weights] \
    OUTPUT_DIR [path_to_outputs]

We provide the model checkpoints here.

model weightsJ&F
Ref-YouTube-VOSmodel64.3

For Ref-YouTube-VOS, submit the result to the online evaluation server.

Ref-DAVIS17

python train_net_ssa.py \
    --config-file configs/mevis/ssa_clip_bs8.yaml \
    --num-gpus 8 --dist-url auto \
    MODEL.WEIGHTS [path_to_weights] \
    OUTPUT_DIR [path_to_outputs]

The model checkpoints are directly adopted from those trained on Ref-YouTube-VOS.

model weightsJ&F
Ref-DAVIS17model67.3

This project saves prediction masks for each object separately, which is not fully compatible with the official DAVIS evaluation script.

An alternative way is using tools/eval_davis.py (modified from eval_mevis.py) to evaluate the predictions directly. The metrics may differ slightly from those obtained with the official evaluation protocol.

The official evaluation code from ReferFormer can be used through a postprocess step to merge prediction masks:

# postprocess step
python tools/postprocess_davis.py  --mask_dir [path_to_outputs] --output_dir [path_to_saves]

# evaluation step 
python tools/eval_davis_official.py --results_path [path_to_saves]/anno_{0,1,2,3}

Training

Following the setup of MeViS,, we employ Mask2Former (CLIP vision encoder as backbone) and train it on COCO dataset as the initialized weights. The checkpoints of the CLIP-based Mask2Former is available here.

Before loading the checkpoint, initialize the CLIP model either with setting open_clip.create_model_and_transforms() or downloading weights from Hugging Face and placing them under open_clip_model/.

MeViS

Training for MeViS, use

python train_net_ssa.py \
    --config-file configs/mevis/ssa_clip_bs8.yaml \
    --num-gpus 8 --dist-url auto \
    MODEL.WEIGHTS [path_to_weights] \
    OUTPUT_DIR [path_to_outputs]

Ref-Youtube-VOS

Training for Ref-Youtube-VOS, use

# pre-training stage
python train_net_ssa.py \
    --config-file configs/coco/ssa_clip_bs8.yaml \
    --num-gpus 8 --dist-url auto \
    MODEL.WEIGHTS [path_to_weights] \
    OUTPUT_DIR [path_to_outputs]

# fine-tuning stage
python train_net_ssa.py \
    --config-file configs/ytvos/ssa_clip_bs8.yaml \
    --num-gpus 8 --dist-url auto \
    MODEL.WEIGHTS [path_to_weights] \
    OUTPUT_DIR [path_to_outputs]

The training settings can be modified from the config-file in./config directory.

Acknowledgement

This project is based on MeViS, FC-CLIP and ReferFormer. Thanks for these great works.