Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
May 9, 2025 · View on GitHub
Paper Link ArXiv.
Install
conda create -n plvl Python=3.8
conda activate plvl
pip install torch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
Data Preparation
1.You can download the images from the original source and place them in ./image_data folder:
- RefCOCO/RefCOCO+/RefCOCOg
- Flickr30K Entities
- Visual Genome
Finally, the ./image_data folder will have the following structure:
|-- ln_data
|-- flickr30k
|-- mscoco/images/train2014/
|-- visual-genome
2.Download data labels here and place them in ./mask_data folder
Pretrained Checkpoints
Download the following checkpoints and place them in the ./checkpoints folder.
Training
-
Training and Evaluation on RefCOCOg.
NGPU=4 PORT=25449 MODEL_NAME=PLVL BACKBONE=ViTDet_Dec DATASET=unc # unc/unc+/gref_umd OUTPUT=outputs/${DATASET}/PLVL # Train python -m torch.distributed.launch --nproc_per_node=${NGPU} --master_port ${PORT} --use_env train.py \ --batch_size 20 \ --device cuda \ --aug_scale --aug_translate --aug_crop \ --is_res --is_rec \ --ca_block_indexes 2 5 8 11 \ --model_name ${MODEL_NAME} \ --backbone ${BACKBONE} \ --dataset ${DATASET} \ --output_dir ${OUTPUT} # Evaluation declare -a items_model=("best_checkpoint.pth" "best_mask_checkpoint.pth") # declare -a items_dataset=("val" "test") for item_model in "${items_model[@]}"; do for item_datatset in "${items_dataset[@]}"; do python -m torch.distributed.launch --nproc_per_node=${NGPU} --master_port ${PORT} --use_env eval.py \ --batch_size 80 \ --is_res --is_rec \ --ca_block_indexes 2 5 8 11 \ --model_name ${MODEL_NAME} \ --backbone ${}BACKBONE} \ --dataset ${DATASET} \ --eval_set ${item_datatset} \ --device cuda \ --eval_model ${OUTPUT}/${item_model} \ --output_dir ${OUTPUT} done donePlease refer to train.sh for training commands on other datasets.
-
For the pretraining result, first use the following command to pretrain model on the mixed dataset. Then use the following command to fine-tune on mixed RefCOCO series datasets.
NGPU=8 PORT=25449 MODEL_NAME=PLVL BACKBONE=ViTDet_Dec # mix pretrain DATASET=mixed_pretrain OUTPUT=outputs/${DATASET}/PLVL_2 python -m torch.distributed.launch --nproc_per_node=$NGPU --master_port $PORT --use_env train.py \ --batch_size 20 \ --device cuda \ --loss_alpha 0.5 \ --epochs 20 \ --aug_scale --aug_translate --aug_crop \ --is_rec --ca_block_indexes 2 5 8 11 \ --model_name $MODEL_NAME \ --backbone $BACKBONE \ --dataset $DATASET \ --output_dir $OUTPUT # coco fine-tune DATASET=mixed_coco OUTPUT=outputs/${DATASET}/PLVL python -m torch.distributed.launch --nproc_per_node=$NGPU --master_port $PORT --use_env train.py \ --batch_size 20 \ --device cuda \ --loss_alpha 0.05 \ --epochs 150 \ --aug_scale --aug_translate --aug_crop \ --is_res --is_rec --ca_block_indexes 2 5 8 11 \ --model_name $MODEL_NAME \ --backbone $BACKBONE \ --dataset $DATASET \ --output_dir $OUTPUT \ --pretrain outputs/mixed_pretrain/PLVL/checkpoint.pthPlease refer to train_mix.sh for training commands on other datasets.
Our checkpoints
Our checkpoints are available at 百度网盘.
Acknowledgement
Our model is related to EEVG, UVLTrack, ViTDet. Thanks for their great work!
Citation
If our work is useful for your research, please consider cite:
@misc{wang2025progressivelanguageguidedvisuallearning,
title={Progressive Language-guided Visual Learning for Multi-Task Visual Grounding},
author={Jingchao Wang and Hong Wang and Wenlong Zhang and Kunhua Ji and Dingjiang Huang and Yefeng Zheng},
year={2025},
eprint={2504.16145},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2504.16145},
}