V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

May 27, 2026 ยท View on GitHub

[๐Ÿ“– Paper] [๐Ÿค— V-Retrver-7B-model] [๐Ÿค—V-Retrver-RFT-model] [๐Ÿค—V-Retrver-SFT-model] [๐Ÿค— V-Retrver-train-data] [๐Ÿค— V-Retrver-eval-data]

๐Ÿ‘€ About V-Retrver

Descriptive alt text

We introduce V-Retrver, an evidence-driven retrieval framework that reformulates multimodal retrieval as an agentic reasoning process grounded in visual inspection. V-Retrver enables an MLLM to selectively acquire visual evidence during reasoning via external visual tools, performing a multimodal interleaved reasoning process that alternates between hypothesis generation and targeted visual verification.

To train such an evidence-gathering retrieval agent, we adopt a curriculum-based learning strategy combining supervised reasoning activation, rejection-based refinement, and reinforcement learning with an evidence-aligned objective.

Experiments across multiple multimodal retrieval benchmarks demonstrate consistent improvements in retrieval accuracy (with 23.0% improvements on average), perception-driven reasoning reliability, and generalization.

All code, models, and data are fully released.

๐Ÿ”ฅ News

  • [2026/2/06] We release the code, model, data of V-Retrver

๐Ÿ“ Features

  • Support Qwen3-VL/Qwen2.5-VL Training
  • Provide full pipeline (dataset, SFT training, RFT training, RL training, evaluation, etc)

๐Ÿ† Performance

V-Retrver-7B demonstrates strong performance across multiple multimodal retrieval benchmarks.

Descriptive alt text

๐ŸŽฅ Reasoning Examples

Some reasoning examples are as follows.

๐Ÿ“ Set up

cd verltool
git submodule update --init --recursive
conda create --name v-retrver python=3.10
conda activate v-retrver
pip install -e verl
pip install -e ".[vllm,acecoder,torl,search_tool]"
pip install "flash-attn==2.8.3" --no-build-isolation

๐Ÿš€ Training

Stage 1: Cold-start Supervised Fine-tuning (SFT)

We recommend to use the popular LLaMA-Factory to perform SFT on our cold-start data.

  1. Install LLaMA-Factory.
  2. Follow the instructions in LLaMA-Factory to configure the cold-start data in data/dataset_info.json, as shown below, then copy the config file sftconfig/qwen2_5vl_retrv_full_sft.yaml into your LLaMA-Factory codebase.
"V-Retrver_SFT": {
  "file_name": "[YOUR_DATASET_FOLDER]/V-Retrver_SFT.json",
  "formatting": "sharegpt",
  "columns": {
    "messages": "conversations",
    "images": "images"
  },
  "tags": {
    "role_tag": "from",
    "content_tag": "value",
    "user_tag": "human",
    "assistant_tag": "gpt",
    "system_tag": "system"
  }
}
  1. Train Cold-start data with the training configs.
llamafactory-cli train sft_configs/qwen2_5vl_retrv_full_sft.yaml

Stage 2: Rejection Sampling Fine-Tuning (RSFT)

In this stage, we improve reasoning reliability through Rejection Sampling.The training process and configurations for this stage are identical to Stage 1 (SFT). You simply need to prepare the RSFT dataset and follow the same training steps described in Stage 1.

Stage 3: Reinforcement Learning (RL)

Training

The reinforcement learning is based on the RSFT model. You could either use the model produced in stage 1, or directly download it from V-Retrver/V-Retrver-RFT-7B.

cd verltool
bash examples/train/v-retrver/train_qwen25vl.sh

It should be able to run under 8 A800 GPUs with 80GB memory. From more details๏ผŒplease refer to verl-tool.

Tips:

  • if output shared memory, try lower the data.dataloader_num_workers
  • if out of cuda memory during vllm rollout, try set actor_rollout_ref.rollout.enforce_eager=True, might be slower.
  • if out of cuda memory during training, try lower the use_dynamic_bs=False.

๐Ÿ”ฎ Inference & Evaluation

We recommend using our provided json files and scripts for easier evaluation.

The json files can be downloaded at: [๐Ÿค— V-Retrver-eval-data].

You can conduct inference on all benchmarks using the following scripts

cd verltool
bash examples/train/AdaTooler-V/eval.sh

Sliding Window Evaluation

Use sliding_window_processor.py for reranking over a large candidate pool. This script splits the input Parquet file into specific windows (default strategy: 30-49, 20-39, etc.), calls the core eval.sh for each window, and merges the results to update the final rankings.

cd verltool
python sliding_window_processor.py \
  --input_parquet /path/to/your/input_data.parquet \
  --output_parquet /path/to/save/reranked_data.parquet \
  --window_size 20

Top-20 Direct Evaluation

Use rerank_top20_processor.py (or the corresponding mode) to directly evaluate and rerank the top-20 candidates, skipping the sliding window logic. This script also wraps the core eval.sh to perform the evaluation.

cd verltool
python rerank_top20_processor.py \
  --input_parquet /path/to/your/input_data.parquet \
  --output_parquet /path/to/save/top20_reranked.parquet \
  --mode top20

Acknowledgements

We sincerely appreciate the contributions of the open-source community. The related projects are as follows: verl-tool, verl, LLaMA-Factory, LamRA

Citations

If you find our work helpful for your research, please consider citing our work.

@article{chen2026v,
  title={V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval},
  author={Chen, Dongyang and Wang, Chaoyang and SU, Dezhao and Xiao, Xi and Zhang, Zeyu and Xiong, Jing and Li, Qing and Shang, Yuzhang and Kan, Shichao},
  journal={arXiv preprint arXiv:2602.06034},
  year={2026}
}