Sherlock Training Guide
May 28, 2025 · View on GitHub
This guide provides detailed instructions for training the Sherlock model.
Prerequisites
- Python 3.10
- 8 x A100 80G GPU
Installation
- Install LLaMA-Factory:
cd train/LLaMA-Factory
pip install -e ".[torch,metrics]" --no-build-isolation
- Install specific versions of required packages:
pip install transformers==4.45.2
pip install trl==0.9.6
- Install vLLM for more efficient data construction:
pip install vllm
- Replace the DPO trainer implementation:
- Locate the
dpo_trainer.pyfile in your Python environment:YOUR_ENV/lib/python3.10/site-packages/trl/trainer/dpo_trainer.py - Replace it with the custom implementation from dpo_trainer.py
- Locate the
Training Pipeline
1. Supervised Fine-Tuning (SFT)
Data Preparation
- Format your LLaVA-CoT data in ShareGPT format:
{
"conversations": [
{
"from": "human",
"value": "<image>{question}"
},
{
"from": "gpt",
"value": "{response}"
}
],
"images": [
"image_path"
]
}
- Create training datasets:
- Randomly sample 10k examples for
- Randomly sample 10k examples for
- Save the ShareGPT format data as JSON files in
./train/LLaMA-Factory/data/
Training Steps
- Train the initial R0 VLM model:
bash ./train/LLaMA-Factory/bash_train/train_sft_ro_vlm.sh
-
Generate Sherlock-SFT dataset:
- Use the trained R0 VLM model
- Run the data generation script:
./train/data_construction/sft/gen_data.py
-
Train the Sherlock SFT model:
bash ./train/LLaMA-Factory/bash_train/train_sft.sh
2. Offline
-
Generate offline preference data:
- Use the script to obtain preference data:
./train/data_construction/offline/gen_data.py - The format of offline data should be:
{ "conversations": [ { "from": "human", "value": "<image>{question}" } ], "chosen": { "from": "gpt", "value": "{chosen_response}" }, "rejected": { "from": "gpt", "value": "{rejected_response}" }, "prefix": "{chosen_prefix}", "prefix_l": "{rejected_prefix}", "weights": "{int}" "images": [ "image_path" ] }
- Use the script to obtain preference data:
-
Train the Sherlock Offline model:
bash ./train/LLaMA-Factory/bash_train/train_offline.sh
Online
-
Sample only question and image from LLaVA-CoT dataset.
-
First generate candidate response with three turn of self-correction using
./train/data_construction/online/gen_candidate.pyfile. -
Running
./train/data_construction/online/gen_chosen_data.pyto filter and obtain 5k chosen reasoning trajectory based on self-consistency. -
Construct online preference data
- Running
./train/data_construction/online/gen_rejected_data.pyfile. - The format of online data should be:
{ "conversations": [ { "from": "human", "value": "<image>{question}" } ], "chosen": { "from": "gpt", "value": "{chosen_response}" }, "rejected": { "from": "gpt", "value": "{rejected_response}" }, "prefix": "{chosen_prefix}", "prefix_l": "{rejected_prefix}", "weights": "{int}" "images": [ "image_path" ] }
- Running
-
Start Online training to obtain Sherlock Iter1 and Sherlock Iter2 model.
bash .train/LLaMA-Factory/bash_train/train_online.sh