UserRL Supervised Fine-Tuning (SFT) Pipeline

September 19, 2025 · View on GitHub

This directory contains the supervised fine-tuning pipeline for training language models on UserRL gym environments using LLaMA-Factory.

Overview

The SFT pipeline trains language models to interact with UserRL environments by fine-tuning on conversation data collected from successful interactions across all nine gym environments.

Pipeline Components

1. Training Data (merged_gym_sft.json)

  • Size: 25MB, ~121K lines
  • Format: ShareGPT conversation format
  • Content: Multi-turn conversations from successful interactions across all UserRL gyms
  • Structure:
    {
      "conversations": [
        {"from": "human", "value": "Environment prompt..."},
        {"from": "gpt", "value": "Model response with tool calls..."},
        {"from": "observation", "value": "Environment feedback..."}
      ],
      "system": "System prompt for the environment",
      "tools": "Tool schema definitions"
    }
    

2. Training Configuration (qwen3_customized.yaml)

  • Model: Qwen3 instruction-tuned models
  • Training: Full fine-tuning with DeepSpeed ZeRO-3
  • Sequence Length: 16,384 tokens
  • Batch Size: 2 per device × 4 gradient accumulation = 8 effective batch size
  • Learning Rate: 1e-5 with cosine scheduling
  • Epochs: 3 training epochs

3. Environment Integration

  • Tool Calling: Supports interact_with_env function calls
  • Multi-Environment: Training data spans all nine UserRL gyms
  • Conversation Format: Maintains proper tool call and observation patterns

Installation & Setup

Step 1: Install LLaMA-Factory

git clone https://github.com/hiyouga/LLaMA-Factory.git
cd LLaMA-Factory
pip install -e .[all]

Step 2: Prepare Training Data

# Copy the training data to LLaMA-Factory data directory
cp merged_gym_sft.json /path/to/LLaMA-Factory/data/

Step 3: Configure Dataset Metadata

Add the following entry to LLaMA-Factory/data/dataset_info.json:

{
  "merged_gym_sft": {
    "file_name": "merged_gym_sft.json",
    "formatting": "sharegpt",
    "columns": {
      "messages": "conversations",
      "system": "system",
      "tools": "tools"
    }
  }
}

Step 4: Configure Training Parameters

# Copy and edit the training configuration
cp qwen3_customized.yaml /path/to/LLaMA-Factory/examples/train_full/

Important: Update the model_name_or_path and output_dir in qwen3_customized.yaml to point to your base model and desired saving directory respectively:

model_name_or_path: /path/to/your/qwen3-model
output_dir: saves/saved_folder_name

Training Execution

cd /path/to/LLaMA-Factory
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
llamafactory-cli train examples/train_full/qwen3_customized.yaml

Training Configuration Details

Model Configuration

  • Base Model: Qwen3 instruction-tuned models
  • Trust Remote Code: Enabled for custom model architectures
  • Fine-tuning Type: Full fine-tuning (all parameters)

Training Parameters

ParameterValueDescription
per_device_train_batch_size2Batch size per GPU
gradient_accumulation_steps4Gradient accumulation steps
learning_rate1e-5Initial learning rate
num_train_epochs3.0Number of training epochs
lr_scheduler_typecosineLearning rate scheduler
warmup_ratio0.1Warmup ratio
cutoff_len16384Maximum sequence length