search_engine: "bing"

July 14, 2026 · View on GitHub

✨ Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents

If you like our project, please give us a star ⭐ on GitHub for the latest update.

Paper Paper Ant Group

📣 Latest News

  • [Apr 21, 2026]: 🔥🔥🔥We have successfully extended IGPO to handle ultra-long-horizon Deep Research tasks (200+ turns). Based on this, we trained DR-Venus, a frontier 4B deep research agent for edge-scale deployment. Empowered by IGPO, the performance of DR-Venus is effectively improved on challenging benchmarks, including BrowseComp and BrowseComp-ZH. For more details, please refer to the DR-Venus technical report DR-Venus DR-Venus, the released models DR-Venus-4B-RL, and the IGPO training codebase within the DR-Venus codebase GitHub.
  • [Feb 01, 2026]: 🔄 Codebase updated with new features. See details
  • [Jan 26, 2026]: 🎉 Our IGPO paper has been accepted at ICLR 2026!
  • [Oct 17, 2025]: 📄 Our IGPO paper is now available on arXiv and Hugging Face daily paper.

🤖Models

DR-Venus-4B-RL DR-Venus-4B-RL-GGUF

💡 Method Overview

We introduce IGPO, a RL algorithm for fine-grained credit assignment in search agent training. By modeling agentic search turns as an incremental information acquisition process, IGPO defines rewards as the marginal gain in the policy's probability of generating the correct answer.

📊 Overall Performance

Updated with the final ICLR 2026 results. GiGPO performance is based on the word-level F1 reproduction reported in BranPO.

🚀 Quick Start

1. Installation

git clone https://github.com/GuoqingWang1/IGPO
cd IGPO

conda create -n igpo python=3.10
conda activate igpo

pip install -r requirements.txt

pip install -e .

2. Configure Web Search API & Prompt Template

Edit tools_server/config.yaml:

# Google Search (via Serper API)
search_engine: "google"
serper_api_key: "your_serper_api_key_here"

# Or Bing Search (via Azure)
# search_engine: "bing"
# azure_bing_search_subscription_key: "your_bing_key_here"

# System prompt template is also defined in this file
# system_prompt: |-
#   ...

💡 Tip: If you don't have access to a search API yet, you can use mock mode for testing:

export IGPO_MOCK_SEARCH=true

Or set mock_mode: true in tools_server/config.yaml. This will return simulated search results without actual API calls.

3. Prepare Data

Place your training data in the data/ directory:

  • data/train.parquet - Training data
  • data/dev.parquet - Validation data
  • data/test.parquet - Evaluation data

Data format: See the provided data for reference.

4. Training

Supported Models: We currently support Qwen series models (e.g., Qwen2.5-7B-Instruct).

Edit train.sh to configure training parameters, then run:

bash train.sh
Key Parameters in train.sh
ParameterDescription
MODEL_PATHPath to your model or Hugging Face model name (e.g., Qwen/Qwen2.5-7B-Instruct)
OUTPUTDirectory for saving checkpoints
EVAL_LOG_PATHDirectory for saving validation results
data.train_filesPath to training data (parquet format)
data.val_filesPath to validation data (parquet format)
algorithm.gammaDiscount factor for reward computation
+algorithm.info_gain_typeInfo gain reward calculation: log_prob_diff (log probability difference) or prob_diff (probability difference)
+algorithm.info_gain_norm_modeReward normalization: separate (normalize info gain and outcome rewards independently) or joint (normalize all rewards together)
+algorithm.use_vectorized_gt_logprobEnable vectorized ground truth log probability computation for faster info gain reward construction
+algorithm.use_curriculumEnable curriculum learning to gradually decay info gain reward weight during training
+algorithm.curriculum_f1_init/finalInitial and final weight for outcome reward in curriculum learning
+algorithm.curriculum_ig_init/finalInitial and final weight for info gain reward in curriculum learning
trainer.save_freqSave checkpoint every N training steps
trainer.test_freqRun validation every N training steps
agent_grpo.nNumber of rollouts per sample (GRPO group size)
max_turnsMaximum number of search turns allowed per episode

5. Evaluation

Edit evaluate.sh to configure evaluation parameters, then run:

bash evaluate.sh
Key Parameters in evaluate.sh
ParameterDescription
MODEL_PATHPath to your trained checkpoint
TEST_FILESPath to test data (parquet format)
OUTPUT_DIRDirectory for saving evaluation results
EVAL_LOG_PATHDirectory for saving evaluation logs
MAX_TURNSMaximum number of search turns allowed per episode

📄 Citation

If you find our code or work useful for your research, please cite our work.

@inproceedings{
wang2026information,
title={Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents},
author={Guoqing Wang and Sunhao Dai and Guangze Ye and Zeyu Gan and Wei Yao and Yong Deng and Xiaofeng Wu and Zhenzhe Ying},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=qkWP6phrvZ}
}

🙏 Acknowledgement

IGPO is inspired by Deepseek-R1, with its implementation built upon veRL, Search-r1, and DeepResearcher. We are grateful to the teams behind these projects for their significant contributions to open-source research and development.

📞 Contact

For any questions or feedback, please reach out to us at guoqingwang905@gmail.com.

📜 License

This project is released under the MIT License.

🌟 Star History

Star History Chart