InfoPO: Information-Driven Policy Optimization for User-Centric Agents
March 16, 2026 · View on GitHub
This repository is the official implementation of InfoPO, an information-driven reinforcement learning algorithm for training user-centric agents in multi-turn interactive scenarios.
News
- 🔥 2026/03: Our paper was accepted to ICLR 2026 Workshop on Lifelong Agents!
InfoPO addresses the sparse reward problem in multi-turn agent training. Traditional GRPO-based methods fail to distinguish between "honorable failures" (correct information gathering, minor execution error) and "complete failures" (entirely off-track behavior) when both receive zero outcome reward. InfoPO introduces turn-level information gain through counterfactual reasoning, providing dense credit assignment for more effective learning.
Overview
InfoPO enhances standard RL training by introducing intrinsic rewards that measure information gain through counterfactual reasoning. Instead of relying solely on sparse terminal rewards, InfoPO quantifies how observations shift the agent's decision-making, providing dense learning signals for efficient exploration.
Key Features:
- Turn-Level Information Gain: Measures how observations influence subsequent actions via KL divergence
- Counterfactual Reasoning: Compares policy distributions with and without observations
- Adaptive Reward Fusion: Combines intrinsic and outcome rewards with variance-gated weighting
- Multi-Turn Credit Assignment: Proper attribution across extended conversations
Method
InfoPO computes turn-level information gain as:
I_t = KL(P(action_{t+1} | context, observation_t) || P(action_{t+1} | context, placeholder))
This measures how much the observation at turn t shifts the agent's action distribution at turn t+1, providing a task-agnostic signal for exploration.
Quick Start
Prerequisites
- Python 3.12
- CUDA-compatible GPU(s)
- OpenAI API key (for user simulation)
Installation
-
Create Environment
conda create -n infopo python=3.12 conda activate infopo -
Install Dependencies
pip install -e .[sglang] pip install flash-attn --no-build-isolation -
Install Gym Environments
bash install_gyms.sh
Training with InfoPO
-
Configure Environment Variables
export CUDA_VISIBLE_DEVICES=0,1,2,3 export OPENAI_API_KEY="your-openai-key" export OPENAI_BASE_URL="https://api.openai.com/v1" export MULTITURN_MODEL_NAME="gpt-4o" -
Update Training Script
Edit
examples/sglang_multiturn/train.sh:PROJECT_DIR="/path/to/your/project" MODEL_PATH="/path/to/your/model" -
Start Training
bash ./examples/sglang_multiturn/train.sh
The training script uses InfoPO by default with:
algorithm.adv_estimator=info_grpoalgorithm.use_intrinsic_reward=True
Configuration
Key Hyperparameters
algorithm:
adv_estimator: info_grpo
use_intrinsic_reward: True
gamma: 0.8
intrinsic_reward:
intrinsic_weight: 0.5 # Weight for intrinsic rewards (β_0)
normalize_intrinsic: True # Normalize with GRPO group baseline
intrinsic_kl_batch_size: 4 # Batch size for KL computation
observation_placeholder: "No information found."
Training Examples
Multi-Turn User Benchmarks:
bash examples/sglang_multiturn/train.sh
Tau2 Benchmarks:
bash examples/tau2/train.sh
ColBench:
bash examples/colbench/train.sh
How It Works
- Rollout Phase: Agent interacts with environments, generating multi-turn conversations
- Intrinsic Reward Computation: For each turn with observation:
- Compute KL divergence between policies with/without observation
- Assign information gain as intrinsic reward to the action that produced the observation
- Advantage Estimation: Combine intrinsic and outcome rewards with adaptive weighting
- Policy Update: Optimize policy using InfoPO objective with variance-gated advantages
Monitoring
Training progress is logged via:
- Console Output: Real-time metrics
- Weights & Biases: Detailed visualizations and metrics
- Checkpoints: Automatic model saving
Citation
If you find this work useful, please cite our paper:
@article{kong2026infopo,
title={InfoPO: Information-Driven Policy Optimization for User-Centric Agents},
author={Kong, Fanqi and Zhang, Jiayi and Deng, Mingyi and Wu, Chenglin and Luo, Yuyu and Liu, Bang},
journal={arXiv preprint arXiv:2603.00656},
year={2026}
}
Acknowledgements
This codebase is built upon UserRL. We thank the authors for their excellent work.
License
This project is licensed under the Apache 2.0 License - see the LICENSE file for details.