README.md
June 20, 2026 ยท View on GitHub
Less Is More: Elevating RAG via Performance-Driven Context Compression
0. Data Download
The data for the distillation and RL training is available at https://drive.google.com/drive/folders/1OMKuh7Mj5_Jf45yrPBtP1Apw8A16LiM6?usp=sharing
1. Distillation for Warm-Start
The implementation for the distillation stage is based on the LLaMA Factory framework. Please refer to their documentation for environment setup.
After completing the setup and placing the downloaded data into the data directory, execute the following commands to begin distillation training:
cd LLaMA-Factory
llamafactory-cli train examples/train_full/qwen_full_sft_nq_lr5e5_epoch2_bs128.yaml
2. Performance-Driven RL Training

Figure: Reward trajectories from two separate training runs. The rewards increase steadily and eventually stabilize, suggesting stable training and convergence.
Our RL training implementation is built upon the verl framework.
-
Deploy the LLM for Reward Calculation:
First, run
reward_llm_serve.shto deploy the QA LLM. This model is responsible for generating answers, which are subsequently used for reward calculation. -
Configure the Reward Service:
Please update the
verl_core/reward.yamlconfiguration file with your specific vLLM deployment IP address and the model name. -
Update the Configuration Path:
Then, navigate to line 105 in the file
verl_core/verl/workers/reward_manager/api_prime.py, where you will find the following line of code:config_path = "verl_core/reward.yaml"Ensure this
config_pathvariable points to the correct location of your modifiedreward.yamlfile. -
Run the Training Script:
Start RL training by running the script below.
cd verl_core/examples/grpo_trainer sh train_compressor_nq.sh