added codes

December 22, 2025 · View on GitHub

TFPI: Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners

arXiv
1Hunyuan, Tencent 
2The Hong Kong University of Science and Techology 
3The University of Hong Kong 

📝 News

Reproduction

This is a temporary repository. Our code and training checkpoints are currently undergoing an audit in accordance with Tencent's regulations. For now, to reproduce our results, you can perform TFPI training using VeRL codebase by modifying utils/datasets/rl_dataset.py:

raw_prompt = self.tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
# added codes
raw_prompt += "<think>\n\n</think>\n\n"

For evaluation, you could refer to DeepScaler and IFEval evaluation.

✏️ Citation

Bib

If you find TFPI useful for your research and applications, please cite using this BibTeX:

@article{xu2025tfpi,
  title={Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners},
  author={Xu, Xin and AI, Cliveb and Yang, Kai and Chen, Tianhao and Wang, Yang and Yang, Saiyong and Yang, Can},
  journal={arXiv preprint arXiv:2509.26226},
  year={2025}
}