added codes
December 22, 2025 · View on GitHub
TFPI: Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
1Hunyuan, Tencent
2The Hong Kong University of Science and Techology
3The University of Hong Kong
2The Hong Kong University of Science and Techology
3The University of Hong Kong
📝 News
- [2025/12/22] We released the official code repository.
- [2025/11/7] We released the model checkpoints.
- [2025/9/30] We released the paper and temporary repo !
Reproduction
This is a temporary repository. Our code and training checkpoints are currently undergoing an audit in accordance with Tencent's regulations.
For now, to reproduce our results, you can perform TFPI training using VeRL codebase by modifying utils/datasets/rl_dataset.py:
raw_prompt = self.tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
# added codes
raw_prompt += "<think>\n\n</think>\n\n"
For evaluation, you could refer to DeepScaler and IFEval evaluation.
✏️ Citation
Bib
If you find TFPI useful for your research and applications, please cite using this BibTeX:
@article{xu2025tfpi,
title={Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners},
author={Xu, Xin and AI, Cliveb and Yang, Kai and Chen, Tianhao and Wang, Yang and Yang, Saiyong and Yang, Can},
journal={arXiv preprint arXiv:2509.26226},
year={2025}
}