CTTS: Collective Test-Time Scaling
August 6, 2025 ยท View on GitHub
We investigate the optimal paradigm of CTTS: (1) single agent to multiple reward models (SA-MR); (2) multiple agents to single reward model (MA-SR); and (3) multiple agents to multiple reward models (MA-MR). Extensive experiments demonstrate that MA-MR consistently achieves the best performance. Based on this, we propose a novel framework named CTTS-MM that effectively leverages both multi-agent and multi-reward-model collaboration for enhanced inference.
News
- [25/08/06] ๐ฅ CTTS-MM is comming! Check our paper. We are currently restructuring our codebase and will release them in a more easy-to-use way.