Dual-Policy-Preference-Optimization

June 30, 2025 ยท View on GitHub

The codebase for the preprint `Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model'