🏆 RM-Bench Leaderboard

July 23, 2025 · View on GitHub

Welcome to the RM-Bench Leaderboard!

First off, a big thank you to the community! Since the release of RM-Bench, we’re excited to see 50+ reward models evaluated using our benchmark. To facilitate further research and model comparison, we are launching this public leaderboard.

🔍 Notes:

  • The results listed here are mainly taken from the original papers or project reports.
  • If you find any mistakes or inconsistencies, please open an issue to let us know.

📢 Contribute Your Results: We welcome everyone to submit their results!
To add your model:

  1. Open a Pull Request (PR).
  2. Include a link to your paper/project.
  3. (Optional but appreciated) Attach your RM-Bench result file for reproducibility.

⚠️ Please note:
The leaderboard is still under construction, and there may be bugs or layout changes as we improve it. Thanks for your understanding!

🎉 Enjoy exploring and contributing to reward model research!

đź“° RM-Bench Leaderboard Update

Hi everyone! We have a few exciting updates and improvements to share with the community. Thank you once again for your enthusiastic support and contributions to RM-Bench!

✨ What’s New:

  1. License Type Column Added
    We've introduced a new column to the leaderboard to indicate each model's license type. This addition aims to make it easier for researchers and practitioners to identify models that are open for further development, research, or deployment.

  2. New Average Metrics: Domain Avg & Difficulty Avg
    Two new columns have been added:

    • Domain Avg: The average score across different domains (Chat, Code, Math, Safety).
    • Difficulty Avg: The average score across different difficulty levels (Easy, Normal, Hard).
  3. Clarification on Average Calculation
    We've noticed that in some cases, the reported Overall score may differ from the computed Domain Avg or Difficulty Avg. This is absolutely understandable — we realized we hadn’t clearly communicated our averaging method before. To clarify:

    We define the Overall score as:

Overall = Domain Avg = Difficulty Avg = (Chat + Code + Math + Safety) / 4

And each difficulty level is computed as:

Easy = (Safety_Easy + Math_Easy + Chat_Easy + Code_Easy) / 4

Normal = (Safety_Normal + Math_Normal + Chat_Normal + Code_Normal) / 4

Hard = (Safety_Hard + Math_Hard + Chat_Hard + Code_Hard) / 4

Following this definition ensures that Domain Avg = Difficulty Avg = Overall.

  1. New Column: Potentially Score Mismatch
    To gently assist with clarity, we’ve added a Potentially Score Mismatch column. This column simply highlights cases where the reported averages might differ slightly from what’s expected under our standard formula.

If your model is flagged here, please don’t worry at all — this is very likely due to us not having communicated the averaging method clearly enough earlier on. It’s not an error on your part, and we truly appreciate the effort that goes into every submission.

If you'd like to recheck your scores or ensure alignment with our formula, we offer a small helper function that can recompute the averages for you. You can find it here: 👉 utils.py from Line 48 to 75

We hope this helps make the leaderboard more transparent and consistent for everyone. Thank you again for your understanding and collaboration! đź’›

Detailed Leaderboard

Model NameModel TypeChatMathCodeSafetyEasyNormalHardDomain AvgDifficulty AvgOverallPotentially Score MismatchLicense Type
Skywork-Reward-V2-Llama-3.1-8B-40MScalar RM----97.696.993.5-96.096Not AvailableOpen Weight
Skywork-Reward-V2-Llama-3.1-8BScalar RM----97.095.086.5-92.892.8Not AvailableOpen Weight
REWARDANYTHING-8BReasoning GenRM76.790.375.290.289.485.384.483.186.486.4TrueOpen Weight
Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual + voting@32GenRM76.393.279.093.592.188.575.985.585.585.5FalseOpen Source
DeepSeek R1Reasoning LLM84.272.793.690.7--80.285.380.285.3Not AvailableOpen Weight
Llama-3.3-Nemotron-Super-49B-GenRM-MultilingualGenRM77.291.974.792.990.786.775.184.284.284.2FalseOpen Source
Llama-3.3-Nemotron-Super-49B-GenRM + voting@32GenRM74.092.777.492.192.687.372.384.084.184FalseOpen Source
RM-R1-DeepSeek-Distilled-Qwen-32BReasoning GenRM74.291.874.195.489.585.476.783.983.983.9FalseOpen Source
Llama-3.3-Nemotron-Super-49B-GenRMGenRM73.791.475.090.691.285.771.282.782.782.7FalseOpen Source
J1-Llama-70BReasoning GenRM---------82.7Not AvailableOpen Weight
Skywork-Reward-V2-Qwen3-8BScalar RM----91.985.770.1-82.682.6Not AvailableOpen Weight
Llama-3.3-Nemotron-70B-Reward-MultilingualScalar RM86.282.466.894.184.384.578.382.482.482.4FalseOpen Source
Qwen-3-Nemotron-32B-RewardScalar RM86.176.170.295.285.183.477.381.981.981.9FalseOpen Source
Skywork-Reward-V2-Qwen3-4BScalar RM----92.184.767.9-81.681.6Not AvailableOpen Weight
RM-R1-DeepSeek-Distilled-Qwen-14BReasoning GenRM71.890.569.594.186.283.674.481.581.481.5FalseOpen Source
Skywork-Reward-V2-Llama-3.2-3BScalar RM----91.584.167.8-81.181.1Not AvailableOpen Weight
Llama-3.3-Nemotron-70B-RewardScalar RM75.484.569.390.492.184.163.579.979.979.9FalseOpen Source
RM-R1-Qwen-Instruct-32BReasoning GenRM75.383.956.293.986.380.570.477.379.179.1TrueOpen Source
Skywork-Reward-V2-Qwen3-1.7BScalar RM----93.083.459.7-78.778.7Not AvailableOpen Weight
Qwen-2.5-Nemotron-32B-RewardScalar RM76.073.966.293.585.680.565.977.477.377.4FalseOpen Source
GPT-4.1LLM79.568.167.393.185.777.069.577.077.477.4FalseProprietary
Skywork-Reward-V2-Llama-3.2-1BScalar RM----91.379.957.8-76.376.3Not AvailableOpen Weight
RM-R1-Qwen-Instruct-14BReasoning GenRM75.675.460.693.682.677.568.876.376.376.1FalseOpen Source
Qwen3-8BLLM66.577.157.084.476.474.374.471.275.075TrueOpen Weight
Skywork-Reward-V2-Qwen3-0.6BScalar RM----90.378.054.8-74.474.4Not AvailableOpen Weight
DeepSeek V3LLM76.365.762.288.380.473.267.373.173.673.6FalseOpen Weight
Skywork-Reward-Llama-3.1-8B-v0.2Scalar RM69.362.153.496.089.375.852.670.272.672.6TrueOpen Source
RM-R1-DeepSeek-Distilled-Qwen-7BReasoning GenRM64.083.956.285.375.973.168.172.472.472.4FalseOpen Source
GRM-Llama3.2-3B-rewardmodel-ftScalar RM68.661.952.895.290.875.949.469.672.072TrueOpen Source
FsfairX-LLLaMA3-RM-v0.1Scalar RM67.362.855.791.887.474.852.869.471.771.7TrueOpen Source
Self-taught-evaluator-llama3.1-70BGenRM73.465.756.390.480.274.559.771.571.571.5FalseOpen Source
INF-ORM-Llama3.1-70BScalar RM66.365.656.894.891.876.144.870.970.970.9FalseOpen Source
Llama-3.1-Nemotron-70B-RewardScalar RM70.764.357.490.392.576.443.170.770.770.7FalseOpen Source
RM-R1-Qwen-Instruct-7BReasoning GenRM66.667.054.692.679.271.759.770.270.270.2FalseOpen Source
Skywork-Reward-Llama-3.1-8BScalar RM69.560.654.595.789.074.746.670.170.170.1FalseOpen Source
URM-LLama-3.1-8BScalar RM71.261.854.193.184.073.253.070.070.170FalseOpen Source
Nemotron-340B-RewardScalar RM71.259.859.487.581.071.456.169.569.569.5FalseOpen Source
Llama-3-OffsetBias-RM-8BScalar RM71.361.953.289.684.672.250.269.069.069FalseOpen Source
internlm2-20b-rewardScalar RM63.166.856.786.582.671.650.768.368.368.3FalseOpen Source
GRM-llama3-8B-sftregScalar RM62.762.557.890.083.572.748.668.268.368.2FalseOpen Source
EvalPlanner-Llama-8BGenRM---------68.1Not AvailableUnknown
ArmoRM-Llama3-8B-v0.1Scalar RM67.857.553.192.482.271.049.867.767.767.7FalseOpen Source
Skywork-Reward-Gemma-2-27BScalar RM69.554.753.291.978.069.254.967.367.467.3FalseOpen Source
internlm2-7b-rewardScalar RM61.771.449.785.585.470.745.167.167.167.1FalseOpen Source
Eurus-RM-7bScalar RM59.960.256.986.587.270.240.265.965.965.9FalseOpen Source
SOLAR-10.7B-Instruct-v1.0GenRM78.652.349.678.957.567.669.464.864.864.8FalseOpen Source
JudgeLRMReasoning GenRM59.959.951.987.373.266.254.864.864.764.7FalseOpen Weight
RM-Mistral-7BScalar RM57.457.052.787.288.667.134.963.663.563.5FalseOpen Source
Mistral-7B-instruct-Unified-FeedbackScalar RM56.558.051.786.887.167.335.363.263.263.2FalseOpen Source
tulu-v2.5-70b-preference-mix-rmScalar RM58.251.455.587.172.865.650.763.063.063FalseOpen Source
stablelm-2-12b-chatLLM67.254.951.665.269.163.546.659.759.759.7FalseOpen Weight