ORPO

April 1, 2024 · View on GitHub

Updates (24.03.25)

 

This is the official repository for ORPO: Monolithic Preference Optimization without Reference Model. The detailed results in the paper can be found in:

Model Checkpoints

Our models trained with ORPO can be found in:

And the corresponding logs for the average log probabilities of chosen/rejected responses during training are reported in:

 

AlpacaEval

Description of the image
Figure 1. AlpacaEval 2.0 score for the models trained with different alignment methods.

 

MT-Bench

Description of the image
Figure 2. MT-Bench result by category.

 

IFEval

IFEval scores are measured with EleutherAI/lm-evaluation-harness by applying the chat template. The scores for Llama-2-Chat (70B), Zephyr-β (7B), and Mixtral-8X7B-Instruct-v0.1 are originally reported in this tweet.

Model TypePrompt-StrictPrompt-LooseInst-StrictInst-Loose
Llama-2-Chat (70B)0.44360.53420.54680.6319
Zephyr-β (7B)0.42330.45470.54920.5767
Mixtral-8X7B-Instruct-v0.10.52130.57120.63430.6823
Mistral-ORPO-⍺ (7B)0.50090.50830.59950.6163
Mistral-ORPO-β (7B)0.52870.55640.63550.6619