YFarmX

Direct preference optimisation (DPO)

AI

Direct preference optimisation (DPO): A method for aligning models directly on human preference pairs, skipping the separate reward model and reinforcement learning loop used in RLHF.

Related terms

Browse the full glossary →