YFarmX

Proximal policy optimisation (PPO)

AI

Proximal policy optimisation (PPO): A reinforcement learning algorithm widely used in RLHF to update a model's behaviour while limiting how far each step changes it.

Related terms

Browse the full glossary →