Proximal policy optimisation (PPO)
AIProximal policy optimisation (PPO): A reinforcement learning algorithm widely used in RLHF to update a model's behaviour while limiting how far each step changes it.
Proximal policy optimisation (PPO): A reinforcement learning algorithm widely used in RLHF to update a model's behaviour while limiting how far each step changes it.