Group relative policy optimisation (GRPO)
AIGroup relative policy optimisation (GRPO): A reinforcement learning method that scores several sampled answers against each other rather than against a learned value model, cutting the cost of training reasoning models.