After each round, the tested hypothesis receives a binary reward where \(\alpha\) increases by one if the edit improves the target's rank over the current best, while \(\beta\) increases by one if the rank remains unchanged or worsens.
Referenced by