After each round, the tested hypothesis receives a binary reward where \(\alpha\) increases by one if the edit improves the target's rank over the current best, while \(\beta\) increases by one if the rank remains unchanged or worsens.