In each round, the loop draws a random sample from the Beta posterior distribution of every candidate edit claim to balance the exploration-exploitation tradeoff, picking the highest-scoring hypothesis for the ideator to apply and test against the competitive re-ranker.

Referenced by