In each round, the loop checks for convergence, samples a hypothesis via the acquisition policy, generates a refined edit, benchmarks it against competitor items using median ranker samples, and updates the claim's Beta distribution based on whether the position improved.