In Stage 1, the system samples a Rank Factor Class weighted by its freshness—penalizing categories that have already undergone heavy testing—to ensure structural diversity across iterations. Once a class is chosen, Stage 2 applies methods like Thompson Sampling or Upper Confidence Bound over individual Beta-distributed hypotheses to select the specific edit with the highest expected reward.