In the Content Optimizer, Thompson sampling models each content hypothesis as a Bernoulli reward distribution, drawing random samples from updated Beta posteriors combined with diversity weighting to select which edit to test next.
Referenced by