Thompson sampling guides the snippet optimization loop by drawing random samples from each hypothesis's Beta posterior distribution and probabilistically selecting edits with the highest expected rank improvement to efficiently balance exploration and exploitation.
Referenced by