Thompson Sampling draws a random success probability from each hypothesis's Beta(alpha, beta) distribution and selects the claim with the highest expected reward theta to balance exploration and exploitation.