Each claim starts with a Beta prior calibrated from honest priors, then undergoes Bernoulli reward updates where alpha increments on a rank improvement and beta increments on failure, shaping the probability distribution sampled by Thompson acquisition.
Referenced by