The loop uses a two-stage acquisition policy to select seeded hypotheses, applies targeted patch edits to candidate text, evaluates rank changes via an LLM judge model, and updates each hypothesis's Beta posterior distribution based on whether the edit achieved a rank improvement.