Re-ranking model preferences is measured by submitting the target snippet alongside competitor snippets to judge models across multi-sample ranking consensus to extract rank positions and per-item rationales.
Referenced by