To smooth out LLM output variance, the system runs N independent ranking calls concurrently (typically 5 passes) and calculates the target item's median rank across N samples, selecting the evaluation closest to the median as the canonical ranking and rationale.
flowchart TD
D[Candidate Draft + Competitors] --> S1[Sample 1 Evaluation]
D --> S2[Sample 2 Evaluation]
D --> S3[Sample 3 Evaluation]
D --> S4[Sample 4 Evaluation]
D --> S5[Sample 5 Evaluation]
S1 --> M[Calculate Median Target Rank]
S2 --> M
S3 --> M
S4 --> M
S5 --> M
M --> C[Select Canonical Sample & Rationale]