Multiple evaluations per round insulate the optimizer from non-deterministic LLM output, aggregating individual scoring passes into a median value that provides a stable, reproducible ranking signal.
Referenced by