The Query Relevance Score (QRS) is computed by evaluating target LLMs with a strict binary prompt—asking directly whether a specific brand is recommended for a given topic—and calculating the percentage of affirmative "yes" responses out of all returned runs. Repeating this probe across multiple stochastic samples (ranging from 1 to 100 iterations per item) isolates statistical confidence against model variance, allowing the system to surface critical recommendation gaps across different providers and categories.
Referenced by