Relevance scores (such as QRS and ERS) are computed by executing \(N\) repeated evaluation samples per active model, prompting the LLM with a strict binary recommendation question such as asking whether it would recommend the brand for a given topic or search intent. The system records the responses in the database, extracts clean "yes" or "no" outcomes, and calculates the final relevance percentage as the proportion of positive recommendations relative to total completed probe samples across queries, entities, and model providers.