The framework executes dual-probe workflows, setting ungrounded queries evaluated via token-level logprobs directly against responses augmented with live web grounding. This setup isolates the divergence between frozen pre-training associations and active retrieval data, enabling analysts to identify hallucinations, unearth blind spots, and reconcile discrepancies between internal model memory and real-world search visibility.
Referenced by