← all concepts

Bayes' Theorem

The rule for updating a belief when new evidence arrives, and the basis of the Rank Signal verdicts behind our Content Optimizer.

Where we use it

Bayes' Theorem is the mechanism behind our Content Optimizer. Each round of optimization tests one content edit against competing passages and asks an AI ranker to pick a winner. Bayes' Theorem turns those win counts into a verdict we can trust, rather than a raw percentage.

We publish the full method on the Rank Signal methodology page. Each signal starts from a Jeffreys prior, \(\mathrm{Beta}(0.5, 0.5)\), the standard prior for an unknown proportion. After \(w\) wins in \(n\) trials the posterior is \(\mathrm{Beta}(w + 0.5,\ n - w + 0.5)\) and we report its mean:

$$\frac{w + 0.5}{n + 1}$$

That mean carries an equal-tailed 95% credible interval alongside it. We call a signal confirmed positive or negative only once the interval clears 50% by a wide enough margin. Otherwise we record it as no effect.

What the prior buys us

The prior is what keeps small samples honest. One signal in our published sample won 0 of 35 trials, a raw rate of 0%. Its posterior mean is 1.4%, because the prior holds open some chance of a win until the evidence rules it out. A signal that is genuinely harmful lands far lower and stays there: the strongest negative in that sample won 6 of 185 trials, for a posterior mean of 3.5% and a credible interval of [1.4%, 6.6%]. Same formula, different weight of evidence, different answer.

The formula

$$P(A \mid B) = \frac{P(B \mid A)\, P(A)}{P(B)}$$

  • \(P(A)\) is the prior, how likely \(A\) was believed to be before the new evidence.
  • \(P(B \mid A)\) is the likelihood, how likely the evidence is if \(A\) is true.
  • \(P(B)\) is the evidence, how likely the evidence is overall, under any hypothesis.
  • \(P(A \mid B)\) is the posterior, the updated probability of \(A\) once the evidence is counted.

The theorem revises a belief in light of new information. It does not replace the prior; it weighs the prior against the evidence and produces a number that reflects both. A strong prior needs strong evidence to shift it. A weak prior shifts easily.

Base rate neglect

Say you run SEO tests on 200 candidate ranking factors, using a test method that is 90% accurate. Ten of the 200 genuinely move rankings. Your tests catch nine of those ten. They also false-flag 10% of the 190 that do nothing, which is another 19 factors. Twenty-eight come back positive and nine of them are real, so a positive test result is right under a third of the time.

The test method is not broken. The base rate is doing the work: most things you can change about a page move nothing, so a small false-positive rate applied to a large pool of duds swamps the true positives from the small pool of factors that matter. Ignoring that is base rate neglect, and it is the same error as reading a raw win rate off 35 trials and calling the signal confirmed.

Related concepts

Concept