Shown two web pages, Gemini picks the first one 92% of the time. Pages rewritten to 31 measured text preferences win 15% from the second slot. Ordinary pages: 0%.
TL;DR. Gemini 3.6 Flash wrote 260 pairs of commercial web pages, 520 in total. One page in each pair came from a plain brief. The other came from the same brief plus 31 text factors measured across 24,445 head-to-head votes, with an instruction to apply whichever positive factors fit and avoid every negative one. The same model then judged 5,200 pairings and picked the guided page in 57.6% of them. That headline figure is inflated by position, because the judge picks whichever page it reads first in 92.4% of votes. Scored only from the second slot, where position works against it, the guided page won 396 of 2,600 votes, or 15.2%. The unguided page won 0 of 2,600.
Rank Signal measures one text factor at a time. It takes a short commercial page, produces a copy of it with a single edit that adds or removes one factor, and asks the model which of the two it would supply as a grounding result for a chat query it imagines behind them. The order of the two texts alternates on every pair. Across 24,445 votes covering 36 factors, 31 carry a 95% posterior interval clear of 50/50, 15 of them positive and 16 negative. Five sit on the line and count as no effect. The whole ledger cost $81.32 across 31 runs. Factors are identified here by number rather than name, and the numbering is stable, so rank factor 010 means the same factor in every table below.
Measuring factors one at a time says nothing about a writer applying all of them at once, so the composite test starts from scratch. Twenty-six business topics, ten pairs each, 520 pages in total. Both arms worked from the same one-line brief and the same 500-word target, and the second arm also received the 31 confirmed factors as a helps-and-hurts list, with the instruction to apply whichever positives suited the page and avoid every negative. Mean length came out at 493 words unguided and 506 guided.
Share of the 260 pages in each arm carrying each factor.
| Rank factor | Direction | Unguided | Guided |
|---|---|---|---|
| 004 | Positive | 0% | 100% |
| 014 | Positive | 3% | 100% |
| 006 | Positive | 0% | 100% |
| 032 | Positive | 15% | 100% |
| 002 | Positive | 72% | 100% |
| 010 | Positive | 32% | 98% |
| 009 | Positive | 100% | 100% |
| 001 | Positive | 98% | 58% |
| 013 | Negative | 100% | 5% |
| 033 | Negative | 9% | 0% |
| 008 | Negative | 100% | 0% |
256 of the 260 guided pages carry all seven positive factors a regular expression can detect, and the remaining four carry six. Unguided pages carry one to four, most often two or three. Rank factor 009 appears in every page in both arms because the generator produces it by default, so that row measures nothing.
Rank factor 001 is the one positive the guided arm dropped rather than added, from 98% of unguided pages to 58% of guided ones.
Rank factor 025, the strongest negative in the ledger at 0.0%, appears in 55% of unguided pages and 58% of guided pages, which reads as a failure to avoid it. The matches differ by arm. In the guided arm they are incidental to the subject: USDA in 23 pages, HVAC in 18, USPS in 14, ASTM in 13. In the unguided arm they are emphasis: SHOP in 45 pages, YOUR in 25, EXPLORE in 16, TAKE in 11.
One vote per pair, with the order alternating between pairs, put the guided page at 59.6%. The same votes show the first of the two texts winning 90.4% of the time, so the counterbalancing holds the average near the truth while censoring it, because one arm sits against a ceiling.
Judging each pair a second time in reverse order separates the causes. A page that wins from both positions did not win on order. 215 of the 260 pairs changed winner when the order changed. Of the 45 that held, the guided page won all 45 and the unguided page won none.
Ten votes in each order per pair, 5,200 judgements in total, measures how much of that flipping is sampling noise. 160 of the 260 pairs returned 0 out of 10 from the second slot, so the judge is close to deterministic on a given pair in a given order, and the flips are position rather than chance.
Every vote above was cast at minimal thinking with an eight-token answer budget, which is how the tool runs. Sixty of the 260 pairs were re-judged with thinking on, stratified so that 20 came from pairs the guided page had already won from the second slot, 20 from the middle, and 20 that had scored zero.
| Judge setting | Guided page, first slot | Guided page, second slot | Unguided page, second slot | First slot overall |
|---|---|---|---|---|
| Minimal, 5,200 votes | 100.0% | 15.2% | 0.0% | 92.4% |
| Thinking, 1,200 votes | 100.0% | 50.7% | 0.0% | 74.7% |
The unguided page won 0 of 2,600 votes from the second slot on reflex and 0 of 600 with thinking on, so its 95% interval reaches 0.42% at the top. The guided page moved from 15.2% to 50.7%. Reweighted to the full 260 pairs, its second-slot rate goes from 16.9% to 33.9%.
The 20 pairs that had scored zero from the second slot averaged 20.5% once the judge could deliberate, and 13 of the 20 moved off zero.
The 24,445 votes behind the ledger were all cast on reflex, so ten factors spread evenly across the win-rate range were re-judged with thinking on, 60 stored pairs each in both orders, 1,200 votes.
| Rank factor | Ledger (24,445 votes, reflex) | Re-judged (1,200 votes, thinking) | Shift |
|---|---|---|---|
| 025 | 0.0% | 0.0% | 0.0 |
| 028 | 14.0% | 12.5% | -1.5 |
| 020 | 29.2% | 21.7% | -7.6 |
| 019 | 35.1% | 31.7% | -3.4 |
| 036 | 47.7% | 55.0% | +7.3 |
| 007 | 50.4% | 48.3% | -2.1 |
| 029 | 62.5% | 67.5% | +5.0 |
| 011 | 80.9% | 81.7% | +0.7 |
| 002 | 93.2% | 93.3% | +0.2 |
| 010 | 98.8% | 100.0% | +1.2 |
None of the ten verdicts changed. The two columns correlate at 0.9935 and sit 2.9 points apart on average. Re-running the reflex judge on the same 60-pair samples put it 3.8 points from the full-ledger figure, so thinking moves the answer less than drawing a different 60 pairs does.
Position bias on these minimal pairs is 69.0% first-slot across all 24,445 stored votes and 66.8% with thinking on, against 92.4% and 74.7% on the unrelated full-length pages. A pair of pages on different businesses gives the judge no discriminator, and a pair differing by one edit gives it one.
How often a pair votes the same way in both orders tracks how far its rate sits from 50%. Rank factors 025 and 010 agree with themselves 60 times out of 60. Rank factor 007 agrees 8 times out of 60 and rank factor 036 agrees 14 times out of 60, which is what a rate near 50% consists of.
One model, Gemini 3.6 Flash, in every role: writing both arms, applying the edits behind the ledger, and judging. The judge prompt asks the model to imagine the chat prompts behind two candidate grounding results and supply one of them, so what is measured is selection by that model at that moment, not ranking, not traffic, and not any other model. Pages are synthetic commercial copy at roughly 500 words on 26 business topics, not live client sites. The five passes described here cost $27.65 in total.
Factors that require inventing a credential, 011 and 010 among them, change the claims a page makes as well as its presentation, which the design does not separate. Rank factor 012 sits at 49.4% and counts as no effect, against rank factor 010 at 98.8%, so the model is not scoring authoritative-sounding additions as a class.