← back
No, AI doesn't prefer Reddit. Search does.

No, AI doesn't prefer Reddit. Search does.

Analysis of citation mining data shows Reddit has a 99.39% rejection rate from OpenAI, zero citations from Anthropic, and heavy usage by Google.

The SEO community is abuzz treating Reddit as the holy grail of AI visibility, but its presence is AI answers is a reflection of its organic search performance, nothing more.

We looked at the last 6 months of our citation mining data and found that Reddit is, in fact, the most rejected website in our entire dataset.

Domain Retrieved Cited Rejected Selection Rejection
Reddit 491,024 3,012 488,012 0.61% 99.39%
Wikipedia 229,879 12,968 216,911 5.64% 94.36%
arXiv 46,700 359 46,341 0.77% 99.23%

OpenAI's models produced only around 3,000 Reddit citations out of the nearly half a million times the domain was supplied as a candidate grounding source. That's a selection rate only 0.6%.

Note: Reddit was among the candidate sources in 76% of OpenAI's searches, and each of those searches carried about eight different Reddit pages, which is how the sample reaches 491,024 Reddit page-suggestions and why 488,012 of them went uncited.

Relevant Read: Grounding Source ≠ Citation ≠ Mention

OpenAI Query Fanouts

Direct use of "reddit" in OpenAI's models over time.

Free Style Prompts

Long form prompts used for qualitative citation mining including 27,351 probes and 103,974 fan-outs.

Month Fan-outs With "reddit" Share
March 79,647 12 0.015%
April 18,904 2 0.011%
May 35,679 2 0.006%
June 31,839 0 0.000%
July 12,268 0 0.000%

Quantized Entity Prompts

Short form entity based prompts used for daily visibility tracking with OpenAI including 75,369 probes and 74,363 fan-outs.

Month Fan-outs With "reddit" Share
April 12,278 0 0.000%
May 26,656 0 0.000%
June 25,456 0 0.000%
July 9,973 0 0.000%

Note: This is "reddit" presence data is biased to our prompts and clients.

Google

Google cites Reddit heavily. Across 697,768 cited sources, 14,127 are Reddit, about 2 percent, second only to YouTube. Google's grounding data reports the sources it cited one-for-one with the answer, and never the wider set it retrieved and set aside. Reddit's selection and rejection rates for Google cannot be measured, only inferred, because the discarded pool is not in the data.

In a separate and dedicated probe on a subset of our citation mining dataset Reddit.com appeared as a result 933 times, across 916 of the 3,450 searches. That's a 26.6% presence rate (about one in four searches), and 3% of all 30,952 results.

Reddit was a candidate source in 76 percent of OpenAI's grounded probes (59,570 of 78,331), against 26.6 percent of the Google searches we sampled.

Extrapolating Google's selection rate

Google never exposes its rejection, so its selection rate has to be inferred rather than measured. The method: a domain's selection rate is its share of citations divided by its share of the candidate pool, read against the engine's overall rate.

OpenAI is the control, because both sides are known there. Reddit is 19.5 percent of OpenAI's candidate pool and 0.93 percent of its citations, a selection rate of 0.048 times the average domain, or 0.61 percent against OpenAI's overall 12.85 percent. Reddit is filtered far harder than the typical source.

For Google only the citation side is measured, at 2.0 percent Reddit, and the candidate side is inferred from the sampled search results, at 3.0 percent Reddit. That puts Reddit's selection at about 0.67 times Google's average domain, roughly 14 times more favorable than OpenAI's treatment of Reddit relative to each engine's own baseline. Google selects Reddit close to the rate it selects everything else.

Turning that factor into an absolute rate requires assuming Google's overall selection rate, which its one-for-one reporting hides. If Google filtered as hard as OpenAI, at 12.85 percent overall, Reddit's rate would be about 8.6 percent. If Google keeps most of what it grounds, at 80 to 90 percent overall, it would be about 54 to 60 percent. The predicted range is roughly 9 to 60 percent, against OpenAI's measured 0.61 percent, so even the low end is about 14 times OpenAI's rate.

These Google figures rest on a proxy. The 3.0 percent candidate share comes from 3,450 sampled DataForSEO searches, not from Google's own grounding pool, which is not in the data. Matching those stored searches to Google's AI-answer citations on the same queries would replace the range with a single measured rate.

Reddit OpenAI (measured) Google (inferred)
Selection rate 0.61% ~9–60%
Rejection rate 99.4% ~40–91%

TL;DR. OpenAI is handed Reddit constantly and rejects almost all of it. Google is handed Reddit less often and keeps it at close to its normal rate. On the measured side, OpenAI cites 0.61 percent of the Reddit pages it retrieves and rejects 99.4 percent. On the inferred side, Google's Reddit selection rate lands between about 9 and 60 percent, so even at its harshest Google keeps Reddit at least 14 times more often than OpenAI.

Google Query Fanouts

Direct use of "reddit" in Google's models over time.

Free Style Prompts

Long form prompts used for qualitative citation mining including 34,237 probes and 121,796 fan-outs.

Month Fan-outs With "reddit" Share
April 26,764 0 0.000%
May 61,460 9 0.015%
June 53,486 5 0.009%
July 22,331 0 0.000%

Quantized Entity Prompts

Short form entity based prompts used for daily visibility tracking with Google including 79,259 probes and 164,041 fan-outs.

Month Fan-outs With "reddit" Share
March 50,117 9 0.018%
April 14,500 0 0.000%
May 41,604 4 0.010%
June 9,425 3 0.032%
July 6,150 3 0.049%

Anthropic

Anthropic cites Reddit zero times. Across 139,601 grounding sources from May to July 2026, Reddit does not appear once. Reddit is not filtered out here the way it is at OpenAI. It is never supplied as a candidate in the first place.

Probe type Probes Total fan-outs "reddit"
Visibility 9,020 12,780 0
Citation mining 7,713 12,021 1


Dan Petrovic · Jul 21, 03:50
“out of the nearly half a million times the domain was supplied as a candidate”

This is fascinating, but how are you retrieving this data? Are you running the fan-out prompts separately to see the citations or are you somehow harvesting from OpenAI backend response file?

Yaacov Goldstein · Questions · · Aug 02, 10:17 ·

Directly with the model and official grounding sources: https://developers.openai.com/api/docs/pricing#:~:text=Web%20search

The API metadata returns both selected and unselected grounding sources, which you can't see by scraping web chat.

Dan Petrovic · Expands · · Aug 02, 13:47

Dis you found arXiv at least as rejected as Reddit? And maybe also as retrieved…

Olivier de Segonzac · QuestionsExpands · · Jul 24, 22:40 ·

Absolutely 99.23% rejection rate but only 10% as a retrieved source compared to Reddit.

Dan Petrovic · SupportsExpands · · Jul 27, 04:50