How Does AI Decide What to Cite?I Tested 125 Keywords Across 4 AI Engines

By Irina Saprykina, working in SEO and digital marketing since 2006.

How does AI decide which sources to cite? My data suggests there are two separate problems: getting into the candidate pool, and being selected once you're there. Google rank is the strongest observable path into that pool, but it doesn't explain most citations on its own. Among pages Google already ranks, page-level links, freshness, query-format match and AI Overview presence help separate cited from uncited pages - and each AI engine weighs those signals differently.

To test that, I pulled the real citations from ChatGPT, Claude, Gemini, Perplexity and Google's AI Overview for 125 keywords across seven topic areas, measured the pages each AI search engine cited against the ones it ignored, then re-ran part of it to see what moved. (My earlier 41-factors guide collected what the research says about AI citations; this is what my own runs show.)

The distinction that runs through everything:

Getting into the candidate pool (mostly a matter of ranking) and getting picked from it (a different, smaller set of signals) are two different jobs - and it's easy to treat them as one.

The short version: what moved citations

If you want the bottom line before the evidence, here is what mattered in my data, strongest first. Each factor is unpacked below, and the full statistical model - with odds ratios and confidence intervals - is at the end.

signal effect on being cited
Ranking in Google's top 10 the gate - the single biggest factor
Being a source in Google's AI Overview strong - and rarely talked about
Page-level backlinks strong, but mostly by helping you rank
Fresh content (updated recently) modest but real
Matching content type to the query's intent matters - editorial for "what/best", product pages for "buy"
Concrete facts and numbers on the page help, but only on data-style queries
Ranking in Bing as well as Google a small extra signal, rarely enough on its own
Blocking AI crawlers (robots) roughly halves the odds
Site-wide Domain Rating small once page-level links are counted
Page length, schema markup, llms.txt little to no effect on their own
Which engine you're aiming at changes the weighting - each one is different

Further down I also dig into the ecommerce picture (reviews vs storefronts), the per-engine differences, and how much citations move from one run to the next.

How I tested this

For each keyword I captured Google's top 40, Bing's top 20, the AI Overview sources, and the live citations from all four assistants (ChatGPT, Claude, Gemini, Perplexity) - then enriched every URL with its page-level backlinks, Domain Rating, word count, fact density, publish date, schema and content type.

Three runs:

That is 130 keyword-run observations representing 125 unique keywords - the SaaS deep-dive re-used the five SaaS keywords from the cross-vertical run. Plus re-runs at day 5 and day 32 to measure churn, two fan-out experiments plus a 160-call observability control, and a repeatability study (12 prompts × 8 identical runs × 3 engines). US, desktop.

Platform giants (Reddit, YouTube, Wikipedia) are excluded from the factor analysis - their scale and platform dynamics make them structural outliers, and less actionable for most sites.

Three scope notes up front:

The article runs in three lenses:

  1. Discovery - what gets into the pool at all (rank, the off-SERP tail, fan-out).
  2. Selection - what separates cited from uncited among ranked pages (links, freshness, format, AI Overview).
  3. Measurement regimes - four engines, two ChatGPT surfaces, sampling noise vs real drift.

The closing model then puts every factor in one place.

I'll grade each finding:

The SaaS findings then replicated on ecommerce - a structurally opposite vertical - which is why I trust the ★★★ ones. Where an outside study agrees (or seems to disagree), I say so.

Part 1 - Getting seen

Google's top 10 is the strongest citation gate I measured

★★★

Rank is the gate. I split Google's top 40 into two bands - the top 10, and positions 11-40 below it (roughly pages two to four). A top-10 page was cited by at least one of the four assistants 78% of the time in SaaS and 69% in ecommerce; an 11-40 page, 27% in both. Across six verticals the top-10 range was 50-87% versus 16-36% below it. (Throughout, "cited" means picked by one of the four AI assistants - ChatGPT, Claude, Gemini or Perplexity; Google's AI Overview is a separate source, analysed later as a supporting signal.) Ahrefs' 863K-SERP study found the same shape - 37.9% of AI Overview citations come from top-10 pages.

But ranking is never enough - in two directions. Depending on the vertical, 13-50% of top-10 pages got no citations at all. And more surprising: while an individual top-10 page has the best odds, most citations don't come from the top 10 at all. Count where every cited page actually ranks and only 19% sit in Google's top 10; 25% are at 11-40, and 56% don't rank in the top 40 at all (64% in ecommerce). Odds favour the top 10; sheer volume doesn't. The same odds-versus-volume gap comes up again - with page formats, and with each engine's habits.

Links separate pages below the top 10

★★★

Links to the exact page (its referring domains) were the most consistent authority signal I measured - but they only tell cited and uncited pages apart below the top 10. In the 11-40 band, the cited page had more links than the uncited one in 89% of keywords (93% in ecommerce) - measured per keyword, so a few heavily-linked topics can't skew it. Inside the top 10, that fell to a coin flip - 50% in both verticals. In other words, links mostly work by getting you ranked; once you're in the top 10, link count stops separating cited from uncited.

How much more linked are cited pages? It depends on the topic (cited vs uncited median referring domains): health 119.5 vs 2, affiliate 117.5 vs 40.5, SEO 15 vs 1, SaaS 5 vs 1 - huge where big institutions and small sites share a SERP, small where everyone is equally linked. "Build links" is still right - just treat it as a ranking lever, not a direct citation trick. (Two caveats: links partly stand in for brand familiarity, which I didn't measure; and inside the top 10 they still predict how many engines cite you, not just whether - more later.)

Domain Rating is mostly a mirage

★★★

Control for page-level links and Domain Rating nearly stops meaning anything: inside each backlink tier, cited and uncited pages have almost the same DR (52/54, 73/71, 88/84) - what looked like a DR effect was mostly its correlation with page links. So a modest site-wide Domain Rating won't hold you back: a well-linked page that fits the query is cited at close to the same rate whatever its domain's DR.

A small leftover effect survives stricter controls (about 1.13× odds per +10 DR), but it is much weaker than the page-level backlink signal. Among ranked pages with zero page links, a DR≥80 was cited 20% and a DR<50 18% - a strong domain does not rescue a link-less page. In affiliate, cited DR 86 = uncited DR 86.

This lines up with Ahrefs' 75K-brand study, which found domain-level link metrics "very weak" for AI visibility. They measured the domain; I measured the page. Same conclusion from both ends: the page is where authority counts.

Ranking in Bing too helps - but Bing alone rarely earns a citation

★★

Ranking in Bing as well as Google helps. Among pages at the same Google rank, those also in Bing's top 10 were cited far more often (11-40 band: 39-67% vs 24-27%). Ranking in two engines, not one, is an extra sign of relevance.

But Bing on its own rarely wins a citation: pages that ranked in Bing but not Google were cited just 3-8% of the time. The old "ChatGPT runs on Bing" shortcut didn't hold either - only 8-12% of ChatGPT's citations matched Bing, versus 23-24% for Google. Where the rest of ChatGPT's citations come from is a separate question, which I come back to below.

Part 2 - Getting quoted

On informational queries, homepages are never quoted

★★★

Zero homepages were cited as a content source - 0 of 109 in SaaS, 0 of 49 in ecommerce. Landing pages 24-29%. What gets quoted is editorial: comparison pages 52%, blog articles 44%, definitions 40%.

You may have seen studies saying homepages get the MOST AI traffic - Similarweb's May 2026 clickstream analysis found that after ChatGPT began surfacing prominent brand links, the share of referrals landing on homepages jumped from roughly 26-32% to around 60%.

Both are true, and they're not in conflict. Those clickstream datasets count referral clicks across all real prompts and include the navigational and branded behavior ("Notion", "HubSpot pricing") that my informational keyword sample deliberately excludes. And assistants often make brand-level, domain-only mentions - Gemini especially - which can render as homepage links. Their own explanation matches mine: the model names a company, the user lands on the homepage to evaluate it.

AI links your homepage as a brand; it quotes your editorial pages as sources. Publish the explainer or comparison for the query you want cited.

Match the format to the query

★★★

Across every measurable query format I tested, format follows intent: "vs" returns comparison pages, "what is" returns definitions, "best" returns roundups, "pricing" returns the vendor's pricing page. And the expected page type performed better within the top-10 rank band. In SaaS, the right type added 12-15 points; it replicated in ecommerce (definition 80% vs 67%, roundup 50% vs 20%, how-to 67% vs 14%, review 82% vs 50%). The "dominant page-type in the top 10" is a direct what-to-produce signal.

Some formats have a low ceiling - "best X" worst of all

★★★

Winning the top 10 pays off differently by format. Validation and definition queries cite their top-10 pages 85-90% of the time; "best X" roundups only ~47-48%, in both verticals. Half your hard-won top-10 roundups get cited, and most "ranked-but-ignored" pages are this format. Ecommerce buy-intent queries sit just as low (47%).

You may have seen studies calling listicles the most-cited format. That's mostly a matter of volume: commercial queries return so many listicles that they add up to a big slice of all citations. But any single top-10 listicle still gets cited only about half the time. A big share of the citations, modest odds per page - share isn't the same as odds.

Concrete facts - but only when the query wants numbers

★★★

For data queries ("average salary", "conversion benchmarks"), cited pages were far fact-denser (15.6 vs 8.7 facts per 1,000 words). For step-by-step queries, no difference at all (4.5 vs 4.2). Within SaaS, only price/spec formats showed a lift. Don't stuff numbers into a how-to; do put them on a benchmarks page. (Self-contained, unhedged claims help too, per the GEO literature - I didn't measure that one directly.)

Freshness matters modestly everywhere

★★★

Freshness is a modest but real factor, and statistically it behaves the same wherever a page sits: holding rank, page-level links and everything else equal, a recently updated page had roughly 1.9× the citation odds of a stale one, both inside the top 10 and below it.

What changes with rank is how much that edge is worth in practice. Below the top 10, citation rates are low, so the same odds boost shows up as a big jump in the actual rate: at positions 11-40, fresh pages were cited 32% of the time vs 18% for pages over a year old, a 14-point gap. Inside the top 10, where most pages get cited anyway, it only widens an already-high rate a little (80% fresh vs 72% stale). Same statistical effect, bigger real-world difference where citations are otherwise scarce.

Two practical notes:

Length, schema and title-match: the proxies and the myths

★★

Length: longer pages were cited a bit more across the whole sample, but the effect disappeared inside the top 10 - so length looks like a proxy for thoroughness, not a lever you pull. Aim for answer density rather than word count (retrieval reads long pages only partially anyway).

Schema presence is flat (Ahrefs' 1,885-page controlled test agrees - no lift). But specific types still correlate after rank control: FAQPage 86% vs 73%, Article 85% vs 69% inside the top 10. I can't fully separate "the markup helps" from "Q&A-shaped content helps and the markup tags along," so treat it as correlation. Boilerplate types (WebSite, Organization) showed no positive association. Add valid Article/FAQ markup as cheap hygiene, not as a lever.

Title-to-query overlap was NOT a citation factor in my data - flat once I controlled for rank. It plausibly helps you rank (and citation follows rank), but adds nothing on top.

A note on nofollow

Many of us nofollow external links by default. Does it help your own page get cited? The data says no: cited pages used plain follow links 89% of the time, and nofollow was slightly more common among uncited pages. This is a weak, suggestive signal, not something to build a firm rule on - but leaving your outbound links as plain follow links clearly doesn't hurt your own chances of being cited. If anything, I'd be especially careful not to nofollow the links to genuinely useful data sources you're citing - those are exactly the ones worth passing real credit to.

AI Overview presence is a strong co-signal

★★★

Among Google-ranked pages, being a source in Google's own AI Overview was one of the strongest co-signals I measured: AIO-source pages were cited ~74% of the time in both verticals, versus ~30-33% for ranked pages outside the AI Overview. Part of that is rank selection - AIO sources rank well - but the effect survives rank control, and it follows the now-familiar pattern: the gap is widest in the 11-40 band (59% vs 25% in SaaS, 56% vs 26% in ecommerce, roughly a 2-3× higher citation rate) and compresses inside the top 10. That makes AI Overview presence a third signal that separates cited from uncited pages most clearly at positions 11-40, alongside page links and freshness.

The AI Overview itself shows up on factual queries and fades on judgment ones (what-is 7/8, is-worth 3/8, when/why 1/8) - and in the query formats where the AIO appeared less often, the assistants tended to cite more sources of their own.

Where the off-SERP citations come from: the fan-out trail

One thread runs through the factors above: page-level links, freshness and AI Overview presence all mattered most in the 11-40 band, not inside the top 10. A likely reason is that these factors help a page rank - just not always for the query you're watching. AI assistants often expand your query into related sub-queries ("fan-outs") and search those too, so a page that ranks nowhere for the main keyword can rank well for a sub-query you never see, and get cited from there. That would also explain the big off-SERP tail from Part 1: many cited pages don't rank for the main query at all.

I tested this directly. Where the sub-queries were visible, cited off-SERP pages ranked for them far more often than matched non-cited controls (same-domain top-20: 30% vs 6%; exact-URL top-10: 28% vs 5%). But visibility is the bottleneck: ChatGPT exposed its fan-out queries in just 2 of 96 calls, so I can't say how much of the off-SERP tail this explains, and absence from the visible fan-outs isn't proof the model didn't use other sub-queries.

Part 3 - ChatGPT, Claude, Gemini and Perplexity cite different pages

The four AI engines barely overlap on what they cite

Take the same keyword and compare what any two of the four AI search engines - ChatGPT, Claude, Gemini and Perplexity - cite: median overlap never exceeded 16% in SaaS, 11% in ecommerce. Four assistants, four different source sets. Outside audits land in the same place - Averi reports ~11% ChatGPT - Perplexity domain overlap, citing a 680M-citation analysis, and similarly low pairwise agreement in other large audits. So "optimize for AI" isn't one target - here it's four, and in general it's as many as there are engines you care about.

Perplexity tracks Google closely; ChatGPT strays furthest off-SERP

Share of each engine's citations that sit in Google's top 10: Perplexity 34% / Claude 31% / Gemini 22% / ChatGPT 17%. Off-SERP share (cited but not ranking in Google's top 40): ChatGPT ~68% / Gemini ~55% / Claude ~46% / Perplexity ~34%. Perplexity is closest to your rankings; ChatGPT strays furthest off-SERP.

One subtlety worth stating carefully: among ranked pages, ChatGPT is actually the most backlink-selective of the four - but most of its citations aren't ranked pages at all.

Each engine rewards a slightly different factor

When I run the citation model on one engine at a time - still holding query, rank, links, freshness and the rest equal - each engine leans on a different factor. (These are four separate per-engine models, not a formal test that the gaps between engines are statistically significant.) The numbers below are odds ratios: an OR of 2 means about twice the citation odds when that factor is present, everything else held equal.

Scope update: this measures one ChatGPT retrieval mode ⚠

One boundary on everything above: it measures a single ChatGPT setup - the fast search-and-answer pipeline (gpt-4o with web search). When I probed a reasoning model instead, the sourcing changed a lot: it read far more pages and then concentrated its citations on primary sources - vendor documentation, independent testing labs, official status pages - while the third-party listicles that dominate the fast pipeline largely dropped out. On one health question it retrieved nothing at all. So read these results as specific to the fast mode; behaviour on other models can differ.

What a page cited by all four engines looks like

Sort pages by how many engines cite them. Pages cited by 0/1/2/3/4 engines have median page-links of 2 → 2 → 5 → 14 → 38, DR 74 → 91, top-10 share 5% → 78%, AIO presence 9% → 50%.

These aren't universal thresholds - they're medians across the specific keywords I checked, so treat the exact numbers as illustrative. But the direction is clear and, frankly, unsurprising: the more engines agree on a page, the more of these signals it tends to have stacked up at once. The portrait of a universal citation magnet: well-linked, top-ranked, present in the AI Overview.

Part 4 - The ecommerce twist

Most of the study so far pools informational and SaaS queries. Shopping queries behave differently enough to deserve their own look, so I ran a separate 36-keyword ecommerce set (~1,600 pages) - things like "best robot vacuum for pets", "running shoes for flat feet" or "air fryer vs oven" - this time letting retailers (Amazon, Walmart and the like) count as competitors alongside editorial pages. Three patterns stood out.

AI prefers reviews to storefronts

★★

The first question was simple: across these product and shopping queries, among pages that rank, which type gets cited more - editorial pages (reviews, guides, comparisons) or retailer/storefront pages? Editorial won comfortably. Storefronts were cited at roughly half the editorial rate - 25% vs 42% among ranked pages. Amazon was cited three times across 36 keywords; Dick's, REI and Home Depot edged it out. SE Ranking's AI-Mode shopping research saw the same thing (Amazon largely missing, eBay leading). AI reaches for reviews and guides much more often than product listings.

Storefronts earn citations mainly on buy and branded intent

★★

Next I split those citations by the intent behind the query, to see when a storefront does get picked. The retailer share of citations was 25% on transactional queries and 14% on branded/review queries - but under 6% on informational ones (what / best / vs). A product page is citable where the query is "buy", rarely cited where it's "best".

Even a brand's own page loses out on review queries

★★

The retailers above are resellers; the manufacturer's own product page (Dyson.com, Purple.com) is a different actor. It doesn't fare much better on judgment queries: on "X review" queries the brand's own page was cited only 2/4. A review pulls in third-party opinion by design, so even the brand's own site rarely wins one - the manufacturer gets cited for facts about its product, not for verdicts on it.

Part 5 - How stable are AI citations?

For this part I stopped asking what gets cited and asked whether a citation is even a stable thing to measure: I ran the same prompts many times over, then re-checked the same keywords across a day and again across a month, to separate real change from noise. Citation measurement itself turned out to be the finding. In repeated identical runs (12 prompts × 8 runs × 3 engines), a single ChatGPT answer exposed only ~29% of the URLs observed across the eight runs - and ten of twelve prompts never returned the same citation set twice - while Perplexity exposed ~89% and its sets repeat.

A 24-hour re-test separated time from that noise: ChatGPT's overlap across a full day (14%) exactly matched its overlap within minutes; what looks like overnight churn is mostly per-response sampling. Perplexity overlapped 90% at both distances - whatever the backend mechanism, its measurements are stable, so its multi-day movements are interpretable as real change.

The month-scale layer reads accordingly. Re-scanning 24 keywords at day 5 and day 32 gave a front-loaded retention curve - 100% → 63% → 57% - with 28% of the "lost" domains cited again by day 32: flicker, not exit. That curve mixes sampling noise with genuine drift and should not be read as pure decay; the stable core is anchored mostly on rank (30% top-10 share vs ~7% in the churn groups).

Practical read: measure ChatGPT and Gemini as distributions over repeated runs, not snapshots; Perplexity is much more defensible to track with single snapshots.

The final model - what holds up with every factor controlled at once

As a closing pass I put every measured factor into one model - query, rank, page links, Domain Rating, freshness, AI Overview presence, robots blocking, schema, length, fact density, all controlled simultaneously across ~2,400 ranked pages, with keyword-clustered confidence intervals. The factors that still stand out when they all compete at once:

factor (unit) odds ratio [95% CI] verdict
Google top-10 (binary) 5.42 [4.2-6.9] the gate - ranking in Google's top 10 is the single biggest factor
AI Overview source (binary) 3.11 [2.2-4.4] being a source in Google's AI Overview - rarely mentioned in GEO advice, yet the second-strongest factor here
Freshness ≤3 months (binary) 1.98 [1.5-2.5] modest but real, inside the top 10 and below it alike
Page-level links (log-transformed) 1.34 [1.24-1.46] per log-unit strong but with diminishing returns; works mainly by helping you rank
Robots AI-block (binary) 0.49 [0.32-0.77] blocking AI crawlers roughly halves the odds (some blocked pages are still cited - mechanism unclear)
Domain Rating (per +10 points) 1.13 [1.08-1.20] only a small effect once page-level links are counted
Raw length (log words) 1.06 [1.00-1.12] adds almost nothing on its own
Schema presence · fact density 0.93 / 0.99 (CIs cross 1) no measurable effect once everything else is held equal

A few notes on reading the table:

As a final check, I re-ran the same all-factors-at-once model from the table above, this time using only citations that resolve to a specific, exact page - dropping the domain-only matches, which are noisier - to make sure the results weren't an artifact of fuzzy URL matching. The main conclusions held (top-10 4.70, page-links 1.34, AI Overview 2.36, robots 0.55), which is the strictest test I could put this data through.

What did NOT hold - the myths I had to drop

Being honest about the popular claims that didn't hold up in my data is part of the point:

So how do you get cited by AI?

If I restrict the recommendations to what this dataset actually measured:

  1. Win the classic ranking for your priority queries, and treat links primarily as the ranking lever they behave like.
  2. Match page type to intent: dedicated editorial pages (explainer / comparison / roundup / guide) for informational queries; product and retailer pages have their clearest citation opportunity on buy-intent.
  3. Put concrete facts where the query wants numbers - higher fact density was associated with citation only on data-style queries.
  4. Be not-stale; expose a real updated date.
  5. Know your engine mix - rank matters most for Perplexity; ChatGPT's ranked selections skew toward domain authority; Claude favors freshness; Gemini leans hardest on Google's AI Overview.

Other GEO literature also points to answer-first structure, self-contained claims and original rather than commodity content; plausible, but I did not test those directly here.

None of it is a trick, and none is a guarantee - only better odds. It's SEO pointed at a new surface, not a separate ritual.

Want to see where a page stands? Run any URL through the free AI Readiness & Citation Checker - it scores a page against most of these factors and shows whether it's already appearing in AI answers.


Method & reliability (appendix)

← All posts