AI Citation Ranking Factors: What My Tests Actually Show
By Irina Saprykina, working in SEO and digital marketing since 2006.
Every week a new checklist promises to fix your AI visibility. Most of it is guesswork, or copied blindly from the last person's post. This guide is different: it's built from the largest public studies on AI citations, plus what I see when I run my own AI Readiness & Citation Checker on real pages.
In short, here is how I built it:
- Studied the research. I gathered the most credible public work on AI citations; the full source list and method are in a detailed note below the factors.
- Added my own findings. A few factors the studies underplayed, from auditing real pages myself.
- Built a checker. A script that scores any page against most of the factors in this list.
- Tested and calibrated. I ran it on real pages and re-weighted the factors against real citations.
One thing up front: none of these factors guarantees a citation, and a citation is not the same as a click. They shift the odds in your favour, nothing more.
Before the list, three things my own runs showed: there is no single recipe, because what matters shifts with the topic and the format of the query; links mostly get a page into the top of Google rather than win the citation once it is there; and a good share of what AI cites never appears in Google's top results at all. Full findings below.
What Counts as "AI Search"?
"AI search" isn't one product. It's the growing set of systems that answer a question directly and cite sources, instead of only listing blue links: Google AI Overviews and Google AI Mode, ChatGPT Search, Perplexity (an answer engine built entirely around citations), Google Gemini and Claude. The factors in this guide are meant to raise your citation probability across all of them, not just one.
AI Ranking Factors
Getting cited takes two things, and they shape the whole list: a page has to rank in classic search, which still feeds the AI, and then hold up on its own once the AI pulls a single passage out of it.
I've split the factors into seven blocks below, listed roughly in order of impact.
1. Can AI Reach and Read Your Page?
Nothing else matters if an AI crawler can't fetch you.
1.1 AI Crawler Accessibility
Crawlability is the most important factor in the data, for a blunt reason: before anything else can help, the assistants have to be able to fetch the page. Each engine sends its own crawler, with its own name in your logs (GPTBot, ClaudeBot, PerplexityBot, and so on). If a page answers any of them with a block, a login wall or a soft 404, it drops out of that engine's candidate pool.
Most of these blocks are accidental. A CDN, WAF or bot-management layer can turn real crawlers away while the page looks perfectly healthy to you. Cloudflare is the one to check first, since it now blocks AI crawlers by default on new domains and fronts a large slice of the web, but Fastly and other bot managers can do the same.
Two checks beat guessing:
- Read robots.txt line by line for each named AI bot, not just the generic
*rule. - Confirm in your server logs that the real crawlers actually arrived and got a 200. A header-spoof test doesn't count; it only proves the page answers a faked user-agent.
1.2 Content Visibility
If your important content only appears after JavaScript runs, or lives inside tabs and accordions, retrieval may never see it. Put the substance in the visible HTML.
1.3 Content Display Permissions
Directives like nosnippet, max-snippet and
data-nosnippet tell search and AI bots how much of a page they may show, and which
parts they may not quote at all.
These directives show up in three places: the X-Robots-Tag HTTP
header, the <meta name="robots"> tag in the page head, or a data-nosnippet attribute
on individual elements in the body. Check whether your pages carry any of them, and make
sure they aren't quietly closing off content you actually want cited.
2. Classic Ranking Still Decides
Before it answers, an AI usually runs a search first, so where you rank in classic results still decides what it even sees. This is the block where traditional SEO carries over almost unchanged: the search engines feeding these systems are the same ones you have always optimised for.
2.1 Organic Ranking
Ranking in Google's organic top 10 for the main query gives you more chances of being cited across AI systems.
In my tests, AI Overview and Perplexity citations overlapped most strongly with the top 10, and Ahrefs' study of 863K SERPs found the same for AI Overviews specifically: about 38% of their citations come from top-10 pages. Claude and Gemini followed the rankings less closely, and ChatGPT least of all, though even there most queries still turned up one or two cited pages from the top results.
Track your organic rankings in Google Search Console; its new generative-AI report, a phased rollout since June 2026, breaks out impressions from AI Overviews and AI Mode (no clicks or queries yet, and not on every account).
2.2 Ranking in Bing
ChatGPT is usually described as running on Bing, and an earlier Seer Interactive study matched 87% of its citations to Bing's top results. That link looks weaker now. Comparing like for like, each engine's own top 20, only 8 to 12% of ChatGPT's citations sat in Bing's against 23 to 24% in Google's, so Google leads by two to three times. OpenAI has been building out its own index, which may explain the drift.
Ranking in Bing is still worth having, just for a different reason than the shortcut it is usually sold as. In my runs, pages that also ranked in Bing's top 10 were cited markedly more often than pages that didn't, even when their Google position was the same, so Bing works as a second, independent read on whether a page really answers the query. The effect is largest below the top 10, which is exactly where you need the help.
Check that Bing has your pages indexed (via Bing Webmaster Tools) and that they rank there too, but don't treat Bing as a back door into ChatGPT.
2.3 Query Wording Match
A page is more likely to be pulled in when its wording sits close to the question, and the title carries the most weight here: AirOps found that titles sharing at least half their words with the query were cited about 2.2x more often. What counts here is literal overlap, not just semantic closeness: a title that means the same thing in different words scores lower than one that reuses the wording people actually type. So lead your title with the main phrase for the topic, and in the body, try to cover the most common questions and angles your readers are likely to have.
2.4 Matching Search Intent
Just as in traditional search, a page is more likely to be cited when its structure and content fit the intent behind the query. If you're not sure what format to use, look at which pages dominate the Google top 10 for that query and follow the pattern. There are exceptions: depending on the query, AI may reach for authoritative sources, research or brand pages instead.
2.5 Fan-Out Coverage (Single Page)
When you ask a question, the AI doesn't run it as a single search. It breaks the topic into several fan-out sub-queries, related questions you never typed, to cover it as widely as possible. Ask Gemini "how to negotiate a salary," for example, and it will also look for pages on things like:
- salary negotiation tips and strategies
- how to prepare for a salary negotiation
- researching salary benchmarks and your market value
- salary negotiation scripts and examples
- how to make a counter-offer
- responding to a job offer
- common salary negotiation mistakes to avoid
The more of these angles a single strong article covers, and ranks for, the more likely that one page is to land in the answer. This isn't long-tail keyword coverage, though: some sub-queries sit close to the original, but others barely touch it ("responding to a job offer" is a different search altogether). You also can't choose them by volume, since AirOps found that roughly 95% of fan-out queries have no search volume at all.
2.6 Topic-Cluster Coverage (Whole Site)
Coverage helps at the site level, not just within a single page. Just as a single page that ranks across the fan-out gets cited more often, a site that covers the topic across more articles, each taking a different angle, has more chances of being cited at all. That breadth is what topical authority actually looks like to a machine. Build real depth around the subject in several well-linked pages, not one thin page per keyword.
3. Content That Actually Gets Cited
Ranking earns a page a look; whether a passage from it actually gets quoted comes down to how the content itself is written.
3.1 AI-Ready Structure
You don't need to "chunk" your content for AI, or rewrite it in some special machine style; whatever the advice going around, Google says directly that neither is needed. What retrieval actually needs is structured content it can slice cleanly: a logical heading hierarchy, short sections, and tables or lists where they fit. Format it for a skim-reader and the machine benefits too.
3.2 Answer Near the Top
AI engines pull only a limited slice of any single URL, and they sample it from the top down. Dan Petrovic's analysis of Google's grounding chunks found that a query gets only a fixed grounding budget, roughly a couple of thousand words shared across all its sources, so the longer your page, the smaller the fraction of it that ever reaches the model: a page over 3,000 words has only about 13% of its content used, against roughly 60% for a short one. An answer buried far down the page may never be read at all. Open the page, and each major section, with the direct answer, then add the detail underneath.
3.3 Self-Contained Passages
Each key claim has to make sense on its own, because retrieval lifts a passage away from everything around it. "It scales better than the alternatives" collapses the moment it's pulled out, because on its own it names neither the thing nor what it beats, while "PostgreSQL handled 10,000 writes per second in this benchmark" still carries its full meaning.
3.4 Concrete, Checkable Facts
Numbers, dates and named specifics are what an engine can lift and attribute; soft generalities give it nothing to hold onto. Wherever you'd reach for a line like "coffee has a lot of caffeine," put the actual figure instead: a brewed cup has about 95mg. This counts most when the question asks for numbers; on step-by-step queries, fact density barely moved anything in my tests.
3.5 Direct, Unhedged Claims
Engines quote confident statements and pass over the ones buried in caveats, so state your position plainly wherever the evidence lets you ("cast iron is the best pan for searing a steak"). Watch hedge density rather than the odd word, though: one "tends to" per 5,000 words is nothing to worry about.
3.6 Links to Sources
Pages that cite sources with credible evidence get cited more. The Princeton GEO study (KDD 2024) measured lifts of up to 40% from its optimization methods, with citing sources, adding quotations and adding statistics among the most effective. Link to authoritative, credible sources near your numbers, not from the footer; pages with cited sources are pulled into answers far more often.
3.7 Content Length
Longer content ranks a little better on average, but the data is mixed and long pages are retrieved whole less often. Optimize for answer density, not word count. If you're unsure what length fits, look at the pages already cited for your query, and pay most attention to those from sites close to yours in size, authority and topic.
3.8 Content Freshness
Freshness correlates with citations, but modestly: Ahrefs measured AI-cited content as 25.7% fresher than organic results across ~17M citations, not the viral "4.3x" figure, which appears in no primary study. Set a sensible refresh cadence for your priority queries and expose a real "last updated" date; more recently updated pages are cited more.
3.9 Language Match
An engine tends to answer in the language it was asked in, and to
reach for sources in that same language, so a page has to be written in your audience's
language to be pulled into their answers. If you serve several markets, give each one its
own localized version and wire them together with hreflang so engines pick the right one.
3.10 Entity Consistency
Describe who you are and what you offer the same way everywhere: not just across your own site, but anywhere you can shape how you're written up, from press mentions and guest posts to profiles and directories. One product name, one clear description of what it does, the same names for the people and places involved. Drifting between phrasings, or trailing off into "it" and "this," makes an engine work to decide whether two mentions are even the same thing, which weakens how confidently it can attribute and quote you.
4. Brand and Off-Page Authority
Off-page signals are where AI search departs most from a pure on-page checklist.
4.1 Unlinked Brand Mentions
Unlinked mentions are about your brand's overall share of voice in AI answers, not any single page. Ahrefs' 75K-brand study measured how often brands get mentioned across ChatGPT, AI Mode and AI Overviews, and unlinked web mentions of a brand tracked that visibility more closely than any link metric: a correlation around 0.66 for AI Overviews, while backlinks and other raw link counts came out, in Ahrefs' words, "very weak." Branded anchor text (0.53) and branded search volume (0.39) outranked them too.
Read it as a correlation, though, not a lever: Ahrefs is explicit that improving these numbers won't automatically lift you, and Google filters obviously inauthentic mentions. Genuine digital PR that gets your brand talked about, linked or not, is still a stronger bet than link building alone.
4.2 YouTube Mentions
Of all the brand signals in Ahrefs' 75K-brand study, mentions of a brand on YouTube, in video titles, descriptions and transcripts, correlated most strongly with its visibility in AI answers, around 0.74, ahead of every other factor. The likely reason: Google and OpenAI both trained their models on YouTube transcripts. Getting your brand named in relevant videos is one of the strongest external moves available, though like every correlation here it's a bet, not a guarantee.
4.3 Third-Party Corroboration
When independent communities like Reddit and Quora discuss your brand, AI systems get confirmation of it from outside your own pages. Most of that discussion carries no link, so its value is as corroboration, not as a backlink.
4.4 Domain Authority
Real, but weaker and far less consistent than most people assume. In some topics domain rating does separate the cited from the ignored; in others both groups sit at the same rating and it tells you nothing. Page-level authority was the steadier signal in every test I ran, so treat domain rating as one input, not the goal.
4.5 Known Source
Not every citation comes from a live search. An engine will sometimes name a page purely because that URL was baked into its training data, without crawling anything in the moment. You can't edit that data, but the pages that land in it tend to be the ones the web has referenced and linked to over years, so steady mentions and links are what earn your place there.
5. Trust Signals (E-E-A-T for Machines)
What the model already knows about you shapes how much it trusts you. A source it recognises as an authority on the topic gets the benefit of the doubt over a nameless site, before either page is even read.
5.1 Author and Expertise
Real author pages, bios with credentials, and Person
schema with sameAs links.
5.2 Experience (First-Hand)
First-person markers like "I tested" or "I measured this myself", plus original data or screenshots, signal genuine first-hand experience. That is what makes a page non-commodity content: something an engine can't find, near-identical, on a hundred other sites.
5.3 Transparency
Clear /about, /contact, an address and a privacy policy: the
basics of a trustworthy site.
5.4 Notability
Presence on third-party authoritative surfaces, a Knowledge Panel, a Wikidata entry.
6. Hygiene, Not Hacks
This block is technical hygiene: worth getting right as good practice, but none of it is a shortcut to more AI citations.
6.1 Structured Data
In a large controlled test, Ahrefs added schema markup to 1,885 pages and measured no lift in AI citations; on AI Overviews it even found a small but statistically significant dip of 4.6%. Worth knowing, but I wouldn't treat one study as the final word: it may have its own blind spots, and this field moves so fast that "no effect six months ago" doesn't mean "no effect today," or that anyone is measuring it the right way yet.
So my take is pragmatic. Valid Article, Organization and Person markup is cheap to add,
and if there's even a small chance it helps an engine parse your page, why not have it? Add
it as hygiene; just don't count on it as a lever. HowTo markup no longer earns a rich result
but stays valid, and a clean HTML step structure does nearly the same job either way.
6.2 Core Web Vitals and Technical Basics
A stable layout, a valid sitemap, semantic HTML and reasonable Core Web Vitals are table stakes; they help every crawler, AI included.
6.3 Content Signals in robots.txt
The emerging ai-train /
search / ai-input content signals let you express how your content may be used.
New, low-cost to add, worth doing.
6.4 llms.txt
A plain-text file at your domain root (/llms.txt) that hands AI tools a
tidy markdown map of your key pages. Google says it doesn't use it, and independent tests
have found no measurable effect on citations, so don't expect it to move the needle. But
it takes minutes to write, does no harm, and AI search reaches well beyond Google, so a
low-effort file that some engines might read is worth having.
7. Agent-Ready Protocol (Emerging)
This block comes from Suganthan Mohanadasan's work on making a site "agent-ready." It matters most if you run a SaaS product or anything AI agents might operate directly, and it's still largely unproven for citations, so treat it as forward-looking hygiene rather than a priority for a content site.
7.1 Discoverable Structure for Agents
Clear AI-bot rules in robots.txt, a valid
sitemap with lastmod, and Link headers whose rel values (api-catalog,
service-desc, service-doc, describedby) surface your key resources to agents that
read headers before HTML.
7.2 Markdown Negotiation
When a client sends Accept: text/markdown, return a
markdown version of the page; it gives agents a cleaner payload and, by Suganthan's
measure, cuts its size by roughly 5x.
7.3 Tool and Agent Manifests
If your site exposes tools, an MCP Server Card
(/.well-known/mcp/server-card.json, per the SEP-1649 proposal) lets agents discover and
call them; WebMCP (navigator.modelContext, a W3C community draft now previewing in
Chrome) does the same from inside the browser, and an A2A Agent Card
(/.well-known/agent.json) is for sites that are themselves agents.
7.4 A Single Discovery Index
The Agentic Resource Discovery standard
(/.well-known/ai-catalog.json, shipped by Google and ten other companies in June 2026)
advertises all your MCP servers, A2A agents and APIs in one place, so an agent finds them
in a single fetch. Adoption is close to zero so far, so this is a bet on where things are
heading.
7.5 Auth for Agent-Callable APIs
If agents hit authenticated endpoints, publish OAuth
Protected Resource metadata (/.well-known/oauth-protected-resource.json, RFC 9728) and
OIDC discovery (/.well-known/openid-configuration) so they can authenticate cleanly.
7.6 Agent Skills
Anthropic's Agent Skills spec, still under active development, lists
the discrete tasks an agent can carry out on your site. The convention points at a
/.well-known/skills.json file, though the path isn't settled yet.
7.7 Open Knowledge Format
OKF, a Google Cloud spec published in June 2026 at v0.1, packages knowledge as a plain directory of markdown files with YAML frontmatter, so an agent can take a curated copy of your content instead of scraping it. Very early, worth watching.
7.8 Verifiable Bot Access (Web Bot Auth)
Signed, verifiable bot identities (HTTP Message Signatures, RFC 9421) instead of blunt user-agent allow-lists, so you can let genuine AI crawlers in without opening the door to spoofers. Whatever your edge provider's defaults are, confirm that real crawlers still get through.
7.9 Commerce Protocols
For stores, a stack of agent-commerce standards is forming, and they sit on different layers rather than competing: UCP (Google) and ACP (Stripe and OpenAI) cover how an agent browses a catalogue and checks out, while x402 (Coinbase) and MPP (Stripe) handle machine-to-machine settlement. Maturity varies: ACP already runs inside ChatGPT's checkout, the others are earlier.
What My Own Tests Showed
To pressure-test the list I ran two experiments: 30 keywords across six verticals (about 1,500 pages), and a deeper single-vertical run of 64 queries (about 3,000 pages). Both used US desktop results, comparing Google's top 40 and everything Bing returned (about 20 results) for each query against what ChatGPT, Claude, Gemini, Perplexity and Google's AI Overview actually cited. I measured only a handful of the factors above, not all of them, so read this as a probe rather than a full audit.
The clearest finding was that there is no single formula: which factors decide a citation depends on both the topic and the type of question.
- The topic changes which factors matter. In health queries like "statins side effects" or "is creatine safe," where bad advice is costly, the assistants leaned hard on authority. Among pages that rank in Google, the cited ones had 60 times more links pointing at the page itself than uncited ones, and a higher median Domain Rating (90.5 against 81). In affiliate queries like "best vpn" or "best credit cards" the same two signals barely separated the groups: a 3-fold link gap, and a median DR that hardly moves (88 against 84.5).
- Query types change which factors matter as well. When a query asks for numbers ("average data scientist salary," "email open rate benchmarks"), cited pages were far more fact-dense than uncited ones: around 15.6 concrete facts per 1,000 words against 8.7. When it asks for steps ("how to boil an egg," "how to tie a tie"), that difference disappeared completely, 4.5 against 4.2. The signal that decides one kind of query is irrelevant to another.
-
Page-level backlinks matter, but how much they matter changes completely by topic. How sharply they separate cited from uncited pages depends on how mixed the authority is in that topic. Counting only pages that rank in Google, the median cited page had 119.5 referring domains against 2 in health, a 60-fold gap; 15 against 1 in SEO; 5 against 1 in SaaS; and 117.5 against 40.5 in affiliate, where the gap narrows to 3-fold. The effect is huge where major institutions sit in the same results as small sites, and it shrinks where every competitor looks alike. Treat it as a range by topic, not one number.
One caveat on how to read that: this run didn't measure brand mentions at all, and page-level backlinks were the nearest proxy I had for how well known something is. So part of what looks like a link effect may really be brand familiarity.
-
Ranking in Google's top 10 raised citation odds by roughly 1.6 to 5.4 times: 50 to 87% of top-10 pages picked up at least one citation, against 16 to 36% at positions 11 to 40. But rank alone was never enough. Depending on the vertical, 13 to 50% of top-10 pages got no citations at all. One page in the SaaS batch sat at #5 on a strong domain (DR 78) with just three links pointing at the page itself, and not one assistant cited it.
- Ranking in Bing adds something the Google position doesn't. Holding the Google position fixed, pages that also ranked in Bing's top 10 were cited far more often: 39 to 67% against 24 to 27% for pages sitting at Google 11 to 40, and 88 to 95% against 76% inside the top 10. Because the Google position is constant across both groups, Bing cannot simply be echoing it: a page carried by both indexes has independent confirmation that it genuinely answers the query. The gap is widest below the top 10, which is exactly where the Google position tells you least.
- Links get a page into the top 10; they decide much less once it is there. In the band from position 11 to 40, cited pages had about 4 times more links to the page than uncited ones. Inside the top 10 that gap shrank to 1.8 times. The other verticals show the same thing more sharply: almost all of health's 60-fold gap comes from positions 11 to 40, where it reaches 158-fold, while inside the top 10 it disappears and even reverses, a median of 65 referring domains for cited pages against 85 for uncited ones. Affiliate behaves the same way, 273 against 364. Links do the work of pulling a page up into the top results; once it is there the field is level, and simply being in the top 10 is what carries the citation.
-
Where an engine looks for sources tells you more than how picky it is. Of everything ChatGPT cited, 68% came from outside Google's top 40. For Gemini it was 55%, for Claude 46%, and for Perplexity 34%. So Perplexity stays closest to the organic results and ChatGPT roams furthest from them.
Among the pages they did take from the search results, ChatGPT asked the most of them: a median of 40.5 referring domains, against 9 for Claude, 7 for Perplexity and 4 for Gemini. Those two groups have to be counted separately. Lump them together and the engine that roams furthest from the rankings looks like the one with the lowest standards, when among ranked pages it is the strictest of the four.
-
Close to half of all the citations — about 47% — landed on pages that showed up in neither Google's top 40 nor the Bing results I pulled. An individual top-10 page still has the best odds of being cited, but there are so many more pages below it that most citations, in aggregate, come from outside the ranked results.
- Pages in the AI Overview tend to get cited too, but the AI Overview is often missing. Cited pages showed up in an AI Overview more often than uncited ones (32% against 21% in the deeper run), which suggests AI Overviews and the assistants draw on a broadly similar pool. But AI Overviews get rarer the less factual the query: they appeared for seven of eight "what is X" keywords, three of eight "is X worth it," and just one of eight "when or why to use X." The rarer the AI Overview, the more the assistants choose sources on their own.
- Domain Rating on its own never decided who got cited. It separated cited from uncited pages in some verticals and barely at all in others, which makes it a weak secondary signal rather than something to chase.
- The format that got cited followed the question, in every vertical without exception. A comparison query ("asana vs trello") returned comparison pages; "what is a crm" returned definitions and explainers; "best project management software" returned listicles; and "notion pricing" put the vendor's own pricing page first, with third-party price breakdowns around it.
- Freshness is a threshold, not a lever. Even in fast-moving topics like SEO, where you would expect recency to count most, the gap was modest: cited pages were fresher on four queries out of five, and on "is link building still worth it" the cited ones were typically 6 months old against 11 for uncited ones. Being current enough not to look stale is what matters; chasing a few weeks of recency is not.
A Practical Action Plan
If you want an audit checklist to start from, work top-down by impact:
- Fix accessibility first: confirm AI crawlers aren't blocked, and that content display permissions and content visibility aren't hiding your best content.
- Win the classic ranking: organic ranking, fan-out coverage and topic-cluster coverage for your priority queries.
- Make each page citable: a direct answer near the top, clear structure, concrete facts and self-contained passages that link to credible sources.
- Build brand signals: unlinked brand mentions, YouTube coverage and third-party corroboration.
- Keep the hygiene clean: structured data, Core Web Vitals, a real refresh cadence.
None of these factors is a trick, and none is a guarantee, only better odds. It's SEO strategy pointed at a new surface: the overlap of SEO and AI, not a separate AI SEO ritual. Earning citations is the natural result of publishing genuinely helpful content.
Want to see where a page stands today? Run any URL, or audit your existing content, through the free AI Readiness & Citation Checker: it scores a page against most of these factors and shows whether it is already appearing in AI answers. That's the honest starting point for optimizing for AI, the core of any AI search optimization plan.
The Sources, and How I Built This
The sources. Each study covers a different slice of "AI search," so I combined them rather than trusting any single one:
- Zyppy Signal, by Cyrus Shepard. A meta-analysis of 54 studies that scores 23 citation factors by how repeatable and well-evidenced each one is. It draws on evidence across ChatGPT, Gemini and Perplexity, plus Google's AI Overviews, which makes it the backbone of this list.
- Google's AI optimization guide. The official word on Google's generative features (AI Overviews and AI Mode), and the best source for what Google explicitly confirms or denies.
- Ahrefs' studies. Large-scale correlation work: 863K SERPs on how top-10 rankings feed AI Overview citations, 17M citations on freshness, 75K brands on brand signals, and a controlled test of schema. Mostly Google AI Overviews, plus brand visibility across AI.
- Princeton's GEO paper. The academic "Generative Engine Optimization" study (KDD 2024), which measured visibility lifts of up to 40% from its methods, with citing sources, quotations and statistics among the best. Generative engines in general.
- AirOps and Seer Interactive. ChatGPT-specific: how retrieval fan-out and overlap with Bing drive what ChatGPT actually cites.
- Suganthan Mohanadasan, Dan Petrovic and Metehan Yesilyurt. The agent and protocol layer, Gemini's retrieval cap, and the reciprocal-rank-fusion view of fan-out.
How I built and calibrated it. I wrote a script (with Claude Code) that audits any URL against most of the factors above. It pulls a page's actual citations from Google's AI Overview, ChatGPT, Perplexity, Gemini and Claude through their APIs, and compares them to the score. Where score and reality diverged, I re-weighted the factors so the score tracks reality more closely: page-level backlinks turned out to matter far more than domain-level authority, for instance, and title-to-query overlap deserved more weight than I first gave it.
How to read this list. Be honest about the sample: the runs behind it cover 30 keywords across six verticals (about 1,500 pages) and one deeper single-vertical run of 64 queries (about 3,000 pages), and they measured only a subset of the factors, so treat the conclusions as directional rather than statistically settled. The script has known blind spots too:
- It checks one URL at a time, so site-wide signals (wasted crawl budget, topical depth, internal linking, content decay) aren't measured.
- It can't see your server logs, Search Console or Bing Webmaster, so it infers crawlability from robots.txt instead of confirming real bot visits.
- Several checks are proxies, not the full factor: brand and entity recall, "known source," and hedging detection are approximations, and some are tuned for English.
- Citations move around. AI Overviews and the assistants can each return different sources for the same query from one run to the next, so a single snapshot can mislead.
And the caveat that matters most: none of these factors guarantees a citation. They are correlations and levers, not switches. Implementing them raises your odds; nothing forces an AI engine to quote you.
Two more things a citation itself won't tell you:
- A citation is not a click. An AI answer can quote you and still send no traffic, because the reader already got what they needed inside the answer.
- A citation carries a context this study doesn't judge. The same brand can be the recommendation for one query and the cautionary counter-example for another, depending on how the question is framed.
I measure whether a page is cited, not how it's cited or whether the citation earns a visit.
Why read the whole list? Because different AI systems weigh these factors differently, and research suggests their importance shifts by vertical: what moves the needle for a SaaS page isn't what moves it for a health article. It's worth scanning the whole list and implementing whatever fits your site, even the lower-priority ones.
