Does Topical Authority Show Up in AI Citations? What 9,471 Questions Reveal

Khadija ZamanAI Search ManagerI'm Khadija Zaman, AI Search Manager at Wellows, where I lead generative and answer engine optimization (GEO/AEO) — building the automated workflows that track brand citations…Read Full Bio- We split each topic’s questions into two sets, measured citation coverage on one, and compared it with citations on the held-out set.
- Coverage correlated more closely with held-out citations than the stored authority score on four of five engines: AI Overviews, AI Mode, Perplexity and ChatGPT. On Gemini the two were level, at 0.20 and 0.21.
- Citation reach on other topics had the strongest correlation of the three measures on every engine, from 0.29 to 0.37.
- Among sites with similar reach, the coverage correlation ranged from 0.208 on AI Overviews down to 0.043 on ChatGPT.
- These are associations inside one six-month window. They show where citations cluster, not what happens when you publish more pages.
Websites that get cited across more questions in a topic also tend to get cited on a separate, held-out set of questions in that same topic. We saw that on all five engines Wellows tracks, across 9,471 questions and 151 topics collected between January and June 2026.
That’s a useful way into the topical authority debate, because it swaps a fuzzy question (does this site own the topic?) for one the data can actually answer: does citation coverage on one set of questions tell you anything about citations on another? It does. How much it tells you depends heavily on the engine, and ChatGPT had the lowest coverage correlation.
The highest correlations came from how often a site was cited on other topics, which changes how the coverage numbers should be read. All of this is first-half 2026 data, collected before the July shifts we covered in our domain-authority study.
Citation coverage isn’t the same as content coverage
Citation coverage is how often a site was cited across a topic’s measurement questions. It’s something the engines do, not something you publish. A site can have deep content on a topic and almost no citation coverage, or pick up citations on questions it never wrote a dedicated page for.
That distinction decides what this study can answer. It tells you whether citation patterns carry over to new questions in the same topic, and it leaves the question of what caused them (more pages, tighter clusters, better internal links) for a different kind of test.
We also measured outside-topic reach: how often a site was cited on the other topics in the dataset. Think of it as citation breadth within the questions we track, rather than backlinks or brand awareness.
How we ran the study
The data comes from Wellows’ citation dataset for January to June 2026, covering ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode. It holds 9,471 distinct English-language questions and 382,176 answer observations. Of those answers, 86% were collected in the United States and the rest across 23 other markets. Questions were weighted equally despite repeated collection. We kept questions where all five engines returned at least one cited website.
Topics are project-assigned labels, pooled where projects used the same label. We kept 151 topics with at least ten distinct questions in each half.
If you measure coverage and outcomes on the same answers, you’re grading a site on a test it has already taken. So within each topic, we assigned every question to one half or the other using a hash of the question text, and repeated answers to the same question stayed on the same side. Coverage came from the first half. The outcome came from the second. Held out means excluded from the coverage calculation, nothing more. Both halves come from the same January to June window.

Each topic’s questions were split in two, so coverage and citations were measured on different questions.
- Coverage: the average share of measurement-half answers citing a website within a topic. For each engine, coverage comes from the other four engines’ answers, so no engine is scored on its own citations.
- Outcome: the share of held-out answers on each engine that cite the website.
- Reach: citation frequency on questions outside the focal topic. Sites fall into three reach groups: niche below 0.05% of other questions, mid-reach between 0.05% and 0.5%, and widely cited at 0.5% or more.
- Authority: the 0–100 domain authority score Wellows stores with each citation, taken as each website’s median across its January to June 2026 observations. 163,156 website–topic pairs, about 82% of the sample, carry a score.
- Statistics: Spearman rank correlations calculated within each topic, with tied values given their average rank, then averaged across the 151 topics. A shuffle of held-out outcomes within topic and reach groups sets the chance baseline.
Scoring each engine on the other four engines’ coverage removes the most obvious self-reference: an engine never gets credit for predicting its own citations. Engines can still share source preferences.
The unit is a website–topic pair: 198,360 of them across 92,112 websites, each cited at least once in the measurement half. So this is a study of sites the engines already cite.
Coverage correlated more closely than authority on four of five engines
We compared three measures against held-out citations: citation coverage, the stored authority score, and outside-topic reach. Here are the first two.

Citation coverage tracked held-out citations more closely than the authority score on four of five AI engines, with Gemini level.
| Engine | Citation coverage | Stored authority score |
|---|---|---|
| Google AI Overviews | 0.29 | 0.22 |
| Google AI Mode | 0.28 | 0.22 |
| Perplexity | 0.25 | 0.15 |
| Gemini | 0.20 | 0.21 |
| ChatGPT | 0.13 | 0.10 |
Coverage correlated more closely on AI Overviews, AI Mode, Perplexity and ChatGPT. The widest gap was on Perplexity, at 0.25 against 0.15. ChatGPT had the lowest values for both measures, and on Gemini the two sat level.
A higher correlation means two measured quantities moved together more closely in this sample. It does not establish a ranking factor or a causal effect. The authority comparison covers the 163,156 website–topic pairs that carry a stored score.
Outside-topic reach had the strongest correlation
Sites cited more often across other topics also tended to receive citations on the held-out questions. Outside-topic reach had the highest correlation of the three measures on every engine.
| Engine | Outside-topic reach | Citation coverage | Stored authority score |
|---|---|---|---|
| Google AI Overviews | 0.33 | 0.29 | 0.22 |
| Google AI Mode | 0.35 | 0.28 | 0.22 |
| Perplexity | 0.29 | 0.25 | 0.15 |
| Gemini | 0.37 | 0.20 | 0.21 |
| ChatGPT | 0.29 | 0.13 | 0.10 |
The gap was widest on ChatGPT and Gemini, where reach ran well ahead of coverage: 0.29 against 0.13 on ChatGPT, and 0.37 against 0.20 on Gemini.
That matters for how you read coverage. A site’s citations inside one topic may be riding on a broader tendency to get picked as a source, so treating the within-topic number as a pure measure of expertise would miss part of the story.
It doesn’t follow that broad publishing, PR or link building would reproduce the effect. We didn’t test any of those, and three separate correlations don’t tell you which combination would predict best in a model.
ChatGPT had the weakest coverage signal among similar sites
To compare like with like, we grouped sites within each topic by their outside-topic reach and recalculated the coverage correlation. We then shuffled the held-out outcomes once within those groups to set a chance baseline.

The observed correlation exceeded the shuffled value on all five engines. ChatGPT had the smallest absolute gap.
| Engine | Observed correlation | Shuffled baseline |
|---|---|---|
| Google AI Overviews | 0.208 | 0.009 |
| Google AI Mode | 0.197 | 0.008 |
| Perplexity | 0.185 | 0.008 |
| Gemini | 0.119 | 0.010 |
| ChatGPT | 0.043 | 0.012 |
The shuffled values ranged from 0.008 to 0.012. Every observed value was higher, with the smallest absolute gap on ChatGPT: 0.043 compared with 0.012.
The data shows where ChatGPT differs, not why. A preference for sources it has already picked is one candidate. Differences in question mix, retrieval and candidate selection are others, and separating them is the next piece of work.
What we’d do with this in a visibility program
If you’re tracking AI citations, the practical move is to keep the measurements apart before you draw a content conclusion from them.
- Track citations by topic and engine. A blended score hides the fact that ChatGPT, Perplexity and Google’s AI surfaces behave differently here.
- Keep citation coverage separate from published coverage. Log which questions you have useful content for, independent of whether an engine cites it. The gap between the two lists is where to dig.
- Monitor a stable question set. Same prompts, locations and collection settings, so a change in how you measure doesn’t show up as a gain or a loss.
- Test content changes over time. Log each change, then compare later citations against a set of pages you didn’t touch.
None of this is a reason to stop building topic coverage for a ChatGPT audience. It says ChatGPT’s citations tracked coverage less closely in this sample, which is a different claim from saying depth doesn’t work there.
Wellows reports citations across ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, and keeping the engine-level view is how you tell a broad shift from one that’s confined to a single platform.
Where this fits with earlier research
Two earlier studies examined related questions. Floyi’s Topical Authority Report 2026 looked at ranked coverage in Google search results, and Kevin Indig’s analysis of Semrush data looked at category ownership in ChatGPT. Their measures differ from citation coverage, so the results are not directly interchangeable.
Wellows’ study adds a different lens: citation coverage measured directly in AI answers, tested on held-out questions, across five engines including Perplexity.
For the named-source side of the same question, our companion study of LLM topic associations looks at which sources models name by topic. It uses different data and a different outcome, so read the two side by side rather than as one result.
The question worth chasing next is what actually moves citation coverage. Answering it takes an intervention, a later collection window and a set of untouched pages to compare against.
Scope
This is an observational comparison within January to June 2026. The held-out split keeps coverage and outcomes on separate questions, and both halves come from the same window.
The sample covers questions where all five engines cited at least one website, and websites already cited in the measurement half. Pooled project labels can combine questions with different intents. The authority comparison uses the website–topic pairs that carry a stored score.
Our domain-authority study picked up shifts in cited-source profiles in July 2026, so a later window is the next test of whether these relationships hold.