TL;DR
  • We split each topic’s questions into two sets, measured citation coverage on one, and compared it with citations on the held-out set.
  • Coverage correlated more strongly with held-out citations than the stored authority score on four of five engines: AI Overviews, AI Mode, Perplexity and ChatGPT. On Gemini, both came in at 0.21.
  • Citation reach on other topics had the strongest correlation of the three measures on every engine.
  • Among sites with similar reach, the coverage correlation ranged from 0.208 on AI Overviews down to 0.043 on ChatGPT.
  • These are associations inside one six-month window. They show where citations cluster, not what happens when you publish more pages.

Websites that get cited across more questions in a topic also tend to get cited on a separate, held-out set of questions in that same topic. We saw that on all five engines Wellows tracks, across 9,471 questions and 151 topics collected between January and June 2026.

That’s a useful way into the topical authority debate, because it swaps a fuzzy question (does this site own the topic?) for one the data can actually answer: does citation coverage on one set of questions tell you anything about citations on another? It does. How much it tells you depends heavily on the engine, and ChatGPT is the outlier.

The strongest signal wasn’t coverage inside the topic, though. It was how widely a site gets cited on other topics, which changes how the coverage numbers should be read. All of this is first-half 2026 data, collected before the July shifts we covered in our domain-authority study.

Citation coverage isn’t the same as content coverage

Citation coverage is how often a site was cited across a topic’s measurement questions. It’s something the engines do, not something you publish. A site can have deep content on a topic and almost no citation coverage, or pick up citations on questions it never wrote a dedicated page for.

That distinction decides what this study can answer. It tells you whether citation patterns carry over to new questions in the same topic, and it leaves the question of what caused them (more pages, tighter clusters, better internal links) for a different kind of test.

We also measured outside-topic reach: how often a site was cited on the other topics in the dataset. Think of it as citation breadth within the questions we track, rather than backlinks or brand awareness.

How we ran the study

The data comes from Wellows’ citation dataset for January to June 2026, covering ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode. It holds 9,471 distinct English-language questions and 382,176 answer observations. Of those answers, 86% were collected in the United States and the rest across 23 other markets. Questions were weighted equally despite repeated collection. We kept questions where all five engines returned at least one cited website.

Topics are project-assigned labels, pooled where projects used the same label. We kept 151 topics with at least ten distinct questions in each half.

If you measure coverage and outcomes on the same answers, you’re grading a site on a test it has already taken. So within each topic, we assigned every question to one half or the other using a hash of the question text, and repeated answers to the same question stayed on the same side. Coverage came from the first half. The outcome came from the second. Held out means excluded from the coverage calculation, nothing more. Both halves come from the same January to June window.

Diagram of the held-out split method, with each topic's questions divided into a measurement half and a held-out half

Each topic’s questions were split in two, so coverage and citations were measured on different questions.

  • Coverage: the average share of measurement-half answers citing a website within a topic.
  • Reach: citation frequency on questions outside the focal topic. A site counted as niche below 0.05% and widely cited at 0.5% or more of other questions.
  • Authority: a stored 0–100 domain-level score, available for about 82% of website–topic pairs.
  • Statistics: within-topic rank correlations, plus a shuffle of held-out outcomes within topic and reach groups as a baseline.

We also ran a leave-one-engine-out version, where each engine’s coverage was calculated from the other four engines only, and the positive pattern held. That removes the most obvious self-reference, though engines can still share source preferences.

The unit is a website–topic pair: 198,360 of them across 92,112 websites, each cited at least once in the measurement half. So this is a study of sites the engines already cite.

Coverage beat the authority score on four of five engines

We compared three measures against held-out citations: citation coverage, the stored authority score, and outside-topic reach. Here are the first two.

Chart comparing how citation coverage and the stored authority score correlate with held-out AI citations on five AI engines

Citation coverage tracked held-out citations more closely than the authority score on four of five AI engines.

Within-topic rank correlations with held-out citation outcomes
Engine Citation coverage Stored authority score
Google AI Overviews 0.30 0.21
Google AI Mode 0.29 0.21
Perplexity 0.28 0.15
Gemini 0.21 0.21
ChatGPT 0.12 0.07

Coverage came out clearly ahead on AI Overviews, AI Mode and Perplexity. ChatGPT was weak on both measures, and Gemini was a tie at two decimal places.

A higher correlation means two numbers moved together more closely in this sample. It isn’t a ranking factor, and it doesn’t mean pushing one moves the other. The authority comparison covers 163,156 website–topic pairs, about 82% of the eligible sample, because not every site had a stored score.

The strongest signal was reach outside the topic

Sites cited widely across other topics also picked up more citations on the held-out questions. Outside-topic reach had the highest correlation of the three measures on every engine.

That matters for how you read coverage. A site’s citations inside one topic may be riding on a broader tendency to get picked as a source, so treating the within-topic number as a pure measure of expertise would miss part of the story.

It doesn’t follow that broad publishing, PR or link building would reproduce the effect. We didn’t test any of those, and three separate correlations don’t tell you which combination would predict best in a model.

ChatGPT had the weakest coverage signal among similar sites

To compare like with like, we grouped sites within each topic by their outside-topic reach and recalculated the coverage correlation. Then we shuffled the held-out outcomes within those same groups to get a baseline.

Chart of observed coverage correlations against a shuffled baseline among sites with similar outside-topic reach, by AI engine

Among sites with similar reach, coverage stayed well above the shuffled baseline everywhere except ChatGPT.

Coverage correlations within topic and reach groups, against a shuffled baseline
Engine Observed correlation Shuffled baseline
Google AI Overviews 0.208 0.009
Google AI Mode 0.197 0.008
Perplexity 0.185 0.008
Gemini 0.119 0.010
ChatGPT 0.043 0.012

Every shuffled value landed around 0.01. The observed values sat well above that on AI Overviews, AI Mode and Perplexity, lower on Gemini, and closest to the baseline on ChatGPT at 0.043.

The data shows where ChatGPT differs, not why. A preference for sources it has already picked is one candidate. Differences in question mix, retrieval and candidate selection are others, and separating them is the next piece of work.

What we’d do with this in a visibility program

If you’re tracking AI citations, the practical move is to keep the measurements apart before you draw a content conclusion from them.

  1. Track citations by topic and engine. A blended score hides the fact that ChatGPT, Perplexity and Google’s AI surfaces behave differently here.
  2. Keep citation coverage separate from published coverage. Log which questions you have useful content for, independent of whether an engine cites it. The gap between the two lists is where to dig.
  3. Monitor a stable question set. Same prompts, locations and collection settings, so a change in how you measure doesn’t show up as a gain or a loss.
  4. Test content changes over time. Log each change, then compare later citations against a set of pages you didn’t touch.

None of this is a reason to stop building topic coverage for a ChatGPT audience. It says ChatGPT’s citations tracked coverage less closely in this sample, which is a different claim from saying depth doesn’t work there.

Wellows reports citations across ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, and keeping the engine-level view is how you tell a broad shift from one that’s confined to a single platform.

Where this fits with earlier research

Two earlier studies set the baseline for this question. Floyi’s Topical Authority Report 2026 looked at ranked coverage in Google search results, and Kevin Indig’s analysis of Semrush data looked at category ownership in ChatGPT.

Wellows’ study adds a different lens: citation coverage measured directly in AI answers, tested on held-out questions, across five engines including Perplexity.

For the named-source side of the same question, our companion study of LLM topic associations looks at which sources models name by topic. It uses different data and a different outcome, so read the two side by side rather than as one result.

The question worth chasing next is what actually moves citation coverage. Answering it takes an intervention, a later collection window and a set of untouched pages to compare against.

Frequently Asked Questions


Sites cited across more of a topic’s measurement questions were also cited more on held-out questions in that topic, on all five engines. That’s evidence about citation patterns in January to June 2026, which is narrower than proof that publishing more pages earns more citations.

On four of five engines, yes: AI Overviews, AI Mode, Perplexity and ChatGPT. On Gemini both measures came in at 0.21.

Outside-topic citation reach, on every engine. Reach here means citations on other topics in our dataset, not visibility across the whole web.

No. We didn’t measure content depth or run a publishing test. ChatGPT had the smallest coverage correlation among sites with similar reach, which tells you where it differs, not what would happen after you improve your content.

No. The split was based on question text, and both halves come from the January to June 2026 window.

It describes the first half of 2026. Our domain-authority study picked up shifts in cited-source profiles in July, so a later collection window is what would confirm whether these relationships held.