TL;DR
  • What we looked at: 22.7 million citations, the sources AI engines point to when they answer a question, across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode, from January to June 2026. All five engines cited on 531,889 of the same questions, and every all-five comparison below runs on those.
  • ChatGPT is the odd one out. On the same question, Perplexity never touches 89.1% of the websites ChatGPT cites.
  • Ask all five the same question and they mostly point at different things. 79.6% of the sources they cite show up on one engine and nowhere else. Only 0.31% show up on all five.
  • But there’s good news. The engines rarely agree on which page to link to (6.8% of the time), yet they agree far more often on which companies to name (30.3%). Brand travels about 4.5x better than pages do.
  • Where agreement does happen, better content isn’t what explains it. It tracks engines sharing the same underlying plumbing, the one thing you can’t change.
  • Bottom line: there is no single “AI visibility”. There are five.

Picture the dashboard your team checks every Monday. One big, confident number: your AI visibility score. It went up two points this week, so someone adds a green arrow to the slide and everyone moves on.

That number is misleading you. Not because the maths is wrong, but because it averages five things that have almost nothing in common.

Definition

First, what a citation actually is

When you ask ChatGPT, Gemini, or Perplexity a question, the answer usually comes with sources attached: links to web pages, or the names of companies mentioned in the text. Those are citations. They’re how a brand gets noticed in AI search, the same way a blue link on Google used to work. This study counted 22.7 million of them.

We pulled six months of these citations from Wellows’ own data: 22.7 million citations, 1.15 million questions, and 441,946 different websites, January through June 2026.

Not every engine answers every question. All five cited at least one website on 531,889 of those questions, covering 13.0 million of the citations and 280,245 websites. Every all-five comparison in this study runs on that set, so no engine is ever judged on a question it sat out. Comparisons between just two engines use every question that pair answered, and we flag the base wherever it changes.

Then we asked a simple thing of it. When five AI engines answer the same question, do they point to the same sources or different ones?

They point to different ones. Overwhelmingly. And the gap is widest exactly where most people assume it’s narrowest.

We compared what each engine cited, question by question, at three levels of detail: the exact web page, the website it sits on, and the brand named in the answer. The full method sits at the end, along with the caveats and the tests we ran to try to break our own finding.

The main finding: the five engines mostly don’t agree

Start with the simplest possible question. Take everything all five engines cited, and count how many engines cited each source. If AI search were one thing, most sources would appear on most engines.

They don’t.

79.6%

of sources are cited by only one of the five AI engines

Those 531,889 questions produce 8,729,964 website-and-question pairings. Exactly one engine cited 6,949,766 of them, and no other engine touched them. Only 0.31%, 27,097 pairings, made it onto all five.

Source: Wellows internal citation dataset, Jan–Jun 2026

Two different measures run through this post, and it is worth separating them now. This 79.6% counts how many of the five engines cited each source. The agreement percentages further down count how much any two engines’ source lists overlap on the same question. Different measures, same direction.

Two-panel chart: 79.6% of AI citations appear on one engine only, plus engine-pair co-citation ranking across ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode


Top: nearly 80% of all cited sources show up on a single engine, and full agreement across all five is a rounding error at 0.31%. Bottom: every engine pair ranked by how often they cite the same sources, with AI Overviews and AI Mode leading at 24.5% and every ChatGPT pairing sitting at the bottom.

A number that big invites a fair challenge: maybe it’s a quirk of one odd month. So we checked each month separately.

The share that only one engine cited held between 77.7% and 81.1% in every single month, a spread of just 3.4 points. This isn’t one strange batch of questions. It’s how the system behaves.

Now look at the lower half of that chart, where the pattern gets specific. Rank all ten possible engine pairings by how often they cite the same sources and the order isn’t random at all:

  • AI Overviews and AI Mode agree most, at 24.5%
  • Gemini and AI Overviews follow at 12.0%, then Gemini and AI Mode at 10.9%
  • Every pairing of two Google engines beats every pairing across different companies
  • Every pairing involving ChatGPT lands in the bottom four, ending with ChatGPT and Gemini at just 6.1%

That ordering is the whole study in miniature, and we unpack why it exists further down.

The practical consequence is immediate. When one of the many AI visibility tools hands you a single blended score, it averages five surfaces that barely overlap. An average of five things that share almost nothing describes none of them.

ChatGPT is the odd one out

One engine stood apart at every stage of this analysis, so it’s worth taking on its own. Flip the question around: of everything each engine cites, how much does no other engine touch for the same question?

Bar chart of per-engine unique citations: ChatGPT at 76.3% cites the most sources no other AI engine touches, Google surfaces lowest near 50%


No other engine touches 76.3% of the sources ChatGPT cites, the highest share of any engine. Google’s two surfaces are the least unique, at around 50%.

ChatGPT leads at 76.3%. More than three-quarters of what it cites, none of the other four engines used for the same question.

Perplexity follows at 69.9%, Gemini at 63.7%, and Google’s two surfaces sit near 50% (AI Overviews 50.4%, AI Mode 49.7%), low precisely because they lean on each other.

Compare ChatGPT against one rival at a time and it gets starker. Of every website ChatGPT cites:

  • Perplexity never touches 89.1% of them on the same question
  • Google AI Overviews never touches 89.2%
  • Google AI Mode never touches 90.2%
  • Gemini never touches 91.5%

No rival engine recovers even one in eight of ChatGPT’s sources.

If you want to see that split on your own website, our Perplexity visibility check shows which of your pages Perplexity cites, so you can set the result against your ChatGPT one.

The takeaway is blunt. Getting cited by ChatGPT is its own separate job. Nothing you learn from doing well in Google’s AI Overviews reliably carries over, which makes it the least mapped surface of the five, at least as far as this data can see.

The good news: brands travel, pages don’t

So far this is a gloomy picture. This next finding is the one that turns it into a plan.

Agreement isn’t a single number. It depends on how specific you’re being. We measured all three levels the same way, on one identical set of questions, and a gap opened up that reframes the whole problem.

Bar chart comparing AI engine agreement at page, website and brand level: 6.8% of exact pages, 10.9% of websites, 30.3% of brands named in the answer


On the same set of questions, measured the same way, agreement climbs from 6.8% of exact pages to 10.9% of websites to 30.3% of brands, roughly a 4.5x jump driven purely by how specific you get.

Read it left to right. The engines share only 6.8% of exact pages, 10.9% of websites, and 30.3% of the brands they name. That’s roughly a 4.5x jump from most specific to most general, on the very same questions.

That comparison narrows to the 236,096 questions where all five engines answered and at least one of them named a brand, so all three bars sit on exactly the same footing. That is why they differ slightly from the whole-dataset figures elsewhere in this post.

Here’s why that matters. “Are you visible in AI search?” was never one question. It’s two, and they have opposite answers:

  • “Do the engines link to my exact pages?” A scattered, engine-by-engine problem. Five separate games with almost no shared board.
  • “Do the engines mention my company?” Much more of a shared signal. Getting named by one engine is real evidence you can get named by the others.

Key finding: the engines agree on brands and split on pages. Brand agreement (30.3%) is about 4.5x page agreement (6.8% of exact pages). If your AI-search work chases specific-page citations, treat every engine as its own channel. If it builds brand presence, you have leverage across all five that the page numbers completely hide.

Why some engines agree: shared plumbing more than better content

If agreement is so rare, what goes with it when it does happen? Our set includes two Google surfaces, AI Overviews and AI Mode, which gives us a rare natural experiment.

Do engines built on the same underlying systems cite more alike than rivals do?

Bar chart showing Google's two AI engines agree on 24.5% of citations versus 7.3% for cross-company engines, a 3.4x shared-infrastructure premium


Google’s two engines agree on 24.5% of the websites they cite. Engines from different companies average 7.3% of websites. Engines that share infrastructure agree 3.4x more often.

Yes, clearly. Measured on the websites they cite, the data splits into three tiers:

  • Two Google engines. AI Overviews and AI Mode agree on 24.5% of websites.
  • The wider Google family. Pair Gemini, one step further from core Search, with either Search surface, and the average drops to 11.4% of websites.
  • Different companies entirely. Pairings across OpenAI, Google, and Perplexity average 7.3% of websites, with individual pairs running from 6.1% to 9.3%.

Engines that share the same plumbing agree 3.4x more often.

Sit with what that means.

The strongest pattern in the data isn’t better content. It isn’t a stronger website. It isn’t a better-optimised page.

It’s whether the two engines were built on the same infrastructure, and that’s the one variable nobody outside those companies can touch.

One honest limit: we can see the pattern, not the cause. Google’s two surfaces share indexes and ranking systems, and they also share a product team and an evaluation set. This data can’t separate those.

Either way, if Google can’t get its own two surfaces past 24.5%, no amount of page tweaking will make OpenAI and Perplexity agree.

Are the engines converging? Yes, agreement rose 31% in six months

Everything so far is a snapshot. Six months of data can also show direction, and the direction is that this split is slowly closing.

Line chart of monthly citation agreement across five AI engines rising from 8.58% in January to 11.25% in June 2026


Average agreement between any two engines rose from 8.58% in January to 11.25% in June 2026, a 31% increase, with a single dip in February.

Take any two engines at random and measure how often they cite the same sources. That number climbed from 8.6% in January to 11.3% in June, dipping once in February and then rising every single month after. That’s a 31% rise in half a year, or about 0.53 percentage points a month in a near-straight line.

Six monthly points is a short run, so read the slope as a direction, not a forecast.

But a completely separate measure points the same way. The share of sources cited by all five engines more than tripled over the same period, from 0.18% to 0.59%. Two different measures, one direction.


If this continues, AI engines will cite much more alike in 2027 than they do today. That cuts two ways. Sources everyone agrees on will pull ahead across all five at once, and today’s window, where a mid-sized site can own a niche one engine found and the others haven’t, is closing. Engine-specific gaps are a shrinking opportunity. Grab them while they’re here.

Your category matters more than the kind of question

Agreement isn’t spread evenly. So where does it pile up?

Most marketers would guess it depends on what the person is trying to do: research, shop around, buy. The data says that barely matters.

Research questions sit at 9.5% agreement, commercial questions at 9.6%, buying questions at 10.6%, and “find me this specific site” questions at 11.8%. A 2.3-point spread that changes no decision you’d ever make.

Sort by category instead and the spread is four times wider.

Horizontal bar chart of cross-engine citation agreement by category: settled categories run 10.1% to 15.5% while emerging AEO/GEO sits lowest at 6.3%


Settled commercial categories agree on 10.1% to 15.5% of cited websites. Emerging AEO/GEO sits at 6.3%, with every category pooled the same way for a like-for-like comparison.

To keep that comparison honest, we pooled every category the same way, adding together all the question topics that belong to it, so we’re never holding one narrow topic up against a broad group.

Settled commercial categories land between 10.1% and 15.5%, topped by data backup and recovery. These areas have a settled, trusted set of sources that every engine already leans on.

Now pool every question about AI visibility, AI SEO, LLM SEO, and generative engine optimisation into one emerging category. Call it AEO/GEO, short for answer engine optimisation and generative engine optimisation.

It sits at 6.3%, under two-thirds of the nearest settled category and about 40% of the highest one. There’s a huge pool of sources, no clear favourites yet, and the five engines each point to largely different writers.

That’s both the opportunity and the risk: easy to break in today, but a real chance today’s citation doesn’t survive as the category settles.

One caveat: these categories come from the question set Wellows tracks for customers, not a survey of every industry.


Category moves agreement by about 2.5 times. The kind of question moves it by two points. To judge how crowded the citation game is for you, how settled your category is matters far more than what your customers are trying to do.

Narrow specialists travel across engines. The biggest sites don’t.

Averages give you the shape of the field. A leaderboard shows you what actually travels.

So we built one. For every website, we counted how many of the five engines cited it on the same question. We withheld the names, since several are sites Wellows tracks for customers, but the categories and the numbers tell the story on their own.

Website Questions cited in (all five engines answered) In all 5 engines Avg engines / question
Product-protection warranty provider 9,343 1,801 2.94
Community discussion platform 272,866 1,420 1.66
Independent online bookstore 2,990 576 3.26
Website-builder platform 12,265 544 1.52
Online encyclopedia 20,097 415 1.36
Ticket-comparison marketplace 1,528 395 2.92

About this table: the six websites all five engines cite together most often, January to June 2026, ranked by the “In all 5 engines” column. The second column counts only the 531,889 questions where every one of the five engines returned sources, so no site gains or loses credit for questions an engine skipped. We withheld the names, and left out Wellows’ own site. This reflects the questions Wellows tracks for its customers, so it is not a ranking of the open web.

The column to watch is the last one, the average number of engines that cite a site whenever it appears at all. A score near 1.0 means it’s almost always cited alone.

Compare it against the second column and the two pull in opposite directions.

The most-cited site in our whole set, a community discussion platform, appears in 272,866 questions, by far the widest reach in the table.

Yet it averages just 1.66 engines. When one engine cites it, the others usually point to different threads on the same site.

The online encyclopedia is even more extreme at 1.36. Engines cite it constantly, and almost always alone.

The sites that genuinely travel are the narrow, trusted specialists: the bookstore at 3.26, the warranty provider at 2.94, the ticket marketplace at 2.92. Small reach, high carry-over.

Reach and cross-engine agreement are two completely different things, and the biggest gathering sites on the web are not the cross-engine magnets you’d assume.

Check your own citations, engine by engine

Every number above argues the same practical point. One blended score can’t tell you where you stand, because the engines aren’t looking at the same web.

The only way to know is to check each engine separately and compare. Both checks below are free, need no signup, and run 40 real buying-intent queries against your website.

Check What it returns Run it
ChatGPT citation check
The engine with the most unique pool (76.3%)
Citation rate, direct citations, brand mentions, and where you appear in the answer ChatGPT Visibility Tracker →
Google AI Overviews citation check
The surface closest to your existing SEO work
Whether your site appears as a direct link, or your brand is named as a source, inside real AI Overviews answers AI Overviews Tracker →

Both checks build their own question set for you. If you’d rather decide exactly which questions get tested, our LLM Query Builder turns your domain into 40 conversational queries, each tagged with the persona and the intent behind it. The Query Fan-Out Generator then takes any single one of those and expands it into the semantically related variations engines tend to split a topic into.

How to read the two results together. If one engine cites you and the other doesn’t, that’s the normal state of this system, not an error to average away. On exact pages the engines agree only 6.8% of the time.

If neither cites you but your brand is named in the answers, you have a page problem, not a brand problem, and the fix is different. If your brand is missing from both, start with brand presence, because that’s the layer that carries across all five engines.

What to do about it, by role

Back to that Monday dashboard. The number wasn’t wrong because the maths failed. It was wrong because it answered a question that has no single answer.

The findings are the same for everyone. What changes is what you do on Monday morning. Find your row, then read the section below it.

Your role What this study changes for you The number to hold onto
In-house SEO Stop reporting one score. Report five, plus a separate brand line. 6.8% pages vs 30.3% brands
Brand or marketing lead Shift budget weight from page optimisation toward brand presence. 4.5x brand agreement over page agreement
Agency Rebuild client reporting per engine, and scope by category maturity. 24.5% Google’s own ceiling

If you’re an SEO

A decade of rank tracking builds the instinct to want one number that says how you’re doing. This data says keep five.

  • Report per engine, not blended. Pages overlap just 6.8% between engines, so a page winning on one and losing on another isn’t noise to average away. It’s the normal state of the system, and each gap needs its own diagnosis.
  • Keep brand mentions on a separate line. They overlap 30.3% and behave like a different channel entirely.
  • Know what transfers and what doesn’t. Within Google’s surfaces, quite a lot: AI Overviews and AI Mode agree 3.4x more with each other than rival engines do. From Google to ChatGPT, almost nothing. No other engine touches 76.3% of what ChatGPT cites, so treat it as a new discipline rather than an extension of your Google playbook.
  • Watch carry-over, not just reach. Borrow the “average engines per question” column from the table above. A site can appear in 272,866 questions and still average 1.66 engines. That’s a carry-over problem, not a visibility problem, and it needs a different fix.

If your current setup reports only a blended figure, our roundup of the Ziptie alternatives compares which platforms break citations out engine by engine.

If you’re a brand

The budget question this study answers is where cross-engine leverage actually lives.

  • Page work pays one engine at a time. Brand work pays across all five. The engines agree on which companies matter about four and a half times more than they agree on which pages to cite.
  • Flip the weighting on digital PR. If you’ve been treating being named, reviewed, and discussed as a nice-to-have next to content production, this data argues for reversing that.
  • Move before the shortlist hardens. Cross-engine agreement rose 31% in six months, and the pool of sources all five engines trust more than tripled. Companies that get on that list early will compound as consensus sets.
  • Win your buyers’ engine first, then budget for ChatGPT separately. With this little overlap you can’t be everywhere at once. This study can’t tell you which engine your buyers use, but it can tell you ChatGPT has both the most unique pool and the least transferable playbook.
  • Ask vendors for per-engine breakdowns. A single blended number can rise while the engine your buyers actually use goes dark.

That timing asymmetry rewards younger challengers most, which is why we’ve set out a separate playbook for early-stage companies building AI visibility from zero.

And if you’re mid-way through choosing a platform, our comparison of the Rankscale alternatives breaks the field down on exactly the per-engine question.

If you’re an agency

This study hands you two things: a reporting structure, and a way to set client expectations before the work starts.

  • Retire the blended score from your decks. Show five scoreboards plus a brand-mention line that cuts across them. It’s more honest, and it’s more defensible.
  • Pre-empt the hardest client question. “We won in AI Overviews, so why doesn’t ChatGPT mention us?” Those two engines cite the same sources just 6.9% of the time, while Google’s own two surfaces reach 24.5%, sharing infrastructure neither of you controls.
  • Scope by category maturity. A settled category at 10 to 15% agreement is a slow, defensible displacement job. The emerging AEO/GEO space at 6.3% is a land grab: easier to enter, harder to hold. Same service, two very different roadmaps, and two very different prices.

Doing it across a full book of clients is its own operational problem, since every account now needs five scoreboards rather than one. Our AI visibility platform for agencies is built around that multi-client, per-engine structure.

We’ve set out what that reporting structure looks like in practice for agencies running AI visibility across a client roster.

None of this is a flaw in AI search. For publishers and challenger brands, it’s the most open playing field since the early days of Google: five chances instead of one gatekeeper, with a brand layer that pays off on all five at once.

But it only works if you can see all five boards. Watching one scoreboard while playing five games is how brands end up confidently invisible.

The full method, and how we stress-tested it

In a new field, a bold claim is only as good as the method behind it. Here it is in full, laid out so you can poke holes in it.

Definition

How we measure agreement

For every question, each engine returns a list of sources it used. To see how much two engines agree, we count the sources they both used, then divide by the total number of different sources the two of them mentioned between them. Share nothing, that’s 0%. Identical lists, that’s 100%. We do this for every question and average the results. We measure at three levels: the exact web page, the website it sits on, and the brand named in the answer. A brand counts when the engine names the company in its answer text, whether or not it also links to that company, and we count named and implied mentions together. Measuring all three the same way, on the same questions, is what lets us compare them fairly. Separately, for the headline number, we counted how many of the five engines cited each individual source. Those are two different measurements, and we keep them apart throughout.


How the numbers narrow, step by step
  1. Six months of citations, January to June 2026: 22,749,707 citations across 1,146,483 questions and 441,946 websites.
  2. Keep only questions where all five engines cited at least one website: 531,889 questions, 13,037,251 citations, 280,245 websites.
  3. Reduce to distinct website-and-question pairings: 8,729,964.
  4. Count how many engines cited each pairing. Exactly one engine: 6,949,766, which is the 79.6% headline. All five: 27,097, or 0.31%.

Every figure in this post traces back to one of those four steps. Where a section uses a narrower base, we say so.


Six things to know before you trust the numbers
  1. The questions lean commercial. They come from the brands Wellows tracks, so this isn’t a random sample of everything people ask an AI engine. The same applies to the categories: they are the topics Wellows tracks for customers, not a survey of every industry.
  2. All-five comparisons use only the 531,889 questions every engine answered, so no engine loses ground for skipping one. Pair-by-pair comparisons use every question both engines in that pair answered, which is a larger base. The page-vs-website-vs-brand comparison narrows to the 236,096 questions where all five answered and at least one named a brand. That’s why the same pair can report slightly different figures in different sections. We keep the base attached to every number rather than blending them.
  3. One website means one website. At website level we treat www.example.com and example.com as the same site rather than two. This matters for the counts above, and barely at all for the percentages: every engine-pair figure in this post moves by less than a tenth of a point either way.
  4. Group figures are averages of pairs. When we say cross-company agreement is 7.3% of websites, that’s the mean of the seven cross-company pairs, which individually run from 6.1% to 9.3%. The Google-family 11.4% is the mean of two pairs.
  5. Exact-page matching runs on tidied-up web addresses, stripping the protocol, www, tracking parameters, in-page anchors, default index files, and trailing slashes. In-page anchors alone move page-level agreement by about three points, because two engines citing the same article via different section links would otherwise look like two different sources.
  6. This is real-world data, not a lab test. We can show how much the engines agree, but we can’t prove exactly why. Where we suggest a reason, we say so.

We tried to break our own headline

Here’s a fair challenge to a number like 79.6%: maybe we picked the settings that flatter it.

The opposite is true, and it’s worth being precise about why. Both choices behind the headline make the engines look more alike, not less:

  • Website-level matching is looser than page-level matching. Two engines citing different articles on the same site count as agreeing. That finds more agreement, so fewer sources come out unique to a single engine.
  • Tidying up web addresses does the same thing. Stripping tracking parameters and in-page anchors makes addresses match that would otherwise have counted as two separate sources.

So 79.6% is the floor, not the ceiling. Tighten either setting and the number climbs.

Here are the same five engines, on the same questions, with only the match getting stricter at each step:

Horizontal bar chart: 79.6%, 85.1% and 93.2% of sources are cited by only one of five AI engines as match strictness increases from website to exact page


Tighten the match and the finding gets stronger, never weaker, rising from 79.6% at the loosest setting to 93.2% at the strictest, across all five engines.

Test conditions Engines Matched at Share cited by one engine only
Headline (loosest match) 5 Website, tidied addresses 79.6%
Tighter 5 Exact page, tidied addresses 85.1%
Tightest 5 Exact page, raw addresses 93.2%

Every version of this test lands higher than the headline. The finding doesn’t weaken under tougher conditions, it strengthens, which is exactly what you want from a result you plan to build a strategy on.

We lead with 79.6% because it is the most conservative number we have: the loosest matching, on the full five-engine set. Every stricter cut lands above it.

Frequently asked questions



For each question, we took the list of items each engine produced (web pages, websites, or named brands), counted how many any two engines shared, and divided by the total number of different items across both, then averaged across all questions. Separately, for the headline number, we counted how many of the five engines cited each individual source. For exact web pages we tidied the addresses first, stripping the protocol, www, tracking parameters, in-page anchors, default index files, and trailing slashes. At website level we treat www.example.com and example.com as one site. At brand level we count every company the engine names in its answer, named and implied together. The study covers January 1 to June 30, 2026: 22,749,707 citations across 1,146,483 questions, from Wellows’ internal citation dataset. All five engines cited on 531,889 of those questions, which produce the 8,729,964 website-and-question pairings behind the headline.


Ranked by how often two engines cite the same source on the same question: Google AI Overviews and Google AI Mode lead at 24.5%, then Gemini with AI Overviews (12.0%) and Gemini with AI Mode (10.9%). Cross-company pairs follow: AI Overviews and Perplexity (9.3%), AI Mode and Perplexity (8.3%), Gemini and Perplexity (7.4%), AI Overviews and ChatGPT (6.9%), AI Mode and ChatGPT (6.5%), ChatGPT and Perplexity (6.5%), and Gemini and ChatGPT last at 6.1%. Every Google-family pairing beats every cross-company pairing, and every ChatGPT pairing sits in the bottom four.


Yes, measured on questions where both engines returned sources. Of every website ChatGPT cited, Perplexity never cited 89.1% on that same question. The figure is similar against every rival: 89.2% against Google AI Overviews, 90.2% against Google AI Mode, 91.5% against Gemini. Compared against all four rivals at once, nobody else touches 76.3% of ChatGPT’s sources.


Very, and it is deliberately the most conservative version of the finding. It comes from 6,949,766 of the 8,729,964 website-and-question pairings across the 531,889 questions all five engines answered. It held in a narrow 77.7% to 81.1% band across all six months, so it isn’t a fluke of one time period. It also uses the loosest matching we have: website level, on tidied addresses. Tighten the match to the exact web page, keeping all five engines, and the share rises to 85.1%, or 93.2% with no address cleanup at all. The finding gets stronger under tighter conditions, not weaker.


It comes down to how specific you get. Many pages can answer a question, but only a handful of brands define a category. So the engines split on which exact page to cite (6.8% overlap) while agreeing far more often on which companies to name (30.3% overlap), measured on the 236,096 questions where all five answered and at least one named a brand. Winning specific pages is an engine-by-engine game; brand presence carries across engines.


When two engines run on the same underlying systems, they cite more alike. Google’s two surfaces (AI Overviews and AI Mode) agree on 24.5% of the websites they cite, versus an average of 7.3% for engines from different companies, a 3.4x difference. It suggests agreement comes more from shared indexes and ranking systems than from content quality, though we can’t observe those systems directly, and Google’s two surfaces share a product team as well as an index.


Yes, slowly. Monthly average agreement rose from 8.58% (January) to 11.25% (June), a 31% rise, dipping once in February before climbing every month after, or about 0.53 percentage points a month in a near-straight line. Six monthly points is a short series, so read it as a direction rather than a forecast. A separate measure backs it up: the share of sources cited by all five engines more than tripled over the same period, from 0.18% to 0.59%.


Barely. Research questions sit at 9.5% agreement, commercial questions at 9.6%, buying questions at 10.6%, and navigational ones at 11.8%, a 2.3-point spread. The category the question sits in matters far more, ranging from 6.3% for emerging AEO/GEO topics to 15.5% for the most settled category in our set. We pooled every category the same way so the comparison is like-for-like.


Five: ChatGPT (OpenAI), Google Gemini, Perplexity, Google AI Overviews, and Google AI Mode. For all-five comparisons we used only the 531,889 questions where every engine returned sources; pair-by-pair comparisons use every question both engines in the pair answered.


Start where your buyers ask questions, and lock down brand presence everywhere, since that carries across engines. Then expand engine by engine at the page level. Because agreement tracks infrastructure more than content, one generic effort aimed at all five at once is exactly the approach this data argues against.

Note on earlier work: others have observed the split of AI citations across engines before, including Kevin Indig in “The Consensus Gap” (Growth Memo, May 2026). We ran this study independently, on Wellows’ own five-engine dataset.