Key Takeaways
  • It measures odds, not a spot: LLM rank tracking runs the same prompts through AI engines many times and reports how often you appear. One screenshot tells you almost nothing.
  • Citations beat mentions: Being named is nice. Being linked as a source is what carries authority and sends clicks, so track the two separately.
  • Five engines, five systems: ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode pull from different sources. A win on one says little about the others.
  • Measure what users see: Answers in the consumer apps can differ from what an API returns. Your tracking should reflect the screen your buyer looks at.
  • Your AI reputation already exists: The sources AI engines cite shape how they describe you. Tracking shows you the exact pages behind that story.
  • Tracking only pays off with action: Every gap you find should map to a page to fix, a source to pitch or a piece to publish.

Short answerLLM rank tracking means running a fixed set of real buyer prompts through ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, again and again, and logging if your brand appears, if it’s cited as a source, where it sits and how it’s described. You get rates to watch, not one position.

If you’ve been trying to work out how to rank in AI Overviews, this is the piece most advice skips: you can’t improve a position you’ve never measured properly. And the audience is large. Google said in May 2026 that AI Overviews has over 2.5 billion monthly active users and AI Mode has passed 1 billion (Google, 2026).

What Is LLM Rank Tracking?

Short answerLLM rank tracking tells you how visible your brand is inside answers written by large language models. You pick the prompts that matter to your business, run them on each AI engine many times, and log presence, citations, position and sentiment. The output is a trend line per prompt and per engine, not a single number.

📘 Definition
LLM rank tracking
The repeated measurement of how often, where and how favorably a brand appears in AI generated answers for a defined set of prompts, broken down by AI engine and by whether the brand is mentioned, cited as a source, or missing.

Why “rank” means something different inside an AI answer

In classic search, position 3 is position 3 for most people at a given moment. In an AI answer, there’s no fixed list. The model writes a fresh response each time, sometimes names you first, sometimes fourth, sometimes not at all.

So the useful question changes. It stops being “where do I rank?” and becomes “out of 20 runs of this prompt, how often was I there, and how often was I the source?”

Then vs. now

Classic rank tracking LLM rank tracking
Unit you track Keyword Prompt (a full question, often long tail)
What you record One position, one URL Presence, citation, position in the answer, sentiment, cited sources
How often it changes Days or weeks Between runs of the same prompt
Samples needed One check per day Several runs per prompt per engine
Where results live One results page Five engines with different source pools
What “winning” means Top 3 blue link Being named and linked as a trusted source

GEO, AEO and AI SEO Rank Tracking: Are They the Same Thing?

Short answerMostly, yes. GEO rank tracking (generative engine optimization), AEO rank tracking (answer engine optimization), AI SEO rank tracking and LLM rank tracking all describe the same job: measuring your visibility in answers written by AI.

The labels lean in slightly different directions, though:

GEO rank tracking
Usually covers the full set of generative engines, including ChatGPT and Perplexity.
AEO rank tracking
Tends to focus on direct answer formats, such as Google AI Overviews and featured answers.
AI SEO rank tracking
The term SEO teams use when they bolt AI visibility onto an existing search report.

Pick whichever word your stakeholders already use. The method below works for all of them. For the wider strategy behind the labels, see our comparison of SEO vs. GEO.

What Should You Track in LLM Rank Tracking?

Short answerTrack five things per prompt and per engine: mention rate, citation rate, position in the answer, sentiment and the source URLs the engine relied on. Then roll them into share of voice against your competitors. Anything less leaves you guessing why a number moved.

The core metrics

Metric What it measures Why it matters Watch out for
Mention rate Share of runs where your brand is named Basic awareness inside AI answers Counts a passing name drop the same as a recommendation
Citation rate Share of runs where your page is linked as a source Authority and referral traffic Some engines hide sources behind a click
Position in answer Where you appear among the brands named First named tends to get the attention Only meaningful as an average across many runs
Sentiment Whether the description is positive, neutral or negative Shows the story the model tells about you Needs the source URL to be fixable
Share of voice Your mentions or citations as a share of all tracked brands Competitive context Changes when you add or drop competitors from the set
AI visibility score A weighted blend of the above A single number for reporting Useless without the parts underneath it

Mentions and citations are different wins

This is where most reports go wrong. A mention means the model said your name. A citation means it pointed to your page as evidence, which is the version that carries authority and can send a visitor your way.

The gap between the two is bigger than most teams expect. In a Wellows study of 20.5 million AI citations across five engines (March to July 2026), 90.4% of cited sources didn’t mention any tracked brand at all, and only 4.5% were a brand’s own page used as the source (Wellows, 2026). If your tracking only counts mentions, you’re missing most of the picture.

What 20.5 million AI citations actually pointed to
No tracked brand named
90.4%
Third party page names a brand
5.1%
Brand’s own page cited
4.5%
Source: Wellows, 2026. Five engines, March to July 2026.

Sentiment and your shadow reputation

Every brand has a shadow reputation in AI: the version of you the models describe when you’re not in the room, built from sources you didn’t write. A review thread from 2023, an outdated comparison page, a forum complaint. If an engine cites it, it shapes the answer.

That’s why sentiment on its own isn’t enough. You need the specific URLs behind a negative description, because a page is something you can respond to, update or outrank. A vague “sentiment dipped” is not.

💡 Did you know?
In the same Wellows study, only 44.3% of sources with a domain authority under 40 that were cited in April to June were still being cited two months later. For domains rated 90 and above, it was 92.0% (Wellows, 2026). Citations from smaller sites come and go, so a single month of data will mislead you.
Share of sources still cited two months later
DA under 40
44.3%
DA 40 to 69
67.0%
DA 70 to 89
86.5%
DA 90 and above
92.0%

Why Do LLM Rankings Change Every Time You Check?

Short answerBecause AI engines generate each answer fresh, and the sources they retrieve shift with the phrasing, the user’s location, the time and the model version. Even with every setting locked, outputs can vary. That’s why a single check is an anecdote and repeated runs are data.

The models themselves vary

Researchers at Thinking Machines Lab ran the same prompt 1,000 times on an open model at temperature zero (the setting meant to remove randomness) and still got 80 different completions (Thinking Machines Lab, 2025). Consumer AI apps add search, personalization and product modules on top of that. Treat this as directional, since it was a lab test on one model, but the lesson holds.

1,000
runs of the same prompt
80
different answers returned
0
temperature setting (meant to remove randomness)

The source pool moves too

Engines also change what they trust. The Wellows domain authority study found ChatGPT’s median cited domain authority jumped 16 points in a single week of July 2026, landing at 51 after a March to June baseline of 35 (Wellows, 2026). Nothing on your site changed, yet your odds of being cited did.

March to June 2026
35
➜
After July 7 to 14
51
ChatGPT median cited domain authority, one week apart

The API isn’t the app

A lot of tracking runs prompts through a model’s API because it’s cheap. The trouble is that the API answer often differs from what a person sees in the ChatGPT or Gemini app, where web search, shopping modules and location all come into play.

If the goal is to know what your buyer sees, measure the screen your buyer sees.

How to Set Up LLM Rank Tracking: 7 Steps

Here’s a setup you can start on Monday. It works for one brand or twenty clients.

  1. 1Build a prompt set of 30 to 100 real questions.Pull them from sales calls, support tickets, People Also Ask and your own search queries. Mix category questions (“best payroll software for a 20 person team”), comparison questions and problem questions. Our guide to finding queries for GEO covers where to source them.
  2. 2Tag every prompt by intent and topic.Group prompts into topic clusters so you can report by theme. Wellows analysis of 2.37 million AI citations found a brand’s AI visibility can vary by more than 4.5x across topics within its own niche (Wellows), so a single blended number hides your weak spots.
  3. 3Pick your engines and treat each one separately.Track all five major engines, but never average them into one number before you’ve read them individually (the table below shows why).
  4. 4Set a sampling plan.Run each prompt several times per engine per cycle and track weekly. . Keep location and language fixed so changes mean something.
  5. 5Capture the answer as a user sees it.Record the full response, every brand named, every source URL and the position of each. Browser based collection gets you closer to the real app experience than API calls.
  6. 6Add three to five competitors.Share of voice only means something against named rivals. Watch where their visibility rises and which URLs are doing the lifting.
  7. 7Turn every gap into a task and review monthly.For each prompt where you’re missing or described badly, write down the fix: a page to update, a third party source to pitch, a new piece to publish. Then check next month whether the rate moved.

How the five engines differ

Engine How it shows sources Median cited DA (Mar to Jun 2026) What to watch
ChatGPT Inline links and a sources panel when it searches the web 35 Whether search triggers at all for your prompt; answers without search carry no citations
Gemini Source links on grounded answers 39 Differences between the app and Google Search surfaces
Perplexity Numbered citations on nearly every answer 34 The most citation friendly engine, so a good early signal
Google AI Overviews Link cards beside the summary 47 Only appears for some queries, so track how often it triggers
Google AI Mode Links throughout a conversational answer 48 Follow up turns can change which sources appear
Median domain authority of cited sources, by engine
Google AI Mode
48
Google AI Overviews
47
Gemini
39
ChatGPT
35
Perplexity
34
Source: Wellows, 2026. March to June 2026 baseline.

Notice the split. The two Google surfaces lean toward higher authority domains, while ChatGPT and Perplexity cite smaller sites far more often.

The same study found 66.5% of domains under DA 40 were cited by only one of the five engines. One engine’s result is not your AI visibility.

How Do You Rank in AI Overviews (and Prove You Did)?

Short answerYou rank in AI Overviews by being the clearest, most trusted answer to the narrow questions Google generates around a query, and you prove it with prompt level tracking. Search Console won’t tell you which competitor got cited when you didn’t, so you need a separate record of every AI Overview for your priority queries.

What should I know about how to rank in AI Overviews?

Three things matter most for measurement:

  • AI Overviews don’t show for every query. In Pew Research Center data from March 2025, about 18% of Google searches produced an AI summary (Pew Research Center, 2025). Track trigger rate first. That share has likely grown since, so check it on your own queries.
  • Google fans out your query. AI Mode and AI Overviews answer related sub questions behind the scenes, so your tracking set should include those follow ups, not only the head term.
  • Authority helps, but it’s not the gate. In the Wellows study, 40.8% of AI Overview citations came from domains under DA 40.

How to show up in AI Overviews with SEO

Solid SEO gets you into the pool Google retrieves from. Structured, direct answers, current facts and third party mentions get you picked from that pool.

For the full workflow, read our guide on how to rank in Google AI Overviews. Then use the tracking setup above to check if it worked.

What Makes the Best LLM Rank Tracking Setup?

People searching for the best AI SEO rank tracking or the best LLM SEO rank tracking tool are usually asking one question: can I trust these numbers enough to act on them? Judge any tool or homegrown setup against these criteria.

Look for
  • ✅All five engines tracked separately, including Google AI Mode
  • ✅Answers captured the way a user sees them in the app
  • ✅Citations and mentions reported as separate metrics
  • ✅The source URLs behind each answer, with sentiment attached
  • ✅Repeated runs per prompt, with rates rather than single positions
  • ✅Competitor tracking on the same prompts and schedule
  • ✅A clear route from gap to action (content, outreach, page updates)
Be wary of
  • ❌A single “AI rank” with no breakdown underneath
  • ❌API only collection with no mention of how it compares to the app
  • ❌Mention counts presented as citations
  • ❌Daily screenshots of one run per prompt
  • ❌Dashboards that show the problem and stop there

What a useful report line looks like

Before
“Brand visibility in ChatGPT: 62. Up 4 points.”
After
“On 40 payroll prompts, ChatGPT named us in 38% of runs and cited our pages in 11%. Competitor A was cited in 24%, mostly via two comparison pages from third party review sites. We’re pitching both publishers this month.”

The second version tells your boss what happened and what you’re doing about it. That’s the standard.

How Wellows helps
This is the job Wellows is built for. It runs your prompts across ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, tells you whether you were cited or only mentioned, and surfaces the exact URLs shaping your shadow reputation. The same scan runs on your competitors, so you can see which sources lift their visibility and where they’re exposed.

Your LLM Rank Tracking Checklist

Copy this into your project tool.

LLM Rank Tracking Checklist
0 of 15 done
Setup 0/5

Every cycle 0/5

Every month 0/5

🎉 All 15 done. Your tracking setup is ready to report on.

If you’d rather audit your current state first, our AI search visibility audit checklist is a good starting point.

What Results Can You Expect From LLM Rank Tracking?

Short answerTracking alone won’t lift a single number. What it gives you is the evidence to decide where to act, and a way to prove the actions worked.

Reporting confidence

The first payoff is usually internal. When leadership asks “are we showing up in ChatGPT?”, you can answer with rates per topic and per engine instead of a screenshot someone took on their phone.

Traffic, with honest expectations

Clicks from AI answers are real but smaller than classic search clicks. Pew found Google users clicked a traditional result in 8% of visits with an AI summary, against 15% without one, and clicked a link inside the summary itself in only 1% of visits (Pew Research Center, 2025). That data is from March 2025, so treat it as a baseline rather than today’s number.

15%
clicked a result on pages without an AI summary
8%
clicked a result on pages with an AI summary
1%
clicked a link inside the AI summary
Source: Pew Research Center, 2025. Browsing data from 900 US adults, March 2025.

What this means in practice: when a click does happen, it usually comes from a citation. Another reason to track citations separately from mentions.

Pipeline and influence

A lot of AI influence never shows up as a click. The buyer reads the answer, remembers the brand and searches for it later.

Watching branded search volume alongside your citation rate is the simplest way to spot this.

📊 Quotable stats
2.5 billion
monthly active users of Google AI Overviews. Google, May 2026
900 million
weekly active users of ChatGPT, as disclosed by OpenAI. TechCrunch, February 2026

How Wellows Handles LLM Rank Tracking

Everything above can be done with spreadsheets and a lot of patience. Wellows runs the same workflow for you, across ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, and ties each gap to something your team can do about it. Here’s how each part maps to the setup steps.

Steps 1 to 4: Prompt TrackingStep 5: LLM CitationsStep 6: AI Visibility ScoreStep 7: Content and Outreach

The dashboard at a glance

Brand Visibility Overview: your whole tracking setup on one screen

Before you dig into any single step, the overview tells you where you stand. It shows your rank, visibility score, total citations split into explicit and implicit, and a trend line against the competitors you track, with a short written brief on what changed since the last scans.

  • A plain language brief on who is gaining ground and where
  • Explicit and implicit citations counted separately
  • Your prompt, topic, competitor and country coverage in one bar, with gaps flagged
Wellows Brand Visibility Overview showing rank, visibility score, explicit and implicit citations, and a visibility trend against competitors
Brand Visibility Overview in Wellows

Steps 1 to 4

Prompt Tracking: see every prompt where you show up (or don’t)

Add the prompts your buyers actually ask and Wellows checks them daily across the major AI engines. You see where your brand appeared, where it was overlooked, and which competitor took the spot instead.

  • Filter results by platform, region, intent and competitor
  • Spot the prompts where you’re mentioned but never cited
  • Keep location and prompt wording fixed, so trends mean something
Wellows Prompt Tracking view showing brand presence by prompt across AI platforms
Prompt Tracking in Wellows

See how Prompt Tracking works ➜

Step 5

LLM Citations: know if you were cited, not only mentioned

This is the part most tracking misses. Wellows separates explicit citations (your own pages used as the source) from implicit ones (third party pages the engines trust that talk about your category), and shows the exact source URL behind each.

  • Track citations, the metric that carries authority and clicks
  • See the exact URLs shaping your shadow reputation, the version of your brand the models describe when you’re not in the room
  • Turn each missing source into an outreach or content task
Wellows citation tracking showing explicit vs. implicit citation distribution and citation score
Explicit and implicit citations, with the source behind each

See how LLM Citations works ➜

Step 6

AI Visibility Score: one number, with the breakdown underneath

The AI Visibility Score is a 0 to 100 measure of the share of citations your brand earns against your competitors. It updates daily and splits by platform, so a strong week in Perplexity can’t hide a weak one in AI Overviews.

  • Competitor leaderboard on the same prompts and schedule
  • Per platform breakdown across all five engines
  • History over time, so you can prove what your fixes did
Wellows AI Visibility Score dashboard with competitor leaderboard and per platform breakdown
AI Visibility Score with competitor leaderboard

The per model view goes one level deeper. It shows how often each brand appears on each engine and calls out your weakest one, with the next moves to close that gap.

Wellows Brand Visibility Across LLMs table comparing brand visibility on Google AI Overviews and ChatGPT, with recommended next steps
Brand Visibility Across LLMs, with the weakest engine flagged

See how the AI Visibility Score works ➜

Step 7

Content Opportunities and Outreach: close the gaps you found

Tracking only pays off when it turns into work. Wellows shows the topics where competitors get cited and you don’t, with suggested titles and intent, then finds the sites where your brand is missing and gives you verified contacts to pitch.

  • Content ideas ranked by likely citation impact
  • Outreach targets on the sources AI engines already trust
  • Verified editor and site owner contacts, ready to export
Wellows Content Opportunities dashboard
Content Opportunities
Wellows Outreach dashboard showing citation score and outreach opportunities
Outreach opportunities

Content Opportunities ➜    Outreach ➜

Conclusion

LLM rank tracking comes down to three habits: sample repeatedly, separate citations from mentions, and read each engine on its own. Get those right and the numbers start telling you what to fix.

Start small. Pick 30 prompts this week, run them across all five engines, and log every cited URL. You’ll have your first real gap list by Friday.

See where AI engines cite you, and where they don’t

Wellows tracks your prompts across ChatGPT, Gemini, Perplexity, AI Overviews and AI Mode, shows you where you’re cited and where your shadow reputation comes from, and turns the gaps into content and outreach work you can act on.

Explore Wellows

`

See how it fits your broader AI visibility measurement plan.


LLM rank tracking checks how often AI tools like ChatGPT, Gemini, and Perplexity mention or cite your brand when people ask questions about your category. You run the same prompts repeatedly, record the results per engine, and watch the rates change over time. It’s the AI search version of keyword rank tracking, built around probability instead of fixed positions.

Regular rank tracking records one position for one keyword on one results page. LLM rank tracking records presence, citations, position within the answer, and sentiment across several runs and several engines, because AI answers change between runs. You also track the source URLs each engine relied on, which tells you why you appeared or didn’t.

Start with 30 to 100 prompts that reflect real buyer questions, grouped by topic. Fewer than 30 makes the numbers jumpy; more than a few hundred gets hard to act on for a small team. Add prompts as you learn which topics drive business, and retire ones nobody asks anymore.

Weekly works for most brands, with several runs per prompt per engine each cycle. Daily tracking makes sense for competitive categories or during a launch. Review trends monthly, since AI engines can shift which sources they trust within a single week, and a monthly view smooths out the noise.

No. A mention means ChatGPT named your brand in the answer. A citation means it linked to your page as a source. Citations carry more authority and are where referral clicks come from, so track them as a separate metric. Plenty of brands get mentioned often and cited rarely.

You can for a handful of prompts, using a clean browser session and a spreadsheet. It breaks down quickly, though. Five engines, several runs per prompt, and dozens of prompts add up to hundreds of checks per cycle, and manual checks miss the cited source URLs that explain each result.

The engines pull from different source pools and weigh authority differently. Industry research shows most smaller domains are cited by only one major engine. Perplexity cites sources on nearly every answer, while ChatGPT only cites when it searches the web. Check which URLs each engine cites for your prompts to find the gap.

It matters, but less than many expect. Analysis of millions of AI citations shows that a large percentage of cited sources have moderate domain authority. Higher authority sites tend to hold citations longer and appear across more engines, so authority helps consistency more than it decides entry.