Your board wants a number for AI search. You open Search Console, open GA4, open whichever AI visibility tool the agency sold you, and get three figures that have nothing to do with each other.
None of them measure the same thing, so nobody can say which is right. Pick the wrong one and you spend a year optimising toward a metric that moves on its own.
There is no rank to track inside an AI answer, so visibility has to be measured as a frequency rather than a position. The working stack in 2026 has five layers: Search Console impressions, Bing citation counts, GA4 referral sessions, server logs, and a sampled prompt panel run often enough to average out the noise.
Key facts
- Google launched a generative AI performance report in Search Console on 3 June 2026 and rolled it out worldwide on 31 August 2026, per the Google Search Central blog.
- That report exposes impressions only — no clicks, no click-through rate and no query dimension — grouped by page, country, device and date, per Google’s Search Console documentation.
- Bing Webmaster Tools opened its AI Performance report in public preview on 10 February 2026, reporting total citations, average cited pages per day and the grounding queries Copilot used to retrieve content.
- Google Analytics 4 now carries AI Assistant as a default channel, covering arrivals from ChatGPT, Gemini, DeepSeek, Copilot and Grok, per the Google Analytics help documentation.
- In research by SparkToro and Gumshoe.ai across 600 volunteers and roughly 2,961 prompt runs in November and December 2025, the odds of ChatGPT, Claude or Google’s AI returning the same list of brands twice were under 1 in 100.
- Conductor’s analysis of 14,000 API calls — 10 industries by seven intent types by four engines, 50 runs each — found brand overlap between two runs ranged from 40% on purchase prompts to 63% on comparison prompts.
This article sets out the five data sources that exist today, what each genuinely measures, where each misleads you, and the check that tells you whether the number in your report is real.
Why can’t you rank-track an AI answer?
Because the answer is regenerated each time and the order inside it is close to random. A large language model samples from a distribution; it was not built to return a stable ordered list, and it does not.
What it looks like: your tool reports you at “position 3” in ChatGPT, and a colleague running the same prompt an hour later cannot find you at all.
Key data: SparkToro and Gumshoe.ai, testing 12 prompts through ChatGPT, Claude and Google’s AI across 600 volunteers in late 2025, put the chance of two identical brand lists at under 1 in 100, and the chance of two identical lists in the same order at roughly 1 in 1,000. Conductor’s larger API study found the same effect graded by intent: 40% brand overlap between runs on purchase prompts, 63% on comparison prompts.
Why it matters: a position inside an AI answer is noise dressed as a measurement, and reporting it will make you look wrong every time the client checks by hand.
Here is the check. Run one priority prompt ten times in a fresh session with history and personalisation switched off, recording whether your brand appears. Six out of ten means 60% visibility for that prompt, and any tool reporting a single position has discarded the only stable signal in the data.
Frequency is what survives the randomness. SparkToro found that although ordering collapsed, leading brands still turned up in 60% to 90% of runs for a given intent. Report the percentage of runs you appear in, over a named run count, and you have something defensible.
The unwelcome part is the sample size. Thirty prompts at ten runs each is 300 queries a month before you have looked at one competitor, and that cost is why cheap tools sample once and call it a rank.

What does Google’s generative AI performance report actually give you?
Impressions, and nothing else. Google announced the report on 3 June 2026 and finished rolling it out to every property on 31 August 2026.
What it looks like: a new Search Console view with a single impressions line, four dimension tabs, and no clicks column anywhere.
Key data: per Google’s documentation for the report, it covers AI Overviews and AI Mode, groups data by pages, countries, dates and devices, filters by text or multimodal search type, excludes Search Labs experiments, and inherits the standard 1,000-row export limit.
Why it matters: this is the only first-party count of how often Google showed your pages inside an AI answer, and it is free, so validate every other source against it.
Open Search Console and find Generative AI performance under the Performance group. Set the range to the last three months, export the Pages tab, then export Pages from the standard Search results report for the same window. The difference tells you which URLs Google reaches for when summarising rather than ranking, and those are rarely the pages your plan was built around.
Two absences matter more than anything the report contains. There is no query dimension, so you cannot learn which questions produced the appearance. And AI Overviews and AI Mode are aggregated together, so a rise could be either surface, or a shift between them, and nothing tells you which.
Does Search Console tell you whether a click came from an AI Overview?
No. Clicks from AI features are folded into the ordinary Search results report under the Web search type, with no flag separating them.
What it looks like: your organic clicks are flat, your generative AI impressions are climbing, and you cannot connect the two.
Key data: Google’s documentation on AI features states there are no additional requirements to appear in AI Overviews or AI Mode and no special optimisations necessary, and that performance data for these features sits inside the Web search type of the standard report.
Why it matters: any agency selling you an “AI Overview click” figure derived from Search Console is deriving it from data that does not exist.
Subtraction gets you close. Look for pages where generative AI impressions rose while Search results clicks for those same URLs fell — that divergence is the nearest thing to a measurable AI Overview effect Google supplies, and it is the signature we describe in the ten reasons rankings hold while clicks keep falling.
The Pew Research Center study is the honest backdrop. Across 68,879 Google searches made by 900 US adults in March 2025, people clicked a traditional result on 8% of visits where an AI summary appeared against 15% where it did not, and clicked a link inside the summary on 1% of visits. Presence in the answer and a visit are close to unrelated events.
What does Bing Webmaster Tools show that Google does not?
Citations and the queries behind them. Microsoft opened AI Performance in Bing Webmaster Tools as a public preview on 10 February 2026.
What it looks like: a report with total citations, average cited pages per day, per-URL citation counts, and a list of grounding queries.
Key data: Microsoft’s announcement covers Copilot, AI-generated summaries in Bing and selected partner integrations, and notes the grounding query list is a sample of citation activity, not a complete record.
Why it matters: grounding queries are the query dimension Google withholds, and Copilot draws on Bing’s index, so the patterns generalise further than Bing’s own market share suggests.
Verify the property in Bing Webmaster Tools — the import from Search Console takes two minutes — then open AI Performance and sort grounding queries by citation count. Compare that list against the queries you target in Google. In most accounts we see, the overlap is smaller than people expect, because retrieval favours pages that state a fact plainly over pages built to rank.
Do not oversell this source. It is a preview, the sampling is undocumented, and Bing’s share of search is a fraction of Google’s. What it gives you is direction, not volume.
How do you see AI assistant referrals in GA4?
There is a channel for it now. Google Analytics 4 added AI Assistant to the default channel group, so the major assistants no longer hide inside Referral.
What it looks like: a new row in Traffic acquisition, usually with a small session count and a suspiciously good conversion rate.
Key data: the Google Analytics default channel group documentation defines AI Assistant as the channel by which users arrive from sources such as ChatGPT, Gemini, DeepSeek, Copilot or Grok, matched either on a medium of exactly ai-assistant or on a referrer in Google’s maintained list.
Why it matters: this is the only layer of the stack that connects an AI surface to revenue, which makes it the number the business will care about most.
Open Reports, Acquisition, Traffic acquisition, and set the primary dimension to Session default channel group. If AI Assistant is missing, switch to Session source and filter for chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com. Then do the thing most people skip: confirm your key events still fire on those sessions, because a channel with no conversions recorded looks identical to one with no value.
The undercount is the catch, and it is large. Native mobile apps frequently send no referrer at all, so a meaningful share of assistant traffic lands in Direct and stays there. Adobe Analytics measured AI-referred traffic to US retail sites growing 393% year on year in the first quarter of 2026 and converting 42% better than non-AI traffic in March 2026. Small, late in the decision, and systematically under-counted. The wider limits of these reports sit in our guide to what GA4 can and cannot answer.
Why do your server logs disagree with your analytics?
Because crawling and referring are different activities, and the gap between them is enormous. Your logs record what AI systems took; GA4 records the fraction they sent back.
What it looks like: tens of thousands of bot requests a month against a few dozen assistant sessions.
Key data: Cloudflare’s crawl-to-refer ratio on Cloudflare Radar, measured across the week of 19 to 26 June 2025, put Anthropic’s ratio at roughly 70,900 HTML requests for every one referral. Cloudflare also notes the ratios may be overstated, because traffic from native apps usually arrives without a referrer header and cannot be counted.
Why it matters: logs are the only source that proves an AI system can reach your content at all, which is the precondition for every other number in this article.
Pull a month of access logs, filter the user-agent field for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended and Bingbot, then group by status code. You want 200s on the pages you care about. A wall of 403s means your CDN or firewall is blocking the crawlers you are paying an agency to court, and that check has resolved more “we are invisible in AI” complaints than any content change we have made.
Screaming Frog, a log analyser or a few lines of grep will do. Run it monthly, and after any CDN, firewall or robots.txt change.
Should you still track Google rankings for AI visibility?
Yes as a leading indicator, no as a substitute. Ranking well makes citation more likely without making it reliable.
What it looks like: pages that hold position 2 and are never cited, alongside a page on the third results page that turns up in answers constantly.
Key data: Ahrefs found 76.10% of AI Overview-cited pages ranked in the top 10, 9.50% ranked between 11 and 100, and 14.40% did not rank in the top 100, across 1.9 million citations from 1 million AI Overviews in July 2025. BrightEdge, measuring differently over 16 months to September 2025, put the share of AI Overview citations that rank organically at 54.5%, up from 32.3% in May 2024.
Why it matters: two credible vendors produce figures twenty points apart on what sounds like one question, which tells you how much weight a single third-party number deserves.
Export your top 100 ranking URLs from Ahrefs or Semrush and cross-reference them against the Pages tab of the generative AI report. Pages in both are working as intended. Pages that rank and are never cited are the interesting set, and the fix is structural rather than editorial — a plain answer near the top, a heading that mirrors the question, figures a model can lift. That pattern is set out in our analysis of what gets cited in AI Overviews and how it changes by query.
Keep rank tracking in the report. Drop the pretence that it measures AI visibility.

How do you judge a third-party AI visibility tool?
By its sampling, not its interface. Every one of these tools is running prompts and counting appearances; the only question that separates them is how many times, and under what conditions.
What it looks like: a dashboard with a share-of-voice percentage, a competitor leaderboard, and no methodology page.
Key data: not disclosed by most vendors, which is itself the finding. Given the SparkToro and Conductor results above, a tool that runs each prompt once a week is reporting a coin flip, and chart design does not fix that.
Why it matters: you will be asked to defend this number in a meeting, and “the tool said so” is not a defence.
Ask four questions before you buy. How many runs per prompt per period? Stateless sessions, or sessions carrying history and personalisation? Which model versions and regions? And is an appearance a citation with a link, or a brand name in prose with no link? A vendor who answers all four plainly is worth paying; one who redirects you to a feature tour is selling rank tracking with new labels.
Run your own control alongside it. Ten manual runs of three priority prompts takes twenty minutes a month, and where the tool disagrees badly with your own count, the tool is wrong until it proves otherwise. That is the standard we apply to any data source in an SEO audit: nothing enters the report without a way to reproduce it.
What should the monthly report contain instead of a rank?
Five numbers, each from a different layer, each reproducible. Presence, citations, referrals, crawlability and outcome.
What it looks like: a report page a sceptical finance director can check without asking you anything.
Key data: Similarweb put ChatGPT’s share of worldwide generative AI web traffic at around 53% by May 2026, down from roughly 76% in June 2025, with Gemini rising to about 27% and Claude to about 9% — so a report built around one assistant is already out of date.
Why it matters: a single blended “AI visibility score” hides which layer moved, and the layers need different work.
Build it as five rows. Generative AI impressions, from Search Console. Citations and top grounding queries, from Bing Webmaster Tools. AI Assistant sessions and key events, from GA4. Crawler requests and status codes, from your logs. And prompt visibility — the share of runs your brand appeared in, with the run count printed beside it — from your own panel.
Only two of the five deserve a target: prompt visibility and AI Assistant key events. Google publishes no coverage figure for AI Overviews or AI Mode, so the denominator behind your impressions can move without notice, and an impressions goal eventually rewards or punishes you for something you did not do.
Matching the question you are asked to the source that answers it
| What you are asked | Layer that answers it | What it cannot tell you | How to confirm the number |
|---|---|---|---|
| “Are we in Google’s AI answers?” | Generative AI performance report, Search Console | Which queries, which surface, how many clicks | Export the Pages tab and compare with the Search results Pages export |
| “What are people asking when we get cited?” | AI Performance, Bing Webmaster Tools | Anything about Google; the grounding sample is partial | Sort grounding queries by citation count and spot-check three in Copilot |
| “Is any of this making money?” | AI Assistant channel, GA4 | Assistant traffic arriving with no referrer | Check key events fire, then cross-check Session source |
| “Can the AI systems read our site?” | Server access logs | Whether anything was cited or shown | Filter by AI user-agent, group by status code, look for 200s |
| “How visible are we against rivals?” | Your own sampled prompt panel | Why the model chose what it chose | Ten stateless runs per prompt; report frequency, not position |
| “Why did the number move?” | All five read together | Nothing reliable from one layer alone | Check whether crawl status, rankings or coverage changed that week |
What to fix first, and what to stop reporting
Start with the logs. Confirm GPTBot, ClaudeBot, PerplexityBot and Bingbot get 200s on the pages that matter, because nothing downstream works if they are turned away at the door. It is an afternoon, not a quarter of content work.
Then build the prompt panel before you buy a tool. Thirty prompts written the way your customers actually ask, ten stateless runs each, recorded with the date and the model version. That gives you a baseline you own, a way to audit any vendor you later hire, and the only metric here stable enough to carry a target. Everything else — Search Console impressions, Bing citations, GA4 sessions, crawler status codes — is already being collected for you and takes an hour a month to pull.
What to stop: reporting a position inside an AI answer, presenting a blended visibility score with no stated method, and attributing organic clicks to AI Overviews, because Search Console does not separate them and no derivation makes that data appear. A number nobody else can reproduce from your written method does not belong in the report — the same standard that keeps a traffic drop diagnosis honest, and the one behind our view of which AI SEO tactics actually work.
If you want this stack set up and audited on your own properties, the measurement side sits inside our AI SEO and generative engine optimisation work, and the content changes that follow are what our blog marketing team builds into a publishing routine.
About the author
Shabir MS leads SEOValley Solutions and has worked in search since 2005. He has spent two years rebuilding client reporting around AI surfaces, and most of it arguing that a number without a stated method is worse than no number. Read more from Shabir MS.
Frequently asked questions
Can you rank-track your brand inside an AI answer?
No. SparkToro and Gumshoe.ai found the odds of two identical brand lists were under 1 in 100, and roughly 1 in 1,000 for the same order. Track appearance frequency instead.
What is the generative AI performance report in Search Console?
A dedicated Search Console view of how often your pages appeared in Google’s AI Overviews and AI Mode. Google announced it on 3 June 2026 and rolled it out worldwide on 31 August 2026.
Does the generative AI report show clicks?
No. It exposes impressions only, grouped by page, country, device and date. Clicks from AI features stay inside the standard Search results report under the Web search type.
Can you see which queries triggered an AI Overview appearance?
Not in Google. The generative AI performance report has no query dimension, so you can see which of your pages appeared but not what was asked to produce the appearance.
What does Bing Webmaster Tools AI Performance report?
Total citations, average cited pages per day, per-URL citation counts and a sample of the grounding queries Copilot used. Microsoft opened it in public preview on 10 February 2026.
How do you see ChatGPT and Gemini traffic in GA4?
Use the AI Assistant default channel in Traffic acquisition. Google defines it as arrivals from sources including ChatGPT, Gemini, DeepSeek, Copilot and Grok, matched on medium or referrer.
Why is AI assistant traffic under-counted?
Native mobile apps frequently send no referrer header, so those sessions land in Direct. Cloudflare makes the same point about its crawl-to-refer ratios being overstated for that reason.
Do Google rankings predict AI Overview citations?
Partly. Ahrefs found 76.10% of cited pages ranked in the top 10, but 14.40% did not rank in the top 100 at all, so ranking raises the odds without guaranteeing anything.
How many runs does a prompt need before the number is usable?
Ten stateless runs per prompt is a workable floor. Report the percentage of runs your brand appeared in and print the run count beside it so the figure can be reproduced.
Which AI metrics should carry a target?
Prompt visibility across a fixed prompt set, and key events on the GA4 AI Assistant channel. Impressions make a poor target because Google publishes no coverage figure for the denominator.
