SourceCited / Phase 4 · Measure

4  Phase 4

Measure whether you're actually getting cited

Rank trackers can't see AI answers. Measure share of voice with a fixed query set, track citation rate over time, and reconcile AI referrals that hide in Direct traffic.

Answer first

Build a fixed set of buyer prompts, run them across ChatGPT, Perplexity, Gemini and AI Mode, and have Claude score whether you're mentioned, linked, or cited. Track the trend — citations churn 40–60% week to week — and build a GA4 channel group for AI referrers, since much AI traffic hides in Direct.

On this page
  1. Baseline share of voice
  2. Catch hidden AI referrals
  3. Scoring rubric
  4. The ghost zone
  5. Measurement cadence

Baseline your share of voice

You can't improve what you don't measure, and traditional rank trackers don't see AI answers. Build a fixed set of the prompts your buyers actually ask, run them across the engines, and record whether — and how — you're mentioned. That baseline is your scoreboard. Because citations churn 40–60% week to week — a University of St. Gallen study measured about 65% of cited sources changing from one day to the next — track the trend, never a single run, and always score per platform rather than blending engines into one number: overlap between them is tiny.

prompt · Build a query set and score answers for share of voice
Help me measure my AI share of voice.
1. From my topic/products (below), generate 25 prompts real buyers would type into ChatGPT,
   Perplexity, Gemini, and AI Mode — a mix of category, comparison, and problem queries.
2. I'll run them and paste the answers back. For each answer, score: is my brand mentioned?
   linked? cited as a source? which competitors appear? what sources did the engine cite?
3. Summarize my visibility %, share of voice vs named competitors, and the 5 sources most often
   cited that I don't yet appear on — those become off-page targets.

Catch AI referrals hiding in your analytics

AI traffic under-reports. Many AI referrals land as Direct or (none), and name-only tracking misses "ghost citations" where an engine cites your URL without naming your brand. Build a referrer channel group for the AI sources and add a self-reported "how did you hear about us" field to triangulate.

MetricHow to get it
Share of voiceYour fixed query set, scored per engine (prompt above)
Citation rate% of your query set where you're cited as a source, tracked over time
AI referral trafficGA4 channel group matching ChatGPT / Perplexity / Gemini / Copilot referrers
Crawler hitsServer/CDN logs — confirm retrieval bots fetch you and get 200s, not 403/CAPTCHA
Freshness on Bing-fed enginesSubmit new/updated URLs via IndexNow / Bing Webmaster — Perplexity and Copilot lean on Bing's index

Score with one rubric, per platform

Free-form notes rot. Score every answer on the same four-point scale, log the same columns, and keep engines separate — with ~11% overlap between platforms, a blended number is fiction.

ScoreMeaning
0Absent — not mentioned, not sourced
1Mentioned, but not linked or cited
2Cited or linked as a source
3Recommended first, or featured as the answer

Alongside the score, log sentiment, an accuracy flag (misquotes, wrong prices), competitors named, and sources cited. Average two or three runs per checkpoint — daily churn is ~65% — and chart each platform on its own line.

prompt · Weekly scoring pass
Here are this week's answers for my fixed prompt set, labeled by platform and prompt. Build a table: platform | prompt | score 0-3 (0 absent, 1 mentioned, 2 cited/linked, 3 recommended first) | sentiment | accuracy issues | competitors named | sources cited. Then compare against the baseline table I'm pasting: the three biggest movers up and down, and any factual errors about my brand I should correct at the source.

Referral numbers are a floor — find the ghost zone

In GA4, a custom channel group catching referrer sources that match chatgpt|openai|perplexity|copilot|gemini|claude|poe captures the visible slice of AI traffic — typically around 1% of sessions, converting several multiples better. Treat it as a floor, not a census: assistant apps often strip referrers, so real usage lands in Direct.

The invisible slice is measurable from the other side. Pull thirty days of server logs, filter for OAI-SearchBot, ChatGPT-User and PerplexityBot hits on your money pages, and compare against referral sessions to those same pages. A big gap means your content is feeding answers without attribution clicks — the 40–60% ghost zone — which is an argument for stronger in-answer branding (named data, named frameworks), not despair.

Set a measurement cadence

  1. Weekly: re-run your top-priority prompts; note movement, expect churn.
  2. Monthly: full query set + GA4 AI-channel review + crawler-log check.
  3. Quarterly: re-verify every stat and crawler name, and re-audit (back to Phase 0).
NW
OSINT researcher and founder; runs a network of ranked content & tools sites
Updated July 21, 2026

Questions people actually ask

Why do I need a special tool to measure AI visibility?

Because classic rank trackers only see the blue links, not what an engine writes in its answer. Measuring AI share of voice means running real prompts and reading the generated answers — which Claude can score in batches.

Why does my AI referral traffic look so small in GA4?

Two reasons: AI referrals often arrive as Direct/(none), and 'ghost citations' name your URL without your brand, so name-only tracking misses them. A referrer channel group plus a self-reported source field closes most of the gap.

How much citation churn is normal?

A lot — studies put week-to-week citation turnover around 40–60%, and only a fraction of brands stay visible across repeated runs. Single spot-checks are nearly useless; measure trends.

How often should I run my prompt set?

Weekly for the priority prompts, monthly for the full set — and average two or three runs each time, because ~65% of cited sources change day to day. Freeze the prompt wording itself: edit the set and you break the trend line, which is the only thing that means anything.

What counts as a good AI share of voice?

There is no universal benchmark. Compute it as your citations divided by all brand citations across your category prompt set, per platform, and judge the trend and the gap to the leader. Early on, moving from absent to intermittent is the real win.