SourceCited / Phase 4 · Measure
4 Phase 4
Rank trackers can't see AI answers. Measure share of voice with a fixed query set, track citation rate over time, and reconcile AI referrals that hide in Direct traffic.
Build a fixed set of buyer prompts, run them across ChatGPT, Perplexity, Gemini and AI Mode, and have Claude score whether you're mentioned, linked, or cited. Track the trend — citations churn 40–60% week to week — and build a GA4 channel group for AI referrers, since much AI traffic hides in Direct.
You can't improve what you don't measure, and traditional rank trackers don't see AI answers. Build a fixed set of the prompts your buyers actually ask, run them across the engines, and record whether — and how — you're mentioned. That baseline is your scoreboard. Because citations churn 40–60% week to week — a University of St. Gallen study measured about 65% of cited sources changing from one day to the next — track the trend, never a single run, and always score per platform rather than blending engines into one number: overlap between them is tiny.
Help me measure my AI share of voice. 1. From my topic/products (below), generate 25 prompts real buyers would type into ChatGPT, Perplexity, Gemini, and AI Mode — a mix of category, comparison, and problem queries. 2. I'll run them and paste the answers back. For each answer, score: is my brand mentioned? linked? cited as a source? which competitors appear? what sources did the engine cite? 3. Summarize my visibility %, share of voice vs named competitors, and the 5 sources most often cited that I don't yet appear on — those become off-page targets.
AI traffic under-reports. Many AI referrals land as Direct or (none), and name-only tracking misses "ghost citations" where an engine cites your URL without naming your brand. Build a referrer channel group for the AI sources and add a self-reported "how did you hear about us" field to triangulate.
| Metric | How to get it |
|---|---|
| Share of voice | Your fixed query set, scored per engine (prompt above) |
| Citation rate | % of your query set where you're cited as a source, tracked over time |
| AI referral traffic | GA4 channel group matching ChatGPT / Perplexity / Gemini / Copilot referrers |
| Crawler hits | Server/CDN logs — confirm retrieval bots fetch you and get 200s, not 403/CAPTCHA |
| Freshness on Bing-fed engines | Submit new/updated URLs via IndexNow / Bing Webmaster — Perplexity and Copilot lean on Bing's index |
Free-form notes rot. Score every answer on the same four-point scale, log the same columns, and keep engines separate — with ~11% overlap between platforms, a blended number is fiction.
| Score | Meaning |
|---|---|
| 0 | Absent — not mentioned, not sourced |
| 1 | Mentioned, but not linked or cited |
| 2 | Cited or linked as a source |
| 3 | Recommended first, or featured as the answer |
Alongside the score, log sentiment, an accuracy flag (misquotes, wrong prices), competitors named, and sources cited. Average two or three runs per checkpoint — daily churn is ~65% — and chart each platform on its own line.
Here are this week's answers for my fixed prompt set, labeled by platform and prompt. Build a table: platform | prompt | score 0-3 (0 absent, 1 mentioned, 2 cited/linked, 3 recommended first) | sentiment | accuracy issues | competitors named | sources cited. Then compare against the baseline table I'm pasting: the three biggest movers up and down, and any factual errors about my brand I should correct at the source.
In GA4, a custom channel group catching referrer sources that match chatgpt|openai|perplexity|copilot|gemini|claude|poe captures the visible slice of AI traffic — typically around 1% of sessions, converting several multiples better. Treat it as a floor, not a census: assistant apps often strip referrers, so real usage lands in Direct.
The invisible slice is measurable from the other side. Pull thirty days of server logs, filter for OAI-SearchBot, ChatGPT-User and PerplexityBot hits on your money pages, and compare against referral sessions to those same pages. A big gap means your content is feeding answers without attribution clicks — the 40–60% ghost zone — which is an argument for stronger in-answer branding (named data, named frameworks), not despair.
Because classic rank trackers only see the blue links, not what an engine writes in its answer. Measuring AI share of voice means running real prompts and reading the generated answers — which Claude can score in batches.
Two reasons: AI referrals often arrive as Direct/(none), and 'ghost citations' name your URL without your brand, so name-only tracking misses them. A referrer channel group plus a self-reported source field closes most of the gap.
A lot — studies put week-to-week citation turnover around 40–60%, and only a fraction of brands stay visible across repeated runs. Single spot-checks are nearly useless; measure trends.
Weekly for the priority prompts, monthly for the full set — and average two or three runs each time, because ~65% of cited sources change day to day. Freeze the prompt wording itself: edit the set and you break the trend line, which is the only thing that means anything.
There is no universal benchmark. Compute it as your citations divided by all brand citations across your category prompt set, per platform, and judge the trend and the gap to the leader. Early on, moving from absent to intermittent is the real win.