Case study: Gumlet turned ChatGPT mentions into 20% of inbound revenue. Read it →
AI Citation Overlap Is ~9%: Why Blended Visibility Scores Lie to B2B SaaS
TL;DR
- Fresh 2026 benchmarks put cross-engine AI citation overlap around 9% to 24% of sources (Foglift Q3 mean Jaccard 0.094). Engines rarely cite the same pages.
- On the same prompts, brand agreement runs much higher: roughly 85% to 98% in Foglift’s five-engine panel. Different shelves, similar verdicts on who to trust.
- A single “AI visibility score” averages ecosystems that do not share inputs. It can look healthy while you are invisible on the engine your buyers actually use.
- August 2026 proved the point: Reddit collapsed in ChatGPT citations while other engines barely moved the same way. Blended Reddit metrics hid a ChatGPT-specific crash.
- DerivateX’s stance: report per engine, separate citation vs brand mention, and chase the brand consensus floor. Start with a free AI visibility audit.
Most AI search dashboards still ship one blended number.
That number is lying to you politely.
Across 2026 studies, AI engines share only a thin slice of cited sources on the same buyer prompts, while they often name the same brands. If your program optimizes for “getting cited somewhere,” you will celebrate the wrong wins and miss engine-specific collapses.
This is DerivateX’s operator read of the newest cross-engine overlap data, what a blended score hides, and how B2B SaaS teams should rebuild reporting this month.
What the September 2026 data says
Nobori’s synthesis of Q3 2026 citation-overlap research centers on a Foglift benchmark: 75 buyer questions, five engines, 375 answers, 1,510 distinct cited domains, mean pairwise Jaccard overlap 0.094.
Other large 2026 datasets land in the same band:
- Writesonic (May to June 2026): pairwise source overlap roughly 0.119 to 0.237 across ChatGPT, Gemini, Perplexity, and Google AI Overviews
- SurfacedBy (~16,400 answers): only 2.7% of cited domains appeared on all five engines; about 70% appeared on exactly one
No serious study puts source overlap comfortably above a quarter.
Brand agreement is the other half of the story. In Foglift’s panel, brand-mention agreement between engine pairs ran 85.5% to 98.4% while source overlap on some pairs fell as low as 0.027. Engines read different shelves and still converge on who belongs on the shortlist.
That split is the useful number. Source wins do not travel. Brand consensus does.
Why citation volume per engine makes blended scores worse
Engines do not even offer the same number of slots.
SurfacedBy’s five-engine sample (summarized in the same Nobori write-up) put approximate sources per answer at:
| Engine | Sources per answer (approx.) |
|---|---|
| Gemini | 11.0 |
| Perplexity | 8.6 |
| Google AI Mode | 7.8 |
| Claude | 6.8 |
| ChatGPT | 3.7 |
A brand cited eighth on Gemini and absent from ChatGPT can still post a “healthy” blended citation rate. You are averaging an 11-slot contest against a 3-to-4-slot contest.
DerivateX already separates recommendation share vs citation share. Overlap data is why that split has to be per engine, not one rollup.
August 2026: the blended-score failure case
In mid-August 2026, Reddit’s share of ChatGPT Search citations collapsed within days (industry tracking summarized via Search Engine Land / Promptwatch coverage in the Nobori piece), while Google surfaces drifted slowly and Perplexity did not mirror the same crash.
If your dashboard said “Reddit share of AI citations” on August 7, it looked stable. By August 14 it was a ChatGPT-specific retrieval change, with fan-out site: queries jumping and government/regulatory pages absorbing attention in ChatGPT answers.
One blended Reddit metric would have told you almost nothing actionable. Five engine columns would have.
What this means for B2B SaaS operators
1. Kill the single AI visibility score
Report four things per engine:
- Citation presence (did they link your domain?)
- Brand mention (did they name you, cited or not?)
- Source-type mix behind answers that include you (owned, review, editorial, community, video, institutional)
- Stability over weeks on a fixed prompt set
Then add one cross-engine layer: share of priority prompts where you appear on at least four of five engines. That is the brand consensus floor the data says is achievable for brands that clear it.
2. Stop treating a Perplexity win as a ChatGPT strategy
ChatGPT sits at the bottom of many pairwise source-overlap tables. Claude behaves like its own channel (near-zero YouTube/Reddit in SurfacedBy’s sample, stronger pull from docs and prestige editorial). Google AI Mode and Perplexity share more web-retrieval DNA with each other than either does with ChatGPT.
Your buyer panel should weight the engines your ICP actually uses, not the engine that is easiest to screenshot.
3. Build brand-layer evidence, not one-source bets
Source-level wins rot when retrieval policy changes. Brand-layer signals (entity consistency, editorial breadth, review depth, constrained comparison evidence) are what travel when engines disagree on pages but agree on brands.
That is the same logic behind DerivateX’s off-site and third-party work, and why directory/list presence still mattered in yesterday’s supplier-prompt study lane: engines need citable evidence, but not the same URL every time.
4. Version every run
Log engine, model, week, and prompt set. OpenAI model defaults and fan-out policy can move inside a fortnight. An unversioned “we were cited last month” slide is not a measurement system.
Worked example: how a SaaS team should use this next week
Assume leadership wants one AI visibility KPI for the board.
Do not give them one KPI.
Give them a five-column table for 25 money prompts:
- ChatGPT / Perplexity / Claude / Gemini / AI Overviews+AI Mode
- Brand mentioned? Domain cited?
- Top source types when you win or lose
Then one summary line: “Present on 4+ engines for X of 25 prompts.” That is the number that survives retrieval churn.
If ChatGPT is empty and Perplexity is full, fix ChatGPT’s evidence paths (owned pricing, security, constrained comparisons, institutional corroboration). Do not celebrate the Perplexity screenshot.
If brand mention is high and citations are low, you have awareness without proof. Ship the pages and third-party evidence the model can footnoted.
Tie movement to CRM the same way you would for any channel. Gumlet is still the public proof that pipeline beats vanity visibility.
How DerivateX turns this into authority
Other agencies will post “AI engines cite different sources.” Fine. The durable DerivateX angle is harsher:
Low source overlap + high brand agreement means your reporting architecture is the product. Blended scores hide ChatGPT crashes, Claude’s doc bias, and Gemini’s slot abundance. The agency that wins is the one that forces per-engine citation and brand metrics into the operating cadence, then builds evidence for the consensus floor.
That is why the 8 GEO metrics we report exists, and why a free AI visibility audit starts with engine-separated baselines. Live options remain on pricing.
FAQ
What is a normal AI citation overlap score?
Published 2026 studies put pairwise Jaccard source overlap roughly between 0.09 and 0.24. Foglift’s Q3 five-engine mean was 0.094.
Do engines agree on brands more than sources?
Yes. Foglift reported brand agreement around 85% to 98% on matched prompts while source overlap stayed low.
Why is a blended AI visibility score dangerous?
It averages engines with different slot counts, indexes, and source personalities. A Gemini win can mask a ChatGPT loss.
What changed in August 2026?
ChatGPT’s Reddit citation share dropped sharply while other engines did not mirror the same crash, exposing how blended community metrics can hide engine-specific retrieval changes.
What should we report instead?
Per-engine citation presence, brand mention, source-type mix, and weekly stability, plus a cross-engine brand consensus rate on a fixed prompt panel.
How is this different from DerivateX’s recommendation vs citation piece?
Recommendation share vs citation share defines the metric split. This piece uses 2026 overlap studies to explain why that split must be engine-separated.












