WashU Audit: 11% of AI Overview Claims Lack Source Support (B2B SaaS Playbook)

Google can cite a credible page and still put a wrong, or missing, fact next to your brand.

That is the operator takeaway from WashU’s Sept 29 release of Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact. Covered same day by WashU The Source and phys.org, the paper is headed to the October 2026 ACM Internet Measurement Conference.

For B2B SaaS teams that treat “we got cited in AI Overviews” as proof the answer is clean, this study separates three layers most dashboards still smash together: activation, source selection, and claim fidelity.


What WashU measured

Haofei Xu, Umar Iqbal, and Jacob M. Montgomery ran a Puppeteer crawler on US Google Trends queries across 19 categories (March 13-April 21, 2026).

For each query they captured:

  1. Whether an AI Overview appeared
  2. Every embedded citation
  3. The co-displayed first-page SERP
  4. Body text and ad presence on cited pages

They then split each overview into atomic claims and checked each claim against the cited pages with an LLM verifier, validated against human annotators (high agreement on failure labels).

Caveat: This is an independent academic audit, not a GEO vendor study. Still read it as a US trending-query sample, not a B2B SaaS commercial-prompt census.


Finding 1: Question buyers live inside AI Overviews

  • All queries (n=55,393): 13.7%
  • Question-form queries: 64.7%
  • Non-question queries: 9.5%
  • How… / Why… (top open interrogatives): 84.3% / 73.4%

Topic rates swing hard too: hobbies/leisure 46.1%, science 39.9%, business and finance 26.2%, politics only 7.5%. Google’s trigger logic is selective and opaque.

Operator read: your buyers do not type “CRM software.” They type full questions. Those are exactly the queries WashU shows AIOs dominate. Classic short-head SEO traffic and AI Overview exposure are different surfaces.


Finding 2: Credible citations, different shortlist than the SERP

Across scored domains, AIO citations averaged higher credibility (PC1 0.732) than co-displayed first-page results (0.645). AIOs also used less unvetted UGC (14.2% of AIO refs vs 41.4% of first-page URLs).

But selection is not “re-rank the blue links”:

  • 29.8% of AIO-cited domains never appear on that query’s first page
  • Median 8 references per overview
  • Off-page AIO refs were more credible on average than on-page AIO refs

That lines up with DerivateX’s own B2B software snapshot in the Google AI Overview vs SERP overlap report (about 35% URL overlap with top-10 organic). WashU’s academic sample is larger and trending-query based; the structural point is the same: winning organic rank is not the same system as winning the AIO citation set.


Finding 3: Citation is not claim fidelity

This is the finding that should change your weekly GEO ritual.

Of 98,020 claims:

  • 89% consistent (clear + vague support)
  • 11% inconsistent: about 7% omitted (not in cited text), about 2.7% incorrect (contradicted), about 1.4% ambiguous (sources conflict)
  • Only 41.9% of overviews were fully grounded on the text collected
  • Source quality and claim fidelity were nearly independent (paper reports r about 0.045)

“Google’s summaries generally draw on credible websites, which is good news. But quality sources are not enough.”

That line, from co-author Jacob Montgomery in the WashU Source write-up, is the board slide.

The authors also warn: excluding social/video bodies and real-time pages (weather, school closings) means 11% is an upper bound. Even under a generous UGC correction, they still put a residual floor around about 5.3%.

This is sharper than the general LLM citation problem DerivateX already covered in how LLMs decide what to cite (SourceCheckup / Nature Communications ranges for chat engines). WashU is Google AI Overviews specifically, at production scale, with claim-level labels.


Finding 4: Publisher economics (why your content partners care)

More than half of AIO-cited pages (50.6%) showed display ads. Google’s own sponsored ads still appeared on some AIO SERPs, occasionally above the overview. The paper does not measure your SaaS demo loss; it documents the incentive structure that makes third-party listicles and docs fragile if clicks dry up.

Pair that with the click story: AI Overview links opening AI Mode already argued citation is not a session. WashU adds citation is not an accurate sentence.


What this means for B2B SaaS GEO

1. Split named, cited, and correct

Track three lines on the same buyer prompts:

  1. Named in the answer (recommendation share)
  2. Cited as a source chip (citation share)
  3. Correct on price bands, limits, ICP, integrations, security claims

Recommendation share vs citation share is incomplete without a claim-check pass. A wrong feature claim next to your logo is worse than silence.

2. Run claim audits on question-form prompts

Build 15-20 prompts the way buyers ask them (How does X compare to Y for mid-market SaaS?, What are the SSO options for…). For each AIO hit:

  • Screenshot the summary sentences about you and rivals
  • Open every citation chip
  • Mark clear / vague / omitted / contradicted against your current product truth
  • File corrections where a third-party page is the broken grounder

3. Harden the pages models actually quote

WashU’s failure mode is mostly omission: the overview asserts something the cited page never said. Defenses that still work under Google’s “GEO is still SEO” guidance:

  • Put constraints, packaging, and not-for statements in plain HTML near the top of money pages
  • Keep comparison tables factual and dated
  • Make third-party docs and partner pages match your current limits (SIGIR-style gatekeepers still matter; see what gets cited)
  • Treat brand hallucinations as a weekly ops queue, not a one-off PR scare

4. Do not confuse WashU’s 13.7% with your category rate

Their corpus is Google Trends heavy (sports alone is half the queries). Your B2B commercial cluster will not match 13.7%. Use their question-form 64.7% as the directional risk for how buyers actually ask.


Worked example: mid-market analytics SaaS

Priority prompts: best product analytics for B2B SaaS; how does Amplitude compare to Mixpanel for PLG; what is the cheapest session replay with SSO.

Weak reaction: AIOs cite Wikipedia and news; we are fine if we rank.

Better reaction this week:

  1. Run those three prompts in a logged-out US Google session; expand every Overview.
  2. Extract every sentence that names you, a rival, or a pricing/limit claim.
  3. Diff each sentence against your pricing page, docs, and the cited third-party URL.
  4. If the Overview invents a plan limit and cites a stale G2 or blog post, fix the grounder and your own canonical table the same day.
  5. Report to the CMO: naming rate, citation rate, claim-error rate. Three lines. No blended AI visibility vanity number.

How DerivateX uses claim-fidelity audits

  1. Treat AI Overviews as an answer surface that can misstate you even when the chip looks premium.
  2. Keep SERP rank, AIO citation, and claim correctness on separate scoreboards (the AIO vs SERP report plus this WashU fidelity layer).
  3. Prefer first-party and third-party pages that state hard constraints in crawlable text, because omission is the dominant failure mode.
  4. Measure Google, ChatGPT, Perplexity, and Claude separately; WashU is Google-only.

If you need a prompt-level baseline across Google AI Overviews, ChatGPT, Perplexity, Gemini, and Claude, start with a free AI visibility audit. Engagement options are on pricing. For proof patterns, see the Gumlet and REsimpli case studies.


FAQ

Is this the same as DerivateX’s 35% AIO-SERP overlap report?

No. That report is a focused B2B software URL-overlap snapshot. WashU is a 55k-query academic audit covering activation, credibility, claim fidelity, and publisher ads. Overlap findings rhyme; claim fidelity is new evidence.

Should we quote that 11% of AI Overviews are wrong?

No. Quote carefully: about 11% of atomic claims were unsupported in their pipeline, as an upper bound, with omission the main mode. Only 41.9% of overviews were fully grounded on collected text. Do not convert that into one in nine Overviews hallucinate.

Does higher domain credibility fix the problem?

Not according to this paper. Credibility and fidelity were largely independent. Better sources help; they do not make claim checking optional.

Is this a GEO vendor study?

No. WashU academic co-authors; preprint for ACM IMC 2026. No product pitch inside the paper. Still: trending US queries are not your ICP prompt set.

What should we change this week?

Add a claim-fidelity pass to your AI Overview monitoring: screenshot answer sentences, check cited pages, fix grounders and owned pages when product facts drift.

Shivanshi Bhatia
Written byCo-founder, DerivateX

Shivanshi Bhatia is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline. She runs operations and delivery, which means every audit, content brief, and published page ships through a system she built. She owns the client relationship from kickoff through reporting, so clients spend their time on decisions instead of chasing updates. She has worked in SaaS since 2019 and reviews client work before it goes live.

Apoorv Sharma
Reviewed byCo-founder, DerivateX

Apoorv Sharma is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline. He is the author of the 2026 AI Visibility Benchmark Report and the Citation Engineering methodology. He's also the brain behind "Found On AI" and has sold 2 of his companies previously