What Gets Cited: Four GEO Gatekeepers From 252,000 SIGIR Tests

Most GEO pitches still sell the same before-and-after: bold the answer, add FAQ schema, “chunk” the page, hope ChatGPT notices.

A SIGIR 2026 paper just stress-tested that story at a scale agencies rarely admit exists. Across 252,000 paired trials, the models cared far more about whether the page could answer the question than whether it looked “AI optimized.”

If you run B2B SaaS into AI search, that is useful only if you also respect what the study did not measure. Citation preference after retrieval is not the same problem as getting into the retrieval set. DerivateX already separates those layers in client work. This paper is strong evidence for the second layer only.


What the SIGIR study actually tested

The paper, What Gets Cited: Competitive GEO in AI Answer Engines, is by Rahul Vishwakarma, Shushant Kumar, and Ratnesh Jamidar, accepted at SIGIR 2026 (Melbourne, July 20–24). Sprinklr also published a company summary on July 24, 2026.

Design in plain terms:

Design choiceWhat they did
Competition shapeExactly two candidate documents injected into a simulated RAG tool response
OutcomeWhich source’s URL appears first in the model’s citations
Factors18 content attributes changed one at a time (relevance, price, specs, hedging, structure, timestamps, position, and more)
Scale1,440 matched scenarios × paraphrases × orders × 5 repeats × 6 models = 252,000 trials
ModelsGemini 2.5 Flash, GPT-5 Nano, GPT-5 Mini, GPT-5.2, Claude 3.5 Sonnet, Kimi K2 Thinking
ControlsBrand/publisher anonymization; counterbalanced source order so position bias is measured separately

They fit logistic mixed-effects models so repeated runs of the same scenario do not pretend to be independent samples. That is why the headline numbers are odds ratios after adjusting for order, not raw “win rates” from a spreadsheet screenshot.

Conflict of interest note: all three authors list Sprinklr affiliations, the paper mentions an internal Sprinklr pilot, and Sprinklr markets an AEO product. The venue is peer-reviewed. Still flag the stake when you brief leadership.


The four gatekeepers (and what “position” really means)

Across all six models, four factors produced very large, consistent first-citation advantages:

  1. Topical relevance — an on-topic page crushed an off-topic alternative. Cosmetic edits do not rescue a page that fails the buyer’s actual question.
  2. List position in the supplied context — the source listed first in the tool response was heavily favored. This is not “rank #1 on Google.” It is order inside the documents the model already received.
  3. Explicit price — when one candidate stated a price and the other did not, the priced page won more often in these product-oriented scenarios.
  4. Recent timestamp — a 2026 date beat a 2019 date under the freshness comparisons they ran.

Secondary (still useful, less universal) gains showed up for specs, comparisons, evidence-backed claims, and deeper coverage. Formatting-only changes (dense vs structured layout, scattered vs organized information) were weak or inconsistent.

Sprinklr’s marketing write-up sharpens the operator insult: “Substance moves citations. Layout does not.” The arXiv paper is more careful, but the directional finding matches.


What the study cannot tell a SaaS marketer

Read the limitations before you rewrite every WP template.

  • No live retrieval. The models never crawled the open web for these trials. Upstream rank, domain trust, embeddings, and index coverage are out of scope.
  • Two sources, not ten. Production answer engines often retrieve larger pools. Pairwise preference is not a full multi-document market share model.
  • First citation ≠ business outcome. Second citations, uncited influence, brand mentions without links, and CRM conversions are different events.
  • Product-review seed corpus. Explicit price helped in commercial review scenarios. Do not bolt a fake price onto a compliance explainer and call it GEO.
  • Some reported odds ratios hit quasi-separation (shown as “>10k”). Treat those as “near-deterministic in this setup,” not as a precise multiplier you can put in a board deck.

DerivateX’s public directories / supplier-prompt citation work and AI Overview vs SERP overlap research already show how often off-site lists win before your owned page gets a fair citation fight. This SIGIR paper starts after that fight has already been narrowed to two pages in context.


Operator playbook for B2B SaaS this month

1. Split retrieval work from citation work

If you are absent from the answer entirely, fix crawlability, indexation, third-party list presence, and category evidence first. Do not spend the quarter rearranging H2s on a page that never enters the context window.

2. Audit commercial URLs for decision-useful facts

For every money prompt you care about (pricing, alternatives, “best X for Y”):

  • Does the page answer the exact constraint in the prompt?
  • Is price, packaging, or plan boundary stated in plain text where it is honest to do so?
  • Are specs, limits, and integration requirements complete enough to support a recommendation?
  • Is the page’s date and substance actually current, not just republished?

Those map to the paper’s gatekeepers without pretending formatting is strategy.

3. Kill formatting theater in the GEO backlog

Bolding “quotable” sentences, adding decorative FAQ blocks, and “LLM-friendly” CSS are easy to sell because the before/after screenshot looks busy. In this controlled setup they barely moved first-citation odds. Keep accessibility and parseability. Stop treating them as citation levers.

4. Keep engine-specific measurement

DerivateX’s read of cross-engine citation overlap (and related measurement posts) still stands: a win in one answer engine is not a win in all of them. This SIGIR study measures citation preference inside a fixed two-doc context across models. It does not give you a single blended visibility score to optimize.


Worked example: a CRM SaaS team

Prompt set: “best CRM for a 40-person SaaS sales team that needs Salesforce migration and SOC 2.”

Wrong sprint: rewrite the homepage with chunked paragraphs and schema flourishes, then ask whether ChatGPT “likes” the new HTML.

Better sprint:

  1. Pull 20 paraphrases of that buyer situation across AI Overviews, AI Mode, ChatGPT, and Perplexity.
  2. Log whether you are retrieved/cited at all. If not, prioritize listicle/directory presence and comparison evidence that engines already trust (same retrieval layer DerivateX already treats as primary for many supplier prompts).
  3. Where you are in the answer but lose to a rival, open both pages. Check topic match to the constraint, explicit packaging/price honesty, migration/SOC proof, and publish dates.
  4. Ship the missing decision facts on the owned URL. Re-test the same prompt panel before celebrating a template change.
  5. Reconcile named/cited sessions to CRM opportunity influence, not only to GSC generative impressions.

That sequence respects both layers the paper forces you to name: get into the context, then win the citation with substance.


How DerivateX reads the paper

Agencies that only sell “AEO formatting packages” just lost their cleanest peer-reviewed excuse.

The durable read:

  1. Citation competition has gatekeepers (relevance, decision facts, freshness, context order).
  2. Context order is mostly an engine/retrieval artifact, not a WordPress theme setting.
  3. Formatting is a weak lever after retrieval.
  4. Live AI search still needs the retrieval layer DerivateX already measures with clients: third-party lists, directories, comparison evidence, and engine-specific prompt panels.

If you want a baseline on whether your commercial pages look like grounding material or decoration, start with a free AI visibility audit. Engagement options stay on pricing. For how we separate the metrics that usually get blended into one vanity score, see the 8 GEO metrics we report.


FAQ

What is the SIGIR “What Gets Cited” study?

A 2026 peer-reviewed paper that ran 252,000 controlled two-document RAG trials across six LLMs to measure which single content factor makes a source more likely to receive the first citation.

What are the four citation gatekeepers?

Topical relevance, position in the supplied context list, explicit price information (in the product scenarios tested), and a recent timestamp. These were consistent across all six models.

Does this mean schema and formatting never matter?

No. The study largely bypassed crawling and retrieval. Structure can still affect accessibility and how systems parse pages before a document enters an LLM context. Inside this post-retrieval testbed, formatting-only edits barely moved first-citation odds.

Is “position” the same as ranking #1 on Google?

No. Here position means which of the two injected documents appeared first in the model’s tool/context list. Upstream ranking systems decide which pages enter that list in production.

Should we trust Sprinklr’s summary as much as the paper?

Use the arXiv paper as the primary source. The Sprinklr blog is a vendor summary by the same authors and ties into a commercial AEO product. Flag that when you brief stakeholders.

How should B2B SaaS teams act on this without overreacting?

Prioritize relevance and decision-useful facts on commercial URLs, keep freshness real, treat retrieval/SEO as a separate workstream, and stop budgeting formatting theater as if it were citation strategy.

Apoorv Sharma
Written byCo-founder, DerivateX

Apoorv Sharma is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline. He is the author of the 2026 AI Visibility Benchmark Report and the Citation Engineering methodology. He's also the brain behind "Found On AI" and has sold 2 of his companies previously

Shivanshi Bhatia
Reviewed byCo-founder, DerivateX

Shivanshi Bhatia is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline. She runs operations and delivery, which means every audit, content brief, and published page ships through a system she built. She owns the client relationship from kickoff through reporting, so clients spend their time on decisions instead of chasing updates. She has worked in SaaS since 2019 and reviews client work before it goes live.