The AI Search Funnel: 6 Metrics From Prompt Impression to Closed-Won Revenue

The AI search funnel is a six-stage measurement model that tracks a buyer from prompt impression to closed-won revenue. DerivateX measures prompt coverage, share of voice, recommendation rate, identified AI visits, AI-evidenced leads, and sourced versus influenced revenue. With 64% of marketing leaders unsure how to measure AI search, each stage needs a fixed denominator and a stated confidence tier.

  • Most AI search reporting fails because the denominator moves. If your prompt set changes month to month, every percentage you report is noise.
  • Only three of the six stages are directly observable. Stages four through six carry known leakage, so they need confidence tiers, not false precision.
  • Stage-to-stage conversion is the useful signal. Whether mentions convert into recommendations, and recommendations into citations, tells you something a raw mention count never will.
  • Self-reported attribution is not a fallback. It is the only source that captures the buyer who read a ChatGPT answer, typed your brand into Google two weeks later and arrived as direct traffic.
  • This model can prove exposure, direction and correlation. It cannot prove causation, and DerivateX does not present it as if it can.

What is the AI search funnel?

The AI search funnel is the chain of measurable events between a buyer typing a prompt into an assistant and revenue landing in your CRM. It has six stages, and DerivateX treats each one as a ratio rather than a count, because counts move with your prompt list and ratios do not.

This matters more for AI search than for classic organic because the surface is sampled, not crawled. You cannot pull a complete impression log from an assistant the way you can from Google Search Console. You construct your own denominator by defining a prompt set, running it on a schedule and recording what comes back. Skip that step and you end up reporting mentions, and a mention count with no denominator is unfalsifiable. The vocabulary shift here is real, and our prompt-based search definition covers why buyer prompts behave differently from keywords.

Here is the full model. Read the numerator and denominator columns carefully, because that is where most B2B SaaS reporting quietly breaks.

StageMetricNumeratorDenominatorConfidence tier
1Prompt coverageAnswer runs that return a named vendor setAll answer runs in the tracked prompt setA: directly observed
2AI share of voiceRuns in which your brand is namedAll vendor-returning runs (or all brand mentions, for share)A: directly observed
3Recommendation rate and citation rateRuns where you sit in the recommended set / runs where your domain is citedAll vendor-returning runsA: directly observed
4Identified AI visitsSessions with an assistant referrer or AI campaign parameterAll sessions in the periodB: observed with leakage
5AI-evidenced leadsLeads with an AI referrer touch or a self-reported AI answerAll identified AI visits, plus self-report-only leadsB and C combined
6AI-sourced and AI-influenced revenueClosed-won revenue with an AI first touch / any AI touchAll closed-won revenue in the periodC: modelled and self-reported

DerivateX reports all six every month for retained accounts, and reports the tier alongside the number. A CMO who sees tier C on the revenue line stops treating it as a rounding-proof figure and starts treating it as a directional read, which is the right posture. We cover the individual reporting metrics in more depth in our breakdown of the 8 GEO metrics we report to clients, so this page stays focused on how the stages connect.


How do you define the numerator and denominator at each stage?

You define them once, version them, and refuse to change them mid-quarter. DerivateX writes the definitions into a one-page spec before the first measurement run, because a model anyone can silently redefine is not auditable.

Stage 1, prompt coverage. The denominator is your prompt set multiplied by engines multiplied by runs. The numerator is the subset of runs where the assistant actually names vendors. A meaningful share of commercial-sounding prompts return an educational answer with no shortlist at all, and ignoring that understates your performance on the prompts that carry buying intent.

Stage 2, AI share of voice. Two ratios, not one. Mention rate is your runs over addressable runs. Share of voice is your mentions over all brand mentions in those same runs. The second is the competitive number and the one your board will ask about. Assistants typically surface 3 to 4 brands per category query, so every point of share you gain comes out of a rival’s slot in a very short list, which is why DerivateX trends share of voice against named competitors rather than against an absolute target.

Stage 3, recommendation rate and citation rate. Being named in passing is not the same as being recommended, and DerivateX scores them separately. Recommendation means you appear in the shortlist with a reason attached. Citation means your domain is linked as a source. These decouple more than people expect: 28% of ChatGPT-cited pages have zero organic Google visibility, and 80% of URLs cited by ChatGPT and Perplexity are not in Google’s top 100. You can be recommended constantly with almost no citations of your own domain, because the recommendation is sourced from third-party pages.

Stage 4, identified AI visits. The numerator is sessions you can positively tie to an assistant. OpenAI documents that ChatGPT search referrals can carry utm_source=chatgpt.com, which gives you a clean parameter to filter on in Google Analytics 4 alongside assistant referrer hostnames. The honest part is what this misses: copied links, mobile app sessions with stripped referrers, and the buyer who reads an answer and searches your brand name later.

Stage 5, AI-evidenced leads. Two paths into the numerator. Path one is a lead whose session history contains an AI referrer inside your lookback window. Path two is a lead who tells you, in an open-text “how did you first hear about us” field, that an assistant recommended you. Path two catches the dark portion of path one, which is why DerivateX insists on the form field before the first month of reporting.

Stage 6, revenue. Numerator is closed-won revenue carrying an AI touch. Denominator is all closed-won revenue. Split it into sourced and influenced, which the next section handles.


Which data sources are needed to measure the AI search funnel?

Five, and each covers a blind spot in another. DerivateX will not start a measurement engagement without at least four of the five in place, because a funnel missing a layer produces a number that cannot be defended in a board review.

  1. Sampled prompt monitoring. Fixed prompt set, four or more runs per prompt per engine, memory and personalization off, fixed locale, same weekly window. Run the set across ChatGPT, Google AI Overviews, Perplexity and Claude, and record the full answer text and every cited URL, not a boolean. Our comparison of the best GEO tools for B2B SaaS covers which platforms handle multi-engine sampling properly.
  2. Google Analytics 4 or your analytics of record. Assistant referral hostnames plus the campaign parameter filter. Build one exploration and one audience, not a separate report per engine.
  3. Server or edge logs. AI crawler and fetcher user agents tell you which pages the retrieval layer is reading. This is a leading indicator that sits before stage 1, not a funnel stage itself.
  4. Your CRM of record. In HubSpot, Salesforce or whatever you run, you need two custom fields on the contact or deal record: an AI touch flag and a self-reported source string. Without fields you cannot roll stage 5 into stage 6.
  5. Self-reported attribution on every form. One open-text question. It is the only source that survives referrer stripping.

A sixth layer is optional and only worth it above a certain reporting complexity: a marketing data platform that centralizes everything for the warehouse. Funnel positions itself as a marketing intelligence platform with 600 or more connectors, and its published plan comparison lists BigQuery and Amazon S3 exports on the Business tier and Snowflake on Enterprise. That is genuinely strong at what it does, and if eleven data sources will not reconcile, it is the right purchase. It does not observe AI answers, and it is not trying to. If your problem is that you do not know whether ChatGPT recommends you, no data platform will tell you, because the answer text was never in your stack to begin with.


What does a worked example look like for a $12M ARR B2B SaaS company?

Every figure below is a rounded hypothetical built to show the shape of the model. None of it is a DerivateX client result, and no client is described by it. Setup: 150 tracked buyer prompts, four engines, four runs each, so 2,400 answer observations per month. The value sits in the conversion column, not the volumes.

StageIllustrative volumeIllustrative rateIllustrative conversion from prior stage
Answer observations2,400DenominatorBaseline
1. Vendor-returning runs~1,900About 80% prompt coverage80%
2. Runs naming the brand~450About 24% mention rate, about 13% share of voice24%
3. Runs recommending the brand~270About 14% recommendation rate60% of mentions
3b. Runs citing the domain~100About 5% citation rate37% of recommendations
4. Identified AI visits~1,200 sessionsRoughly 2% of all sessionsNot a clean ratio, see note
5. AI-evidenced leads~60 (45 referrer, 15 self-report only)About 5% of identified AI visitsRoughly a third more than referrer data alone
6. AI-influenced pipeline~20 opportunities, about $600,000 at a $30,000 ACVAbout a third of leads to opportunity~12 sourced, about $360,000
6b. AI-influenced closed-won~4 deals, about $120,000About a 20% win rateTier C confidence

Read the interesting rows. Mention to recommendation converts at 60% in this model: when the company gets named, it usually gets endorsed. Recommendation to citation converts at 37%, which tells you the assistants are corroborating this brand through third-party pages rather than its own site. That is a content and off-site instruction, not a technical one.

Stage 4 deliberately has no clean conversion ratio from stage 3, and DerivateX says so out loud rather than inventing one. You cannot divide sessions by answer runs, because one run is not one buyer and you have no impression volume. What you can do is trend both lines side by side over six months and check whether they move together. Correlation across two independent data sources is a genuine finding. A ratio between them is arithmetic theater.

One approved number is worth keeping in view: AI-sourced visitors convert at 4.4x the rate of other traffic. Judged on volume, that channel looks negligible next to organic. Judged on conversion it is the strongest acquisition path in the model, which is the argument you take into the budget conversation.


How do you separate AI-sourced pipeline from AI-influenced pipeline?

Sourced means the AI touch is the earliest identified touch on the record. Influenced means an AI touch appears anywhere in the deal timeline before closed-won. DerivateX reports both, always as two separate lines, because collapsing them is the most common way AI search reporting loses credibility with a CFO.

DefinitionRuleUse it forKnown weakness
AI-sourcedFirst identified touch is an assistant referrer or the self-reported source names an assistantChannel investment decisions, new logo acquisitionUnderstates reality, since the earliest AI touch is usually invisible
AI-influencedAny AI touch on any contact on the deal within the lookback windowContent and citation strategy, showing consideration-stage impactOverstates causation, since presence is not persuasion
AI-corroboratedSelf-reported source names an assistant AND a referrer touch existsThe high-confidence subset you quote when challengedSmall sample, so it moves erratically month to month

DerivateX added the third row after too many meetings where a marketing leader was asked to defend the influenced number and had nothing narrower to fall back on. AI-corroborated deals are the ones where two independent sources agree. There will be few of them. Their job is to establish that the mechanism is real, after which the influenced number becomes credible as a scale estimate rather than a claim.

Set a lookback window and write it down. DerivateX typically sees a 45 to 90 day sales cycle across its accounts in the $5M to $50M ARR band, and defaults to a 120-day window on that basis. Shorter windows undercount assistant-led research, which happens early. Longer windows start capturing coincidence. The number matters less than the fact that it is fixed and disclosed.

Gumlet shows what the top of this table looks like when the work compounds: more than 20% of monthly inbound revenue attributed to AI discovery. That is an attributed figure from a running measurement model, not a projection, and it is why DerivateX treats attribution plumbing as part of the engagement rather than a reporting afterthought. Our write-up on how a GEO agency connects citations to revenue walks through the CRM side in more detail.


What can this model prove, and what can it not prove?

DerivateX puts this table in the appendix of every quarterly review. It is the fastest way to end a debate about whether the numbers are real, because it concedes the limits before anyone has to attack them.

This model can proveThis model cannot prove
Your exposure across a defined prompt set, at a stated sample sizeTotal AI impression volume, since no engine publishes it
Directional change in mention, recommendation and citation rates over timeThat a specific citation caused a specific deal
Which competitors occupy the recommended set, and where they are corroborated fromWhy an engine changed its answer in a given week
That identified AI traffic exists, converts, and at what rateThe size of the unidentified portion, only that it exists
Correlation between prompt-level gains and pipeline movement across quartersCausation between them, at any sample size you will realistically have
Whether your reported figures reconcile with CRM recordsIncrementality without a holdout, which is rarely practical here

One nuance on volatility. Assistant answers vary run to run, so a single-run measurement is close to worthless. DerivateX reports the median across runs and flags any prompt whose result changed in more than half of runs as unstable rather than improved. A tracking tool reporting one run per prompt per week will show you movement that is sampling variance, and that is how measurement programs lose the room.


How do you report the AI search funnel to leadership?

One slide, six numbers, three confidence tiers, and a stated denominator. DerivateX builds the reporting pack around what a board member can interrogate in ninety seconds, not around what a dashboard can render.

The slide carries: prompt set size and version, share of voice with last quarter’s figure beside it, recommendation rate, identified AI visits, AI-evidenced leads, and the sourced and influenced revenue lines with tiers marked. Then one sentence naming the biggest prompt cluster you do not yet appear in, and what is being built against it this quarter.

Two internal frameworks earn their place here. The AI Visibility Score, or AVS, is the weighted composite DerivateX uses to roll stages 1 through 3 into a single trackable index, weighted toward recommendation and away from passing mentions. The Citation Surface Map is the inventory of third-party pages, listicles, review sites and community threads that assistants actually pull from for your category, which is what makes stage 3 actionable instead of merely observable. AVS is the number you trend on a slide; the Citation Surface Map is the document that tells your team where to work next. Reddit accounts for 46.7% of Perplexity’s top sources, which is usually the first surprise a map surfaces for software firms who have only ever invested in their own blog. Our B2B SaaS AI citation study documents those source patterns across categories.

Cadence beats depth. Monthly funnel report, quarterly model review where definitions can be changed and versioned, annual prompt set rebuild. If you change the prompt set in month five, restate the prior months against both sets or your trend line is fiction.


Where this model is the wrong tool

DerivateX would rather tell you this up front than three months into a retainer. The six-stage funnel assumes a category where buyers actually ask assistants for vendor recommendations, and enough deal volume for stage 6 to produce more than single digits.

Three situations where you should not build it yet. First, as a DerivateX rule of thumb, if you close fewer than about twenty deals a quarter, stages 5 and 6 will be too sparse to trend and you should stop at stage 3. Second, if your category returns educational answers rather than vendor shortlists, which in DerivateX’s experience shows up as prompt coverage below roughly 40%, your problem is category definition and not measurement. Third, if your analytics and CRM are not reconciled today, fix that first, because an AI funnel bolted onto broken plumbing will produce a new set of numbers nobody trusts.

There is also a genuine case for buying software instead of an agency. If you have a data team, a defined prompt set and someone who will own the weekly runs, a tracking platform plus internal analyst time is cheaper than a retainer and will get you stages 1 through 4. For plenty of SaaS teams that is the right first move. DerivateX is the better call when the measurement model needs to be built and the citation surface needs to move, because reporting the gap every week and changing nothing is the failure mode most teams have already lived through. Citation Engineering is the methodology DerivateX uses to make assistants recommend a brand deliberately, and it is the half of the engagement that measurement alone cannot substitute for.

On cost, DerivateX publishes its numbers. On current published pricing, the entry engagement is a $5,000 retainer plus $1,000 to $1,200 of off-site budget, so $6,000 to $6,200 all in, on a 90-day pilot with no lock-in afterward, and full details sit on the DerivateX pricing page. Own Your Category runs $9,500 to $10,000 all in, and a standalone two-week Diagnostic is $3,500, credited in full against month one if you convert within thirty days. Below roughly $5M ARR the maths on a retainer usually does not work, and we say so.


Frequently asked questions

How do I measure the AI search funnel end to end?

Define a fixed prompt set, run it weekly across engines with memory off, and record answer text and cited URLs. Then filter assistant referrers and campaign parameters in analytics, add an AI flag and a self-reported source field in your CRM, and report six ratios with confidence tiers rather than raw mention counts.

What is AI share of voice and what counts as a good number?

AI share of voice is your brand mentions divided by all brand mentions across the same set of answer runs. There is no universal threshold. Since assistants typically name 3 to 4 brands per category query, judge it against the named rivals in your own tracked set and against your own prior quarter, not an absolute figure.

Can I track ChatGPT referrals in GA4?

You can, but only partly. OpenAI documents that ChatGPT search referrals can carry utm_source=chatgpt.com, so you can filter on that parameter plus assistant referrer hostnames. That captures clicked links only. App sessions with stripped referrers, copied URLs and later brand searches will not appear, which is why self-reported attribution sits alongside it.

What is the difference between AI-sourced and AI-influenced pipeline?

AI-sourced means an assistant touch is the earliest identified touch on the record. AI-influenced means any assistant touch appears in the deal timeline within your lookback window. DerivateX reports both separately, plus an AI-corroborated subset where a referrer touch and a self-reported answer agree, which is the figure to quote when challenged.

How long before AI search measurement shows pipeline movement?

Stages 1 through 3 move inside 60 to 90 days when citation work lands. REsimpli became the most cited and recommended real estate CRM for investors in ChatGPT within 90 days. Stage 6 lags by one full sales cycle after that, so plan on two quarters before revenue lines are readable. These are targets, not guarantees.


Start with the denominator, not the dashboard

The mental model worth keeping: AI search visibility is not a metric you read, it is a ratio you construct. Every argument about whether AI search can be measured rigorously collapses once you write down the prompt set, the run count and the confidence tier, because at that point the number is falsifiable, and a falsifiable number is the only kind that survives a board review. Verito is a useful example of what the top of the funnel looks like when it moves properly: from an average position of 40 on Google to first page, and cited and recommended on ChatGPT and Google AI Overviews for 40 of their commercial hosting queries. That gain is only legible because the prompt set was fixed before the work started.

Request the free AI visibility audit and you will get back, within 48 hours, your current mention and recommendation rates across a sampled prompt set for your category, with the competitors currently occupying the recommended set named.

Ayush Sharma
Written byVP, SEO & AI Search, DerivateX

VP, SEO & AI Search at DerivateX. We're a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline.

Apoorv Sharma
Reviewed byCo-founder, DerivateX

Apoorv Sharma is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline. He is the author of the 2026 AI Visibility Benchmark Report and the Citation Engineering methodology. He's also the brain behind "Found On AI" and has sold 2 of his companies previously