Case study: Gumlet turned ChatGPT mentions into 20% of inbound revenue. Read it →
Prompt-to-Pipeline Attribution in HubSpot: 7 Steps for B2B SaaS
DerivateX measures AI search attribution in HubSpot with three evidence tiers rather than one last-click source. Confirmed contacts arrive with an AI referrer or UTM, corroborated contacts name an assistant in a self-reported field, and inferred contacts match a cited page and prompt theme. Weight them 1.0, 0.6 and 0.3, then report confidence-weighted pipeline.
- AI search attribution is measurable in HubSpot, but only if you define the numerator and the denominator before you build a single property. Most teams skip that and end up arguing about the number instead of using it.
- Three data layers are required: prompt-side sampling, referral and session data, and CRM evidence including self-reported attribution. None of the three is sufficient alone.
- Machine-readable evidence exists. ChatGPT can append utm_source=chatgpt.com to outbound links, per OpenAI’s publishers and developers FAQ, so a portion of AI-sourced traffic is directly identifiable.
- Most AI discovery leaves no referrer at all, which is why a self-reported field and a confidence weight beat pretending the referrer data is complete.
- Report three numbers to leadership: AI-sourced pipeline, AI-influenced pipeline, and confidence-weighted AI pipeline, each with the sample size and the weights printed underneath.
- Gumlet attributes more than 20% of monthly inbound revenue to AI discovery, which is the kind of number this model is built to produce and defend.
How do I measure AI search attribution in HubSpot?
You measure it by classifying every inbound contact into one of four evidence tiers, stamping that tier onto the deal, and reporting pipeline three ways: sourced, influenced, and confidence-weighted. DerivateX uses this structure because AI search produces partial evidence by design, and a model that only counts perfect evidence will undercount reality by a wide margin.
Here is the problem in plain terms. A buyer asks ChatGPT which tools solve their problem, reads three or four recommendations, does not click anything, opens a new tab two days later and types your brand name into Google. In HubSpot that contact looks like branded organic search or direct traffic. The AI conversation that created the demand is invisible to every default report you own.
That is not a tracking bug you can fix with better UTMs. It is a structural property of how answer engines work. Roughly 40% of Google queries now show AI Overviews, and click-through rates drop by about 61% when they appear, which means the ratio of influence to clicks keeps moving against you. Meanwhile 64% of marketing leaders say they are unsure how to measure AI search at all, which is why so many decks show a citation count where a pipeline number should be.
The fix is not more precision. It is honest classification with a stated confidence level, the same way a finance team handles pipeline coverage: you do not refuse to forecast because deals are uncertain, you apply a weight and disclose it. DerivateX builds AI search attribution in HubSpot the same way, and the resulting number survives a CFO asking how it was calculated.
Which data sources does AI search attribution need?
Three, and DerivateX treats them as a set rather than alternatives. Drop any one of them and the model breaks in a specific, predictable way.
Prompt-side sampling tells you whether you are present in the conversations that matter. This is a controlled panel of buyer prompts run on a fixed schedule against ChatGPT, Perplexity, Google AI Overviews, Gemini and Claude, with the outputs logged. It produces prompt coverage, citation rate, recommendation rate and AI share of voice. It cannot tell you anything about revenue.
Analytics and referral data tells you who arrived with a traceable AI footprint. ChatGPT search referrals can carry utm_source=chatgpt.com on outbound links, and other assistants pass identifiable referrer domains with varying consistency. This layer is accurate but incomplete, and you should assume it understates AI-driven demand rather than overstating it.
CRM and self-reported evidence is the layer most software companies skip, and it is the one that carries the most signal. A required “how did you first hear about us” field with an AI assistant option, followed by a free-text “what did you ask it” field, produces the only data that connects a specific prompt to a specific deal. It is self-reported and therefore imperfect. It is also the only bridge between the prompt layer and the pipeline layer.
If you want the search-engine side of this picture in more depth, our walkthrough of Search Console generative AI reporting for B2B SaaS measurement covers what Google exposes and what it withholds. This article stays inside HubSpot.
What are the 7 steps to build prompt-to-pipeline attribution in HubSpot?
The build takes a RevOps person about two working days, plus a form change that needs sign-off. DerivateX runs these seven steps in order, because steps three onward depend on the definitions set in step one.
Step 1: Define the numerator and the denominator in writing
Write four sentences before you open HubSpot. What counts as AI-sourced. What counts as AI-influenced. What time window applies. What the denominator is for every rate you plan to report. A citation rate of 34% is meaningless unless the reader knows it means “34 of the 100 prompts in our tracked panel returned a response containing at least one link to our domain, sampled three times each, logged out, US region.”
DerivateX defines AI-sourced as first touch, meaning the earliest recorded interaction carries AI evidence. AI-influenced means any AI evidence exists on the contact or account at any point before the deal closed. Those two definitions are not interchangeable and mixing them is the single most common way these reports lose credibility in a board meeting.
Step 2: Capture the raw signals at the form
Add hidden fields to every conversion form so the raw evidence lands in the CRM rather than dying in a session. At minimum, capture the document referrer, the landing page URL, and the full query string on first touch, stored in first-touch-only properties so a later visit cannot overwrite them.
Do not rely on a default source drop-down to carry an AI value. Build your own properties so the classification logic is yours, visible, and editable when a new assistant appears. HubSpot serves 306,000+ customers worldwide across its marketing, sales and Smart CRM products, and its own homepage now lists an AEO beta for seeing where a brand shows up in AI results, but pipeline-grade attribution still needs fields you control.
Step 3: Create the property set
DerivateX uses this exact property structure. Names are suggestions, the shape is the point.
| Object | Property | Type | Purpose |
|---|---|---|---|
| Contact | ai_first_touch_surface | Dropdown | ChatGPT, Perplexity, Google AI Overviews, Gemini, Claude, Copilot, Other AI, None |
| Contact | ai_evidence_tier | Dropdown | Confirmed, Corroborated, Inferred, None |
| Contact | ai_self_reported_source | Dropdown | Answer to “how did you first hear about us” |
| Contact | ai_reported_prompt | Multi-line text | Free text: what the buyer actually asked |
| Contact | first_referrer_raw | Single-line text | Hidden field, first touch only |
| Contact | first_utm_source_raw | Single-line text | Hidden field, first touch only |
| Contact | ai_evidence_date | Date | When the tier was first stamped |
| Deal | ai_sourced | Boolean | First touch carried AI evidence |
| Deal | ai_influenced | Boolean | Any AI evidence before close date |
| Deal | ai_confidence_weight | Number | 1.0, 0.6, 0.3 or 0 |
| Deal | ai_weighted_amount | Calculation | Deal amount multiplied by ai_confidence_weight |
| Deal | ai_prompt_theme | Dropdown | Category, alternatives, comparison, problem, integration, pricing |
The prompt theme field earns its place faster than people expect. A comparison prompt and a problem prompt produce very different win rates, and once you have ninety days of data you can see which prompt themes are worth building assets against.
Step 4: Write the evidence tier rules
This is the auditable core of the model. DerivateX writes the rules as a workflow with no ambiguity and no manual judgment at the contact level.
- Confirmed, weight 1.0. first_referrer_raw or first_utm_source_raw matches a maintained list of AI surfaces, for example chatgpt.com, perplexity.ai, gemini.google.com, claude.ai, copilot.microsoft.com.
- Corroborated, weight 0.6. No machine evidence, but the contact selected an AI assistant in ai_self_reported_source, or a rep logged an AI assistant as the discovery route on a call and recorded a prompt.
- Inferred, weight 0.3. No machine evidence and no self-report, but first touch was direct or branded organic onto a page that appeared as a cited source in your tracked prompt panel during the same month, and the deal maps to one of your prompt themes.
- None, weight 0. Everything else, including everything you are not sure about. Ambiguity defaults down, never up.
Two rules keep this honest. A contact can only move up a tier if new evidence arrives, and every tier change writes a timestamp. If someone later asks why a deal is counted, the record answers.
Step 5: Add self-reported attribution and the prompt question
Make “how did you first hear about us” required on demo request and trial signup forms, with these options: search engine, AI assistant such as ChatGPT or Perplexity, peer or colleague, review site such as G2, social, event, ad, other. When a contact selects the AI assistant option, show a conditional free-text field asking what they asked it.
That second field is the highest-value data point in the entire build. It gives you the buyer’s actual prompt phrasing, in their words, which almost never matches the keyword list your SEO team maintains. DerivateX feeds these strings straight back into the tracked prompt panel, which is how the measurement loop stops being a dashboard and starts driving what gets built. Our note on how B2B SaaS buyers use ChatGPT to evaluate vendors covers the prompt patterns that show up most often in this field.
Step 6: Stamp the deal and calculate weighted pipeline
On deal creation, copy ai_evidence_tier from the associated contact to the deal, set ai_confidence_weight from the tier, and let the calculation property produce ai_weighted_amount. Set ai_sourced from the first-touch record and ai_influenced from any AI evidence on any associated contact in the buying group, which matters because B2B SaaS deals routinely involve four or more people and the one who found you is rarely the one who signs.
Confidence-weighted AI pipeline is then a single sum: confirmed amount plus 0.6 times corroborated amount plus 0.3 times inferred amount. DerivateX reports that figure alongside the unweighted sourced number so nobody can accuse the model of inflating or hiding.
Step 7: Build two reports and a monthly review
Report one is prompt-side: prompt coverage, citation rate, recommendation rate, AI share of voice, trended monthly. Report two is pipeline-side: AI-sourced pipeline, AI-influenced pipeline, confidence-weighted pipeline, plus win rate and average deal size split by evidence tier. Review them together once a month, never separately, because the only interesting question is whether movement in the first report precedes movement in the second.
What is the right way to calculate AI share of voice, citation rate and recommendation rate?
Every one of these metrics lives or dies on its denominator. DerivateX defines them as follows and prints the definition on the report itself, because a rate without a denominator is a decoration.
| Metric | Numerator | Denominator | What it tells you |
|---|---|---|---|
| Prompt coverage | Prompts in your tracked panel | Total mapped buyer prompts in the category | Whether you are even watching the right conversations |
| Citation rate | Prompt runs returning a link to your domain | Total prompt runs in the sample | Whether your pages are being used as evidence |
| Recommendation rate | Prompt runs naming you as a recommended option | Total prompt runs in the sample | Whether you make the shortlist, which is the commercial event |
| AI share of voice | Your brand mentions across the sample | All vendor brand mentions across the same sample | Your position relative to the named competitive set |
| Confidence-weighted AI pipeline | Confirmed + (0.6 x corroborated) + (0.3 x inferred) | Not a rate, an absolute dollar figure | What you take to the board |
Sampling control matters more than sample size. Run the same prompt panel on the same weekday, logged out, from a fixed region, three runs per prompt, and record the model version. Answer engines vary run to run, so a single run is an anecdote. DerivateX treats a prompt as covered only when it has been sampled at least three times in the period, and reports recommendation rate against prompt runs rather than prompts, because that is the honest denominator.
One benchmark worth holding in your head: AI category queries typically surface three to four recommended brands. That is the shelf. If your recommendation rate is 20% in a category with a well-established set of incumbents, you are on the shelf in one conversation out of five, and the realistic goal is to move that number, not to reach 100%. Teams choosing tooling for the prompt-side layer can compare options in our review of LLM visibility trackers for B2B SaaS.
What does this look like for a real B2B SaaS company?
Here is a worked example using an illustrative software firm at roughly $12M ARR with a $24,000 average contract value and 42 inbound deals created in a quarter. The numbers below are constructed to show the arithmetic, not drawn from a client account.
| Evidence tier | Deals | Pipeline value | Weight | Weighted value |
|---|---|---|---|---|
| Confirmed (AI referrer or UTM) | 5 | $118,000 | 1.0 | $118,000 |
| Corroborated (self-reported assistant) | 9 | $221,000 | 0.6 | $132,600 |
| Inferred (cited page plus prompt theme) | 7 | $164,000 | 0.3 | $49,200 |
| None | 21 | $505,000 | 0 | $0 |
| Total | 42 | $1,008,000 | $299,800 |
Three things stand out, and DerivateX points to all three in the first review with a client. First, the confirmed tier alone would have reported $118,000 and understated AI influence by more than half. Second, the corroborated tier is the biggest contributor, which is an argument for the form field, not for better UTM hygiene. Third, the honest headline is “$300,000 of confidence-weighted AI pipeline from 21 of 42 deals carrying some AI evidence,” not “AI drove a million dollars.”
Track win rate by tier as well. If AI-sourced deals close at a higher rate or at a higher average contract value than the rest, that is worth knowing and worth funding. Published research puts conversion from AI-sourced visitors at around 4.4x that of traditional organic visitors, but you should measure your own delta rather than quoting anyone’s benchmark to your CFO. For the broader frame on tying acquisition work to revenue, see our piece on connecting B2B SaaS SEO to pipeline.
What can this model prove, and what can it not prove?
DerivateX puts this table directly into client reporting, because stating the limits up front is what stops the number from being challenged later.
| This model can prove | This model cannot prove |
|---|---|
| That a measurable share of pipeline arrived with AI evidence attached | That AI search caused the deal, as opposed to being one touch among several |
| Which prompt themes produce deals that close | Which specific answer the buyer read, or how you were described in it |
| Whether recommendation rate moved in a tracked prompt panel | Whether recommendation rate moved across all prompts, including ones you never sampled |
| Whether AI-sourced deals convert differently from the rest of your pipeline | Statistical significance at typical B2B SaaS deal volumes, in the first two quarters |
| Directional correlation between visibility work and pipeline over time | Causation, which would require holdouts almost no company will accept |
Two structural blind spots deserve naming. Around 28% of pages cited by ChatGPT have zero organic Google visibility, and roughly 80% of URLs cited by ChatGPT and Perplexity do not appear in Google’s top 100, so the pages driving your AI presence may not be the pages your SEO reporting watches. And with roughly 11% domain overlap between ChatGPT and Perplexity citations, per-engine measurement is not optional. A single blended number hides which engine you are losing.
How do I report AI search attribution to leadership?
One slide, three numbers, one footnote. DerivateX uses this format with every client reporting to a board or an executive team, and it has survived more finance scrutiny than any dashboard we have built.
- AI-sourced pipeline, unweighted, with the deal count. This is the conservative floor.
- AI-influenced pipeline, unweighted, with the deal count. This is the ceiling.
- Confidence-weighted AI pipeline, the number you actually plan against.
The footnote states the weights, the sample size of the prompt panel, the number of runs per prompt, and the sentence “this is directional, not causal.” Including that sentence increases trust rather than reducing it, because every senior finance person already knows marketing attribution is directional and is waiting to see whether you know it too.
Alongside the pipeline slide, DerivateX reports an AI Visibility Score. The AI Visibility Score, or AVS, is a composite of recommendation rate, citation rate and share of voice across the tracked prompt panel, expressed as a single trended figure so leadership can see direction without reading five charts. Underneath it sits the Citation Surface Map, which is the inventory of every source an engine actually pulls from when answering your category prompts, including review sites, communities and comparison pages you do not own. That map explains the metric: if 46.7% of Perplexity’s top sources come from Reddit, then your Reddit presence is a measurement input, not a side project.
For evidence that this converts into revenue rather than reporting, Gumlet attributes more than 20% of monthly inbound revenue to AI discovery, and REsimpli became the most cited and recommended real estate CRM for investors in ChatGPT within 90 days. Neither outcome would have been visible, let alone defensible, without the tier structure described above.
When is HubSpot the wrong place to build this?
Three situations, and DerivateX says so before quoting anything.
If Salesforce is your system of record, build the identical tier structure there instead. Same four picklist values on Lead and Contact, same weight field on Opportunity, same self-reported field on the web-to-lead form. The model is CRM-agnostic. Only the field syntax changes, and running it in two systems at once guarantees two different numbers.
If you have fewer than roughly 20 inbound deals per quarter, skip the weighting math entirely. At that volume the confidence weights add false precision to what is really a list you can read in ten minutes. Read the self-reported field every month, log the prompts, and revisit the full build when volume supports it.
If the argument in your company is about the revenue definitions themselves, ARR calculation, renewal treatment, what counts as pipeline coverage, then attribution is the wrong first purchase. A revenue data platform such as Discern, which positions itself as an AI data layer for B2B SaaS with audit-ready metric definitions and transparent calculation logic, solves a problem attribution modeling cannot touch. Fix the denominator of the business before you refine the numerator of one channel.
There is also a tooling boundary worth stating. HubSpot’s homepage lists more than 2,000 integrations and an AEO beta for brand visibility in AI results, and native visibility tooling inside the CRM will keep improving. If all you need is a rough read on whether your brand appears in AI answers, native and low-cost tools may be enough. The tier model in this article exists for teams that have to defend a pipeline number, not just observe a mention count.
Frequently asked questions
How do I track ChatGPT traffic in HubSpot?
Capture the raw referrer and full query string as hidden first-touch fields on every form, then classify contacts whose referrer or UTM matches chatgpt.com. ChatGPT can append utm_source=chatgpt.com to outbound links per OpenAI’s publishers FAQ, but most AI discovery leaves no referrer, so pair this with a self-reported source field.
What is a good AI share of voice for a B2B SaaS company?
There is no universal benchmark, and anyone quoting one is guessing. AI category queries typically surface three to four recommended brands, so in a category with five serious vendors an even split is roughly 20%. Judge your figure against your own trend line and your named competitive set, not an industry average.
Can HubSpot attribution reports show AI search as a source?
Not reliably through default source categories, which is why DerivateX builds custom contact and deal properties instead. HubSpot’s homepage lists an AEO beta for seeing where a brand appears in AI results, but connecting that visibility to closed revenue requires your own evidence tier field, weight field and calculation property.
Is self-reported attribution reliable enough to report to a board?
It is reliable enough at 0.6 weight, which is exactly why the weight exists. Buyers misremember and sometimes skip the field, so treat it as corroborating evidence rather than proof. Reported alongside confirmed referral data and a stated sample size, it is the strongest signal available for demand that leaves no click trail.
How long before AI search attribution shows movement in pipeline?
Expect prompt-side metrics such as citation rate and recommendation rate to move first, usually within 60 to 90 days of consistent work. Pipeline evidence lags by a full sales cycle after that. DerivateX commits to the measurement cadence and the reporting format, not to a guaranteed citation or revenue outcome.
The model that survives the next budget review
The mental model is this: AI search attribution is a confidence problem, not a tracking problem. You will never see every conversation, so stop building toward completeness and start building toward defensibility. Four tiers, three weights, one written definition of sourced versus influenced, and a footnote that says directional. That structure holds up when a CFO asks how you got the number, and it keeps working as new assistants appear, because the classification logic is yours rather than a vendor’s.
DerivateX builds this model inside client engagements rather than selling it as a standalone dashboard, because a measurement layer with nothing changing underneath it just reports the gap more precisely every week. As of September 2026, the entry engagement is $5,000 retainer plus $1,000 to $1,200 of off-site budget, so $6,000 to $6,200 all in on a 90-day pilot with no lock-in after, and there is a $3,500 diagnostic that credits in full against month one, both detailed on the DerivateX pricing page. If you want to see how the prompt panel feeds the model before any of that, our analysis of modeling the revenue value of individual buyer prompts is the companion piece to this one.
Request the free AI visibility audit and DerivateX will return, within 48 hours, your recommendation rate across a sampled panel of buyer prompts, the engines where you are absent, and the competitors being named in your place: https://derivatex.agency/free-ai-visibility-audit/.













