The Prompt P&L: 7 Inputs That Put a Dollar Value on an AI Buying Prompt

AI prompt value is the expected pipeline a single buying prompt can generate, calculated from prompt demand, intent weight, recommendation share, referral transfer, opportunity conversion, deal value and a confidence discount. DerivateX models this across seven inputs, sampling each prompt at least 25 times per engine before assigning any dollar figure to it.

  • A prompt is worth money only when it produces a recommendation set, your brand is in that set, and the recommendation moves someone into your CRM. Each of those three steps has its own measurable rate.
  • The formula DerivateX uses is Prompt Value = Volume x Intent Weight x Recommendation Share x Transfer Rate x Opportunity Rate x Opportunity Value x Confidence Factor.
  • Recommendation share is meaningless until you define the numerator and the denominator. DerivateX counts a run as valid only when the engine returns a vendor list at all.
  • The model is directional, not precise. At 25 runs, any measured share carries a wide confidence interval, so DerivateX reports ranges and confidence tiers rather than a single clean number.
  • Its real job is prioritization. Not every commercial prompt is worth working on, and the dollar value tells you which 15 out of 200 deserve budget this quarter.

What is AI prompt value, and why measure it per prompt instead of in aggregate?

AI prompt value is the estimated dollar contribution of one buying prompt to your pipeline over a defined period. DerivateX treats a prompt the way a paid team treats a keyword in an ad account: as a unit with its own demand, its own competitive set, its own conversion behavior and its own cost to win. Aggregate metrics hide all of that.

Most B2B SaaS teams currently report AI search at the portfolio level. Share of voice is up, mentions are up, citations are up. Those numbers move without telling anyone what to do next, because a few points of gain spread across 200 prompts is invisible in pipeline, while the same gain on the two prompts your buyers actually type before a demo request is not.

The reason per-prompt matters more in AI search than it did in classic search is structural. Google returns ten blue links and a buyer can scroll. A language model typically returns 3 to 4 brands per category query, and there is no page two. Being fourth on a prompt with 4,000 monthly instances is worth more than being first on a prompt with 40, and only a per-prompt model lets you see the difference before you spend against it. AI-sourced visitors also convert at 4.4x the rate of other visitors, so averaging them into a site-wide conversion rate destroys the signal you are trying to find. For the funnel view that sits above this model, DerivateX maintains a separate AI search funnel measurement model covering stage definitions and cohort logic.


How do you calculate AI prompt value? The seven inputs

DerivateX calculates prompt value as a chain of seven multipliers, each with a named data source. The chain matters: if any one input is zero, the prompt is worth zero, and the model tells you exactly which link is broken.

InputWhat it measuresWhere the number comes fromTypical ranges DerivateX uses as starting assumptions
1. Prompt volume (V)Estimated monthly instances of the prompt and its close paraphrasesSearch demand for the equivalent question, sales call language, self-reported buyer wording50 to 8,000 per month
2. Intent weight (I)How close the prompt sits to a purchase decisionPrompt class taxonomy, fixed and documented in advance0.1 to 1.0
3. Recommendation share (R)Share of valid sampled runs where your brand appears in the recommended setScheduled prompt sampling across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews0 to 0.6
4. Transfer rate (T)Share of recommendation events that produce a measurable visit or brand searchAnalytics referral data, branded search lift, self-reported attribution0.02 to 0.12
5. Opportunity rate (Q)Share of AI-attributed visits that become a qualified opportunityCRM, filtered to the AI-sourced cohort0.04 to 0.20
6. Opportunity value (OV)Average new deal value for that segmentCRM closed-won data, trailing 12 monthsYour ACV
7. Confidence factor (F)Haircut for sampling noise and attribution gapsRun counts, CRM field completion, referral parameter coverage0.3 to 0.9

Monthly prompt value = V x I x R x T x Q x OV x F. Run the same chain without OV and you get expected opportunity count. Run it with average contract value instead of win-rate-adjusted value and you get pipeline rather than revenue. DerivateX reports all three lines, because finance will ask which one you handed them.

Intent weight is the input teams argue about, so fix it before you calculate anything. The taxonomy DerivateX uses assigns 0.8 to 1.0 to vendor comparison, alternatives and “best tool for” prompts, 0.7 to 0.9 to pricing and cost prompts, 0.5 to 0.7 to integration and implementation prompts, 0.3 to 0.5 to problem-framing prompts, and 0.1 to 0.2 to definitional prompts. Publish the table, then stop adjusting it per prompt. A weighting that moves whenever the result is inconvenient is not a measurement system.


How do you define the numerator and denominator for recommendation share?

Recommendation share is the number of sampled runs where your brand appears inside the recommendation set, divided by the number of runs where the engine returned a recommendation set at all. DerivateX defines it that way deliberately, and the second half of that sentence is where most reporting quietly breaks.

Ask a language model “what is the best inventory management software for mid-sized distributors” and you reliably get a vendor list. Ask it “how do I reduce stockouts” and much of the time you get a process answer with no vendors named. If you divide by all runs, your share looks terrible on prompts where nobody could have won. If you divide by valid runs only, you get a number that reflects competitive position rather than prompt type. On problem-framing prompts, the gap between those two denominators is routinely large enough to reverse the conclusion you would draw from the report.

The numerator needs rules too. DerivateX counts three separate states rather than one:

  1. Recommended: named inside the answer’s vendor list, with a description. This is the only state that feeds R.
  2. Cited: your domain appears as a source link but your brand is not in the recommendation set. Tracked separately, because it predicts future recommendation but does not convert today.
  3. Mentioned: named in passing, in a caveat, or in a comparison against the recommended option. Logged, never counted as a win.

Sample size decides how much you can say. With 25 runs and 7 appearances, measured share is 28%, but the 95% Wilson score interval, the standard small-sample method for estimating a proportion reviewed in Brown, Cai and DasGupta, Statistical Science, 2001, runs from roughly 11% to 44%. Applying the same Wilson calculation at 50 runs and the same share tightens the interval to roughly 17% to 42%. That is still wide, which is the honest answer to whether AI search can be measured precisely at the prompt level: it cannot, but it can be measured directionally, and directionally is enough to allocate budget. DerivateX reports the range alongside the point estimate and refuses to publish a single figure below 25 runs per engine.

Run the same prompt across engines separately, never blended. Only 11% of domains are cited by both ChatGPT and Perplexity, and 80% of URLs cited by those two engines do not appear in Google’s top 100. A blended share number averages away the exact differences you need to act on. DerivateX published a comparison of ChatGPT, Claude, Gemini and Perplexity citation behavior for B2B SaaS if you want the engine-by-engine breakdown.


Which data sources do you actually need?

DerivateX treats three of the following five sources as the minimum for building a prompt value model at all. Below three, the confidence factor sits under 0.3 and the output stops being a decision tool.

  • Scheduled prompt sampling. A fixed prompt set, run on a fixed cadence, across five engines, with responses stored. Manual spot checks do not work here because model outputs vary run to run. Profound and Peec AI both position themselves as AI search visibility platforms that run this sampling at scale and report share over time. If sampling infrastructure and a weekly read are genuinely your only gap, one of those subscriptions on its own is the better purchase than any retainer, including ours, and DerivateX says that on discovery calls. The DerivateX review of GEO tools for B2B SaaS compares the current set.
  • Referral parameters in analytics. OpenAI’s publishers and developers FAQ confirms that clicks from ChatGPT search can arrive carrying utm_source=chatgpt.com, which makes a segment in Google Analytics 4 straightforward to build from standard campaign URL parameters. Perplexity referrals show up as their own referral host. Build the segment once, then treat it as a floor rather than the whole number, since a buyer who reads a recommendation and then types your name into a browser leaves no referral trace at all.
  • Self-reported attribution in the form and on the call. One required open-text field on the demo form asking how the buyer first heard about you, plus a scripted discovery question. Self-report over-credits whatever the buyer remembers most vividly, so it inflates. It is still the only instrument that catches the untracked path, and 64% of marketing leaders say they are unsure how to measure AI search, largely because they skipped this step.
  • A CRM opportunity source field. A picklist in HubSpot or Salesforce with an explicit AI search value, filled in on at least 80% of opportunities. If sales does not complete it, your opportunity rate is a guess wearing a decimal point.
  • Sales call transcripts. Search them, in Gong or whatever your team records with, for engine names and for the literal phrasing buyers use. This is where you discover the prompts that matter, which rarely match the keyword list you inherited. DerivateX has written separately on how B2B SaaS buyers use ChatGPT to evaluate vendors, and the prompts in there consistently surprise SEO teams with strong Google rankings.

Context for why this belongs in the reporting stack: 40% of Google queries now show Google AI Overviews, and there are more than 200 million weekly ChatGPT Search users. OpenAI also sells ChatGPT Business seats to teams of 2 to 200 employees from $20 per standard seat per month billed annually, per OpenAI’s business pricing page checked at the time of writing and subject to change. Buying committees at software companies are inside these tools during working hours, on company accounts.


A worked example: valuing one prompt for a $12M ARR B2B SaaS company

Every figure in this section is invented for illustration on a hypothetical software company at $12M ARR selling into operations teams. It is a model showing how the arithmetic behaves, not a case study, and no DerivateX client result should be read from it. The worksheet format below is the one used with clients so that every cell traces to a source someone can audit.

Input (illustrative)ValueSourceRunning total
Prompt volume (V)2,000 / monthSearch demand proxy plus transcript frequency2,000 prompt instances
Intent weight (I)0.9Vendor comparison class1,800 commercial instances
Recommendation share (R)0.28 (7 of 25 valid runs)Sampled runs, ChatGPT504 recommendation events
Transfer rate (T)0.06GA4 referral segment plus self-report30.2 AI-attributed visits
Opportunity rate (Q)0.12CRM, AI-sourced cohort3.63 opportunities
Average contract value$18,000CRM, trailing 12 months$65,340 pipeline
Win rate22%CRM, same cohort$14,375 expected revenue
Confidence factor (F)0.625 runs, partial CRM completion$8,625 / month

On these illustrative inputs, the prompt carries a discounted expected revenue contribution of about $8,625 per month, or roughly $103,500 annualized, at a recommendation share of 28%. The marginal case is the number that decides whether work is worth doing. Moving recommendation share from 28% to 45% adds 306 recommendation events, 18.4 visits, 2.2 opportunities and roughly $5,235 per month in discounted expected revenue, or about $62,800 annualized. Those figures come only from multiplying the hypothetical inputs in the table above, not from any external benchmark or client account.

Set that against cost honestly. A DerivateX Own Your Category engagement runs $8,000 retainer plus $1,500 to $2,000 in off-site budget, so $9,500 to $10,000 all in per month, with a 90-day pilot and no lock-in after, and full current figures sit on the DerivateX pricing page. One prompt moving 17 points does not cover that on its own. Twelve prompts of similar weight, each moving 8 to 20 points, do. That is the real arithmetic of an AI search program, and any agency that will not put it on a page is asking you to buy on faith.


How do you prioritize prompts once every one has a dollar value?

Rank by marginal value per unit of effort, not by absolute value. DerivateX scores each prompt on the dollar gain from a realistic share improvement, then divides by the estimated work required to get there, which produces a very different list from ranking by volume.

Effort is not uniform across prompts, and that is the whole point. A prompt where you are already cited but not recommended usually needs framing and corroboration rather than new content, which is cheap. A prompt where the recommendation set is locked to three well-reviewed incumbents needs third-party evidence on the surfaces the engines pull from, which is slow. Reddit accounts for 46.7% of Perplexity’s top sources, so a prompt whose answer is built from community threads cannot be fixed by publishing another page on your own domain.

The Citation Surface Map is the artifact DerivateX builds for this: a per-prompt inventory of which domains, threads, reviews and listicles the engines actually pulled from across sampled runs, ranked by how often each source appeared. It converts “we need better content” into a named list of surfaces where corroboration is missing. Three prioritization rules hold across most software firms in the $5M to $50M ARR range:

  1. Fund prompts where you are cited but not recommended first. The engine already trusts your domain. Closing that gap is the cheapest share gain available and often lands inside one quarter.
  2. Deprioritize high-volume definitional prompts. An intent weight of 0.15 crushes even large volume, and winning them produces traffic your sales team will never see.
  3. Treat prompts with zero valid runs as a category problem. If the engine never returns vendors, no amount of citation work helps. Reframe toward the adjacent prompt that does return a list.

How do you report AI prompt value to leadership?

DerivateX reports four numbers to a board or executive team, in this order, and nothing else on the first slide.

  1. Recommendation share across the commercial prompt set, stated as a range with run counts, for example 31% across 40 prompts at 50 runs each.
  2. Modeled prompt value for the top 15 prompts, as a discounted annual figure with the confidence tier printed next to it.
  3. AI-sourced pipeline and AI-influenced pipeline, pulled straight from the CRM, on separate lines with opportunity counts alongside the dollar values and never summed into one figure.
  4. Movement since last period on the same prompt set, with a plain note on what changed in the market that quarter. Hold the prompt set constant for at least two quarters, because swapping prompts between reports produces movement that has nothing to do with performance.

Confidence tiers do more work than any chart. The DerivateX tiering runs Tier A for 50 or more runs per prompt per engine with above 80% CRM field completion and referral parameters present, supporting a confidence factor of 0.8 to 0.9; Tier B for 25 runs and partial CRM hygiene, at 0.5 to 0.7; and Tier C for anything below that, at 0.3 to 0.4, labeled as an estimate that should not enter a forecast.

Context helps the room understand why this reporting exists. Around 73% of B2B sites lost significant traffic between 2024 and 2025, click-through drops by 61% when AI Overviews appear, and 28% of pages cited by ChatGPT have zero organic Google visibility. Traffic decline in that environment is a market condition, not a failure of anyone’s SEO program, and framing it that way keeps the conversation on reallocation rather than blame. The eight GEO metrics DerivateX reports to clients covers the full stack around these four headline numbers.


When is prompt-level valuation the wrong tool?

DerivateX will tell a prospect to skip this model in four situations, and each one is a real disqualification rather than a soft objection.

Low deal volume. If you close fewer than roughly 30 opportunities per quarter, your opportunity rate and win rate are built on too few events to survive being multiplied by five other estimates. Track recommendation share on its own, watch the trend, and revisit the dollar model when volume supports it.

Self-serve motion under $5,000 ACV. Product-led SaaS companies with high-velocity signups get better answers from cohort analysis on signup source than from a seven-input pipeline chain. The chain was built for considered purchases with a sales conversation in the middle.

You already have the content engine and just need measurement. If you have an in-house SEO team producing quality work and clean CRM hygiene, a tracking platform subscription plus a quarter of internal analyst time will get you most of this model without a retainer. That is genuinely the better purchase, and it is the honest recommendation. A $3,500 DerivateX diagnostic, delivered in two weeks and credited in full against month one if you convert within 30 days, is the right size of commitment when you want the model built once and then run in house.

You want one agency for classic SaaS SEO, content and paid together. SimpleTiger positions itself as a SaaS-focused SEO and paid search agency with a long track record in software marketing, and if your priority is a broad organic and paid program under one roof rather than a prompt-level citation program, that is the stronger fit and DerivateX will say so. The DerivateX comparison of SimpleTiger and DerivateX for B2B SaaS AI search sets out where each one wins.

Where DerivateX does fit is companies at $5M to $50M ARR that need the prompts identified, the sampling run, the model built and the citation work executed against the priority list in the same engagement. That combination is why REsimpli became the most cited and recommended real estate CRM for investors in ChatGPT within 90 days, and why Gumlet now attributes more than 20% of monthly inbound revenue to AI discovery. Those are outcomes from executed programs, and they are targets DerivateX works toward rather than guarantees anyone can make.


Frequently asked questions

How do I measure AI prompt value?

Multiply seven inputs: estimated prompt volume, intent weight, recommendation share from sampled runs, transfer rate from analytics, opportunity rate from your CRM, average deal value, and a confidence factor between 0.3 and 0.9. DerivateX requires at least 25 sampled runs per prompt per engine before publishing any dollar figure from that chain.

What is the difference between recommended, cited and mentioned?

Recommended means your brand sits inside the answer’s vendor list with a description, and only that state feeds recommendation share. Cited means your domain appears as a source link without your brand being recommended. Mentioned means named in passing. DerivateX logs all three separately and counts only recommendations as wins.

Can I track ChatGPT referrals in GA4?

Yes, though only partially, because OpenAI’s publishers and developers FAQ confirms clicks from ChatGPT search can carry utm_source=chatgpt.com, which you can segment in GA4. Treat that segment as a floor rather than a total, since buyers who read a recommendation and then type your brand name directly leave no referral trace.

What is a good AI recommendation rate for B2B SaaS?

Language models typically return 3 to 4 brands per category query, so a recommendation share above 30% across a commercial prompt set puts you inside the consideration list more often than not. DerivateX treats 30% as a working benchmark, adjusted for how many credible vendors compete in your category.

What is the difference between AI-sourced and AI-influenced pipeline?

AI-sourced pipeline means the first identifiable touch was an AI search referral or a self-report with no prior touch recorded. AI-influenced pipeline means an AI touch appears anywhere in the journey. DerivateX reports both on separate lines with opportunity counts, and never sums them into a single figure.

How much does an AI search measurement and GEO program cost?

DerivateX pricing starts at $6,000 to $6,200 all in per month for Rank and Get Found, $9,500 to $10,000 for Own Your Category, and $14,000 to $15,000 for Market Leader with a six-month minimum. A one-time diagnostic is $3,500, delivered in two weeks, credited against month one if you convert within 30 days.


The mental model worth keeping

Treat every buying prompt as a small P&L line rather than a ranking. It has revenue potential you can estimate, a competitive position you can sample, a conversion path you can instrument, and a cost to improve that varies enormously from one prompt to the next. Once those four things sit on one row, the argument about whether AI search can be measured ends, and the real argument starts: which fifteen prompts get funded this quarter.

Citation Engineering is the methodology DerivateX uses to move recommendation share on a chosen prompt deliberately, by building the evidence and third-party corroboration the engines pull from rather than editing a page and hoping. The measurement model above is what tells you which prompts deserve that work. For teams still building the case internally, the DerivateX guide to connecting B2B SaaS SEO to pipeline covers the reporting foundation this model sits on.

The free DerivateX AI visibility audit returns your recommendation share across your commercial prompt set, engine by engine, within 48 hours.

Ayush Sharma
Written byVP, SEO & AI Search, DerivateX

VP, SEO & AI Search at DerivateX. We're a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline.

Apoorv Sharma
Reviewed byCo-founder, DerivateX

Apoorv Sharma is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline. He is the author of the 2026 AI Visibility Benchmark Report and the Citation Engineering methodology. He's also the brain behind "Found On AI" and has sold 2 of his companies previously