Case study: Gumlet turned ChatGPT mentions into 20% of inbound revenue. Read it →
How AI Describes Your Brand: The BUT Audit

How AI describes your brand decides more deals than whether AI mentions your brand at all. Google AI Overviews attach a boundary to almost every recommendation, and DerivateX found one on 44% of the 445 brand recommendations we studied. The BUT Audit is the method for reading that boundary and fixing what sits behind it.
Most B2B SaaS teams stop measuring the moment their name appears. They screenshot the mention, drop it in Slack, and move on. The sentence that decides whether the recommendation survives is the one they never read.
TL;DR
- Across 100 Google AI Overview answers covering 20 B2B software categories, DerivateX counted 445 brand recommendations, and 44% of them carried a qualifying clause in the sentence directly after the brand name.
- The most common qualifier is not a criticism. A segment lock, phrased as “best for” or “ideal for”, followed roughly a quarter of all recommendations, and a company-size gate followed another 16%.
- Actual complaints are rare in AI Overviews. A capability gap followed 2.9% of recommendations, and a cost objection followed 2.0%.
- Sentiment scoring cannot see any of this, because a scope statement reads as praise while quietly removing the brand from every query outside that scope.
- The BUT Audit takes about 30 minutes, runs only on prompts where your brand already appears, and the repair is missing evidence rather than more published articles.
What the BUT Audit Is
The BUT Audit is a diagnostic that reads the sentence immediately after your brand appears in an AI answer, classifies the boundary that sentence places on the recommendation, and counts which boundaries repeat across your buyer prompts. DerivateX built it because a brand can be named in an answer and still lose the deal inside the same paragraph.
The name comes from the connector that used to signal the problem: “X is great, but…”. Real AI Overview answers use that structure less often than people assume. They use a friendlier one that does the same damage.
Think about what “Best for freelancers on a tight budget” actually does to a company selling into 200-person teams. The line is complimentary, accurate, and completely disqualifying. A buyer running a mid-market evaluation reads it and moves down the list.
This is why the BUT Audit is not sentiment analysis. Sentiment tools sort language into positive, neutral, and negative, and a scope statement lands in the positive bucket every time.
Your brand is a clean, well-designed option covering invoicing and expense tracking. Best for solo consultants and very small service businesses.
Complimentary adjectives, no negation, no complaint. Scores at the top of the range and clears the audit.
- Tone: favourable
- Complaint detected: none
- Action flagged: none
The clause is a filter. A 200-person finance team reads it, rules the brand out, and moves to the next name on the list.
Mid-market shortlistFinance team queriesMulti-entity comparisons
In the DerivateX State of AI Visibility benchmark, nearly nine in ten of the B2B SaaS companies we tested scored at the very top of the sentiment range, and plenty of them still could not win a shortlist. Being liked was never the constraint.
What Google AI Overviews Actually Say After They Name You
In a DerivateX study of 100 Google AI Overview answers run across 20 software categories in June 2026, 44% of the 445 individual brand recommendations carried a qualifier in the next sentence. We logged every answer while signed out, pulled each recommended product, and tagged the clause that followed the brand’s first appearance.
Here is how those qualifiers break down across all 445 recommendations.
| Qualifier class | How it sounds in the answer | Share of recommendations |
|---|---|---|
| Segment lock | “Best for”, “ideal for”, “designed for” | 24.7% |
| Size gate | “small teams”, “freelancers”, “enterprise environments” | 15.7% |
| Effort flag | “complex”, “requires technical setup”, “learning curve” | 9.0% |
| Capability gap | “lacks”, “limited”, “fewer”, “does not include” | 2.9% |
| Cost flag | “expensive”, “per-user pricing”, “premium” | 2.0% |
| Any qualifier present | 44% |
Three numbers from that sample are worth holding onto.
The phrase “best for” appeared 245 times across the 100 answers, averaging two and a half explicit segment locks per answer. Google AI Overviews are not writing a ranked list. They are writing a sorting exercise, and every brand gets filed into a drawer.
Every category in the matrix sorts brands on a dominant axis. Identify which axis yours runs on before you decide what evidence to build.
Also: AI time tracking, iPaaS, CRM
What to build: proof you serve the adjacent use case, written in that segment’s own words.
Also: help desk, ITSM, business intelligence
What to build: named customers at the size you want, with seat counts and outcomes attached.
Also: tech pack, project management, video hosting
What to build: not much here yet. Spend the quarter on getting named rather than on widening scope.
Roughly half the answers, 49 out of 100, ended by asking the reader a qualifying question back. Questions like “How large is your team?” and “What is your annual revenue?” show the model narrowing a shortlist by segment before the buyer has clicked anything. Your drawer label is what decides whether you survive that narrowing.
Qualification also varies sharply by category. In our sample, compliance automation queries qualified 70% of recommended brands, while HR software queries qualified 18%. Categories where buyers carry hard constraints get sorted hardest, which means regulated and technical categories need this audit more than most.
One limitation belongs in the same breath as the finding. These figures describe Google AI Overviews in one week of June 2026, across commercial “best software” queries in 20 categories. They are a strong signal about how AI describes brands in that surface, and they are not a measurement of ChatGPT, Perplexity, Gemini, or Claude, which DerivateX tracks separately.
The Five Qualifier Classes and What Each One Is Telling You
Every qualifier is a confession. The model is telling you which piece of evidence it could not find, which is far more useful than knowing your sentiment score.
| Qualifier class | What the model is really saying | The evidence that is missing |
|---|---|---|
| Segment lock | “I can only place you in one drawer.” | Proof you serve the adjacent segment: customers, use cases, and outcomes described in that segment’s own language |
| Size gate | “I have only seen you succeed at one company size.” | Named examples at the company size you want, with employee counts, seat counts, or deal volumes attached |
| Effort flag | “I cannot tell how long implementation takes.” | Published onboarding timelines, migration documentation, and a stated time to first value |
| Capability gap | “I found no confirmation that you do this.” | A documented feature page, integration listing, or third-party write-up that states the capability plainly |
| Cost flag | “Your pricing is opaque, so I will warn people.” | Public pricing, what each tier includes, and an honest note on what costs extra |
Notice what none of these ask for. Not one of them is solved by publishing more blog posts. Each one is solved by putting a specific, checkable fact somewhere a model can read it.
How to Run the BUT Audit
Then loop back to step one after 30 days. The audit is a cadence rather than a one-off, and the signal you are watching for is the boundary changing while your mention count stays flat.
Set aside 30 minutes. The BUT Audit runs on prompts where your brand already appears, so pull the prompt set you already track rather than writing new ones.
- Pick 20 money prompts. Use the commercial questions your buyers actually type, phrased the way a person types them. Start with the same 20 prompts you use for visibility tracking so the two datasets line up.
- Run each prompt in Google and capture the AI Overview. Use a logged-out window and keep the same day and time each round, because retrieval shifts with recency.
- Find your brand, then copy the next sentence verbatim. Do not paraphrase it, and do not copy the sentence your brand sits in. The boundary lives in what comes after.
- Tag each captured sentence to one of the five classes. Segment lock, size gate, effort flag, capability gap, or cost flag. Mark it clean when nothing qualifies the recommendation.
- Count the repeats. One instance is noise. The same class appearing across five or more of your 20 prompts is a pattern, and the pattern is your priority.
- Fix the evidence behind the top three, then rerun in 30 days. Compare the tagged sentences side by side rather than comparing whether you were mentioned.
What a clean result looks like: your brand is named, the following sentence describes what you do without attaching a condition, and the same pattern repeats across most prompts. What a scoped result looks like: your brand is named, and the next sentence hands you a drawer you did not choose.
What to Fix First, and Why Volume Will Not Do It
One class repeats across five or more of your 20 prompts. That is the pattern, not the one-off.
Example“Best suited to small teams” appeared on 9 of 12 mentions.
Translate the class into the specific fact the model could not find. Do not guess at a content topic.
ExampleNo customer above 50 employees documented anywhere.
State the fact in extractable words on your own site, then get one independent source saying the same thing.
ExampleThree mid-market customers with seat counts, plus one listing.
Rerun the same prompts after 30 days and compare the tagged sentences rather than the mention count.
Still measuringEffect size and time to change are not yet established.
Why stage four is dashed. DerivateX has measured how often boundaries appear and which classes dominate. We are still building the evidence on how quickly a repaired boundary changes the answer, so treat the audit as a diagnosis you act on rather than a guaranteed lever.
Fix the qualifier that repeats most often, and fix it with evidence rather than with content. A boundary appears because a model could not find a fact, so publishing another article about your category adds nothing it was looking for.
Take the most common case in our sample, the segment lock. A brand that keeps getting filed as “best for small teams” does not need a thought leadership piece about scaling. It needs three named mid-market customers, their seat counts, and their outcomes, published somewhere a model reads and repeated somewhere it trusts.
Order the work this way. Start with the qualifier that appears across the most prompts, because breadth of repetition beats severity.
Move next to the qualifier attached to your highest-intent prompts, which are usually comparison and alternatives queries. Leave one-off qualifiers alone until they repeat.
Two habits protect the result. Put the new evidence on your own site and get it corroborated on at least one independent source, because agreement across sources is what raises a model’s confidence. Then rerun the audit on a fixed cadence, since your qualifier can move without your mention rate changing at all.
Where the BUT Audit Does Not Apply
The BUT Audit is the wrong tool when your brand is not being named in the first place. It reads the sentence after a mention, so a brand with no mentions has nothing to read. Absence is a different diagnosis with a different fix.
Each DerivateX framework owns exactly one question. Running them out of order wastes the quarter, because a boundary cannot be fixed on a brand a model cannot name or explain.
What limit does AI attach once it names us?
You are hereRun these in order. If AI answers do not name you, start with the AI Visibility Score to establish a baseline, and check whether a model can explain you at all using the 4 C’s of being explainable to AI. Once you are being named, the BUT Audit tells you what the mention is costing you.
The audit also will not fix a product that genuinely does not serve the segment you want. When a model files you as best for small teams and you have no mid-market customers, the qualifier is accurate. That is a positioning and proof problem before it is an AI search problem.
FAQ
How is the BUT Audit different from AI sentiment analysis?
Sentiment analysis sorts language into positive, neutral, and negative. The BUT Audit ignores tone and classifies scope instead. A line like “best for freelancers” scores positive on any sentiment tool while removing the brand from every enterprise query, which is precisely the failure sentiment scoring is built to miss.
How often should I run the BUT Audit?
Monthly is enough for most B2B SaaS companies, and quarterly works for slower categories. Run it on the same 20 prompts each time, on the same weekday, so the comparison stays clean. Rerun sooner after publishing evidence intended to move a specific qualifier.
Why does Google AI Overviews say my software is only for small businesses?
In most cases the model found proof of small-business use and no comparable proof at larger sizes. Company-size gates followed roughly 16% of the brand recommendations in our sample. The repair is named customers at the size you want, with seat counts and outcomes stated plainly on pages a model can read.
Can I run the BUT Audit without a tool?
Yes. The audit needs a browser, a spreadsheet with five tag columns, and about 30 minutes. DerivateX designed it to run manually because the value sits in reading the sentences carefully, not in automating the collection.
The One Thing to Take Away
How AI describes your brand is a sorting decision, and the sort happens in one sentence you are probably not reading. Google AI Overviews filed 44% of the brands in our study into a drawer, and the label was usually a compliment.
Read the next sentence. Then go find the fact that would change it.
If you want the scored version run for you across Google AI Overviews, ChatGPT, Perplexity, Gemini, and Claude, with your repeating qualifiers ranked and the evidence gaps named, that is exactly what our free AI visibility audit delivers.









