How AI Describes Your Brand: The BUT Audit

How AI describes your brand decides more deals than whether AI mentions your brand at all. Google AI Overviews attach a boundary to almost every recommendation, and DerivateX found one on 44% of the 445 brand recommendations we studied. The BUT Audit is the method for reading that boundary and fixing what sits behind it.

Most B2B SaaS teams stop measuring the moment their name appears. They screenshot the mention, drop it in Slack, and move on. The sentence that decides whether the recommendation survives is the one they never read.

What the BUT Audit Is

The BUT Audit is a diagnostic that reads the sentence immediately after your brand appears in an AI answer, classifies the boundary that sentence places on the recommendation, and counts which boundaries repeat across your buyer prompts. DerivateX built it because a brand can be named in an answer and still lose the deal inside the same paragraph.

The name comes from the connector that used to signal the problem: “X is great, but…”. Real AI Overview answers use that structure less often than people assume. They use a friendlier one that does the same damage.

Think about what “Best for freelancers on a tight budget” actually does to a company selling into 200-person teams. The line is complimentary, accurate, and completely disqualifying. A buyer running a mid-market evaluation reads it and moves down the list.

This is why the BUT Audit is not sentiment analysis. Sentiment tools sort language into positive, neutral, and negative, and a scope statement lands in the positive bucket every time.

One sentence, two readings Prompt: best accounting software for a growing business

Your brand is a clean, well-designed option covering invoicing and expense tracking. Best for solo consultants and very small service businesses.

How a sentiment tool reads it
Positive

Complimentary adjectives, no negation, no complaint. Scores at the top of the range and clears the audit.

  • Tone: favourable
  • Complaint detected: none
  • Action flagged: none
How a buyer reads it
Not for us

The clause is a filter. A 200-person finance team reads it, rules the brand out, and moves to the next name on the list.

  • Mid-market shortlist
  • Finance team queries
  • Multi-entity comparisons
The blind spot. Sentiment analysis measures tone. The BUT Audit measures scope. Across the 445 brand recommendations in the DerivateX study, roughly 44% carried a scope clause like this one and only about 5% carried an actual complaint.

In the DerivateX State of AI Visibility benchmark, nearly nine in ten of the B2B SaaS companies we tested scored at the very top of the sentiment range, and plenty of them still could not win a shortlist. Being liked was never the constraint.

What Google AI Overviews Actually Say After They Name You

In a DerivateX study of 100 Google AI Overview answers run across 20 software categories in June 2026, 44% of the 445 individual brand recommendations carried a qualifier in the next sentence. We logged every answer while signed out, pulled each recommended product, and tagged the clause that followed the brand’s first appearance.

Here is how those qualifiers break down across all 445 recommendations.

Qualifier classHow it sounds in the answerShare of recommendations
Segment lock“Best for”, “ideal for”, “designed for”24.7%
Size gate“small teams”, “freelancers”, “enterprise environments”15.7%
Effort flag“complex”, “requires technical setup”, “learning curve”9.0%
Capability gap“lacks”, “limited”, “fewer”, “does not include”2.9%
Cost flag“expensive”, “per-user pricing”, “premium”2.0%
Any qualifier present44%

Three numbers from that sample are worth holding onto.

The phrase “best for” appeared 245 times across the 100 answers, averaging two and a half explicit segment locks per answer. Google AI Overviews are not writing a ranked list. They are writing a sorting exercise, and every brand gets filed into a drawer.

Three archetypes: find yours in four seconds

Every category in the matrix sorts brands on a dominant axis. Identify which axis yours runs on before you decide what evidence to build.

Sorted by who you serve
Compliance automation

Also: AI time tracking, iPaaS, CRM

Segment lock45
Size gate35
Effort flag20

What to build: proof you serve the adjacent use case, written in that segment’s own words.

Sorted by company size
Accounting

Also: help desk, ITSM, business intelligence

Size gate38
Segment lock19
Cost flag5

What to build: named customers at the size you want, with seat counts and outcomes attached.

Barely sorted at all
HR software

Also: tech pack, project management, video hosting

Segment lock14
Effort flag5
Size gate0

What to build: not much here yet. Spend the quarter on getting named rather than on widening scope.

Read as: percentages are the share of recommended brands in that category whose next sentence carried that boundary class. Base: 445 brand recommendations across 100 Google AI Overview answers, 20 B2B software categories, June 2026. Category samples run between 19 and 36 recommendations, so treat these as directional.

Roughly half the answers, 49 out of 100, ended by asking the reader a qualifying question back. Questions like “How large is your team?” and “What is your annual revenue?” show the model narrowing a shortlist by segment before the buyer has clicked anything. Your drawer label is what decides whether you survive that narrowing.

Qualification also varies sharply by category. In our sample, compliance automation queries qualified 70% of recommended brands, while HR software queries qualified 18%. Categories where buyers carry hard constraints get sorted hardest, which means regulated and technical categories need this audit more than most.

One limitation belongs in the same breath as the finding. These figures describe Google AI Overviews in one week of June 2026, across commercial “best software” queries in 20 categories. They are a strong signal about how AI describes brands in that surface, and they are not a measurement of ChatGPT, Perplexity, Gemini, or Claude, which DerivateX tracks separately.

The Five Qualifier Classes and What Each One Is Telling You

Every qualifier is a confession. The model is telling you which piece of evidence it could not find, which is far more useful than knowing your sentiment score.

Qualifier classWhat the model is really sayingThe evidence that is missing
Segment lock“I can only place you in one drawer.”Proof you serve the adjacent segment: customers, use cases, and outcomes described in that segment’s own language
Size gate“I have only seen you succeed at one company size.”Named examples at the company size you want, with employee counts, seat counts, or deal volumes attached
Effort flag“I cannot tell how long implementation takes.”Published onboarding timelines, migration documentation, and a stated time to first value
Capability gap“I found no confirmation that you do this.”A documented feature page, integration listing, or third-party write-up that states the capability plainly
Cost flag“Your pricing is opaque, so I will warn people.”Public pricing, what each tier includes, and an honest note on what costs extra

Notice what none of these ask for. Not one of them is solved by publishing more blog posts. Each one is solved by putting a specific, checkable fact somewhere a model can read it.

How to Run the BUT Audit

The BUT Audit loop 30 minutes per round
Step 01 Find Run your 20 fixed buyer prompts in a logged-out window. 20 answers captured
Step 02 Read Copy the sentence that follows your brand, word for word. Verbatim sentences
Step 03 Tag Assign each sentence to one of five boundary classes, or mark it clean. Tagged log
Step 04 Count Total the tags. Five or more of one class is a pattern worth acting on. BUT Score
Step 05 Fix Publish the missing evidence, then corroborate it off your own domain. Evidence live

Then loop back to step one after 30 days. The audit is a cadence rather than a one-off, and the signal you are watching for is the boundary changing while your mention count stays flat.

Keep the prompt set fixed for at least 90 days. Changing prompts between rounds makes the boundary trend meaningless, because you are no longer measuring the same questions.

Set aside 30 minutes. The BUT Audit runs on prompts where your brand already appears, so pull the prompt set you already track rather than writing new ones.

  1. Pick 20 money prompts. Use the commercial questions your buyers actually type, phrased the way a person types them. Start with the same 20 prompts you use for visibility tracking so the two datasets line up.
  2. Run each prompt in Google and capture the AI Overview. Use a logged-out window and keep the same day and time each round, because retrieval shifts with recency.
  3. Find your brand, then copy the next sentence verbatim. Do not paraphrase it, and do not copy the sentence your brand sits in. The boundary lives in what comes after.
  4. Tag each captured sentence to one of the five classes. Segment lock, size gate, effort flag, capability gap, or cost flag. Mark it clean when nothing qualifies the recommendation.
  5. Count the repeats. One instance is noise. The same class appearing across five or more of your 20 prompts is a pattern, and the pattern is your priority.
  6. Fix the evidence behind the top three, then rerun in 30 days. Compare the tagged sentences side by side rather than comparing whether you were mentioned.

What a clean result looks like: your brand is named, the following sentence describes what you do without attaching a condition, and the same pattern repeats across most prompts. What a scoped result looks like: your brand is named, and the next sentence hands you a drawer you did not choose.

What to Fix First, and Why Volume Will Not Do It

From boundary to repair Worked on a repeating size gate
1
Boundary detected

One class repeats across five or more of your 20 prompts. That is the pattern, not the one-off.

Example“Best suited to small teams” appeared on 9 of 12 mentions.

2
Missing evidence named

Translate the class into the specific fact the model could not find. Do not guess at a content topic.

ExampleNo customer above 50 employees documented anywhere.

3
Published and corroborated

State the fact in extractable words on your own site, then get one independent source saying the same thing.

ExampleThree mid-market customers with seat counts, plus one listing.

4
Boundary moves

Rerun the same prompts after 30 days and compare the tagged sentences rather than the mention count.

Still measuringEffect size and time to change are not yet established.

Why stage four is dashed. DerivateX has measured how often boundaries appear and which classes dominate. We are still building the evidence on how quickly a repaired boundary changes the answer, so treat the audit as a diagnosis you act on rather than a guaranteed lever.

Fix the qualifier that repeats most often, and fix it with evidence rather than with content. A boundary appears because a model could not find a fact, so publishing another article about your category adds nothing it was looking for.

Take the most common case in our sample, the segment lock. A brand that keeps getting filed as “best for small teams” does not need a thought leadership piece about scaling. It needs three named mid-market customers, their seat counts, and their outcomes, published somewhere a model reads and repeated somewhere it trusts.

Order the work this way. Start with the qualifier that appears across the most prompts, because breadth of repetition beats severity.

Move next to the qualifier attached to your highest-intent prompts, which are usually comparison and alternatives queries. Leave one-off qualifiers alone until they repeat.

Two habits protect the result. Put the new evidence on your own site and get it corroborated on at least one independent source, because agreement across sources is what raises a model’s confidence. Then rerun the audit on a fixed cadence, since your qualifier can move without your mention rate changing at all.

Where the BUT Audit Does Not Apply

The BUT Audit is the wrong tool when your brand is not being named in the first place. It reads the sentence after a mention, so a brand with no mentions has nothing to read. Absence is a different diagnosis with a different fix.

Where the BUT Audit sits Run the layers bottom to top

Each DerivateX framework owns exactly one question. Running them out of order wastes the quarter, because a boundary cannot be fixed on a brand a model cannot name or explain.

04
The BUT Audit

What limit does AI attach once it names us?

You are here
03

How do we get recommended on purpose?

Execution
02

Are we named at all, and how prominently?

Measurement
01

Can a model explain us in one clean line?

Foundation
Read as: the 4 C’s decide whether a model can describe you, the AI Visibility Score decides whether it does, Citation Engineering builds the sources that make it happen, and the BUT Audit measures the limit attached to the result. Skip a layer and the one above it has nothing to work with.

Run these in order. If AI answers do not name you, start with the AI Visibility Score to establish a baseline, and check whether a model can explain you at all using the 4 C’s of being explainable to AI. Once you are being named, the BUT Audit tells you what the mention is costing you.

The audit also will not fix a product that genuinely does not serve the segment you want. When a model files you as best for small teams and you have no mid-market customers, the qualifier is accurate. That is a positioning and proof problem before it is an AI search problem.

FAQ

How is the BUT Audit different from AI sentiment analysis?

Sentiment analysis sorts language into positive, neutral, and negative. The BUT Audit ignores tone and classifies scope instead. A line like “best for freelancers” scores positive on any sentiment tool while removing the brand from every enterprise query, which is precisely the failure sentiment scoring is built to miss.

How often should I run the BUT Audit?

Monthly is enough for most B2B SaaS companies, and quarterly works for slower categories. Run it on the same 20 prompts each time, on the same weekday, so the comparison stays clean. Rerun sooner after publishing evidence intended to move a specific qualifier.

Why does Google AI Overviews say my software is only for small businesses?

In most cases the model found proof of small-business use and no comparable proof at larger sizes. Company-size gates followed roughly 16% of the brand recommendations in our sample. The repair is named customers at the size you want, with seat counts and outcomes stated plainly on pages a model can read.

Can I run the BUT Audit without a tool?

Yes. The audit needs a browser, a spreadsheet with five tag columns, and about 30 minutes. DerivateX designed it to run manually because the value sits in reading the sentences carefully, not in automating the collection.

The One Thing to Take Away

How AI describes your brand is a sorting decision, and the sort happens in one sentence you are probably not reading. Google AI Overviews filed 44% of the brands in our study into a drawer, and the label was usually a compliment.

Read the next sentence. Then go find the fact that would change it.

If you want the scored version run for you across Google AI Overviews, ChatGPT, Perplexity, Gemini, and Claude, with your repeating qualifiers ranked and the evidence gaps named, that is exactly what our free AI visibility audit delivers.

Apoorv Sharma
Apoorv Sharma

Apoorv Sharma is the co-founder of DerivateX, a B2B SaaS SEO and Generative Engine Optimization agency that engineers AI citations in ChatGPT, Perplexity, Claude, and Gemini and connects them to demo bookings and revenue pipeline.

He is the author of the 2026 AI Visibility Benchmark Report and the Citation Engineering methodology.

He's also the brain behind "Found On AI" and has sold 2 of his companies previously