Home / Blog / How Prompt Selection Shapes Your AI Visibility Score
How Prompt Selection Shapes Your AI Visibility Score
Published: July 31, 2026
Share on LinkedIn Share on Twitter Share on Facebook Click to print Click to copy url
Contents Overview
Somebody chose the questions behind your AI visibility report. Maybe a strategist wrote them. Maybe a platform suggested them and a strategist signed off. Either way, that score isn’t measuring your brand on its own. It’s measuring your brand against a list someone decided should stand in for the market.
Keyword research had the same selection problem. But it also had an outside number to argue with: estimated search volume. You could dispute the estimate, the seed terms, or the intent labels. You couldn’t claim the number came from your own team.
Prompt volume tools are starting to fill that gap, and they’re genuinely useful for sizing a subject like “ecommerce,” “digital PR” or “AI visibility.” What they can’t tell you is whether a buyer asks for a “GEO agency” or spends two sentences describing a company, a competitor and a problem. Both sit inside the same topic, but that doesn’t mean they produce the same answer.
That matters because the same need reaches an assistant in dozens of forms. One person knows the category term. Another only knows that competitors keep showing up in recommendations. A third adds company size, channel, platform, and everything they’ve already tried.
So the list has to be built by hand. That isn’t manipulation. It’s a measurement decision. But it’s a decision that should be visible, because whoever writes the questions helps set the score.

The list gets written by one team. The number gets defended by another.
We tested this on our own brand across ChatGPT, Google AI Overviews and Perplexity, from July 20 through July 26, 2026. Our experiment held 1,554 completed responses across 73 unique prompt wordings. Under one kind of language the brand looked strong. Under another it was nearly absent.
What Happens After Someone Asks for a Recommendation
Often the system doesn’t search using the sentence a person typed. Google describes a “query fan-out“ process in which AI features may issue multiple related searches across subtopics and data sources before assembling a response.
Two platforms in our export exposed their search logs. They didn’t behave the same way.
Here’s a 24-word prompt:
“Looking for a GEO agency that has actually worked with B2B SaaS companies and can show real case studies. Who are the top options?”
| Platform | What the log showed for this one prompt |
| ChatGPT | 13 search queries across seven runs; all 13 were different. |
| Perplexity | The full prompt was searched verbatim in four runs; the same shorter rewrite was used in the other three. |
Across every visible query string, longer prompts were cut down harder. The table below compares search-query length against prompt length. It doesn’t prove the system ignored the missing context. It shows how much less of that context stayed explicit in the query text.
| Original prompt length | Visible query strings | Average reduction in word count |
| Under 12 words | 747 | 11.2% shorter |
| 12 to 17 words | 188 | 27.1% shorter |
| 18 words or more | 186 | 56.1% shorter |
Microsoft now shows a version of this from the publisher side. Bing Webmaster Tools’ AI Performance report, launched in public preview, includes a sampled view of grounding queries: the key phrases Microsoft’s AI systems used to retrieve content that was later cited. Useful for understanding retrieval. Not a stand-in for the language a buyer actually used.
So the front of the workflow has changed. You still choose what to test. The system may then rewrite, split or expand it after you hit send. Prompt construction isn’t the whole visibility problem. It’s the first measurement decision.
How We Ran the Analysis
The data supports a strong observation. It doesn’t support a clean causal experiment. The categories differ in wording, length and specificity, so here are the boundaries up front.
| Element | Method used |
| Study window | July 20-26, 2026 |
| Market | United States |
| Platforms | ChatGPT, Google AI Overviews and Perplexity |
| Response count | 1,554 completed responses, 518 per platform |
| Question set | 73 unique prompt wordings. Seventy-two ran seven times per platform; one identical wording ran fourteen times per platform. |
| Primary outcome | Whether the answer named Go Fish Digital. Citation presence and recommendation position were not treated as equivalent to a brand mention. |
| Weighting | Each unique prompt carries equal weight in category-level rates, so the duplicated wording doesn’t count twice. |
| Industry language | Prompts containing “GEO,” “generative engine optimization,” “AEO” or “answer engine optimization.” |
| Problem-led language | Prompts without those industry labels and without the full-situation “realistic buyer” tag. |
| Full buyer situations | Realistic buyer prompts without an industry label. Two realistic prompts that used “GEO” stayed in the industry-language group. |
What This Design Can and Cannot Establish
The study can show that reported visibility moved sharply across different kinds of prompts in this dataset. It can’t prove that one vocabulary choice caused the move. Full-situation prompts also ran longer, more specific and more constrained, and any of that may affect retrieval and recommendation behavior alongside the terminology.
That’s why what follows says “associated with,” “sensitive to” and “diagnostic.” None of it treats personas, prompt length or category terms as ranking factors.
How Much Question Framing Changed Our Visibility
Give each unique prompt equal weight and Go Fish Digital was named in 23.6% of responses to category-led prompts, against 3.6% of responses to full buyer situations written without the industry label. In this sample, that’s a 6.6-times difference.
| How the question was framed | Unique prompts | Observed runs | Prompt-balanced mention rate |
| Category-led: used GEO/AEO terminology | 35 | 756 | 23.6% |
| Problem-led: no industry label | 30 | 630 | 7.6% |
| Full buyer situation: no industry label | 8 | 168 | 3.6% |
The obvious objection is intent. Category language usually signals a buyer who already knows what service they want. So we narrowed to the 49 prompts tagged commercial, which holds buying intent roughly steady and leaves vocabulary as the main thing that differs. Inside that narrower group the gap held: 24.0% with the industry label, 8.6% without.
That removes one source of difference. It doesn’t leave vocabulary as the only variable. Prompt length, vertical, specificity and requested outcome all still vary. Read it as evidence that prompt-set composition matters, not as proof that a single term creates visibility.
The pattern appeared on all three platforms:
| Platform | Category-led prompts | Full buyer situations |
| ChatGPT | 31.6% | 3.6% |
| Google AI Overviews | 28.4% | 7.1% |
| Perplexity | 10.8% | 0.0% |
The absolute numbers moved a lot by platform, which is exactly why one blended score hides more than it shows. The direction didn’t move at all. Our brand turned up far more often when the question already used the category’s professional vocabulary.
Wording Sensitivity Reveals What the Score is Measuring
The useful question isn’t which agency appears most often. It’s what each type of prompt actually tests.
Category-led prompts test whether a brand is tied to the professional name of a service. Situation-led prompts test whether that connection holds when a buyer knows the problem but not yet what the solution is called.
Those are two different measurements. A score built mostly from category terminology can describe visibility among category-aware buyers accurately and still tell you very little about buyers who are describing symptoms, constraints and business problems.
A separate Rankshift study points the same way. Across seven near-identical CRM prompts, an established brand held steady while lower-visibility brands moved more as small wording choices changed. That doesn’t turn wording sensitivity into a brand-ranking metric. It suggests prompt variation can show how consistently a brand is connected to the problem being discussed.
The Growth Opportunity Sits Outside Category Vocabulary
Buyers who say “GEO agency” already know the category exists. A company can be visible to that group and still show up inconsistently when a buyer describes the problem in ordinary language: competitors appear in ChatGPT, product recommendations skip the brand, or AI Overviews absorb the discovery traffic.
Which is why an overall mention rate isn’t enough on its own.
The distance between category-led, problem-led and situation-led prompts is a useful diagnostic, though prompt length, specificity and retrieval behavior may also feed into it. We treat our own gap as a workstream across positioning, content, digital PR and third-party coverage. Not as a verdict about who’s winning the category.
How to Measure AI Visibility Without Fooling Yourself
None of this is an argument against prompt tracking. Every number in this article came out of prompt tracking, and without it the vocabulary gap would have stayed invisible while the headline score looked fine. It’s the only method that surfaces any of this. The problem isn’t the tracking. It’s what goes into it, and what gets reported out.
You can’t take the judgment out of this. The goal is to make the judgment visible, repeatable, and tied to a buyer population you’ve actually named.
1. Decide what the score is supposed to represent. A category-demand benchmark, a problem-aware buyer benchmark and an executive composite answer three different questions. Name the population before you pick the prompts. A prompt set is only inflated when it overweights language the population you named wouldn’t use.
2. Write the same buyer need at three distances from your vocabulary. Build a category-led version, a problem-led version and a full situation-led version. Hold the underlying need as steady as you can. You end up with one question written three ways, not three unrelated questions.
3. Start with people for the situation-led version. Use sales-call transcripts, support tickets, win-loss interviews and free-text form responses. That’s where buyers explain the problem before anyone hands them the company’s preferred terminology. AI tools can help cluster and expand that language. They shouldn’t invent the buyer’s voice from scratch.
4. Keep buyer prompts and grounding queries in separate datasets. Buyer language tells you what to test. Grounding queries tell you how the system retrieved. Feed machine-written retrieval phrases back into your buyer prompts and the measurement starts reflecting the system’s vocabulary instead of the market’s.
5. Report the distribution, not just the average. Break results out by platform, prompt family and run frequency. A composite is fine for executive reporting, as long as it sits on top of the breakdown rather than replacing it.
Our Answer is Not a Measurement
Every standard prompt-platform combination ran seven times. That repetition is what exposed how unstable a single screenshot can be.
| Go Fish mentions across seven runs | Prompt-platform combinations |
| 0 of 7 | 148 |
| 1 of 7 | 27 |
| 2 of 7 | 8 |
| 3 of 7 | 8 |
| 4 of 7 | 5 |
| 5 of 7 | 4 |
| 6 of 7 | 8 |
| 7 of 7 | 8 |
Across all 216 unique prompt-platform combinations, 68 named Go Fish Digital at least once. Only eight named it in every run. A screenshot proves an answer happened. It can’t show you how representative that answer was.
Never Accept a Blended Score Without the Platform Breakdown
Platform differences aren’t a rounding issue. Ahrefs analyzed 730,000 AI Overview and AI Mode response pairs for semantic similarity and 540,000 pairs for citation overlap. It reported 86% average semantic similarity but only 13.7% overlap among cited URLs. Two Google surfaces can reach similar conclusions off substantially different source sets.
Our own numbers said the same thing. Category-led visibility ran from 10.8% on Perplexity to 31.6% on ChatGPT. A composite score can summarize that spread, but only if the weighting is disclosed and the engine-level results stay available.
What This Data Cannot Tell You
The result is useful because its boundaries are clear. Eight limitations matter most.
- It isn’t a causal wording experiment. The groups differ in length, specificity, industry context and constraints as well as terminology.
- It covers one brand and one service category. Don’t generalize the size of the gap to every company or market.
- It measures three surfaces, not the full AI ecosystem. Gemini, Claude, Copilot, Grok and Google AI Mode aren’t in the dataset.
- It’s a United States snapshot. Location, language and market availability can change both retrieval and recommendations.
- Session conditions can move the result. Account history, memory, personalization and product configuration may all affect an answer.
- Visible search logs are incomplete. Search-query text was missing for some runs, and what’s there may not expose the full retrieval or ranking process.
- A mention isn’t a recommendation. The analysis doesn’t score sentiment, prominence, citation quality or conversion impact.
- The systems keep changing. Every percentage here belongs to the study window. Treat it as a dated benchmark.
The Measurement Still Starts With a Human Judgement
Keyword research hasn’t gone anywhere. Neither has the need for technical accessibility, useful content, third-party authority and a clear brand position. What’s gone is the certainty we used to borrow from a visible query and an estimated volume.
But the fundamentals holding doesn’t mean the work is the same. The unit of research is a question set you build rather than a keyword list you look up. The measurement is mention and citation rates per platform with a range attached, rather than a ranking position. And retrieval is non-deterministic, so the same question can answer differently an hour later. That’s a different discipline running on familiar foundations, and treating it as a checkbox on an SEO retainer is how you end up with a number nobody can defend.
AI visibility measurement starts with a question set somebody designed. The platform may then rewrite those questions, search different sources, and hand back a different answer on the next run. That makes methodology part of the result.
The question isn’t whether an AI visibility score is correct. It’s which buyer questions that score was built to represent.
We ran one brand through different kinds of language and got dramatically different estimates of how visible it is. No tracking tool can decide which set represents your market. That takes knowing the buyer, writing the assumptions down, and being willing to show the question list.
Get a clearer view of your AI visibility with a free GEO Scorecard. Our experts will assess your LLM performance, surface competitors taking traffic and attention, and recommend a path to stronger visibility.
Key Takeaways
- AI visibility scores are shaped by the question set. The score doesn’t measure a brand in isolation. It measures that brand against prompts someone picked to stand in for the market.
- Prompt framing can change the result dramatically. In this study, Go Fish Digital was mentioned in 23.6% of category-led prompts but only 3.6% of full buyer-situation prompts. A 6.6-times difference.
- AI platforms often rewrite what users ask. Longer prompts were cut down harder in visible search queries, so much of the buyer’s context may not stay explicit during retrieval.
- One answer or one blended score isn’t enough. Results moved substantially by platform and across repeated runs. Report AI visibility by platform, prompt family and frequency, not as a single unsupported average.
- The biggest growth opportunity usually sits outside industry language. Measure whether you appear when buyers describe their problem in ordinary words, not only when they already know terms like “GEO agency.” That takes real buyer research, transparent methodology, and access to the underlying question list.
Proprietary dataset
The analysis used 1,554 completed responses from July 20-26, 2026, after removing one blank row from the supplied CSV export.
About Ashish Jacob
MORE TO EXPLORE
Related Insights
More advice and inspiration from our blog
Your Reviews Are Now GEO: What the Yelp and ChatGPT Deal Means for Your Brand
One of the biggest assumptions in GEO just broke. Most brands...
Tiffany Crockett| July 29, 2026
Data Report: Why Organic Traffic Benchmarks Mislead You in the GEO Era
Ahrefs released median organic traffic data from more than 344,956 real...
Shane White| July 24, 2026
SEO vs GEO vs AEO: What They Are, Differences & How To Use Them
Search is changing, and so is the way brands earn visibility....





