Here's a question I'd ask any local-services marketer right now: when someone asks ChatGPT to recommend a plumber in your city, do you come up? You probably don't know. Most people don't. They have a strong opinion about how they should show up, and no data on how they actually do.
That's a problem, because the instinct is to skip straight to fixing it. Add some schema markup, write a few more pages, chase whatever the latest "get cited by AI" advice happens to be. But you can't tell whether any of that worked if you never measured where you started. You'd be running tests with no baseline — adding effort and hoping, which is just a more expensive version of guessing.
So before you change anything, measure. Here's a concrete method you can run this afternoon. It takes about an hour and costs nothing but your time.
The 5-prompt baseline method
The idea is simple. Pick the handful of questions a real prospect would type, ask them across the major AI assistants, and write down exactly what comes back. You're not optimizing anything yet. You're taking a photograph of where you stand today so you have something to compare against later.
Step 1: Write five prompts the way a customer would.
Not the way a marketer would. A marketer types "best HVAC company Austin." A customer types things like:
- "Who are the best HVAC companies in Austin?"
- "I need emergency AC repair in Austin tonight, who should I call?"
- "Most reliable HVAC installation near me in Austin" (with location enabled)
- "Compare the top-rated heating and cooling companies in Austin"
- "Affordable HVAC repair Austin with good reviews"
Five is enough to start. You want a mix: a general "best of" question, an urgent one, a comparison one, and one anchored on price or reviews. These map to how people actually search now — in full sentences, with intent baked in.
Step 2: Run each prompt across all five assistants.
The ones worth checking are ChatGPT, Claude, Gemini, Perplexity, and Google AI. Don't just check one. They pull from different sources and they disagree more than you'd think. I've seen a business get named in three of five assistants and be completely invisible in the other two. If you only checked the two where you appear, you'd walk away thinking you're in great shape.
That's five prompts times five assistants — 25 answers. It sounds like a lot. It goes fast once you've got the prompts written.
Step 3: Record three things for every answer.
For each of the 25 responses, write down:
- Did you get mentioned? Yes or no. Nothing fancy.
- Where in the list? First name out of the gate, buried at position six, or only mentioned if the user asks a follow-up. Position matters. Being the fourth of seven names is not the same as being the recommendation.
- Who else got named? This is the part people skip, and it's the most useful part. The competitors that show up are your actual competitive set in AI answers — which may be different from the competitors you think about. Write down every name.
A simple spreadsheet does the job. Prompts down the side, assistants across the top, and in each cell: mentioned (Y/N), position, and competitors named.
What the numbers tell you
Once the grid is filled in, you can compute a rough share of voice. Count the cells where you appear, divide by 25. If you show up in 8 of 25, that's 32%. Write that number down and date it. That's your baseline.
Then look at the competitor names. Tally how often each one appears. You'll usually find two or three businesses that show up far more than the rest. Those are the ones AI assistants currently treat as the default answer in your market. If a competitor appears in 20 of 25 cells and you appear in 8, that gap is the real story — not your absolute number.
I'd pay special attention to the disagreements between assistants. When one names you first and another doesn't name you at all, that tells you the underlying signals aren't consistent. That's diagnostic. It usually means your presence is thin in whatever sources that particular assistant leans on.
Why you do this before anything else
Two reasons.
First, you can't measure improvement without a starting point. If you go make changes and re-run the grid in 90 days and your share of voice went from 32% to 44%, you've learned something real. If you never ran the first grid, you've learned nothing — you just have a number with no context. This is the same discipline as any acquisition test: establish the baseline, then test against it.
Second, the baseline tells you where to aim. Maybe you appear fine in the "best of" prompts but vanish on the emergency ones. Maybe you're strong in ChatGPT and absent in Perplexity. Maybe a competitor you'd never worried about is the default name everywhere. You don't know any of that until you measure, and each of those findings points to a different next move.
The mistake is treating AI visibility as a thing you fix on instinct. It isn't. It's a thing you measure, and the measurement comes first.
Doing it more than once
The hour-long manual version is perfect for a one-time snapshot. The catch is that these answers change. Re-run the same grid in a month and you'll get different results — sometimes meaningfully different — because the assistants update constantly. So a single audit is a photograph, not a trend. To see whether you're actually gaining or losing ground, you need the same prompts run on a schedule and tracked over time. That's the part worth automating, and it's what we built LLMClarity to do: track what the five assistants say about you over time. But you don't need a tool to start. You need a spreadsheet and an honest hour.
Run the five prompts. Fill in the grid. Write down your share of voice and the date. Then, and only then, start thinking about what to change.
