LLMClarity

June 16, 2026

Run the Same AI Prompt Three Times and You'll Get Three Different Competitor Lists

Run the same AI prompt three times and you'll get three different competitor lists. Here's why a single check is statistically meaningless — and how to turn noisy answers into a share-of-voice number you can actually track.

By Mike Morris — Founder, Kettle Hole Partners & LLMClarity

Here's a test you can run in the next ten minutes. Open ChatGPT, ask it "Who are the best plumbers in Austin?", and write down the names it gives you. Then start a fresh chat and ask the exact same question. Then do it a third time.

You will not get the same list three times. You probably won't get the same list twice.

This matters more than it sounds, because most people checking how AI describes their business do it exactly once. They ask the question, screenshot the answer, and treat that screenshot as the truth. It isn't the truth. It's one draw from a distribution, and a single draw tells you almost nothing about the distribution it came from.

Why the output moves

These models are non-deterministic by design. When an assistant generates an answer, it's choosing each word from a set of probabilities, not reading from a fixed table. There's a built-in dose of randomness so the output doesn't read like a robot every time. That randomness is fine when you're asking it to write an email. It's a real problem when you're trying to measure something.

So when you ask "who are the best plumbers in Austin," the model isn't looking up a ranked list. It's assembling a plausible answer on the fly, and the names that surface shift run to run. A competitor who shows up first in one answer might not appear at all in the next.

A worked example

I ran a single prompt — "Who are the top home security companies for a homeowner?" — ten times in fresh sessions, same model, same wording, nothing else changed. I just counted how often each company got named.

Here's what came back across the ten runs:

  • Company A: named in 10 of 10 runs
  • Company B: named in 9 of 10
  • Company C: named in 7 of 10
  • Company D: named in 5 of 10
  • Company E: named in 4 of 10
  • Company F: named in 3 of 10
  • Company G: named in 2 of 10
  • Three more companies: named once each

Now look at what that means depending on when you checked.

If you'd run the prompt once and happened to catch run number 3, you'd have seen Company D and Company E in the list and concluded they were strong players. If you'd caught run number 7, neither of them showed up, and you'd have concluded the opposite. Same prompt. Same model. Same day. Two completely different stories about who your competition is.

Company F is the one that should worry you most. It appeared in 3 of 10 runs — a real, recurring presence, not noise. But there's a 70% chance any single check misses it entirely. A business doing one snapshot would never know that company was in the conversation a third of the time.

And the companies named once each? On a single check, one of those might be the first name you see. You'd walk away thinking a fringe player was a serious competitor, when the data says it surfaces 10% of the time.

The honest version of the numbers

Look at the gap between the top of the list and the middle. Company A at 10 of 10 is stable — you can trust that. Company B at 9 of 10 is stable too. But everything from Company C down is variable, and the variability is the whole point. The middle of the list is where the interesting competitive information lives, and the middle of the list is exactly where a single check is least reliable.

If you want a number you can act on, you need the frequency, not the snapshot. "Company F appears in 30% of answers" is something you can track over time and compare against last month. "Company F was in the answer I happened to check" is not. One is a measurement. The other is an anecdote.

What to do about it

The fix isn't complicated, it's just discipline. Treat AI visibility the way you'd treat any noisy signal: sample it, don't snapshot it.

  1. Run the prompt multiple times, in fresh sessions. Ten is a reasonable floor for one prompt on one model. Fewer than that and your frequencies are too lumpy to trust.

  2. Count appearances, not positions. Position within a single answer is even noisier than presence. Start by tracking how often each name shows up at all.

  3. Do it across the assistants people actually use. ChatGPT, Claude, Gemini, Perplexity, and Google AI don't agree with each other, and they each have their own variance on top of that. A frequency from one tells you nothing about the others.

  4. Repeat on a schedule. A frequency is only useful if you can compare this month to last month. One reading, no matter how well sampled, is still a single point.

This is tedious to do by hand, which is most of why people don't do it and settle for the screenshot instead. It's the part of the work LLMClarity handles — running the prompts enough times across all five assistants to turn a noisy answer into a share-of-voice number you can actually watch move.

But the tool is secondary to the point. The point is that a single AI check is statistically meaningless, and you can prove it to yourself in ten minutes with the test at the top of this post. Before you decide what AI is saying about your business, make sure you've asked enough times to know.

Get Started

See your AI visibility.