Konabos

Stop Pretending AI Has Rankings: Measuring AI Visibility

Akshay Sura - Partner

10 Aug 2026

Share on social media

A client asked me last month where they rank in ChatGPT. Fair question. I have spent about 20 years in the CMS and search world, and my honest answer probably was not what they wanted to hear. You do not rank in ChatGPT, at least not the way you rank in Google, and anyone selling you that number is selling you a screenshot.

Let me be careful, because the loud version of this take is wrong too. I am not saying AI visibility cannot be measured. It can. I am saying there is a difference between measuring AI visibility and pretending AI has rankings. One is real work. The other is precision theater. Most of what is being sold right now blurs the two, usually on purpose.

Google gives you something real. AI answers mostly do not.

I will be fair to old-school SEO first. Google was never perfectly stable. Location, device, personalization, and time all move the results, and Search Console reports average position, not some universal rank carved in stone. But there is still something concrete underneath it: an ordered results page, plus impressions, clicks, and position data straight from the search engine. You can screenshot it and spend against it.

AI answer engines do not work like that, and the wording matters. Depending on the product and the mode, an AI assistant might answer straight from what the model already knows, or it might go search the web. ChatGPT Search, Perplexity, Gemini, Copilot, and Google's AI Mode retrieve, rerank, personalize, and then generate an answer. Some of them show source cards. None of them hand a publisher a stable, trackable position the way a results page does.

There is also no Search Console for ChatGPT, Claude, or Perplexity. No first-party feed telling publishers where they showed up. Google and Microsoft are starting to expose a little of this, and honestly it helps my case. Microsoft's new AI Performance report in Bing Webmaster Tools shows how often your pages get cited in AI answers, and Microsoft says right there in the product that this does not indicate ranking, authority, or a page's role in an individual answer. Sit with that for a second. Microsoft has the actual data and still will not call citation visibility a ranking. So why would I take that label from a third-party tracker that does not have the data?

So two different questions are getting mixed up. Search rank asks where your URL sat in a defined results set. AI visibility asks how often your brand got mentioned, cited, or recommended across a set of prompts and conditions. Different questions. Different answers. Treating the second like the first is the whole mistake.

The data says it does not repeat.

Ask a chatbot once and you have learned almost nothing, because the machinery moves before it even writes a word. When ChatGPT Search does go to the web, OpenAI says it rewrites your prompt into one or more targeted queries, runs follow-up searches after seeing the first results, uses your rough location from your IP, and, if Memory is on, pulls in things it already knows about you when it rewrites. Two people with the same goal, or the same person on two different days, can kick off different queries and get different answers.

I want to be precise, because the sloppy version of this gets picked apart. These systems do rank things internally. OpenAI says outright that ranking in ChatGPT Search is based on a number of factors, and that there is no way to guarantee top placement. And inside a single answer, your brand can show up first or get the first citation. Fine. Call that first-mention rate. What does not exist is a durable, repeatable, publisher-facing position you can track over time. One answer's order is a data point, not a standing.

This is not just my hunch. People have measured it, and they had no reason to want this result.

Rand Fishkin and the SparkToro team ran a real experiment. 600 volunteers, three tools, about 3,000 runs. Ask an AI for brand recommendations 100 times and you get the same list fewer than 1 in 100 tries. The same list in the same order? Fewer than 1 in 1,000. Fishkin's word for any tool selling a "ranking position in AI" was baloney. I will give him credit for changing his own mind by the end, which I will get to.

A 2026 preprint out of St. Gallen, "Don't Measure Once," measured the same instability from the source side. Across repeated samples of the same prompt, the cited sources overlapped only 32 to 43 percent. From one day to the next, source overlap sat around 34 to 42 percent and brand overlap around 45 to 59 percent. The point is not that nothing sticks. It is that one answer is a weak stand-in for the real visibility distribution. Their advice: treat visibility as a distribution, not a single number.

Two more 2026 preprints, both from Ronald Sielinski, put real statistics on it. The first shows citation counts follow a power-law shape with big run-to-run swings, and once you put confidence intervals around them, a lot of the differences between domains disappear into the noise. A single run gives you a number that looks precise and is not. His follow-up is the part the dashboard vendors will not love. There is no universal "run every prompt 100 times" rule. How much data you need depends on the platform and the topic, you have to show your measurement actually settled down, and when two competitors sit close together, the data may not support declaring one meaningfully ahead of the other.

So "you are number three in ChatGPT" is one pull of a slot machine reported as a leaderboard.

It is measurable. It is just a sampling problem, twice.

Here is where the loud skeptics overshoot, and where I get off the bus. The lazy take is that these tools are useless. They are not.

Fishkin started out sure the whole thing was a boondoggle and came around halfway. A durable rank is nonsense, but aggregate visibility, how often you show up across a lot of prompts run a lot of times, holds up. In his data, City of Hope showed in 97 percent of ChatGPT's cancer-hospital answers even though it was the top name only about a third of the time. The frequency meant something. The position did not.

The structure underneath is measurable too. That St. Gallen study found a small group of domains taking most of the citations, a Gini coefficient around 0.715. That concentration is real even when the exact sources bounce around answer to answer. And a controlled study at SIGIR 2026, "What Gets Cited," ran 252,000 trials across six models. The biggest drivers of getting cited were boringly familiar: how well the page matched the topic, and where it landed in the retrieval list. Stating specifics like price helped, and so did a recent timestamp. Completeness and evidence gave smaller bumps. Formatting tricks barely moved the needle. The levers you actually control are mostly just good content, which is what we should have been building anyway.

But running the same prompt a lot only fixes half the problem, and it is the half most tools ignore. The other half is picking prompts real buyers actually use. SparkToro asked 142 people to write a prompt for the same intent and found an average semantic similarity of just 0.081 between them. That is basically nothing. So you have two sampling problems. Run each prompt enough times to see its spread. And choose prompts people actually type. A tool can run its made-up prompt 400 times and hand you a gorgeous confidence interval on a question nobody asks. That is measuring the wrong thing, very precisely.

The right way to think about it is not rank tracking. Rank tracking reads a result off a page. AI visibility estimates a distribution, which is why polling is the better comparison. You design a sample, run it over and over, and report a number with a margin of error. "Candidate A: 51 percent" is not a fact. It is an estimate, and the method behind it is the whole game.

Show me your math.

Say a tool shows you a big number. AI Visibility Score, 74. Before that means anything, it has to answer a pile of questions. Which prompts, and who decided they are what your buyers ask? How many runs, and how do you know that was enough? Which model and version? Search on or off? Logged in or out? Which country and language? Was memory in play? Did visible mean mentioned, cited, recommended, or linked? How much did it swing between runs? And the one nobody wants to answer: what is the denominator? A 74 with no answers is not a measurement. It is a vibe with a decimal point.

This is a governance problem, the same way AI anywhere in a build pipeline is. You do not point a black box at a dashboard, read a number, and call it truth. Somebody has to design the prompt set so it reflects real buyers, keep the conditions constant, save the raw prompts and raw answers so the whole thing is auditable, judge whether the measurement actually settled, and decide what matters. You cannot hand that judgment to a score. It is the same thing I ask for when AI writes code or drafts content: reproducibility, an audit trail, and a person who actually reads the output. A visibility number with no saved prompts, no run count, and no variance is the measurement version of shipping AI-generated code straight to prod without a review. The tech is useful. The discipline around it is what makes it trustworthy.

How we built RankReady, and how to check any tool.

We built RankReady on this, and I will hold it to the same bar I am asking you to hold everyone else to. It reports rates, not ranks. It runs prompts over and over instead of once. It keeps ChatGPT, Claude, Perplexity, Gemini, and Google's AI features separate instead of mushing them into one flattering number. It keeps the raw answers so you can audit the claim. And it works to connect visibility back to something you own, like referral sessions and conversions, instead of a score that floats off on its own.

Whatever tool you are looking at, including ours, get straight answers to a few things:

  • Does it report a rate against a stated sample, or just a rank?
  • How did it decide the sample was big enough, beyond "we run it 100 times"?
  • Are the conditions written down and held constant: model and version, search mode, country, language, logged-in state, date?
  • Does it tell a mention from a citation from a recommendation, and track sentiment and whether you own the source?
  • Does it say whether it tested the public product, an API, a logged-in account, or a controlled browser session?
  • Can you export the raw data and read the actual answers?

Here is what an honest result reads like. Across 30 buyer-relevant prompts, sampled until the numbers settled, logged out, US English, we showed up as a cited source in 42 percent of Perplexity answers, with the confidence interval reported right next to it. That is measurement. "We rank number two in AI" is theater.

I am not anti-measurement. I am anti-pretending. AI visibility is real, earning citations is genuine work, and the things that earn them are mostly the good content we should be building anyway. What I will not do is call a probabilistic observation a deterministic rank.

In SEO, rank is a position. In AI search, visibility is a probability. Confusing the two is where measurement quietly turns into marketing. So make them show their math, and make sure you can inspect the evidence yourself. A number you cannot audit has not earned your trust.

Sources

Share on social media

Akshay Sura

Akshay Sura

Akshay is a ten-time Sitecore MVP and a two-time Kontent.ai. In addition to his work as a solution architect, Akshay is also one of the founders of SUGCON North America 2015, SUGCON India 2018 & 2019, Unofficial Sitecore Training, and Sitecore Slack.

Akshay founded and continues to run the Sitecore Hackathon. As one of the founding partners of Konabos Consulting, Akshay will continue to work with clients, leading projects and mentoring their existing teams.


Subscribe to newsletter