AI Visibility Rankings Aren’t Stable – New Research Shows It’s Mostly Statistical Noise
New research finds that brands' positions in AI-generated answers shift too randomly to be gamed, making consistent content quality more important than visibility tactics.
On this page
Your AI-visibility dashboard is probably showing you noise, not a ranking
A paper published April 11 by IQRush co-founder Ron Sielinski, covered by Search Engine Journal on July 11, tested how stable AI citation rankings actually are across repeated queries to platforms like SearchGPT, Gemini and Perplexity. In an earlier test on running gear, Tom's Guide picked up roughly 9.5% of SearchGPT citations against Runner's World's 6.0% — a gap that looked decisive on a single dashboard reading but sat inside the margin of error once repeated sampling was accounted for. Some platform-topic combinations still hadn't produced a trustworthy ranking after 125 separate questions. Search Engine Journal also ties this to Rand Fishkin's SparkToro finding that AI tools return a different set of recommended brands more than 99% of the time the same question is asked.
Why do the same query and the same brand get different citation results each time?
Generative answer engines deliberately inject randomness into each response, so any single citation is one draw from a wide range of possible sources, not a fixed fact.
That's a different failure mode from what this site usually flags with detectors — false positives on human writing — but the underlying problem is the same shape: a tool reports a confident-looking number from a process that is inherently variable, and the dashboard doesn't tell you the variability exists. A single AI Overview or SearchGPT answer citing your competitor over you is not evidence of anything on its own.
How many responses does it actually take before a ranking is trustworthy?
There's no universal cutoff — it depends on the platform and topic, and some competitive pairs never separate cleanly even after 125 sampled answers.
Sielinski's method requires two things to both be true: the order of sites has to stop reshuffling as more answers come in, and the gap between the top sites has to be bigger than each one's margin of error. Neither condition alone is enough. On SearchGPT in particular, some topic pairs stayed too close together to ever produce a reliable ranking, no matter how many questions were asked.
What should buyers ask their AI-visibility vendor before paying for it?
Ask how many queries per platform-topic pair the ranking is built from, and whether the vendor reports a margin of error at all.
Fishkin's advice, as quoted by Search Engine Journal, is blunt: before spending money on AI visibility tracking, make the provider "show their math." That's a reasonable bar for any tool in this space — AI detectors included — where a single score is routinely presented as fact when the underlying signal is probabilistic. If your GEO or AI-visibility vendor can't say how many samples back a citation-share number, treat that number as directional at best.
Frequently asked questions
Does this mean AI visibility tracking is worthless?
No — it means single-snapshot readings are unreliable. Repeated sampling with a stated margin of error, as IQRush's method proposes, can still produce a meaningful ranking once enough answers are collected.
Is this the same instability problem as AI detectors giving inconsistent scores?
Not the same mechanism, but the same lesson: any tool built on a generative model's variable output needs to disclose sample size and error margins, not just a headline percentage.
Which platform was hardest to get a stable ranking on?
SearchGPT, where some competing sites remained too close together to separate even after 125 questions, according to the IQRush research reported by Search Engine Journal.