All posts
Analysis · 7 min read · 2026-09-07

What Your AI Visibility Score Is Actually Measuring

Got a score from an AI visibility audit and noticed the numbers feel inconsistent -- or fixed your directories and expected the score to move but it didn't? Two research papers published in 2026 explain the gap. Neither questions whether your audit found real problems. They clarify what an AI citation score is actually measuring, and what it isn't.

The score is a sample from a probability distribution

In April 2026, a team of researchers published a paper with an unusually direct title: "Don't Measure Once: Methodological Implications of LLM Non-Determinism for GEO Research" (Schulte et al., arXiv:2604.07585). The core finding: AI language models are non-deterministic. Two identical queries to the same platform, submitted seconds apart, can produce different citation outputs.

Our July 8, 2026 methodology investigation of that paper documented the practical implication for AI audits. A business that appears in 1 out of 3 runs of the same query isn't "cited" or "not cited" -- it has a citation probability of roughly 33%. The number your audit reports is a sample from that distribution, not a measurement of a fixed state.

The businesses most affected are those near the citation threshold -- where the AI platform is on the margin about whether to include them. A business with a complete Foursquare record, an active Yelp profile with 80+ reviews, and consistent directory NAP data will appear in nearly every run. A business that just claimed its first Foursquare listing may flip between cited and not-cited run to run. Both audit results are accurate. They're just different samples from different probability distributions.

The practical implication: a single-run score is useful for identifying structural gaps (where are you absent from platform data?) and less useful for comparing your position across time. A score of 6.2 today vs 6.8 three months ago might reflect real improvement. It might also reflect run-to-run sampling variance. You need trend direction across multiple audits to see through the noise.

The citation pool rotates faster than an audit cycle

SISTRIX published a longitudinal citation drift study in April 2026: 82,619 prompts, 1,548,213 snapshots, 17 weeks of tracking across three platforms and six countries. In our July 30, 2026 methodology review of that data, two numbers stood out.

Google AI Mode rotates 56% of its cited sources every week. ChatGPT Search rotates 74% of its sources weekly. The citation pool is in constant motion. Not all of it -- brand domains held consistent across all 17 weeks in 43% of cases, making them the stable anchor category. But the surrounding co-citation pool turns over almost completely.

That clarifies what "not cited on audit day" actually means. If your audit ran and you didn't appear in Google AI Mode, there's a 56% chance you wouldn't have appeared that week even if your infrastructure was correct -- because the pool that week was already occupied by different sources. The audit isn't inaccurate; it's a slice of a rotating system.

The path to reliable presence is to become part of the stable anchor category, not to show up every time an auditor checks. Brand domain anchors are earned through sustained infrastructure signals: directory presence across Foursquare, Yelp, and BBB; consistent NAP data; schema establishing your entity identity; GBP completeness feeding Google's own AI surfaces. These signals move you from the volatile 57% toward the stable 43%.

The SISTRIX data also documents what model launches do to citation pools. When GPT-5.5 launched, 47% of ChatGPT's cited domains shifted within 48 hours. When Gemini 3 launched in January 2026, 42% of AI Overviews' cited domains changed overnight. An audit taken immediately after a major model release captures a temporarily disrupted state. If your audit timing overlapped with a model launch, that's worth accounting for when interpreting the results.

Are you measuring the queries where AI Overviews actually appear?

A third measurement gap is harder to spot without knowing what to look for.

Whitespark's 2026 Guide to Google AI Mode for Local Businesses, which we reviewed in our July 15, 2026 methodology session, quantified AI Overview appearance rates by query type across a 540-query study:

- Local-intent queries ("plumber near me," "find a contractor in [city]"): AI Overviews appear in 15% of results - Informational queries ("how to unclog a drain," "what to look for in a licensed HVAC contractor"): 92% - Hybrid queries ("best plumber for emergency water heater repair"): 97%

If a plumbing company audits its AI visibility using local-intent queries and finds low AI Overview presence, that's largely expected behavior. AI Overviews are uncommon in local-intent queries regardless of how well-optimized the business is. Google Maps handles local-intent. The AI visibility problem -- if there is one -- shows up on informational and hybrid queries, where AI Overviews are nearly universal.

The fix plan differs depending on which gap you have. Local-intent gaps are primarily GBP and review optimization problems. Informational and hybrid AI visibility gaps are primarily on-page content, directory presence, and structured data problems. Whitespark's weighting data shows on-page content carries 24% of AI search visibility weight vs just 12% for GBP signals -- the inverse of local pack rankings, where GBP dominates at 32%. An audit that doesn't specify query type can leave you optimizing for the wrong layer.

What to use an AI audit for

None of this makes AI visibility audits unreliable as a category. The structural gaps they find are real. An unclaimed Foursquare record, a 3.6-star Yelp profile, and missing sameAs links in your schema are real problems that an audit finds correctly regardless of the sampling noise around the final score.

What the score doesn't reliably communicate is your precise position in the citation hierarchy, or whether a specific number change from one audit to the next reflects genuine progress. For that, you need trend direction tracked across multiple measurements, 60-90 days apart -- long enough to average through weekly rotation noise and run-to-run variance.

The re-audit result worth looking for isn't a higher number. It's consistent presence across repeat checks: appearing in 4 of 5 runs instead of 1 of 5, or holding across multiple audit dates rather than flickering. That's the stable anchor signal the SISTRIX data shows is achievable. It takes sustained infrastructure, not a one-time optimization pass.

Sourcepull's Signal Check shows your current citation status across ChatGPT, Perplexity, and Google. Run it as a recurring check over time rather than treating the first result as a fixed truth. The score tells you where to start. The trend tells you whether it's working.

See how your business scores on AI platforms.

Check your score — free