All posts
Deep Dive · 7 min read · 2026-08-28

The AI Citation Threshold: Why Some Businesses Get Cited Everywhere

The posts about why you're invisible to AI are everywhere. Most say the same things: fix your schema, clean up your directories, claim your GBP, write FAQs. These are not wrong. But they describe individual fixes without answering the question that actually matters to a business owner: what is the bar I need to clear before AI models start citing me reliably?

Three 2026 research papers -- all documented in our research vault over the past two weeks -- put numbers to that bar for the first time.

There is a measurable threshold, and the number is 12 out of 16

In our August 24, 2026 review of arXiv:2509.10762, Scout documented a study by Arlen Kumar and Leanid Palkhouski (Wrodium and UC Berkeley SkyDeck, submitted September 2025). The paper ran 70 prompts, analyzed 1,702 citations across Brave Summary, Google AI Overviews, and Perplexity, and scored each cited page against a 16-pillar quality framework called GEO-16.

The finding: pages that scored 0.70 or above on the framework AND hit 12 or more of the 16 pillars achieved a 78% cross-engine citation rate.

That is the threshold. Hit 12 of 16 quality signals and AI models cite you on roughly three out of four queries across three major platforms.

The three pillars most correlated with citation, in order: 1. Metadata and freshness (title tags, publication dates, last-modified signals that AI retrieval systems actively read) 2. Semantic HTML (proper heading hierarchy, structured content flow that makes content parseable) 3. Structured data (schema markup, JSON-LD entities)

The important word in this finding is "simultaneously." You need to be consistently correct across many signals -- not exceptional at one. A perfect LocalBusiness schema on a page with no semantic heading structure and stale metadata is still below the threshold.

Two caveats worth naming. First, the study covers B2B SaaS, not home services. The 78% figure may shift somewhat in a local services context. Second, the study covers Brave Summary, Google AI Overviews, and Perplexity -- not ChatGPT or Claude. Applicability to those platforms requires separate evidence.

Perplexity plays by different rules

The GEO-16 paper also published the mean quality score of pages actually cited by each engine:

- Brave Summary: 0.727 - Google AI Overviews: 0.687 - Perplexity: 0.300

Perplexity cites significantly lower-quality pages -- by GEO-16 metrics -- than the other two engines.

This is not noise. Our research from earlier in the summer identified the explanation in a separate dataset: Perplexity weights format-matching over content quality. A listicle query produces listicle citations regardless of the page's schema integrity or heading structure. Reddit -- which scores poorly on semantic HTML and structured data -- accounts for roughly 46.7% of Perplexity's citation share (5W State of AI Search 2026) precisely because it matches the format of recommendation queries.

The practical read: optimizing toward the GEO-16 threshold improves your position on Google AI Overviews and Brave. For Perplexity, third-party presence (Yelp, directories, relevant forum mentions) and page format alignment matter more than pillar scores. These are separate tracks, not one optimization effort.

The structure fix that doesn't require rewriting anything

The second piece of research came from Scout's August 18, 2026 review of arXiv:2603.29979 by Junwei Yu and colleagues (submitted March 31, 2026). This paper tested what happens when you change the organization of a page without changing the words.

The result: a 17.3% average citation rate lift across 6 generative AI search engines from structural changes alone. Semantic content was held constant.

The framework they used organizes content into three levels:

- **Macro-structure:** the overall document architecture (guide format vs. FAQ format vs. list format) - **Meso-structure:** information chunking (paragraph length, subheading density, logical grouping) - **Micro-structure:** visual emphasis (bold phrases, numbered steps, callout formatting)

This is directly actionable for service businesses. Most local service pages are organized around what the business wants to communicate: credentials first, then services briefly listed, then testimonials, then a contact form. That's a conversion-oriented layout, not a citation-optimized one.

A page organized around explicit subheadings (one per service, in plain language), short answer-first paragraphs, and a clear FAQ block at the bottom is the same content -- just restructured for how AI models extract information. According to the GEO-SFE finding, that restructure alone is worth more than most tactical content additions.

Extraction failure: the problem most businesses don't know they have

The third paper, reviewed in our August 19, 2026 research session, is arXiv:2603.09296, "Diagnosing and Repairing Citation Failures in Generative Engine Optimization."

The paper's core argument is that existing advice makes a measurement error: most GEO guidance optimizes for contribution (how much a document influences a response) rather than citation (whether the document earns a traceable attribution). These are different outcomes, and optimizing for one doesn't guarantee the other.

The paper identifies extraction failure as one of the most common citation failure modes. What this means: the AI found your page, read it, and could not extract a clean, usable answer. So it used the content without attributing it, or skipped the page entirely.

The study tested two repair approaches: - Generic GEO rewriting (the standard approach): 25% improvement in citation rates - Targeted, diagnosis-first repair (identify the specific failure mode, then fix only that): more than 40% improvement -- while modifying only 5% of the content

The 5% figure is the striking part. Businesses that approach citation gaps by rewriting large portions of their content are doing more work for less result than businesses that find the specific extraction barrier and remove it.

In practice, extraction failure looks like a page where the AI cannot locate your service area, cannot identify which specific services you offer, or cannot connect your business name to a clear category. These are usually fixable with a handful of targeted sentences -- not a content overhaul.

What this means if you've already done the basics

If you've fixed your schema, claimed your directories, and built out your GBP -- and you're still not showing up consistently -- this research points to where to look next.

Check structure before adding content. Does each page have a clear macro-structure matched to the query type? Are subheadings explicit service names rather than marketing phrases? Can AI pull a usable answer from the first two paragraphs? If the structure is flat, reorganizing is faster than writing more content and the evidence for impact is now direct.

Run diagnosis before rewriting. A specific page that isn't cited is failing at a specific stage -- discovery, extraction, or competition for the limited citation slots. The extraction failure research is clear that generic rewrites are less efficient than targeted repairs. Find what's failing before adding more words.

Expect platform-specific results. The 78% cross-engine rate from GEO-16 is an average across Brave, Google AIO, and Perplexity. For Perplexity specifically, that threshold matters less than format alignment and directory presence. For Google AI Overviews and Brave, the 12-pillar multi-signal approach is the right target.

Our Signal Check audit reports which of these pillars are present and which are missing on your pages. For businesses that have already addressed the infrastructure basics, it's the structural and extraction gaps that typically explain what's left.

See how your business scores on AI platforms.

Check your score — free