Why Schema Markup Improved AI Citations 1,500% in One Study and Zero in Another
In 2025, OtterlyAI published a case study showing that adding structured schema markup to a set of uncited pages led to a 1,500% increase in AI Overview appearances. That number circulated widely in AEO circles and became the go-to evidence for schema as a citation lever.
Ahrefs then ran a large-scale study across hundreds of thousands of pages and found no meaningful correlation between schema implementation and AI citation rates.
Both numbers are real. Both came from genuine measurement. And they appear to directly contradict each other.
In our investigation of schema markup effects on AI citations (knowledge/schema-markup-effects.md, Scout sessions 80 through 99, updated August 2026), we tracked this conflict for months. In August 2026, the analysis resolved it. The short version: both studies measured real effects, but on structurally different populations. That difference explains everything.
OtterlyAI Measured Businesses Starting From Zero
The OtterlyAI study tested schema against pages that were receiving no AI citations before the intervention. Adding LocalBusiness markup -- complete address block, sameAs links to claimed directory profiles, attribute-rich FAQPage -- produced a dramatic improvement starting from zero.
This makes structural sense. When an AI system has no prior entity signal for a business -- no training-data footprint, no directory presence, no GBP-anchored record -- schema markup functions as an introduction to the system. It tells AI platforms: here is a classifiable entity, here are its core attributes, here is how it connects to verifiable external records. A system that couldn't reliably classify the business now can.
Going from unclassifiable to classified is a large change. A 1,500% increase starting from zero is a large number with a small denominator. Both can be true at the same time.
Ahrefs Measured Businesses Already in the Citation Pool
The Ahrefs study sampled from a large web population and measured schema presence against AI citation rates. A large-scale web sample of pages that receive any AI citations is dominated by businesses with some established entity record -- some training-data footprint, some directory presence, some history of being cited.
For that population, adding schema shows no measurable citation-rate lift. The entity is already classified. Schema does not increase how often an already-known business gets cited; it helps an unknown entity enter the citation pool in the first place.
Ahrefs' own full-population data actually contains the signal: their 6-million-URL full sample showed a 3x schema correlation before they narrowed to already-cited pages. Once they filtered to pages already appearing in AI citations, the correlation collapsed. That is not a null result across the board -- it is a threshold effect that disappears inside a filtered sample.
Our August 6, 2026 methodology rec (2026-08-06-schema-threshold-fix-plan-stratification.md, Scout session 99) describes the working hypothesis directly: schema helps uncited pages enter the citation pool, but does not increase citation frequency for pages that are already being cited. No controlled split-population study has confirmed this mechanism, but the convergent evidence -- OtterlyAI's case study, Ahrefs' full-population data, and independent practitioner analyses -- supports it as a working framework for structuring fix plans.
What This Means for a Fix Plan
The reconciliation has a direct practical implication: schema recommendations need to be calibrated to where a business sits in the citation baseline.
**If your business has few or zero AI citations across all platforms:** Schema is foundational infrastructure. LocalBusiness markup with correct address fields, sameAs links to your claimed Yelp, Foursquare, and BBB listings, and a well-formed @id tells AI systems how to classify and connect you. The goal is entity establishment -- not boosting citation frequency, but becoming classifiable in the first place. Schema belongs early in the fix sequence.
**If your business already receives AI citations but has gaps on specific platforms:** Schema is almost certainly not the explanation for the gap, and fixing it is unlikely to close the gap. Platform-specific levers dominate at this stage: crawlability and freshness for Perplexity, third-party directory data quality for ChatGPT, GBP completeness for Google AI Mode and AI Overviews. Schema work belongs in a secondary pass, after those platform-specific fixes are in place.
The cost of treating schema as a universal first fix is sending already-cited businesses on a structured-data project that produces no measurable change in their actual citation gaps -- while the platform-specific work waits.
One Exception to Schema-First for Uncited Home Service Businesses
For home services businesses with zero AI citations across ChatGPT, Claude, and Gemini, the most effective Phase 1 action is not schema -- it's claiming profiles on the platforms that now feed those AI systems directly through data agreements.
As of 2026, Yelp data feeds ChatGPT via a 330M-review licensing agreement (confirmed July 23, 2026). Thumbtack professionals are surfaced directly in ChatGPT via API (October 2025), in Claude via API (April 2026), and in Gemini via Connected App (August 2026). Foursquare's POI database feeds ChatGPT local results.
Our August 2026 methodology rec on these integrations (2026-08-02-yelp-chatgpt-thumbtack-claude-partnerships.md, Scout sessions 95 through 110) is direct on what this means for Phase 1 fix plans: a business absent from Yelp, Thumbtack, or Foursquare is not underweighted in a training corpus -- it is absent from a direct data feed. Schema markup on the brand website cannot compensate for that structural absence, because ChatGPT, Claude, and Gemini are not reading your website for their local response layer. They are reading the data they licensed.
The updated Phase 1 sequence for home services: claim and complete Foursquare, Yelp, and Thumbtack profiles first. Schema complements those directory signals once the data feed presence is established, not before.
The Pattern Behind the Contradiction
The OtterlyAI-versus-Ahrefs gap is one instance of a recurring measurement problem in AEO research: two studies sample from different populations, both produce accurate measurements of those populations, and both get cited as universal claims that seem to contradict each other.
The practical lesson is to treat schema as a stratified recommendation, not a universal one. Before adding structured data to a fix plan, identify where the business sits: is it uncited and needs entity establishment, or is it already cited somewhere and needs platform-specific gap work?
A Sourcepull Signal Check breaks out scores per platform and surfaces which situation you're in. The answer is often different from what business owners expect -- and it usually determines whether schema or directory work is the right first step.
See how your business scores on AI platforms.
Check your score — free