Page Structure Is an AI Citation Signal. Here's the Research.
Most AEO advice focuses on what you say. Get the right keywords. Cover the right topics. Answer the questions your customers ask. That's all necessary. It's not sufficient.
A March 2026 study published on arXiv (2603.29979) tested whether document structure -- independent of content -- affected AI citation rates. The researchers made structural edits only: changed heading hierarchy, adjusted paragraph length, moved data into tables, added visual emphasis. They did not change the underlying facts or claims.
Citation rates improved by 17.3% on average across six AI search engines. Subjective answer quality improved 18.5% in human evaluation.
That number is worth sitting with. Structure alone -- not new content, not new facts, not new claims -- moved citation rates more than many content-level interventions do.
What "structure" means in this context
The study (Yu, Yang, Ding, Sato -- Scout's content-format investigation, filed 2026-06-18) breaks document structure into three levels:
**Macro-structure** is document architecture: does the page have a clear heading hierarchy? Do sections flow logically from question to answer? A page with a coherent heading spine is easier for an LLM to navigate than one with walls of text broken up by decorative headings.
**Meso-structure** is information chunking: how are claims organized within sections? The study found the strongest citation improvements at this level. Shorter paragraphs with one claim per block, data in tables rather than inline, comparison grids for multi-option analysis. Pages with three or more data tables earned 25.7% more citations than structurally equivalent pages without tables, on the same content.
**Micro-structure** is visual emphasis: bold text, lists, table cells. These act as parsing shortcuts -- they tell retrieval systems where the high-density information is.
The practical implication: a competitor with weaker content but better structure may get cited more reliably than you do. Structure is competing on the same axis as content quality.
The FAQ problem is a calibration problem
Most FAQ sections on business websites fail on evidence density. The questions look fine. The answers are the problem.
"Do you offer emergency plumbing?" "Yes, we offer 24/7 emergency plumbing services across the Greater Toronto Area."
That answer is not wrong. It will not get cited by a RAG system. It contains no specific data point, no measurement, no timeframe. A retrieval system evaluating whether to cite you has no reason to prefer your answer over any other business with the same boilerplate.
Our content-format investigation (Scout, 2026-06-18, session 51) found the working specification for FAQ sections that produce citations:
- **4-8 questions per page.** Fewer is thin; more dilutes keyword density per question. - **40-80 words per answer.** Below 40 is too thin. Above 80 and RAG retrieval systems may split the answer or use only the first chunk. - **Direct answer in the first sentence.** Preamble before the answer is a cost -- retrieval systems often weight the first sentence heavily. - **One specific data point per answer.** A timeframe, a measurement, a cost range, a service area. "Most water heaters need replacement after 8-12 years" is citable. "Water heaters eventually wear out" is not.
Google removed FAQ rich results from search snippets in May 2026, which confused some practitioners into treating FAQ schema as obsolete. The format still matters for AI extraction -- the LLM reads the content, not the rich results. What drove citation was always the evidence density inside the answer, not the markup around it.
Why Perplexity's architecture makes structure a prerequisite
In our directory research (Scout, 2026-04-22), we found that Perplexity fetches approximately 10 pages per query and cites 3-4. The selection from those 10 is not random.
Perplexity is a real-time RAG system. When it fetches a page, it needs to extract a self-contained passage that directly answers the user's query. A page that buries its answer in a 600-word flowing paragraph is technically fetchable -- but the relevant 80 words are embedded in noise. A page with clearly bounded 100-300 word sections gives the system a clean extraction target.
The xSeek analysis of one million queries found that 62% of Google AI Overview citations come from passages in the 100-300 word range. Sources longer or shorter are cited proportionally less. That range is not a coincidence -- it aligns with the retrieval chunk size that RAG systems are typically configured around. Content that fits inside a single retrieval chunk gets cited. Content that doesn't, even if it's relevant, gets passed over for the page that does.
Being in Perplexity's top 10 fetched pages is roughly correlated with being in Google's top 10 organic results (60% overlap, per the same directory research). Getting from "fetched" to "cited" is a second gate -- and structure is what determines which of the 10 fetched pages clears it.
What to check on your existing pages
The test is simple. Pick your most important service page and ask: can you identify a discrete 100-300 word passage that directly answers the most common question someone would ask an AI about your business?
The passage should state what you do, where, and something specific (a timeframe, a price range, a credential, a service detail). It should stand alone -- a retrieval system pulling that passage out of context should still get a useful answer.
If you can't identify that passage, a retrieval system will have the same trouble. The content may be accurate, complete, and well-written -- and still fail at the extraction stage.
Structure problems are fixable faster than authority problems. You don't need new facts or new pages. You need the facts you already have in a form that an AI system can find and use.
If you want to see where your pages stand in actual AI platform queries -- not what you think they say but what AI systems are actually returning when someone asks about your business -- Sourcepull's Signal Check runs live queries across ChatGPT, Perplexity, and Gemini and scores what comes back.
See how your business scores on AI platforms.
Check your score — free