AI Citations Drive Branded Search: The First Causal Study Has Numbers
The argument for AI visibility has run on observation: businesses that show up in ChatGPT or Perplexity recommendations report more calls, more web traffic, more brand awareness. The problem is that observation is correlational. A business investing in AI visibility also tends to invest more broadly. The business case has been assumed, not measured.
Three studies published in June and August 2026 change that. One uses a randomized controlled design -- the first in the field. The findings make the AI citation argument simpler and harder to dismiss.
The causal study: +4.3 percentage points of branded search
Our August 15, 2026 Scout session (session 108) documented arXiv:2606.10907, "From Prompt to Purchase." The paper uses a randomized controlled design to test whether AI recommendation causes subsequent consumer search behavior.
The setup: users were shown AI responses that included brand recommendations for previously-unknown brands. A control group received AI responses without those brand recommendations. Researchers then measured Google search behavior for the mentioned brands in both groups.
The finding: being recommended by AI increased branded Google search volume by +4.3 percentage points. The causal design separates this from correlation. The AI mention drove the additional searches -- not a confounding factor like the brand's ad spend or category seasonality.
For a local service business, that number compounds. A business cited in 15-20 AI sessions per week accumulates a meaningful and persistent branded search signal over time. Each session generates incremental brand awareness that converts to direct searches. The mechanism is direct: someone asks AI who to call, gets a name, and searches for it to find the phone number or read reviews before calling.
Every prior study connecting AI citations to traffic was observational. This one isn't. That distinction matters for how confidently businesses can treat AI visibility as an investment with expected return.
The incumbent advantage and how to close it
The same Scout session documented arXiv:2606.17443, which tests what happens when unknown brands compete with established brands in AI recommendation environments under controlled conditions.
The result is stark: known brands were recommended 100% of the time. Unknown brands -- with identical prices, specifications, and review counts -- were recommended at near-zero rates.
This is not about brand recognition in the everyday sense. It is about data availability in AI training and retrieval systems. A national chain has years of directory coverage, third-party mentions, earned media, and review volume accumulated in the training corpus. A local HVAC company two years old may have almost none of that. The AI system defaults to the entity it knows -- not because it is explicitly programmed to favor the chain, but because its confidence in the chain is high and its confidence in the local business is near-zero.
The paper tests whether third-party content can overcome this advantage. It can. Injecting third-party signals -- reviews, citations, earned media mentions -- for previously-unknown brands measurably increased their recommendation rates to competitive levels.
The mechanism the paper frames as a "vulnerability" -- AI recommendations can be shifted by building external content signals -- is the same mechanism that legitimate entity-building exploits. Claiming and completing directory profiles, building review volume on Yelp and BBB, earning third-party mentions in local press and niche directories: these are the signals that close the data gap between a local business and a national brand in an AI training corpus. The paper confirms they work.
This is the grounding for why Phase 1 of any fix plan starts with entity infrastructure rather than website changes. Schema on a website that AI systems have no existing confidence in accomplishes less than establishing presence on the third-party platforms AI systems already treat as trusted sources.
The multi-platform problem: 41.6% agreement across four platforms
Our session 108 investigation also documented arXiv:2606.23057, which measured recommendation consistency across ChatGPT, Gemini, Perplexity, and Claude simultaneously for identical queries.
The finding: only 41.6% of top recommendations are consistent across all four platforms. More than half of what any individual platform recommends is not recommended by the other three.
For a business working on AI visibility, this reframes what "good" looks like. Strong ChatGPT citation rates don't transfer to Perplexity. Gemini citations don't transfer to Claude. Each platform draws from a distinct combination of data sources and applies different retrieval logic -- ChatGPT for home services pulls from Foursquare, Thumbtack, and Yelp via data partnerships; Perplexity retrieves live directory and web pages; Gemini weights GBP completeness and the Google index; Claude for home services surfaces Thumbtack profiles directly.
A business that invests in one platform's signals and calls it done is optimized for 41.6% of recommendation scenarios at best. The remaining 58.4% of citations, across the three other major platforms, come from different sources and require different infrastructure.
This is the empirical basis for why a multi-platform audit produces different recommendations for each platform rather than a single unified fix list. The platforms agree on less than half of who to recommend. The infrastructure gaps are genuinely distinct, not redundant.
What the three papers together say
The business case for AI citation investment now has a causal study behind it -- not just observation. Being recommended by AI causes measurably more people to search for your brand. The incumbent advantage that keeps unknown businesses invisible is real, but it is closeable through legitimate entity-building. And closing it on one platform does not close it on others.
The practical implication for prioritization: start with the platform where the gap is largest and the fix is most accessible, typically Perplexity for businesses with solid directory presence but thin content structure, or ChatGPT for businesses missing from Foursquare and Yelp. Run through the full multi-platform picture before deciding where infrastructure work will compound fastest.
A Signal Check at sourcepull.ca shows your current score across ChatGPT, Perplexity, Claude, and Gemini in about 60 seconds -- which is the fastest way to know whether your business is in the incumbent gap or the citation pool.
See how your business scores on AI platforms.
Check your score — free