ACADEMY
All posts

Schema Markup Won't Get You Cited by AI: What Ahrefs' 1,885-Page Study Found

Ahrefs tracked 1,885 pages that added JSON-LD schema against 4,000 control pages and found AI citations barely moved, and even declined on Google AI Overviews. Here's what the data says actually earns a citation.

If you've read any AI-search advice published in the last year, you've been told to add FAQ schema, HowTo schema, and Organization markup because it "helps AI understand your content" and get you cited by ChatGPT or Google's AI Overviews. It's repeated so often it reads like settled fact. Then, in May 2026, Ahrefs actually tested it against real pages over an eight-month window, and the advice didn't survive contact with the data.

What Ahrefs Actually Tested

Ahrefs started from six million URLs to find pages that added JSON-LD schema between August 2025 and March 2026, then narrowed to 1,885 pages that added schema during that window and matched them against 4,000 control pages that didn't. Every page in both groups was already receiving 100+ Google AI Overview citations as of February 2025, so this wasn't a study of obscure pages hoping to get noticed for the first time, it was a study of pages already in the game. Researchers Louise Linehan and Xibeijia Guan ran four separate statistical checks on the 30-day windows before and after schema was added: a two-sample t-test, a difference-in-differences analysis to strip out platform-wide trends, an event-study plot of the weekly progression, and a symmetrical window analysis to exclude recrawl noise. That's a meaningfully more rigorous setup than the single before/after screenshot that usually passes for "proof" in SEO content.

The Results Didn't Move the Way the Advice Promised

Across the three platforms measured, adding schema produced no citation lift on two of them and a small, statistically significant decline on the third:

  • **Google AI Overviews: -4.6%**, a small but statistically significant decline, working out to roughly 12 fewer daily citations per page in samples where most pages were already receiving hundreds
  • **Google AI Mode: +2.4%**, indistinguishable from zero
  • **ChatGPT: +2.2%**, indistinguishable from zero
Change in AI citation rate after adding JSON-LD schema, by platform. Source: Ahrefs, "We Tracked 1,885 Pages Adding Schema," May 2026
Change in AI citation rate after adding JSON-LD schema, by platform. Source: Ahrefs, "We Tracked 1,885 Pages Adding Schema," May 2026

The confusing part, and the reason this myth took hold in the first place, is that the correlation clearly exists: Ahrefs' initial analysis found AI-cited pages were almost three times more likely to carry JSON-LD than non-cited pages. That's a real pattern. It's just not a causal one. Pages that already do well tend to be run by teams that also implement schema, use clean HTML, and publish authoritative content, schema is a symptom of a well-built page, not the cause of its citations. When you isolate schema as the only variable, as this experiment did, the lift disappears.

Why the Mechanism Doesn't Support the Myth

The result stops being surprising once you look at how these systems actually retrieve a page at answer time. A separate real-time test cited in Ahrefs' own writeup, run by SEO researcher searchVIU, watched five AI systems, ChatGPT, Claude, Perplexity, Gemini, and Google AI Mode, fetch pages live to answer a query. None of them parsed JSON-LD, hidden Microdata, or hidden RDFa during that direct retrieval step. Every system extracted only the visible HTML content a human reader would see. Schema is metadata sitting alongside the content, not embedded in it, and if the retrieval step never opens that metadata, no amount of it can change what the model reads. This is the same "look at what's actually happening under the hood, not what the marketing copy claims" gap we cover on the retrieval side of RAG pipelines in RAG vs fine-tuning: how to actually decide, where the failure mode is identical: a component everyone assumes is doing work turns out not to be touched at the moment that matters.

What's Actually Correlated With Getting Cited

If schema isn't the lever, something else is, and the visible-content angle points at what. A separate analysis by Evertune Research, covered by Search Engine Land, reviewed close to 400 million citations across roughly 25,000 of the most-cited URLs on six models (ChatGPT, Copilot, Gemini, Google AI Mode, Google AI Overview, and Perplexity) collected in March and April 2026. Listicles accounted for 63% of all citations across every model, ranging from 40-65% depending on the model, and within that group, ranked numbered lists made up 71-86% of the listicles cited, far ahead of unranked lists or institutional rankings. That's not a coincidence, a numbered list with a clear heading per item is already structured the way an answer engine wants to quote it: a discrete, extractable claim with a label attached, no parsing of buried metadata required.

Ahrefs' own AI Overviews research points at the same visible-content mechanism from a different angle: pages structured with question-form H2 headings, a declarative answer in the first two sentences of each section, and the exact named entity repeated (the actual company or product name, not "it" or "the platform") are easier for a model to lift a clean answer from than a page that reads well to a human but makes the model do inference work to find the claim.

What Still Makes Schema Worth Keeping

None of this makes schema worthless, it just narrows what it's for. Structured data still earns rich results in classic Google search, feeds voice assistant responses, and helps knowledge graphs and entity-recognition systems resolve who or what a page is about. Ahrefs is explicit that their finding is scoped to already well-cited pages; they note schema "might still play a role in helping pages get crawled, parsed, or indexed" for pages that aren't showing up in AI answers at all yet, which is a different problem than the one this study measured. Keep it for those reasons. Just stop budgeting it as your AI-citation strategy, because the pages in this study already had both schema and citations, and adding more of the former didn't add any of the latter.

What to Actually Do This Quarter

If your GEO checklist currently leads with "add FAQ schema to every page," reorder it. Rewrite section openings so the first sentence states the answer, not the setup, before any supporting detail. Use the exact entity name every time instead of a pronoun, since that's what a model quotes back. Convert dense prose sections into a ranked, numbered list wherever the content is genuinely a list, not because it's a trendy format but because that's the shape citation engines are already pulling from at a 63% rate. And measure your own before/after the way Ahrefs did, with a control group and more than one week of data, before crediting any single tactic. This is the same content-structure discipline we build out fully in GEO vs SEO: what actually changes, and it's the kind of test-before-you-templatize thinking we teach across The AI Marketing Engine, because a marketing team that ships changes based on a real experiment beats one repeating advice nobody has actually checked.

Go deeper

The AI Marketing Engine