Schema Markup for GEO: Which Structured Data Actually Gets Cited

Schema Markup for GEO: Which Structured Data Actually Gets Cited

Here is the sentence most GEO vendors won't put in a pitch deck: adding schema markup to a page does not, on its own, get you cited more by AI. The largest controlled test to date says so plainly. Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026, matched them against 4,000 control pages, and measured citation changes across Google AI Overviews, AI Mode, and ChatGPT. The result was not a lift. On Google AI Overviews, citations for schema-added pages actually fell.

That finding sits awkwardly next to the marketing claims — "2.5x more AI citations," "schema is the language AI speaks" — that fill most schema-for-GEO guides. Both things can be partly true, and the gap between them is the whole point of this article. Schema is real infrastructure. It is not a citation lever you pull. Below is what the 2026 evidence actually supports: which structured data types influence AI citations, which never did, and how schema interacts with the thing that actually decides citations — passage extraction. For the foundational entity work schema does do, see my guide on connecting entities with schema markup for the Knowledge Graph. This piece is about the narrower, noisier question: does markup make AI prefer you as a source?

The headline finding: adding schema doesn't move AI citations

The cleanest data we have shows schema has no measurable citation lift — and Google now says the same. In the Ahrefs study, a matched difference-in-differences analysis produced Google AI Overviews down 4.6% and statistically significant, Google AI Mode up 2.4% and not statistically significant, and ChatGPT up 2.2% and not statistically significant. In plain terms: the only statistically significant movement went the wrong way, and the two positive numbers were indistinguishable from noise.

The academic work agrees on direction. A February 2026 cross-platform study by Kurt Fischman found that the initial pooled analysis produced a significant negative association between schema presence and AI citation, which proved to be a methodological artifact — Google's ranking algorithm systematically enriches top-10 organic results for schema-bearing pages, inflating schema prevalence. Once you control for that, the real driver appears: the dominant predictor of AI citation was Google organic rank position, with position-1 pages cited in 43% of queries in which they appeared, declining to 5% at position 7.

And Google itself has now closed the door on the "required schema" myth. Google published its first official generative AI search guide on May 15, 2026, and said flatly that structured data isn't required for AI Overviews or AI Mode, and there's no special schema.org markup you need to add for them.

So why does every second GEO guide claim otherwise? Because schema correlates with citations without causing them. Schema markup tends to live on better-maintained, more technically sophisticated sites, and those same sites publish stronger content, build more authority, and earn more links — schema could be doing real work, but it could also just be riding the wave of every other signal. The 2.5x figure that circulates everywhere traces back to correlational 2025 data, not a controlled test.

Why schema barely moves the needle: AI cites passages, not markup

Citations are decided at the passage level, and schema doesn't live in the passage. This is the mechanical reason the studies land where they do. AI retrieval reads rendered content and asks a specific question: can the engine extract a clean, standalone answer from the visible page? Schema answers a different question — what is this page about — that AI engines already resolve through their own classification systems.

There's a second, more surprising mechanism. AI systems don't parse your JSON-LD the way you think. In a February 2026 test, Mark Williams-Cook created a fake company, put its address only inside made-up JSON-LD schema and nowhere in the visible text, then prompted both ChatGPT and Perplexity — and both read the fake schema to find the address. Because the schema was invalid yet still worked, his conclusion was blunt: the LLM is simply picking up whatever you list in the HTML — it does not matter if it is valid schema; if the system interprets the text as relevant, it's included, which indicates schema is not being used in the explicit sense it was designed for.

That cuts both ways. It means a well-formed sameAs or a clean Organization block can reinforce a fact the model reads elsewhere on the page — but it also means every property in a JSON-LD block is readable as plain text, so accurate fields reinforce visible HTML while inaccurate or orphaned fields introduce contradictions the model may surface as hallucinations. Schema-as-text is a benefit only when it agrees with everything else. This is exactly why your facts must match everywhere.

And where content lives only in markup, extraction collapses. An October 2025 SearchVIU test across five engines found that pages with data only in JSON-LD, Microdata, or RDFa saw near-zero extraction, while visible well-structured HTML achieved consistent success — schema-only content failed across the board. If a fact matters for citation, it must be visible on the page. Markup is a reinforcement, never a hiding place. I cover the extraction side in depth in how to make your content citable by AI, and the retrieval mechanics in how Google AI Overviews choose citations via query fan-out.

The schema types by citation influence

No schema type has proven direct citation lift, but they differ sharply in strategic value. The honest ranking is by what each type does — entity clarity, fact reinforcement, or nothing — not by a citation multiplier no one has demonstrated.

Schema type Real job in GEO Citation influence (2026 evidence)
Organization Establishes your entity: name, sameAs, identifier, logo, founding date Indirect but the most valuable — feeds entity resolution, not passages
Person / Author Attributes expertise to a named entity Indirect — supports author authority AI systems already weigh
Article / BlogPosting Marks author, publish/modified dates, publisher Indirect — dates reinforce recency, a real citation signal
Product / Offer / Review Specs, price, aggregate ratings for comparison answers Modest — helps shopping/comparison summaries stay accurate
FAQPage Labels Q&A structure None proven; rich result gone; keep only if it mirrors visible Q&A
HowTo Labels step sequences None proven; rich result deprecated in 2023
LocalBusiness NAP consistency for local entities Indirect — must match GBP and citations exactly

Organization and Person: the only schema worth prioritizing

Entity schema is the one category where markup earns its keep, because it does work AI systems can't fully reconstruct from prose. Organization with accurate sameAs links to your Wikidata, LinkedIn, and Crunchbase profiles, plus identifier for your KGMID, tells engines which entity a page belongs to. This is disambiguation infrastructure, and it matters most for founders and SaaS brands still building recognition. Wellows' synthesis of 2026 research is consistent with this: across major AI platforms, brand search demand, citation overlap, and entity signals now outweigh backlink volume or standalone domain authority. Entity schema supports the signals that actually correlate — it does not replace them. Start with building your entity and, if you don't have one yet, a Knowledge Panel.

Article and BlogPosting: recency is the only real lever here

Article schema's quiet value is the date fields, because recency is a demonstrated citation signal. Ahrefs' analysis of 17 million citations found a strong bias toward recently updated content, and separate 2026 research found 65% of AI bots access pages updated within the past year. Article and BlogPosting let you state datePublished and dateModified in a machine-readable form that mirrors visible dates. Use full ISO 8601 dates, and make sure they match what's on the page — a modified date in schema that contradicts a stale visible byline is exactly the contradiction that hurts you.

FAQPage and HowTo: deprecated theater

FAQ schema is the clearest example of a tactic outliving its usefulness. Google deprecated FAQ rich results on May 7, 2026, ending the expandable Q&A dropdowns, and the rich results disappeared from Search. The reporting infrastructure follows: Search Console reporting and Rich Results Test support end in June 2026, and Search Console API support ends in August 2026. For years FAQ structured data was the single most recommended AEO tactic.

Here is the sober read. The markup type is not dead — FAQPage as a Schema.org type is still valid, and the markup can stay on your pages without causing problems. But the claim that it improves AI citations is unproven: the claim that FAQ schema improves your odds of getting cited in ChatGPT, Perplexity, or AI Overviews isn't confirmed by Google or any AI vendor — treat it as unproven, not a fact. The value was never in the tag. As The HOTH put it after the deprecation, the change made one thing obvious: the schema was never doing the work. Well-built, visible Q&A content still earns citations. The <script> block around it does not.

What actually drives AI citations instead

If schema isn't the lever, spend the budget on what the data says is. Three factors dominate the 2026 evidence, and none of them is markup.

Organic rank and retrievability. Fischman's finding that position-1 pages are cited in 43% of appearing queries versus 5% at position 7 means the classic SEO fundamentals still gate AI citation. And for ChatGPT specifically, you must be in Bing's index at all — see why ChatGPT can't see your website.

Third-party trust and consensus. The single most striking 2026 number comes from Trustpilot's analysis of more than 800,000 AI responses: brands with no active review profile were cited in only 1% of answers, while brands that actively collected and responded to feedback were cited in 75.3% of answers — a 75x gap. In the same sample, review and trust sites now account for 14% of all AI citations, second only to general brand websites. This is why AI keeps citing Reddit and review sites instead of your own content, and why Reddit belongs in every serious GEO plan.

Cross-engine agreement and clean passages. Research from the GEO-16 framework applied to B2B SaaS found that cross-engine citations — pages cited by ChatGPT, Perplexity, and Google simultaneously — exhibit 71% higher quality scores than single-engine citations. Passage structure is the practical route there. Kevin Indig's analysis of 1.2 million ChatGPT answers found 44.2% of citations come from the first 30% of content, and heavily cited text averaged 20.6% entity density, three to four times normal English. Lead with the answer. Pack it with named entities. That does more than any tag.

The verdict: schema is hygiene, not strategy

Treat schema as a hygiene factor, not a growth channel. Add Organization and Person because entity clarity compounds. Keep Article dates accurate because recency is a proven signal. Make sure every schema property mirrors visible content, because AI reads it as text and punishes contradictions. Then stop there, and put the rest of your effort into rank, third-party consensus, and answer-first passages.

The cleanest framing I've seen came from a German practitioner after the FAQ deprecation: schema is useful for LLMs — not as a magic citation lever, but as a clean Q&A format that any embedding pipeline, any Graph-RAG index, and any future crawler processes without friction. That's the right expectation. Or in Williams-Cook's words: Schema, good. Repackaging the basics as some magical new GEO formula, bad. If a vendor is selling you the second thing, you now have the studies to push back. Want the full sequence? Start with the GEO guide, then run a free AI visibility report to see where you actually stand.

Frequently asked questions

Does schema markup increase AI citations?

The largest controlled test — Ahrefs tracking 1,885 pages that added JSON-LD against 4,000 control pages — found no meaningful lift, with Google AI Overviews citations actually falling 4.6% (statistically significant) and AI Mode and ChatGPT changes indistinguishable from noise. Google's own May 2026 AI search guide states structured data isn't required for AI Overviews or AI Mode. Schema correlates with citations because it lives on stronger sites, but it doesn't cause them.

Which schema types matter most for GEO?

Organization and Person schema are the only types worth prioritizing, because they do entity-resolution work AI can't fully reconstruct from prose — accurate sameAs, identifier, and author attribution. Article/BlogPosting is useful mainly for machine-readable dates, since recency is a proven citation signal. Product/Offer/Review helps keep comparison answers accurate. FAQPage and HowTo have no proven citation influence, and their Google rich results are deprecated.

Do ChatGPT and Perplexity actually parse JSON-LD?

Not as structured data. A February 2026 test by Mark Williams-Cook placed a fake address only inside invalid JSON-LD, and both ChatGPT and Perplexity extracted it — meaning they tokenize the script block as plain text. This means schema fields reinforce facts when they match visible content, but introduce contradictions and possible hallucinations when they don't. Content that lives only in markup, never in visible HTML, sees near-zero extraction.

Is FAQ schema dead in 2026?

The FAQ rich result is dead — Google removed it from Search on May 7, 2026, with reporting and API support ending through June and August 2026. But FAQPage is still a valid Schema.org type you can safely keep. The key point: no vendor or engine has confirmed FAQ schema improves AI citations, so treat that claim as unproven. Well-built, visible Q&A content still earns citations; the tag around it never did the work.

If not schema, what actually drives AI citations?

Three factors dominate the 2026 data. First, organic rank and retrievability — position-1 pages are cited in 43% of queries where they appear versus 5% at position 7. Second, third-party trust: Trustpilot found brands actively managing reviews were cited in 75.3% of answers versus 1% for those with no review profile. Third, answer-first passages with high entity density, since 44.2% of ChatGPT citations come from the first 30% of content.

References

  1. Ahrefs — We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.
  2. Search Engine Roundtable — ChatGPT & Perplexity Treat Structured Data As Text On A Page
  3. Quattr — FAQ Schema in 2026: What's Confirmed, What's Not & What to Do
  4. Passionfruit — FAQ Rich Results Deprecated: Google's May 2026 Change
  5. Search Engine Land — How schema markup fits into AI search — without the hype
  6. AuthorityTech — Does Schema Markup Help AI Citations? Ahrefs Tested 1,885
  7. Passionfruit — How LLMs Search for Citations: What They Find [2026 Data]
Cory Maki
About the author

Cory Maki is an AI search strategist based in Taichung, Taiwan, specializing in GEO, AI reputation management, and AI branding for SaaS founders. Author of Reddit, AI Overviews & GEO and creator of the ARC Method. Read more →