Schema Markup

Schema markup is structured data added to a web page, usually as a JSON-LD block using the schema.org vocabulary, that states in machine-readable form what the page is about: that this is a product with a price, a recipe with a cooking time, an article with an author and a modification date. Google defines structured data as "a standardized format for providing information about a page and classifying the page content." It remains worth having for classic search, but the 2026 evidence does not support the popular claim that it raises a page's chances of being cited by AI answer engines.

Schema markup's established job is eligibility for rich results: the star ratings, prices, recipe cards, event listings and video key moments that make a listing larger and more clickable. Google recommends JSON-LD as the format that is easiest to implement and maintain, and its documentation cites publisher case studies such as a 25% higher click-through rate for Rotten Tomatoes pages with structured data and an 82% higher click-through rate for Nestlé pages shown as rich results. Those figures are examples selected by Google, not a controlled estimate, and Google is explicit that correct markup makes a page eligible for enhanced display without guaranteeing it. Beyond rich results, structured data feeds entity understanding and knowledge graphs, which is where its indirect value for AI systems is usually argued to lie.

The Controlled Evidence: No AI Citation Lift

The strongest test to date is an Ahrefs study published May 11, 2026. It tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched control pages, using a difference-in-differences design. The measured effects were −4.6% for citations in AI Overviews (small but statistically significant), +2.4% in Google AI Mode and +2.2% in ChatGPT, the latter two indistinguishable from zero. Its conclusion: "adding schema produced no major uplift in citations on any platform."

The limitations matter. The sample was restricted to pages already receiving 100 or more AI Overview citations, so it says little about pages trying to enter the consideration set for the first time. The post-treatment window was 30 days, only JSON-LD was tested, and schema changes could not be fully separated from other edits made at the same time. Google's own guidance is consistent with the result: its AI features documentation says "there's also no special schema.org structured data that you need to add" to appear in AI Overviews or AI Mode.

The Conflicting Lab Evidence

Laboratory work points in a more favourable direction, at a different stage of the pipeline. SAGEO Arena (KDD 2026) built a full retrieve, rerank and generate pipeline over roughly 171,000 web documents that keep their structural fields, which the authors define as title, meta description, headings and schema/JSON-LD markup. Optimising only those structural fields raised retrieval hit rate by 22%, whereas rewriting body text alone, the standard generative engine optimization approach, degraded visibility at every stage. At the generation stage the benefit largely disappeared, because the great majority of citations were drawn from body text.

Two cautions apply before treating this as a rebuttal. "Structural information" in that study bundles schema with titles and headings, so the gain cannot be attributed to JSON-LD specifically. And the pipeline is a research system with a keyword-based retriever as its primary configuration, not a production engine. The two findings are compatible: structure may help a document get retrieved in systems that index it, while adding JSON-LD to a page that is already being retrieved does not make an engine cite it more.

Treat It as Hygiene

The reasonable position is that schema markup is maintenance, not a growth lever. It should be present, valid and consistent with the visible page, because rich results, entity recognition and accurate dates all depend on it and because the cost is low. It should not be sold or bought as a route to AI citations, and a 2026 survey of 45 GEO studies (Martinez) found no reviewed technique with a stable, cross-platform causal effect on discoverability. Since the markup is derived from fields a content system already holds, it is best generated rather than hand-written; LightCMS, the CMS serving this site, emits schema.org JSON-LD automatically for each page.

Further Reading