Visora
Sign up
← Back to blog

· Visora

SchemaGEOStructured Data

JSON-LD vs microdata vs visible HTML: which format do AI engines actually parse?

Structured data tells search engines and AI assistants what a page means, but the format you choose matters. Google officially supports three: JSON-LD, microdata, and RDFa. Yet for AI-generated answers, JSON-LD has quietly become the one that actually gets read.

The short answer: if your data is inside JSON-LD script tags, an AI engine can usually find and parse it. If it lives only in microdata attributes scattered through your HTML, many LLM-based systems will miss it entirely. And if it exists only as visible text with no markup, extraction is possible but slower and more error-prone.

Why JSON-LD wins for AI engines

JSON-LD keeps all the structured information together in one or two script blocks, separate from your visible content. That clean separation makes it the easiest format for a crawler or a large language model to isolate. Google says it prefers JSON-LD, and most schema generators and ecommerce platforms emit it by default. AI assistants that read a page's raw HTML tend to pull JSON-LD first because it is predictable to locate.

Microdata and RDFa are embedded directly inside your HTML elements with attribute tags. They are valid and Google still parses them, but they are noisier for a model that is not running a full schema validator. A model trying to answer a shopping question wants to quickly find price, availability, and shipping data. JSON-LD hands it to them in one place.

The realistic three-format ranking

In practice, the parsing success rate looks roughly like this:

1. JSON-LD in the head: highest success rate across Google, ChatGPT, and Perplexity. 2. Plain visible HTML text: reliably found, but the model must interpret it, so wording matters. 3. Microdata or RDFa attributes: valid markup, but the least consistently picked up by LLM-based engines that read raw HTML without running schema tooling.

A useful mental model: JSON-LD is the machine-readable answer, and visible text is the fallback the model reads when structured data is missing or unclear. They work best together, saying the same thing.

How to move your key pages to JSON-LD

1. Find which format your platform emits. Many Shopify themes and WordPress GEO plugins already output JSON-LD, so you may only need to check, not rebuild. 2. For every product page, confirm Price, Availability, and SKU live in a JSON-LD Product or Offer block, not just in microdata. 3. Verify the structured data and the visible page tell the same story. If the markup says in stock and the page text says backordered, an engine may distrust both. 4. Test with a validator or a free scan that shows what a model can extract, then fix the fields it reports as missing.

When visible HTML is the real fallback

Even a clean JSON-LD block loses value if the visible page hides the facts. Some engines, especially agentic shopping assistants, read the rendered text and ignore markup they cannot trust. So the two work as a pair: JSON-LD for fast reliable extraction, and plain, complete sentences on the page as insurance for models that only read text.

This is why the same lesson keeps coming back in GEO work: match your structured data to readable content, and do not rely on markup alone.

FAQ

*Do I need to add all three formats?*

No. JSON-LD plus clear visible text is enough. Adding microdata on top adds maintenance with little extra benefit, and can create conflicting signals if the three copies drift apart.

*What about WordPress plugins and Shopify apps?*

Both ecosystems increasingly ship JSON-LD by default. You are usually better off enabling a maintained plugin that outputs valid JSON-LD than hand-writing microdata, because a plugin updates its schema with schema.org as the vocabulary evolves.

*How do I know the AI engines are actually reading my structured data?*

The most direct check is to ask an assistant what it can extract from your product URL, or run a scan that shows the extractable fields. If price or stock are empty in the result, the engine is not reading your JSON-LD and you need to check the markup.

Visora's free scan at geovisora.com/audit reads your product URLs and shows which structured-data fields a model can extract today, so you can see whether your JSON-LD is actually making it through — and what to fix first. If you are unsure whether markup or visible text is the weak link, /faq walks through the common failure points before you make changes.

Put this into practice

Audit your PDP or category page with Visora, then fix schema and FAQ gaps that block AI citations.

Run a free GEO audit →

https://geovisora.com/en/blog/structured-data-formats-json-ld-microdata-html-ai