Guide

Crawlers don't see what you see.

Open your site in a browser and it looks finished. A crawler opening the same URL can get a blank shell, a redirect, or a wall. The gap between those two views is where structured data and AI visibility quietly break.

A crawler is not a browser

Your browser runs JavaScript, keeps cookies, follows redirects, and waits for the page to settle. A crawler usually does the cheap thing: one HTTP request, read the HTML that comes back, move on. If the price, the headings, or the JSON-LD only appear after client-side rendering, the crawler never sees them. It read the page before your page existed.

That is not a hypothetical. In our Ghost Price study a naive shopping agent could read a machine-readable price on just 4 of 50 major US retail sites. The other 46 were blocked, unreachable, or served a page with no readable price at all. Read the full breakdown in Ghost Price.

What a crawler cannot read

Three things trip most sites. Content injected after load, because the first response is empty. Content behind a bot wall, because the request is refused before it returns anything. And content that is present but unlabeled, so a machine sees text but cannot tell a price from a phone number. The first two are access problems. The third is a markup problem, and it is the one that structured data exists to solve.

Test what a crawler actually gets

Do not trust the browser view. Fetch your own URL the way a crawler would, with plain HTTP and no JavaScript, and read the raw HTML. If the facts you care about are missing, a crawler is missing them too. Then check who is actually reaching you: a lot of crawler traffic is spoofed, so a name in your logs is not proof. Bot or Not verifies a crawler's claimed identity against the operator's own published IP ranges, so you know which agents are real before you decide what to do about them.

Decide your robots.txt on evidence

Once you know which crawlers are genuine, the next question is what each one claims to respect. Vendors state their rules in documentation that rarely matches the folklore. Crawler Verdict tracks what each AI crawler says it respects in robots.txt, checked against the vendor's own docs, so your allow and disallow rules rest on primary sources rather than guesses. Block what you mean to block, allow what you want read, and verify that the crawler honoring the rule is the one you think it is.