AI crawler source registry: what each operator actually publishes
26 source URLs, re-checked on a schedule. Last checked 2026-08-03. 11 of them do not do what a reader of the operator's documentation would expect.
A verifier that checks HTTP status codes and not response bodies will accept every row below as a successful fetch, including the one that returns an HTML page where JSON is declared. That is the failure this table exists to make visible.
- 15 Verified
- 8 Moved
- 2 Client rendered only
- 1 Wrong content type
Findings
https://developer.amazon.com/amazonbot/ip-addresses/
zero CIDR blocks present without executing JavaScript
https://developer.amazon.com/amazonbot/live-ip-addresses/
zero CIDR blocks present without executing JavaScript
https://docs.claude.com/claudebot.json
declared JSON, served text/html (735035 bytes). A verifier fetching this gets 200 and no ranges.
Widely cited as Anthropic's published range file. It returns HTTP 200 and serves the documentation site's HTML application shell, not JSON. A verifier that checks status codes rather than parsing the body treats this as a successful fetch of an empty range list and fails open: every request claiming to be ClaudeBot passes or every one fails, depending on which way the code errs. Anthropic's actual list is at claude.com/crawling/bots.json, named only in the body of a support article.
Documented location: https://claude.com/crawling/bots.json
https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
redirects to https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
redirects to https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
https://developers.google.com/static/search/apis/ipranges/googlebot.json
redirects to https://developers.google.com/static/crawling/ipranges/common-crawlers.json
The legacy path. It still serves valid JSON, so allowlists pinned here keep working, but Google's current documentation points at /static/crawling/ipranges/ and treats that as canonical. Anything pinned to the old path is depending on a URL Google no longer documents.
Documented location: https://developers.google.com/static/crawling/ipranges/common-crawlers.json
https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/
redirects to https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
https://platform.openai.com/docs/bots
redirects to https://developers.openai.com/api/docs/bots
https://docs.perplexity.ai/guides/bots
redirects to https://docs.perplexity.ai/docs/resources/perplexity-crawlers
https://www.perplexity.com/perplexity-user.json
redirects to https://www.perplexity.ai/perplexity-user.json
https://www.perplexity.com/perplexitybot.json
redirects to https://www.perplexity.ai/perplexitybot.json
All sources
| Operator | Covers | URL | Status | Served | CIDRs | Classification |
|---|---|---|---|---|---|---|
| Amazon | Amzn-SearchBot, Amzn-User, Amazonbot | developer.amazon.com/amazonbot | 200 | text/html | n/a | Verified |
| Amazon | Amazonbot | developer.amazon.com/amazonbot/ip-addresses/ | 200 | text/html | 0 | Client rendered only |
| Amazon | Amzn-User | developer.amazon.com/amazonbot/live-ip-addresses/ | 200 | text/html | 0 | Client rendered only |
| Amazon | Amzn-SearchBot | developer.amazon.com/amazonbot/searchbot-ip-addresses/ | 200 | text/html | 512 | Verified |
| Anthropic | ClaudeBot, Claude-SearchBot, Claude-User | claude.com/crawling/bots.json | 200 | application/json | 20 | Verified |
| Anthropic | ClaudeBot, Claude-SearchBot, Claude-User | docs.claude.com/claudebot.json | 200 | text/html | 1 | Wrong content type |
| Anthropic | ClaudeBot, Claude-SearchBot, Claude-User | support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler | 200 | text/html | n/a | Moved |
| Apple | Applebot | search.developer.apple.com/applebot.json | 200 | application/json | 12 | Verified |
| Apple | Applebot, Applebot-Extended | support.apple.com/en-us/119829 | 200 | text/html | n/a | Verified |
| Common Crawl Foundation | CCBot | commoncrawl.org/ccbot | 200 | text/html | n/a | Verified |
| Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot | developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests | 200 | text/html | n/a | Verified | |
| Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot, Google-Extended | developers.google.com/search/docs/crawling-indexing/google-common-crawlers | 200 | text/html | n/a | Moved | |
| Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot | developers.google.com/static/crawling/ipranges/common-crawlers.json | 200 | application/json | 169 | Verified | |
| Googlebot | developers.google.com/static/search/apis/ipranges/googlebot.json | 200 | application/json | 169 | Moved | |
| Meta | Meta-ExternalAgent, Meta-ExternalFetcher, Meta-WebIndexer, Meta-ExternalAds, FacebookExternalHit | developers.facebook.com/docs/sharing/webmasters/web-crawlers/ | 200 | text/html | n/a | Moved |
| Microsoft | Bingbot | www.bing.com/toolbox/bingbot.json | 200 | application/json | 28 | Verified |
| Microsoft | Bingbot, AdIdxBot | www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0 | 200 | text/html | n/a | Verified |
| OpenAI | OAI-AdsBot | openai.com/adsbot.json | 200 | application/json | 2 | Verified |
| OpenAI | ChatGPT-User | openai.com/chatgpt-user.json | 200 | application/json | 322 | Verified |
| OpenAI | GPTBot | openai.com/gptbot.json | 200 | application/json | 21 | Verified |
| OpenAI | OAI-SearchBot | openai.com/searchbot.json | 200 | application/json | 35 | Verified |
| OpenAI | GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot | platform.openai.com/docs/bots | 200 | text/html | n/a | Moved |
| Perplexity | PerplexityBot, Perplexity-User | docs.perplexity.ai/guides/bots | 200 | text/html | n/a | Moved |
| Perplexity | PerplexityBot | www.perplexity.ai/perplexitybot.json | 200 | application/json | 8 | Verified |
| Perplexity | Perplexity-User | www.perplexity.com/perplexity-user.json | 200 | application/json | 4 | Moved |
| Perplexity | PerplexityBot | www.perplexity.com/perplexitybot.json | 200 | application/json | 8 | Moved |
Detected changes
No changes detected since tracking began on 2026-08-03. Each scheduled run compares against the last committed result and records anything that moved, broke, or changed contents. Full change history.
Method
Each URL is fetched with redirects followed. The response body is parsed rather than trusted:
a source declared as a machine-readable range file must parse as JSON and expose a
prefixes array, and a documentation page said to list addresses must contain
CIDR blocks in the HTML actually served. Locale redirects are normalised away, so
moved means the host or path changed, not the language.