AI crawler source registry: what each operator actually publishes

26 source URLs, re-checked on a schedule. Last checked 2026-08-03. 11 of them do not do what a reader of the operator's documentation would expect.

A verifier that checks HTTP status codes and not response bodies will accept every row below as a successful fetch, including the one that returns an HTML page where JSON is declared. That is the failure this table exists to make visible.

Findings

https://developer.amazon.com/amazonbot/ip-addresses/

Client rendered only Amazon · covers Amazonbot · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

zero CIDR blocks present without executing JavaScript

https://developer.amazon.com/amazonbot/live-ip-addresses/

Client rendered only Amazon · covers Amzn-User · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

zero CIDR blocks present without executing JavaScript

https://docs.claude.com/claudebot.json

Wrong content type Anthropic · covers ClaudeBot, Claude-SearchBot, Claude-User · HTTP 200 · declared application/json, served text/html · 1 CIDR parsed

declared JSON, served text/html (735035 bytes). A verifier fetching this gets 200 and no ranges.

Widely cited as Anthropic's published range file. It returns HTTP 200 and serves the documentation site's HTML application shell, not JSON. A verifier that checks status codes rather than parsing the body treats this as a successful fetch of an empty range list and fails open: every request claiming to be ClaudeBot passes or every one fails, depending on which way the code errs. Anthropic's actual list is at claude.com/crawling/bots.json, named only in the body of a support article.

Documented location: https://claude.com/crawling/bots.json

https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler

Moved Anthropic · covers ClaudeBot, Claude-SearchBot, Claude-User · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

redirects to https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler

https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers

Moved Google · covers Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot, Google-Extended · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

redirects to https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers

https://developers.google.com/static/search/apis/ipranges/googlebot.json

Moved Google · covers Googlebot · HTTP 200 · declared application/json, served application/json · 169 CIDRs parsed

redirects to https://developers.google.com/static/crawling/ipranges/common-crawlers.json

The legacy path. It still serves valid JSON, so allowlists pinned here keep working, but Google's current documentation points at /static/crawling/ipranges/ and treats that as canonical. Anything pinned to the old path is depending on a URL Google no longer documents.

Documented location: https://developers.google.com/static/crawling/ipranges/common-crawlers.json

https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/

Moved Meta · covers Meta-ExternalAgent, Meta-ExternalFetcher, Meta-WebIndexer, Meta-ExternalAds, FacebookExternalHit · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

redirects to https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers

https://platform.openai.com/docs/bots

Moved OpenAI · covers GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

redirects to https://developers.openai.com/api/docs/bots

https://docs.perplexity.ai/guides/bots

Moved Perplexity · covers PerplexityBot, Perplexity-User · HTTP 200 · declared text/html, served text/html · 0 CIDRs parsed

redirects to https://docs.perplexity.ai/docs/resources/perplexity-crawlers

https://www.perplexity.com/perplexity-user.json

Moved Perplexity · covers Perplexity-User · HTTP 200 · declared application/json, served application/json · 4 CIDRs parsed

redirects to https://www.perplexity.ai/perplexity-user.json

https://www.perplexity.com/perplexitybot.json

Moved Perplexity · covers PerplexityBot · HTTP 200 · declared application/json, served application/json · 8 CIDRs parsed

redirects to https://www.perplexity.ai/perplexitybot.json

All sources

OperatorCoversURLStatus ServedCIDRsClassification
Amazon Amzn-SearchBot, Amzn-User, Amazonbot developer.amazon.com/amazonbot 200 text/html n/a Verified
Amazon Amazonbot developer.amazon.com/amazonbot/ip-addresses/ 200 text/html 0 Client rendered only
Amazon Amzn-User developer.amazon.com/amazonbot/live-ip-addresses/ 200 text/html 0 Client rendered only
Amazon Amzn-SearchBot developer.amazon.com/amazonbot/searchbot-ip-addresses/ 200 text/html 512 Verified
Anthropic ClaudeBot, Claude-SearchBot, Claude-User claude.com/crawling/bots.json 200 application/json 20 Verified
Anthropic ClaudeBot, Claude-SearchBot, Claude-User docs.claude.com/claudebot.json 200 text/html 1 Wrong content type
Anthropic ClaudeBot, Claude-SearchBot, Claude-User support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler 200 text/html n/a Moved
Apple Applebot search.developer.apple.com/applebot.json 200 application/json 12 Verified
Apple Applebot, Applebot-Extended support.apple.com/en-us/119829 200 text/html n/a Verified
Common Crawl Foundation CCBot commoncrawl.org/ccbot 200 text/html n/a Verified
Google Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests 200 text/html n/a Verified
Google Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot, Google-Extended developers.google.com/search/docs/crawling-indexing/google-common-crawlers 200 text/html n/a Moved
Google Googlebot, GoogleOther, Googlebot News, GoogleOther-Image, GoogleOther-Video, Google-CloudVertexBot developers.google.com/static/crawling/ipranges/common-crawlers.json 200 application/json 169 Verified
Google Googlebot developers.google.com/static/search/apis/ipranges/googlebot.json 200 application/json 169 Moved
Meta Meta-ExternalAgent, Meta-ExternalFetcher, Meta-WebIndexer, Meta-ExternalAds, FacebookExternalHit developers.facebook.com/docs/sharing/webmasters/web-crawlers/ 200 text/html n/a Moved
Microsoft Bingbot www.bing.com/toolbox/bingbot.json 200 application/json 28 Verified
Microsoft Bingbot, AdIdxBot www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0 200 text/html n/a Verified
OpenAI OAI-AdsBot openai.com/adsbot.json 200 application/json 2 Verified
OpenAI ChatGPT-User openai.com/chatgpt-user.json 200 application/json 322 Verified
OpenAI GPTBot openai.com/gptbot.json 200 application/json 21 Verified
OpenAI OAI-SearchBot openai.com/searchbot.json 200 application/json 35 Verified
OpenAI GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot platform.openai.com/docs/bots 200 text/html n/a Moved
Perplexity PerplexityBot, Perplexity-User docs.perplexity.ai/guides/bots 200 text/html n/a Moved
Perplexity PerplexityBot www.perplexity.ai/perplexitybot.json 200 application/json 8 Verified
Perplexity Perplexity-User www.perplexity.com/perplexity-user.json 200 application/json 4 Moved
Perplexity PerplexityBot www.perplexity.com/perplexitybot.json 200 application/json 8 Moved

Detected changes

No changes detected since tracking began on 2026-08-03. Each scheduled run compares against the last committed result and records anything that moved, broke, or changed contents. Full change history.

Method

Each URL is fetched with redirects followed. The response body is parsed rather than trusted: a source declared as a machine-readable range file must parse as JSON and expose a prefixes array, and a documentation page said to list addresses must contain CIDR blocks in the HTML actually served. Locale redirects are normalised away, so moved means the host or path changed, not the language.

All crawlers · Machine-readable dataset