AI crawlers: user agents, IP ranges and verification methods
29 crawlers and control tokens from 9 operators. Every IP range file listed here is fetched live at build time and every documentation URL is checked for a 200 response before this page ships. Last verified 2026-08-02.
| GPTBot | OpenAI | Training | GPTBot | Published CIDR list | 21 |
| OAI-SearchBot | OpenAI | Search | OAI-SearchBot | Published CIDR list | 35 |
| ChatGPT-User | OpenAI | Not a crawler | ChatGPT-User | Published CIDR list | 322 |
| OAI-AdsBot | OpenAI | Retrieval | OAI-AdsBot | Published CIDR list | 2 |
| ClaudeBot | Anthropic | Training | ClaudeBot | Published CIDR list | 20 |
| Claude-SearchBot | Anthropic | Search | Claude-SearchBot | Published CIDR list | 20 |
| Claude-User | Anthropic | Not a crawler | Claude-User | Published CIDR list | 20 |
| PerplexityBot | Perplexity | Search | PerplexityBot | Published CIDR list | 8 |
| Perplexity-User | Perplexity | Not a crawler | Perplexity-User | Published CIDR list | 4 |
| Googlebot | Search | Googlebot | Published CIDR list and reverse DNS | 169 | |
| GoogleOther | Retrieval | GoogleOther | Published CIDR list and reverse DNS | 169 | |
| Googlebot News | Search | Googlebot-News | Published CIDR list and reverse DNS | 169 | |
| GoogleOther-Image | Retrieval | GoogleOther-Image | Published CIDR list and reverse DNS | 169 | |
| GoogleOther-Video | Retrieval | GoogleOther-Video | Published CIDR list and reverse DNS | 169 | |
| Google-CloudVertexBot | Retrieval | Google-CloudVertexBot | Published CIDR list and reverse DNS | 169 | |
| Google-Extended | Control token | Google-Extended | None available | none | |
| Bingbot | Microsoft | Search | bingbot | Published CIDR list | 28 |
| AdIdxBot | Microsoft | Retrieval | adidxbot | None available | none |
| Applebot | Apple | Search | Applebot | Published CIDR list and reverse DNS | 12 |
| Applebot-Extended | Apple | Control token | Applebot-Extended | None available | none |
| Meta-ExternalAgent | Meta | Training | Meta-ExternalAgent | None available | none |
| Meta-ExternalFetcher | Meta | Not a crawler | Meta-ExternalFetcher | None available | none |
| Meta-WebIndexer | Meta | Search | Meta-WebIndexer | None available | none |
| Meta-ExternalAds | Meta | Retrieval | Meta-ExternalAds | None available | none |
| FacebookExternalHit | Meta | Retrieval | facebookexternalhit | None available | none |
| Amzn-SearchBot | Amazon | Search | Amzn-SearchBot | Published, but not machine-readable | none |
| Amzn-User | Amazon | Not a crawler | Amzn-User | Published, but not machine-readable | none |
| Amazonbot | Amazon | Training | Amazonbot | Published, but not machine-readable | none |
| CCBot | Common Crawl Foundation | Training | CCBot | Reverse DNS only | none |
Not included
These are commonly listed elsewhere. They are held back because no primary source could be fetched, and this reference does not publish claims it cannot trace to the operator.
- Bytespider (ByteDance) - No ByteDance-operated documentation page for Bytespider could be fetched. Every available description is third-party. Held until a primary source exists.
- Diffbot (Diffbot) - docs.diffbot.com serves API documentation, not crawler identification or verification guidance. No published user-agent or IP source found.
- Timpibot (Timpi) - timpi.io has no crawler documentation page and docs.timpi.io does not resolve. No primary source.