# Crawl Nine, https://crawlnine.com # # Everything is allowed, to everyone. The named blocks below grant nothing that # the wildcard has not already granted; they exist because a crawler operator # checking whether they are welcome should be able to see their own token in # this file, and because an explicit Allow survives a future tightening of the # wildcard rule. # # Training crawlers are allowed too. Long-term familiarity in model weights is # worth more to a brand nobody knows yet than withheld training data. # # A 200 here is necessary but not sufficient: Cloudflare can still return 403 at # the edge regardless of what this file says. Verify with scripts/verify-crawlers.sh. User-agent: * Allow: / # --- OpenAI --- User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-AdsBot Allow: / User-agent: GPTBot Allow: / # --- Perplexity --- User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # --- Anthropic --- User-agent: ClaudeBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # --- Google --- User-agent: Googlebot Allow: / User-agent: Google-Extended Allow: / # --- Microsoft --- # Bing's index feeds ChatGPT Search and Copilot. A page missing from Bing cannot # be cited by ChatGPT regardless of its Google rank. User-agent: Bingbot Allow: / # --- Apple --- User-agent: Applebot Allow: / # --- DuckDuckGo --- User-agent: DuckAssistBot Allow: / # --- Meta --- User-agent: Meta-ExternalFetcher Allow: / # --- Amazon --- User-agent: Amazonbot Allow: / # --- Common Crawl --- User-agent: CCBot Allow: / Sitemap: https://crawlnine.com/sitemap-index.xml