Machine layer
CrawlNineBot
Published · Last updated
If you found CrawlNineBot in your server logs, this page tells you what it is, why it
came, and how to stop it. There is nothing to sign up for and nothing to agree to.
What is CrawlNineBot?
CrawlNineBot is the automated visitor behind the two free checking tools on this site. It fetches a web page, looks at what came back, and reports whether AI assistants can read that page. It runs from Crawl Nine, an AI visibility and web performance agency, and it belongs to nobody else.
It does not collect content to train anything. It does not build a search index. It reads a page, measures a handful of things about the response, and stops.
Why did it visit my site?
Because somebody typed your web address into one of our free tools and pressed the button. That is the only thing that makes CrawlNineBot visit anything. There is no schedule, no crawl queue, and no list of sites we work through.
Most often the person who asked was you, or somebody who looks after your website. It can also be somebody comparing your site with their own, the tools are free and open to anyone, exactly like every other “check a website” tool on the internet. We think that is worth saying plainly rather than implying every visit was invited by the owner.
What does it actually do?
One page, one time, on request. Nothing recursive.
- It asks for
/robots.txtfirst, and for the page whose address was typed in. - It asks for
/sitemap.xmland/sitemap-index.xml, to see whether a list of your pages is published. - It asks for that same page once for each AI assistant we test, so the site can answer each one differently if it wants to.
- If the page names a different address on the same site as its main one, it asks for that address once, without following it anywhere, to see whether it answers or redirects.
- That is at most 22 requests in total, over a few seconds, and then it is finished.
- It reads the page itself once, to check what is on it, and at most the first part of every other response, enough to tell a real page from a security challenge. It never downloads images, video, scripts or stylesheets.
- It does not fill in forms, follow links to other pages, or sign in to anything.
The requests come from Cloudflare’s network, so the address in your logs will be a Cloudflare one rather than one of ours.
Does it send my web address anywhere else?
Yes, to two places, and neither of them touches your site. When somebody checks a site, the grader also asks:
- Wikidata, a free, open database that anyone can edit, whether it has an entry that names your web address as its official website. It sends Wikidata your web address and nothing else.
- Google’s Chrome UX Report, how fast your site has been for real visitors on phones over the last 28 days. It sends Google your web address and nothing else.
Both are lookups of public data. Neither makes any request to your server, and neither
is made at all when your robots.txt asks CrawlNineBot to stay out. Each is sent with
the same identity as every other request.
Does it obey robots.txt?
Yes. If your robots.txt tells an assistant not to read a page, CrawlNineBot does not
request that page under that assistant’s name. It reports the answer from your file
instead, because your file has already given it. Its own requests, for the page, the
sitemaps and any address the page names as its main one, follow the rules your file sets
for CrawlNineBot.
This is not a courtesy. We sell the argument that robots.txt is how a business controls
what reads its website, and a tool that ignored the rules it found there would make that
argument impossible to keep making. Until 5 September 2026 this tool did ignore them,
and that was wrong.
What does it look like in my logs?
Every request identifies itself. The name of the assistant being tested comes first, so that your rules see what they would normally see, and our identity follows it:
OAI-SearchBot CrawlNineBot/1.0 (+https://crawlnine.com/bot)
ClaudeBot CrawlNineBot/1.0 (+https://crawlnine.com/bot)
Googlebot CrawlNineBot/1.0 (+https://crawlnine.com/bot)
One further request is sent with an ordinary browser user-agent and the same identity on the end. That one exists to tell “this site turns assistants away” apart from “this site is down”, without it, a site that is simply offline looks like a site that is blocking.
CrawlNineBot is never the real crawler whose name it carries. A request from us that
says Googlebot is our request, testing what your server does with that name. It comes
from a Cloudflare address and it will not pass Google’s own reverse-DNS check. If you
verify crawlers properly, you should see it fail that check, and that is correct.
How do I block it?
Add this to your robots.txt and we will stop, on every page, on the next check anyone
runs:
User-agent: CrawlNineBot
Disallow: /
To allow the checks but keep one area out of them, disallow only that area:
User-agent: CrawlNineBot
Disallow: /members/
Disallow: /checkout/
Two things worth knowing before you block it. Blocking CrawlNineBot has no effect on
whether ChatGPT, Google, Perplexity or Claude can read your site, it only stops people
finding out. And if you would rather not rely on a file, a firewall rule matching
CrawlNineBot in the user-agent works too; we send it on every single request, which is
the whole point of sending it.
How do I get in touch?
Write to hello@crawlnine.com. If CrawlNineBot has caused a problem on your server, say
so and we will look at it the same day, tell us the web address and roughly when, and we
can tell you what was asked for.
If you want the requests to stop immediately and do not want to edit a file, ask us and we will refuse checks for your domain at our end.