See what your site's robots.txt actually serves — including whether

Cloudflare's "AI Crawl Control" feature is silently appending its own rules

on top of yours.

While deploying a few small sites behind Cloudflare, we noticed our own

robots.txt files weren't what was actually being served. Cloudflare

injects a managed block (look for # BEGIN Cloudflare Managed content) that

adds its own AI-crawler-blocking rules — real search engines are usually

left alone, but plenty of site owners have no idea this is happening at all,

and it's easy to assume the file in your repo/origin is the one being served

when it isn't.

robots-check fetches the live file over HTTPS and tells you, in plain

terms:

- whether it's Cloudflare-managed (and where to check/adjust it)

- whether real search engines (Googlebot, Bingbot, etc.) are actually allowed to crawl

- which AI training/crawling bots (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, and others) are blocked vs. allowed

- whether a sitemap is declared at all

No API keys, no signup, no dependencies — just Node's built-in https

module.

npx github:janibert1/robots-check example.comOr clone it and run directly:

git clone https://github.com/janibert1/robots-check

cd robots-check

node robots-check.js example.comRequires Node 18+.

example.com — 47 non-blank lines

⚠ This robots.txt is (at least partly) Cloudflare-managed.

Cloudflare's "AI Crawl Control" feature injects its own block into what's

actually served — your origin server's own robots.txt file may say

something different from what's shown below. Check your Cloudflare

dashboard under Bots → AI Crawl Control if this doesn't match your intent.

Real search engines:

(any not listed above default to the same as User-agent: * → allowed)

AI training/crawling bots:

gptbot blocked

google-extended blocked

ccbot blocked

...

Sitemap: https://example.com/sitemap.xml

It can't see your origin server's own robots.txt file directly — only

what's actually served over HTTPS to a real request, which is what matters

for crawling anyway. If Cloudflare (or another CDN/WAF) sits in front of

your site, that's the version to trust regardless of what's in your repo.

MIT