Edit your robots.txt file to list the exact AI crawler user-agents you want to allow or block, then confirm your CDN or WAF won’t silently override those rules. This combination handles most cases: Google and OpenAI both publish crawler names you can target directly, and Cloudflare offers a one-click training opt-out. After deployment, watch your server logs for a week to make sure the rules behave as intended.
TL;DR:
- Allow OpenAI’s search crawler while blocking GPTBot if you want ChatGPT search visibility without permitting model training; Google’s generative AI controls remain separate from rankings.
- Robots.txt only requests access rules; check CDN and firewall settings, since they can block permitted crawlers, and use verified crawler exceptions when needed.
- User agent strings are easy to spoof, so verify crawler requests against vendors’ published IP ranges and use reverse DNS or TLS checks for added confidence.
- Use X-Robots-Tag for PDFs, images, and API responses, then check logs after deployment and recheck any block or rate limit within 24 to 48 hours.
Table of Contents
- Quick allow/block checklist
- Robots.txt: exact directives and examples
- WAF, CDN, and bot management: allowlists and exceptions
- User-agent strings, published IP lists, and verification
- HTTP headers and meta directives: X-Robots-Tag, nosnippet, noindex, and training opt-out
- Platform controls: OpenAI, Cloudflare, and Google guidance
- Monitoring and troubleshooting abusive or misbehaving crawlers
- Deployment examples: Next.js, WordPress/WooCommerce, and Edge functions
- How Stellor helps measure and manage AI citation and crawler behavior
- Balancing discoverability and training opt-outs
- Stellor as an alternative to manual monitoring and rule maintenance
- FAQ
- Sources
Quick allow/block checklist
Before touching any config file, decide what outcome you actually want. “Allow everything” and “block AI training but keep search visibility” require different moves.
- Choose your policy: full access, search-only access, or a complete block.
- Update robots.txt with explicit
User-agentlines for each crawler, adding a Disallow AI Training directive where your host supports it. - Check your CDN or WAF dashboard. Many bot-management systems block crawlers that robots.txt technically allows.
- Add page-level
X-Robots-Tagor meta directives for assets robots.txt can’t cover, like PDFs or API responses. - Watch your logs for a week and add temporary rate limits if any crawler behaves badly.
Each step closes a gap the previous one leaves open. Robots.txt alone is a request, not an enforcement mechanism, and that distinction drives most of the configuration below.
Robots.txt: exact directives and examples
Robots.txt is a plain text file at your domain root that crawlers check before fetching pages. According to MDN’s practical implementation guide, formatting mistakes like missing blank lines between blocks or incorrect path casing are among the most common reasons rules get ignored. Some crawlers also cache the file, so changes don’t always take effect instantly.
A typical setup that allows OpenAI’s search crawler while blocking its training crawler looks like this:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
OpenAI’s crawler documentation lists OAI-SearchBot and GPTBot as separate agents with separate purposes, which is why you can permit one and block the other in the same file. Where your host supports it, a Disallow AI Training style directive lets you draw that same line without listing every bot by name, a pattern Cloudflare has built directly into its dashboard.
Watch out for these common pitfalls:
- A robots.txt that returns a 404 or redirects is often treated as “no rules exist,” not as a full block.
- Trailing slashes matter.
/blogand/blog/can match differently depending on the crawler. - Caching means your fix might not show up in logs for several hours.
WAF, CDN, and bot management: allowlists and exceptions
Robots.txt is a request. Your firewall doesn’t have to honor it. A CDN or WAF can flag an AI crawler as suspicious traffic and block it entirely, regardless of what your robots.txt says, which is the single most common reason “allowed” crawlers never show up in your analytics.
- Check your bot management rules before assuming robots.txt is the problem.
- In Cloudflare, you can create a custom rule that uses a specific Bot Detection ID to Skip or Allow a verified crawler, which is documented in Cloudflare’s AI Search configuration guide.
- User-agent allowlisting is easy to set up but easy to spoof. IP range allowlisting is more reliable but requires you to keep the list current.
- Apply rate limits or challenge pages to traffic that claims to be a known crawler but behaves erratically.
Pro Tip: Test your allowlist with a tool that spoofs the user-agent string before trusting it in production; a rule that only checks UA will pass that test even when it shouldn’t.
User-agent strings, published IP lists, and verification
Anyone can set their user-agent string to “GPTBot.” That’s why serious verification goes beyond the header.
- OpenAI and Cloudflare both publish JSON endpoints listing the IP ranges their crawlers use, referenced in OpenAI’s crawler overview, letting you cross-check a request’s source IP against the official range.
- Reverse DNS lookups and TLS certificate checks add a second layer of confidence beyond the user-agent string alone.
- Behavior checks matter too: a legitimate crawler respects crawl-delay and doesn’t hammer the same URL in a loop.
- Automate a weekly pull of published IP lists so your allowlist doesn’t go stale.
HTTP headers and meta directives: X-Robots-Tag, nosnippet, noindex, and training opt-out
Robots.txt controls whether a crawler fetches a page. It doesn’t control what happens after that, which is where headers and meta tags take over.
- Use the
X-Robots-Tagheader for PDFs, images, and API responses that have no HTML to carry a meta tag, as described in MDN’s X-Robots-Tag reference. nosnippetandnoarchivesuppress preview text and cached copies, which also affects what generative features can quote.- Support varies by vendor. Google and Bing both respect
noindexbroadly, but training-specific opt-outs are newer and less standardized. - Example:
X-Robots-Tag: noindex, nosnippetsent as a response header, or<meta name="robots" content="noindex">in the page head.
Platform controls: OpenAI, Cloudflare, and Google guidance
Each major platform handles permissions a little differently, and mixing them up is the fastest way to get an unintended result.
- OpenAI: allow OAI-SearchBot if you want to appear in ChatGPT search results, and separately disallow GPTBot if you want to opt out of training data collection. Verify both against the published IP lists.
- Cloudflare: enable Disallow AI Training when you want search crawlers to keep working while training crawlers get blocked, and add a Bot Detection ID exception when a legitimate crawler is caught by general bot filtering, per Cloudflare’s configuration docs.
- Google: the Search generative AI controls let you include or exclude your content from generative AI features without touching your regular search ranking settings, a distinction worth remembering since the two toggles are separate.
Monitoring and troubleshooting abusive or misbehaving crawlers
Even a correctly configured crawler can misbehave. Site owners have reported GPTBot ignoring robots.txt entirely and repeatedly hitting a single URL with different query-string variants, according to a report on OpenAI’s developer community, which drove up server load until the site owner applied a temporary block.
- Watch for spikes in request rate, repeated parameter variants on the same path, or unusual status codes in your logs.
- Apply a short-term IP block, rate limit, or challenge page while you investigate.
- If the pattern continues, report it to the crawler’s operator with timestamps, request paths, and your robots.txt content.
- Recheck your logs 24 to 48 hours after any change to confirm the new rule actually reduced the traffic.
Pro Tip: Keep a saved log sample from before and after your fix. It’s the fastest way to prove a rule worked when you’re troubleshooting later.
Deployment examples: Next.js, WordPress/WooCommerce, and Edge functions
Implementation details shift depending on your stack, but the goal stays the same: serve the right rules consistently.
- Next.js: drop a static robots.txt in the
publicfolder, or generate one dynamically with the metadata robots helper when rules need to change by environment. - WordPress/WooCommerce: a plugin-generated robots.txt can conflict with a root-level file, so pick one source of truth, and for high-traffic abuse, add server-level rules through your host rather than relying on the plugin alone.
- Edge functions: a lightweight function can inspect the user-agent or IP at the edge and return a 403 or a throttled response before the request ever reaches your origin server.
A rule that isn’t tested in production is just a guess about what your server will do.
After deploying, fetch your own robots.txt from a browser, check response headers with a tool like curl, and confirm the crawler shows up correctly (or not at all) in your logs within a day.
How Stellor helps measure and manage AI citation and crawler behavior
Weekly audits inside Stellor check robots.txt configuration, llms.txt readiness, and schema completeness alongside the technical factors Google has always cared about. It also tracks whether your business gets cited across ChatGPT, Claude, Perplexity, and Gemini, and a free onboarding AI Visibility Audit delivers a 90-day action plan showing exactly what to fix first. For a deeper look at how AI citation affects visibility, see how AI search is changing local business discovery.

Balancing discoverability and training opt-outs
Most site owners treat this as an all-or-nothing choice, and that’s the wrong frame. If organic search traffic matters to your business, allow the search-specific crawler user-agents and pair that with a training opt-out where your host supports it. If your content is the product itself, the training opt-out deserves more weight than the discoverability upside. Vendor behavior shifts often enough that a setting worth trusting in January can need a second look by summer, so put a quarterly review on the calendar rather than treating this as a one-time setup.
— Cole
Stellor as an alternative to manual monitoring and rule maintenance
Chasing crawler updates across OpenAI, Cloudflare, and Google by hand gets old fast, especially when each vendor changes its rules on its own schedule. We built Stellor to handle that ongoing work instead of leaving it on your plate.

- Weekly audits flag robots.txt and crawler configuration issues with one-click fixes.
- LLM visibility checks across ChatGPT, Claude, Perplexity, and Gemini show exactly where your business gets cited or missed.
- A free AI Visibility Audit during onboarding gives you a 90-day plan before you commit to anything.
Readers in home services who want a sector-specific breakdown can check our GEO and SEO guidance for home service businesses. If broader AI content tactics are also on your radar, our partner Solaya covers AI imagery use cases for marketers worth a look. Start your 3-day free trial, no card required, and see your first audit within 48 hours.
FAQ
Does blocking GPTBot hurt my Google ranking?
No. Google’s regular search index is separate from OpenAI’s crawlers, so disallowing GPTBot in robots.txt has no effect on your position in Google’s results. It only affects whether OpenAI can use your content for model training.
How long does a robots.txt change take to apply?
It depends on the crawler, but some vendors can take roughly a day to reflect changes, according to OpenAI’s crawler documentation. Always verify with server logs instead of assuming the change is live immediately.
Can a crawler ignore my robots.txt file entirely?
Yes. Robots.txt is a voluntary request, not an enforcement mechanism, and misbehaving bots have been reported ignoring it and repeatedly hitting the same URL. A CDN or WAF rule is the only way to actually enforce a block.
What’s the difference between blocking training and blocking search indexing?
Blocking training stops a model from learning from your content, while blocking search indexing stops a crawler from showing your pages in results or answers. Cloudflare’s Disallow AI Training feature and Google’s generative AI controls let you manage these separately.
Should I allowlist by user-agent or by IP address?
User-agent allowlisting is simpler but easy to spoof, so pair it with the published IP ranges vendors like OpenAI release for stronger verification. For most sites, checking both together catches impersonators that a user-agent check alone would miss.

