Block Training Bots if You Want — Don’t Accidentally Block Being Cited
Block GPTBot for training if you want — keep OAI-SearchBot allowed for ChatGPT search citations. Google-Extended does not control Google Search indexing.
William Spurlock Founder — Spurlock Studios Updated 22 MIN
You can block training crawlers if your policy says so — but do not treat every AI user-agent as “ChatGPT.” OpenAI’s GPTBot (training), OAI-SearchBot (ChatGPT search / citation index), and ChatGPT-User (user-initiated fetch) are different agents with different consequences. Blocking the wrong one removes you from being cited while you congratulate yourself for “opting out of AI.”
Google-Extended is a third category again: a robots.txt product token for Gemini training and grounding uses, not a search crawler and not a ranking lever. Google Search still rides on Googlebot.
This spoke is the robots.txt decision layer of the Answer Engine Optimization playbook. Pair it with llms.txt for brands so the file you publish matches the agents you actually allow.
The short answer
GPTBot≠ ChatGPT search. BlockingGPTBotopts out of OpenAI training use per OpenAI’s crawler overview. It does not, by itself, opt you out of ChatGPT search surfacing.OAI-SearchBotis the search crawler. Sites opted out “will not be shown in ChatGPT search answers,” though they may still appear as navigational links.ChatGPT-Useris a user-triggered fetch, not automatic crawl. OpenAI says robots.txt rules may not apply, and it is not the Search opt-out control.Google-Extendedis a product token for Gemini training and grounding uses — not the crawler that controls Google Search indexing. Google states it does not affect Search inclusion or ranking.- Default for most brands that want citations: allow search and retrieval bots; decide training bots as a policy call.
GPTBot vs search bots vs Google-Extended
Three different jobs. Three different files. One muddy “block AI” decision breaks the wrong one.
| Control | Vendor | What it is | What Disallow does | What it does not do |
|---|---|---|---|---|
GPTBot | OpenAI | Automated crawler for foundation-model training | Signals content should not be used to train OpenAI generative AI foundation models | Does not remove you from ChatGPT search. Does not touch Google. |
OAI-SearchBot | OpenAI | Automated crawler for ChatGPT search features | Site is not shown in ChatGPT search answers (may still appear as navigational links) | Does not opt you out of training. Does not control Google. |
ChatGPT-User | OpenAI | User-initiated fetch in ChatGPT / Custom GPTs | Live user-directed fetches may fail; OpenAI notes robots.txt may not apply | Not the Search inclusion control. Not a training crawler. |
Googlebot | Common crawler for Google Search (and related Search surfaces) | Blocks the Search crawl/index path for those URLs | Does not set Gemini training policy. That is Google-Extended. | |
Google-Extended | robots.txt product token — no separate HTTP user-agent | Opts crawled content out of listed Gemini training and grounding uses | Does not control Search inclusion, ranking, or AI Overviews eligibility |
OpenAI states these settings are independent. You can allow OAI-SearchBot while disallowing GPTBot. Google states Google-Extended is independent of Search. Treat “AI bot” as three questions, not one switch.
Decision list:
- Do we want ChatGPT search citations? →
OAI-SearchBot - Do we allow OpenAI foundation-model training? →
GPTBot - Do we want Google Search crawl/index? →
Googlebot(and do not confuse it with the next line) - Do we allow Gemini training / listed grounding uses? →
Google-Extended
Should I block GPTBot in robots.txt?
Only if you intend to opt out of OpenAI foundation-model training collection — not because you “don’t want to show up in ChatGPT.”
Per OpenAI’s crawler overview (developers.openai.com/api/docs/bots, checked 2026-08-16):
| User agent | Job | If you Disallow |
|---|---|---|
GPTBot | Crawl content that may be used to train OpenAI generative AI foundation models | Signals content should not be used in that training |
OAI-SearchBot | Surface sites in ChatGPT search features | Not shown in ChatGPT search answers (may still appear as navigational links) |
ChatGPT-User | User actions in ChatGPT / Custom GPTs fetch a page | Live user-directed fetches may fail; robots.txt may not apply; not the Search control |
OAI-AdsBot | Validate landing pages submitted as ChatGPT ads | Only visits ad landing pages; data is not used for foundation-model training |
OpenAI also notes that if a site allows both GPTBot and OAI-SearchBot, they may use results from a single crawl for both use cases to avoid duplicate crawling. That is an efficiency note, not a reason to treat the tokens as one bot.
Publisher-side confirmation lives in OpenAI’s Publishers and Developers FAQ: for content to be included in summaries and snippets in ChatGPT, do not block OAI-SearchBot. Referral traffic from ChatGPT search can carry utm_source=chatgpt.com.
- Legal/policy has a written training stance (allow or disallow
GPTBot) - Marketing has a written citation stance (allow or disallow
OAI-SearchBot) - Those two stances are not collapsed into one “AI” checkbox
-
OAI-AdsBotis left alone unless you actually run ChatGPT ads and need to restrict landing-page checks
What does each OpenAI agent look like in logs?
Do not match on a vibe. Match on the token and then on IP.
OpenAI publishes example user-agent strings and JSON IP lists on the same bots page. Version numbers change. Match the product token, not a pinned GPTBot/1.4 string.
| Agent | Example token in the UA | Published IPs | robots.txt behavior |
|---|---|---|---|
GPTBot | compatible; GPTBot/…; +https://openai.com/gptbot | openai.com/gptbot.json | Honors User-agent: GPTBot |
OAI-SearchBot | compatible; OAI-SearchBot/…; +https://openai.com/searchbot | openai.com/searchbot.json | Honors User-agent: OAI-SearchBot |
ChatGPT-User | compatible; ChatGPT-User/…; +https://openai.com/bot | openai.com/chatgpt-user.json | User-initiated; robots.txt may not apply |
OAI-AdsBot | compatible; OAI-AdsBot/…; +https://openai.com/adsbot | openai.com/adsbot.json | Only ad landing pages; not training |
When OpenAI fetches robots.txt itself, the UA may include an extra robots.txt marker so you can tell a rules fetch from a content fetch in logs that omit paths. That marker is documentation hygiene, not a fifth bot.
| Log line | Treat as | Next action |
|---|---|---|
UA has GPTBot and IP is in gptbot.json | Real training crawl | Honor your GPTBot Allow/Disallow |
UA has OAI-SearchBot and IP is in searchbot.json | Real search crawl | This is your ChatGPT citation path |
UA has ChatGPT-User and IP is in chatgpt-user.json | User-triggered fetch | Do not use this as a Search opt-out |
| UA claims any of the above, IP is not in the matching JSON | Impostor | Rate-limit or block the IP; do not rewrite policy |
What does Google-Extended actually control?
Google-Extended is not a crawler you will see as its own HTTP user-agent. Google’s common crawlers list is explicit: crawling is done with existing Google user-agent strings; the robots.txt token is used in a control capacity.
What the token manages, in Google’s words: whether content Google crawls from your site may be used for training future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and for grounding (providing content from the Google Search index to the model at prompt time) in Gemini Apps and Grounding with Google Search on Vertex AI.
What Google also says, in the same entry: Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.
| Token | Separate HTTP UA? | Controls Search ranking? | Controls ChatGPT search? | Typical policy use |
|---|---|---|---|---|
Googlebot | Yes (Googlebot/2.1) | Yes — crawl/index path for Search surfaces | No | Search visibility |
Google-Extended | No | No (per Google) | No | Gemini training / listed grounding opt-out |
GoogleOther | Yes | No specific product | No | Generic Google R&D fetches; not Search |
GPTBot | Yes | No | No (training) | OpenAI training opt-out |
OAI-SearchBot | Yes | No | Yes | ChatGPT search eligibility |
Google’s crawling overview splits agents into common crawlers (always respect robots.txt on automatic crawls), special-case crawlers, and user-triggered fetchers (ignore robots.txt because a person asked). Google-Extended sits on the common-crawlers page as a token, not as a fourth fetch class.
- You have a
User-agent: Google-Extendedblock if legal wants a Gemini training/grounding stance - You did not Disallow
Googlebotthinking you only touched “AI” - You did not expect
Google-Extendedto hide you from AI Overviews
Does Google-Extended turn off AI Overviews?
No. AI Overviews and AI Mode are Search features. They ride the Search index and snippet controls, not the Gemini training token.
Google’s robots meta tag spec states that nosnippet applies to all forms of search results — including AI Overviews and AI Mode — and prevents the content from being used as a direct input for those features. max-snippet likewise applies to those surfaces. Google-Extended is not in that list.
| Goal | Control that actually applies | Control that does not |
|---|---|---|
| Stay in Google Search | Allow Googlebot; do not noindex the URL | Google-Extended Allow/Disallow |
| Limit or block snippets / AI Overview input | nosnippet or max-snippet (and related snippet rules) | User-agent: Google-Extended Disallow: / |
| Opt out of listed Gemini training / grounding | User-agent: Google-Extended + Disallow: / | Blocking GPTBot |
| Stay in ChatGPT search answers | Allow OAI-SearchBot | Blocking Google-Extended |
If marketing wants “no AI Overviews” and legal wants “no Gemini training,” those are two tickets. Write both. Do not ship one token and assume it covers the other.
What about ClaudeBot and Claude-SearchBot?
Anthropic documents three robots on its crawler help article. Same split as OpenAI: training vs search vs user fetch. Anthropic states its bots honor robots.txt, including the user-initiated agent — that is a difference from OpenAI’s ChatGPT-User note.
| Bot | Job (Anthropic’s wording) | If you restrict it |
|---|---|---|
ClaudeBot | Collect web content that could contribute to model training | Signals future materials should be excluded from training datasets |
Claude-SearchBot | Navigate the web to improve search result quality | Prevents indexing for search optimization; may reduce visibility and accuracy in user search results |
Claude-User | Fetch a site when a person asks Claude a question | Prevents retrieval in response to a user query |
Anthropic also warns that IP-blocking their bots can prevent them from reading robots.txt at all, which breaks a clean opt-out. Prefer robots.txt signals over silent IP bans when you want the vendor to honor a Disallow.
| Goal | OpenAI control | Anthropic control | Google control |
|---|---|---|---|
| Training opt-out | Disallow GPTBot | Disallow ClaudeBot | Disallow Google-Extended |
| Search / citation eligibility | Allow OAI-SearchBot | Allow Claude-SearchBot | Allow Googlebot (Search); snippet rules for Overviews |
| User-initiated fetch | Understand ChatGPT-User limits | Claude-User (honors robots.txt per Anthropic) | User-triggered fetchers ignore robots.txt |
Confirm on the vendor page before you ship. Names and scopes change. This table is a worksheet, not a forever spec.
What robots.txt pattern should most brands ship?
For brands that want answer-engine citations and are willing to decide training separately:
# Citation / search retrieval — keep allowed
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
# Training — policy decision (example: opt out)
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
# Gemini training/grounding token — policy decision
User-agent: Google-Extended
Disallow: /
# Search crawl stays on Googlebot via User-agent: * or an explicit Allow
# Do not blanket-ban everything unknown with User-agent: * Disallow: /
# unless you also Allow the bots you need.
Adjust the training lines to Allow: / if legal wants training inclusion. The important part is splitting the decisions.
Google’s robots.txt interpretation requires the file at the site root on a supported protocol. Rules without a path are ignored. User-agent groups are matched by token. A Disallow: / under User-agent: * will also hit any bot that falls through to * — including search crawlers you meant to keep.
| Pattern | What happens | Use when |
|---|---|---|
| Split groups: search Allow, training Disallow | Independent policy, as vendors document | Default for citation-seeking brands |
User-agent: * + Disallow: / only | You opted out of almost everything that honors * | Rare; regulated or private sites |
One group listing every AI name under * | Easy to miss a token; easy to hit Googlebot by accident | Do not |
| Training Disallow, no search Allow | Fine if * still allows crawl — until someone tightens * | Add explicit search Allows anyway |
Decision list before you paste:
- ChatGPT search citations? → explicit
OAI-SearchBotAllow - Claude search-style indexing? → explicit
Claude-SearchBotAllow - Foundation-model training? →
GPTBot/ClaudeBot/Google-Extendedper vendor - Did a WAF or CDN already block these at the edge? → fix that next
- Is
/robots.txtitself fetchable anonymously? → if no, no bot can honor it
Which robots.txt syntax mistakes silently opt you out?
Most “we blocked ChatGPT by accident” incidents are not policy. They are a * group, a missing path, or a file that is not the file the crawler fetched.
Google’s robots.txt spec is the reference for how Google reads the file. OpenAI’s bots page is the reference for which tokens OpenAI honors. A valid file that names the wrong token is still a miss.
| Mistake | What the crawler sees | Fix |
|---|---|---|
User-agent: * + Disallow: / and no explicit search Allows | Search bots that fall through to * are opted out | Add User-agent: OAI-SearchBot / Googlebot Allows above the catch-all, or drop the catch-all |
Disallow: with no path | Rule is ignored (Google: rules without a path do not restrict) | Write Disallow: / if you mean the whole site |
User-agent: GPT-Bot or OAI_SearchBot | Token does not match; group is skipped | Use the exact tokens from the vendor page |
File at /static/robots.txt or a CMS “SEO plugin” URL | Crawlers fetch https://host/robots.txt only | Serve the file at the host root |
| HTML error page with status 200 | Crawler receives markup, not groups | Return text/plain and 200 with real groups |
robots.txt on www empty, apex has the rules (or the reverse) | Each host has its own file | Duplicate the worksheet on every host you serve |
Path-level example when legal wants training off the docs set but marketing still wants the About page cited:
User-agent: GPTBot
Disallow: /docs/
Allow: /about/
User-agent: OAI-SearchBot
Allow: /
That is still two decisions. Do not “simplify” it back into one * block after the first incident ticket.
-
curl -sIthe exact URLhttps://<host>/robots.txtfor apex andwww - Diff the served bytes against the worksheet, not against last month’s gist
- Confirm no plugin is injecting a second file that origin never sees
What about www, staging, and other hosts?
robots.txt is per host. A clean production file does not cover staging.example.com, www.example.com if that host answers separately, or a marketing subdomain that still ranks.
| Host | Typical failure | What to ship |
|---|---|---|
Apex example.com | File exists; www 301s after a bot already fetched an empty www file | Same groups on both hosts, or one host + a consistent redirect before the fetch |
www | CDN serves a default deny | Explicit search Allows on the host users and crawlers actually hit |
| Staging / preview | User-agent: * Disallow: / copied to production in a merge | Staging deny is fine; production must not inherit it |
| Docs subdomain | Training Disallow on apex only | Repeat GPTBot / Google-Extended on the host that holds the corpus |
Googlebot is documented on the Googlebot page as the Search crawler for that host’s URLs. OpenAI’s search crawler is the same idea: it fetches the host you allowed. Do not assume a Disallow on the marketing site covers the help center.
- Inventory every public host that serves indexable HTML
- Each host has a reachable
/robots.txt - Staging deny is not the production file
- Subdomains that hold the real answers have the same search Allows as the homepage
Failure mode: the “block all AI” CDN default
What breaks: a security dashboard enables “block AI crawlers” globally. robots.txt says Allow for OAI-SearchBot, but the edge returns 403 first. ChatGPT search never indexes you. You spend a quarter “doing AEO” on content that cannot be fetched.
What it costs: zero citations despite answer-first pages.
What you do instead:
- Fetch
https://yoursite.com/robots.txtanonymously — expect 200, not a challenge page - Confirm separate
User-agentblocks exist (not one muddy*) - Check CDN / WAF bot scores against OpenAI’s published ranges (gptbot.json, searchbot.json, chatgpt-user.json)
-
curl -A "OAI-SearchBot"(and vendor equivalents) on About, offer, and top answer pages — expect 200 - Re-test after every security policy change
- Confirm the WAF exception is documented, not a one-off ticket that dies in Slack
Google’s common crawlers always obey robots.txt on automatic crawls, per the common crawlers page. A 403 in front of robots.txt is still a 403. The file cannot save a fetch that never reaches origin.
This is the same fetchability check that belongs in an AEO audit: robots, then edge, then the page.
Will blocking GPTBot hurt Google rankings?
No — not as a Google ranking lever. GPTBot is OpenAI’s training crawler. Google Search crawl and index is Googlebot. These are different systems, documented on different vendor pages.
What is Googlebot is the Search crawler reference. Blocking GPTBot does not tell Google to demote you. Allowing GPTBot does not boost Google rank. Blocking Google-Extended does not demote you either — Google says the token is not a ranking signal.
| Action | Google Search effect (per Google / OpenAI docs) | ChatGPT search effect (per OpenAI) |
|---|---|---|
Disallow GPTBot | None | None (training only) |
Disallow OAI-SearchBot | None | Not shown in ChatGPT search answers |
Disallow Google-Extended | None (not a ranking signal) | None |
Disallow Googlebot | You blocked Search crawl | None |
| CDN 403 on “AI bots” | Depends whether Googlebot is in that rule | Often 403s OAI-SearchBot too |
Do not conflate “AI” into one SEO myth. I have been SEO certified since 2021. The AEO version of that work still starts with “which agent, which product.”
How fast do robots.txt changes take effect?
OpenAI documents that for search results, it can take about 24 hours from a site’s robots.txt update for their systems to adjust (bots overview). Other vendors differ. Assume hours to a few days for automated crawlers, then re-verify with log lines and live answer tests.
| Change | Expect |
|---|---|
Allow OAI-SearchBot after accidental block | ~24h for OpenAI search systems to adjust (per OpenAI), then longer for re-crawl of key URLs |
Disallow GPTBot | Future training collection should respect the signal; already-trained model memory is a separate, slower clock |
Disallow Google-Extended | Future Gemini training / listed grounding uses; not a Search recrawl event |
| CDN unblock | Immediate for new fetches; old index residue clears on re-crawl |
nosnippet / noindex | Requires the page to remain crawlable so the tag can be read (robots meta spec) |
Training residue in model weights is not cleared by robots.txt. robots.txt governs future crawl and use signals — not a memory erase.
OpenAI’s publisher FAQ adds a second clock: if they obtain a disallowed URL from a third-party search provider or by crawling other pages, they may still surface just the link and page title. If you do not want that, use noindex — and the crawler must be allowed to read the tag.
How to verify bots can fetch key pages
- Confirm
robots.txtallows the search agent on the path. - Confirm CDN/WAF allows the vendor’s published IPs / verified bot.
- Confirm the page returns 200 without login.
- Confirm the page is not
noindexif you also care about Google Search / AI Overviews. - Re-run five ChatGPT search prompts that should cite you; log whether your URL returns.
- robots.txt split training vs search vs
Google-Extended - WAF exceptions documented
- About + offer + top 5 answer URLs fetch clean with
OAI-SearchBotUA - Same URLs fetch clean with a normal browser UA (no accidental auth wall)
- Prompt panel archived post-change
-
/llms.txtis consistent with what you allow crawlers to read — see llms.txt for brands
Google verification is a different procedure from OpenAI’s JSON lists. Google’s verify crawler requests page gives two methods: reverse-then-forward DNS (googlebot.com / google.com / googleusercontent.com), or match the source IP to the published common-crawler / special-crawler / user-triggered-fetcher ranges.
| Check | Command / source | Pass |
|---|---|---|
| robots.txt reachable | curl -sI https://yoursite.com/robots.txt | 200, text/plain |
| Search bot fetch | curl -A "OAI-SearchBot" -sI https://yoursite.com/about/ | 200, not 403/401/challenge |
| OpenAI IP | Compare request IP to searchbot.json | Address in published ranges |
| Googlebot IP | Reverse + forward DNS per Google’s verify doc | Hostname and IP match |
Impostor bots and log hygiene
User-agent strings are trivial to spoof. Before you panic about “GPTBot ignoring robots.txt,” match the request IP to the vendor’s published ranges. Impostors wearing the UA are common.
| Check | Action |
|---|---|
| UA says GPTBot, IP not in gptbot.json | Treat as impostor; do not rewrite policy on fakes |
| UA + IP match, hits a Disallow path | File a vendor report if persistent; verify your robots.txt is reachable |
| robots.txt itself blocked by WAF | Fix that first — crawlers that cannot read rules cannot honor them |
UA says Googlebot, reverse DNS is not googlebot.com | Impostor; Google documents this spoof pattern on the Googlebot page |
Google common crawlers typically reverse-resolve to crawl-***-***-***-***.googlebot.com or geo-crawl-***-***-***-***.geo.googlebot.com. Special-case crawlers use a different mask (rate-limited-proxy-…google.com) and may or may not honor robots.txt. Do not apply a Googlebot mental model to every Google UA.
Anthropic has warned that IP-blocking their bots can prevent them from reading robots.txt at all. Prefer robots.txt signals over silent IP bans when you want a clean opt-out.
- Log pipeline stores UA and source IP
- Weekly job diffs claimed-OpenAI IPs against the three JSON files
- Googlebot claims go through reverse+forward DNS before a block
- Security and marketing share one incident channel so a fake UA does not become a policy rewrite
Policy worksheet for legal + marketing
Fill this once. Store it next to the deploy checklist. Re-open it when a vendor changes their crawler page — not when a Slack thread says “AI is stealing our site.”
| Decision | Options | Owner | Notes |
|---|---|---|---|
OpenAI training (GPTBot) | Allow / Disallow | Legal | Future collection only |
ChatGPT search (OAI-SearchBot) | Allow / Disallow | Marketing | Citation eligibility |
ChatGPT user fetch (ChatGPT-User) | Allow / monitor / edge-restrict | Eng | robots.txt may not apply |
Anthropic training (ClaudeBot) | Allow / Disallow | Legal | Confirm current help article |
Anthropic search (Claude-SearchBot) | Allow / Disallow | Marketing | Visibility in Claude search-style results |
Gemini training/grounding (Google-Extended) | Allow / Disallow | Legal | Not Search ranking |
Google Search (Googlebot) | Allow (default) | Marketing + SEO | Do not mix with Extended |
| Review cadence | Quarterly or on vendor-doc change | Named pair | Date the last check |
- Worksheet signed off
- robots.txt matches worksheet
- CDN rules match worksheet
- Change log entry with date
-
/llms.txtdoes not promise access the edge then denies
If marketing wants citations and legal wants training opt-out, that is a normal, supported split — not a conflict that requires blocking everything. OpenAI’s bots page exists to make that split possible. Google’s Google-Extended entry exists for the same reason on the Gemini side.
What robots.txt cannot do
Use this as the “stop asking the file to do that” list.
| Ask | robots.txt answer | Actual control |
|---|---|---|
| Erase us from a model already trained | No | Weights are not a crawl cache |
| Remove us from Google AI Overviews | No (Google-Extended will not) | Snippet rules / nosnippet / noindex |
| Stop a person from pasting our URL into ChatGPT | Not reliably | ChatGPT-User may ignore robots.txt |
| Prove a UA is the real vendor | No | Published IP lists + DNS |
| Survive a WAF 403 | No | Edge allowlist |
| Replace a briefing file | No | llms.txt is a different artifact |
Google’s robots.txt spec also reminds you that user-triggered fetchers and safety crawlers are outside the Robots Exclusion Protocol. A Disallow is a signal to automated crawlers that honor it. It is not a legal lock, a DRM layer, or a ranking boost.
Lane overview for the rest of the visibility stack: /visibility.
FAQ
What is OAI-SearchBot vs GPTBot?
GPTBot crawls for OpenAI foundation-model training. OAI-SearchBot crawls to surface sites in ChatGPT search features. OpenAI treats them as independent robots.txt controls on the official bots page. Blocking training does not equal blocking search — and blocking search is how you disappear from ChatGPT search answers.
What about ClaudeBot and Google-Extended?
ClaudeBot is Anthropic’s training-oriented crawler; use Claude-SearchBot when the question is Claude search indexing. Google-Extended is Google’s product token for certain Gemini training and grounding uses — it does not control Google Search indexing or ranking. Do not block Googlebot thinking you only touched “AI.”
Do WAFs silently block AI bots?
Yes, often. CDN “AI scraper” defaults and bot-fight scores can 403 OAI-SearchBot while your robots.txt looks fine. Verify with user-agent curls and vendor IP lists after every security change. A 200 on the file and a 403 on the page is still a failed citation path.
Will blocking GPTBot hurt Google rankings?
No. Google rankings depend on Google’s crawlers and ranking systems, not on whether OpenAI’s training bot can fetch you. Keep Googlebot decisions separate from GPTBot decisions, and keep Google-Extended separate from both.
How fast do robots.txt changes take effect?
OpenAI notes roughly 24 hours for search systems to adjust after a robots.txt update. Plan for re-crawl lag beyond that. Training opt-outs affect future collection; they do not rewrite model memory overnight.
How do I verify bots can fetch key pages?
Allow the right agents in robots.txt, allow them at the CDN/WAF, confirm HTTP 200 on canonical answer URLs with the bot user-agent, match IPs to OpenAI and Google published lists, then re-test live ChatGPT search prompts and log citations.
CTA
Split training policy from citation eligibility — then prove the fetch with a 200, not a vibes-based robots.txt screenshot.
Lane overview: /visibility. Next step: a visibility audit.
What questions does this article answer?
- What is OAI-SearchBot vs GPTBot?
- `GPTBot` crawls for OpenAI foundation-model training. `OAI-SearchBot` crawls to surface sites in ChatGPT search features. OpenAI treats them as independent robots.txt controls on the [official bots page](https://developers.openai.com/api/docs/bots). Blocking training does not equal blocking search — and blocking search is how you disappear from ChatGPT search answers.
- What about ClaudeBot and Google-Extended?
- `ClaudeBot` is Anthropic’s training-oriented crawler; use `Claude-SearchBot` when the question is Claude search indexing. `Google-Extended` is Google’s product token for certain Gemini training and grounding uses — it does not control Google Search indexing or ranking. Do not block `Googlebot` thinking you only touched “AI.”
- Do WAFs silently block AI bots?
- Yes, often. CDN “AI scraper” defaults and bot-fight scores can 403 `OAI-SearchBot` while your robots.txt looks fine. Verify with user-agent curls and vendor IP lists after every security change. A 200 on the file and a 403 on the page is still a failed citation path.
- Will blocking GPTBot hurt Google rankings?
- No. Google rankings depend on Google’s crawlers and ranking systems, not on whether OpenAI’s training bot can fetch you. Keep `Googlebot` decisions separate from `GPTBot` decisions, and keep `Google-Extended` separate from both.
- How fast do robots.txt changes take effect?
- OpenAI notes roughly 24 hours for search systems to adjust after a robots.txt update. Plan for re-crawl lag beyond that. Training opt-outs affect future collection; they do not rewrite model memory overnight.
- How do I verify bots can fetch key pages?
- Allow the right agents in robots.txt, allow them at the CDN/WAF, confirm HTTP 200 on canonical answer URLs with the bot user-agent, match IPs to [OpenAI](https://openai.com/searchbot.json) and [Google](https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests) published lists, then re-test live ChatGPT search prompts and log citations.
- developers.openai.com
- help.openai.com
- openai.com
- openai.com
- openai.com
- openai.com
- developers.google.com
- developers.google.com
- developers.google.com
- support.claude.com
- developers.google.com
- developers.google.com
- developers.google.com
- openai.com
- openai.com
- openai.com
- openai.com
- host
- yoursite.com
- yoursite.com
Last reviewed — OpenAI crawler overview, publisher FAQ, Google common crawlers, Google-Extended token, robots.txt spec, Googlebot, and Anthropic crawler docs checked 2026-08-16.
AI Visibility
AI Visibility Cannabis visibility when the ad accounts are banned
Google and Meta will not take the usual spend. The models still answer dispensary, cultivator, and brand questions — if the site can be read and the cart can clear a 21+ order.
AI Visibility When ChatGPT names the franchise, not your shop
Run the best-HVAC-near-me prompt panel. If the model names a national franchise, fix corroboration and entity facts — not another blog calendar.
AI Visibility How do I get cited by Perplexity specifically
Allow PerplexityBot, put a liftable answer and unique numbers in HTML, then log numbered sources on a frozen prompt panel. There is no bought citation rate.
AI Visibility What belongs in an AI visibility monthly retainer vs a one-time audit
A one-time audit is the baseline plus prioritized fixes. A monthly retainer is prompt-panel tracking, entity hygiene, page jobs, and citation recovery.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.