Spurlock Studios
Contact
Share LinkedIn X
A lime beam hitting a small brass nameplate. Thesis: BLOCK TRAINING BOTS IF WANT.

You can block training crawlers if your policy says so — but do not treat every AI user-agent as “ChatGPT.” OpenAI’s GPTBot (training), OAI-SearchBot (ChatGPT search / citation index), and ChatGPT-User (user-initiated fetch) are different agents with different consequences. Blocking the wrong one removes you from being cited while you congratulate yourself for “opting out of AI.”

Google-Extended is a third category again: a robots.txt product token for Gemini training and grounding uses, not a search crawler and not a ranking lever. Google Search still rides on Googlebot.

This spoke is the robots.txt decision layer of the Answer Engine Optimization playbook. Pair it with llms.txt for brands so the file you publish matches the agents you actually allow.

The short answer

  • GPTBot ≠ ChatGPT search. Blocking GPTBot opts out of OpenAI training use per OpenAI’s crawler overview. It does not, by itself, opt you out of ChatGPT search surfacing.
  • OAI-SearchBot is the search crawler. Sites opted out “will not be shown in ChatGPT search answers,” though they may still appear as navigational links.
  • ChatGPT-User is a user-triggered fetch, not automatic crawl. OpenAI says robots.txt rules may not apply, and it is not the Search opt-out control.
  • Google-Extended is a product token for Gemini training and grounding uses — not the crawler that controls Google Search indexing. Google states it does not affect Search inclusion or ranking.
  • Default for most brands that want citations: allow search and retrieval bots; decide training bots as a policy call.

GPTBot vs search bots vs Google-Extended

Three different jobs. Three different files. One muddy “block AI” decision breaks the wrong one.

ControlVendorWhat it isWhat Disallow doesWhat it does not do
GPTBotOpenAIAutomated crawler for foundation-model trainingSignals content should not be used to train OpenAI generative AI foundation modelsDoes not remove you from ChatGPT search. Does not touch Google.
OAI-SearchBotOpenAIAutomated crawler for ChatGPT search featuresSite is not shown in ChatGPT search answers (may still appear as navigational links)Does not opt you out of training. Does not control Google.
ChatGPT-UserOpenAIUser-initiated fetch in ChatGPT / Custom GPTsLive user-directed fetches may fail; OpenAI notes robots.txt may not applyNot the Search inclusion control. Not a training crawler.
GooglebotGoogleCommon crawler for Google Search (and related Search surfaces)Blocks the Search crawl/index path for those URLsDoes not set Gemini training policy. That is Google-Extended.
Google-ExtendedGooglerobots.txt product token — no separate HTTP user-agentOpts crawled content out of listed Gemini training and grounding usesDoes not control Search inclusion, ranking, or AI Overviews eligibility

OpenAI states these settings are independent. You can allow OAI-SearchBot while disallowing GPTBot. Google states Google-Extended is independent of Search. Treat “AI bot” as three questions, not one switch.

Decision list:

  1. Do we want ChatGPT search citations? → OAI-SearchBot
  2. Do we allow OpenAI foundation-model training? → GPTBot
  3. Do we want Google Search crawl/index? → Googlebot (and do not confuse it with the next line)
  4. Do we allow Gemini training / listed grounding uses? → Google-Extended

Should I block GPTBot in robots.txt?

Only if you intend to opt out of OpenAI foundation-model training collection — not because you “don’t want to show up in ChatGPT.”

Per OpenAI’s crawler overview (developers.openai.com/api/docs/bots, checked 2026-08-16):

User agentJobIf you Disallow
GPTBotCrawl content that may be used to train OpenAI generative AI foundation modelsSignals content should not be used in that training
OAI-SearchBotSurface sites in ChatGPT search featuresNot shown in ChatGPT search answers (may still appear as navigational links)
ChatGPT-UserUser actions in ChatGPT / Custom GPTs fetch a pageLive user-directed fetches may fail; robots.txt may not apply; not the Search control
OAI-AdsBotValidate landing pages submitted as ChatGPT adsOnly visits ad landing pages; data is not used for foundation-model training

OpenAI also notes that if a site allows both GPTBot and OAI-SearchBot, they may use results from a single crawl for both use cases to avoid duplicate crawling. That is an efficiency note, not a reason to treat the tokens as one bot.

Publisher-side confirmation lives in OpenAI’s Publishers and Developers FAQ: for content to be included in summaries and snippets in ChatGPT, do not block OAI-SearchBot. Referral traffic from ChatGPT search can carry utm_source=chatgpt.com.

  • Legal/policy has a written training stance (allow or disallow GPTBot)
  • Marketing has a written citation stance (allow or disallow OAI-SearchBot)
  • Those two stances are not collapsed into one “AI” checkbox
  • OAI-AdsBot is left alone unless you actually run ChatGPT ads and need to restrict landing-page checks

What does each OpenAI agent look like in logs?

Do not match on a vibe. Match on the token and then on IP.

OpenAI publishes example user-agent strings and JSON IP lists on the same bots page. Version numbers change. Match the product token, not a pinned GPTBot/1.4 string.

AgentExample token in the UAPublished IPsrobots.txt behavior
GPTBotcompatible; GPTBot/…; +https://openai.com/gptbotopenai.com/gptbot.jsonHonors User-agent: GPTBot
OAI-SearchBotcompatible; OAI-SearchBot/…; +https://openai.com/searchbotopenai.com/searchbot.jsonHonors User-agent: OAI-SearchBot
ChatGPT-Usercompatible; ChatGPT-User/…; +https://openai.com/botopenai.com/chatgpt-user.jsonUser-initiated; robots.txt may not apply
OAI-AdsBotcompatible; OAI-AdsBot/…; +https://openai.com/adsbotopenai.com/adsbot.jsonOnly ad landing pages; not training

When OpenAI fetches robots.txt itself, the UA may include an extra robots.txt marker so you can tell a rules fetch from a content fetch in logs that omit paths. That marker is documentation hygiene, not a fifth bot.

Log lineTreat asNext action
UA has GPTBot and IP is in gptbot.jsonReal training crawlHonor your GPTBot Allow/Disallow
UA has OAI-SearchBot and IP is in searchbot.jsonReal search crawlThis is your ChatGPT citation path
UA has ChatGPT-User and IP is in chatgpt-user.jsonUser-triggered fetchDo not use this as a Search opt-out
UA claims any of the above, IP is not in the matching JSONImpostorRate-limit or block the IP; do not rewrite policy

What does Google-Extended actually control?

Google-Extended is not a crawler you will see as its own HTTP user-agent. Google’s common crawlers list is explicit: crawling is done with existing Google user-agent strings; the robots.txt token is used in a control capacity.

What the token manages, in Google’s words: whether content Google crawls from your site may be used for training future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and for grounding (providing content from the Google Search index to the model at prompt time) in Gemini Apps and Grounding with Google Search on Vertex AI.

What Google also says, in the same entry: Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.

TokenSeparate HTTP UA?Controls Search ranking?Controls ChatGPT search?Typical policy use
GooglebotYes (Googlebot/2.1)Yes — crawl/index path for Search surfacesNoSearch visibility
Google-ExtendedNoNo (per Google)NoGemini training / listed grounding opt-out
GoogleOtherYesNo specific productNoGeneric Google R&D fetches; not Search
GPTBotYesNoNo (training)OpenAI training opt-out
OAI-SearchBotYesNoYesChatGPT search eligibility

Google’s crawling overview splits agents into common crawlers (always respect robots.txt on automatic crawls), special-case crawlers, and user-triggered fetchers (ignore robots.txt because a person asked). Google-Extended sits on the common-crawlers page as a token, not as a fourth fetch class.

  • You have a User-agent: Google-Extended block if legal wants a Gemini training/grounding stance
  • You did not Disallow Googlebot thinking you only touched “AI”
  • You did not expect Google-Extended to hide you from AI Overviews

Does Google-Extended turn off AI Overviews?

No. AI Overviews and AI Mode are Search features. They ride the Search index and snippet controls, not the Gemini training token.

Google’s robots meta tag spec states that nosnippet applies to all forms of search results — including AI Overviews and AI Mode — and prevents the content from being used as a direct input for those features. max-snippet likewise applies to those surfaces. Google-Extended is not in that list.

GoalControl that actually appliesControl that does not
Stay in Google SearchAllow Googlebot; do not noindex the URLGoogle-Extended Allow/Disallow
Limit or block snippets / AI Overview inputnosnippet or max-snippet (and related snippet rules)User-agent: Google-Extended Disallow: /
Opt out of listed Gemini training / groundingUser-agent: Google-Extended + Disallow: /Blocking GPTBot
Stay in ChatGPT search answersAllow OAI-SearchBotBlocking Google-Extended

If marketing wants “no AI Overviews” and legal wants “no Gemini training,” those are two tickets. Write both. Do not ship one token and assume it covers the other.

What about ClaudeBot and Claude-SearchBot?

Anthropic documents three robots on its crawler help article. Same split as OpenAI: training vs search vs user fetch. Anthropic states its bots honor robots.txt, including the user-initiated agent — that is a difference from OpenAI’s ChatGPT-User note.

BotJob (Anthropic’s wording)If you restrict it
ClaudeBotCollect web content that could contribute to model trainingSignals future materials should be excluded from training datasets
Claude-SearchBotNavigate the web to improve search result qualityPrevents indexing for search optimization; may reduce visibility and accuracy in user search results
Claude-UserFetch a site when a person asks Claude a questionPrevents retrieval in response to a user query

Anthropic also warns that IP-blocking their bots can prevent them from reading robots.txt at all, which breaks a clean opt-out. Prefer robots.txt signals over silent IP bans when you want the vendor to honor a Disallow.

GoalOpenAI controlAnthropic controlGoogle control
Training opt-outDisallow GPTBotDisallow ClaudeBotDisallow Google-Extended
Search / citation eligibilityAllow OAI-SearchBotAllow Claude-SearchBotAllow Googlebot (Search); snippet rules for Overviews
User-initiated fetchUnderstand ChatGPT-User limitsClaude-User (honors robots.txt per Anthropic)User-triggered fetchers ignore robots.txt

Confirm on the vendor page before you ship. Names and scopes change. This table is a worksheet, not a forever spec.

What robots.txt pattern should most brands ship?

For brands that want answer-engine citations and are willing to decide training separately:

# Citation / search retrieval — keep allowed
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

# Training — policy decision (example: opt out)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

# Gemini training/grounding token — policy decision
User-agent: Google-Extended
Disallow: /

# Search crawl stays on Googlebot via User-agent: * or an explicit Allow
# Do not blanket-ban everything unknown with User-agent: * Disallow: /
# unless you also Allow the bots you need.

Adjust the training lines to Allow: / if legal wants training inclusion. The important part is splitting the decisions.

Google’s robots.txt interpretation requires the file at the site root on a supported protocol. Rules without a path are ignored. User-agent groups are matched by token. A Disallow: / under User-agent: * will also hit any bot that falls through to * — including search crawlers you meant to keep.

PatternWhat happensUse when
Split groups: search Allow, training DisallowIndependent policy, as vendors documentDefault for citation-seeking brands
User-agent: * + Disallow: / onlyYou opted out of almost everything that honors *Rare; regulated or private sites
One group listing every AI name under *Easy to miss a token; easy to hit Googlebot by accidentDo not
Training Disallow, no search AllowFine if * still allows crawl — until someone tightens *Add explicit search Allows anyway

Decision list before you paste:

  1. ChatGPT search citations? → explicit OAI-SearchBot Allow
  2. Claude search-style indexing? → explicit Claude-SearchBot Allow
  3. Foundation-model training? → GPTBot / ClaudeBot / Google-Extended per vendor
  4. Did a WAF or CDN already block these at the edge? → fix that next
  5. Is /robots.txt itself fetchable anonymously? → if no, no bot can honor it

Which robots.txt syntax mistakes silently opt you out?

Most “we blocked ChatGPT by accident” incidents are not policy. They are a * group, a missing path, or a file that is not the file the crawler fetched.

Google’s robots.txt spec is the reference for how Google reads the file. OpenAI’s bots page is the reference for which tokens OpenAI honors. A valid file that names the wrong token is still a miss.

MistakeWhat the crawler seesFix
User-agent: * + Disallow: / and no explicit search AllowsSearch bots that fall through to * are opted outAdd User-agent: OAI-SearchBot / Googlebot Allows above the catch-all, or drop the catch-all
Disallow: with no pathRule is ignored (Google: rules without a path do not restrict)Write Disallow: / if you mean the whole site
User-agent: GPT-Bot or OAI_SearchBotToken does not match; group is skippedUse the exact tokens from the vendor page
File at /static/robots.txt or a CMS “SEO plugin” URLCrawlers fetch https://host/robots.txt onlyServe the file at the host root
HTML error page with status 200Crawler receives markup, not groupsReturn text/plain and 200 with real groups
robots.txt on www empty, apex has the rules (or the reverse)Each host has its own fileDuplicate the worksheet on every host you serve

Path-level example when legal wants training off the docs set but marketing still wants the About page cited:

User-agent: GPTBot
Disallow: /docs/
Allow: /about/

User-agent: OAI-SearchBot
Allow: /

That is still two decisions. Do not “simplify” it back into one * block after the first incident ticket.

  • curl -sI the exact URL https://<host>/robots.txt for apex and www
  • Diff the served bytes against the worksheet, not against last month’s gist
  • Confirm no plugin is injecting a second file that origin never sees

What about www, staging, and other hosts?

robots.txt is per host. A clean production file does not cover staging.example.com, www.example.com if that host answers separately, or a marketing subdomain that still ranks.

HostTypical failureWhat to ship
Apex example.comFile exists; www 301s after a bot already fetched an empty www fileSame groups on both hosts, or one host + a consistent redirect before the fetch
wwwCDN serves a default denyExplicit search Allows on the host users and crawlers actually hit
Staging / previewUser-agent: * Disallow: / copied to production in a mergeStaging deny is fine; production must not inherit it
Docs subdomainTraining Disallow on apex onlyRepeat GPTBot / Google-Extended on the host that holds the corpus

Googlebot is documented on the Googlebot page as the Search crawler for that host’s URLs. OpenAI’s search crawler is the same idea: it fetches the host you allowed. Do not assume a Disallow on the marketing site covers the help center.

  • Inventory every public host that serves indexable HTML
  • Each host has a reachable /robots.txt
  • Staging deny is not the production file
  • Subdomains that hold the real answers have the same search Allows as the homepage

Failure mode: the “block all AI” CDN default

What breaks: a security dashboard enables “block AI crawlers” globally. robots.txt says Allow for OAI-SearchBot, but the edge returns 403 first. ChatGPT search never indexes you. You spend a quarter “doing AEO” on content that cannot be fetched.

What it costs: zero citations despite answer-first pages.

What you do instead:

  • Fetch https://yoursite.com/robots.txt anonymously — expect 200, not a challenge page
  • Confirm separate User-agent blocks exist (not one muddy *)
  • Check CDN / WAF bot scores against OpenAI’s published ranges (gptbot.json, searchbot.json, chatgpt-user.json)
  • curl -A "OAI-SearchBot" (and vendor equivalents) on About, offer, and top answer pages — expect 200
  • Re-test after every security policy change
  • Confirm the WAF exception is documented, not a one-off ticket that dies in Slack

Google’s common crawlers always obey robots.txt on automatic crawls, per the common crawlers page. A 403 in front of robots.txt is still a 403. The file cannot save a fetch that never reaches origin.

This is the same fetchability check that belongs in an AEO audit: robots, then edge, then the page.

Will blocking GPTBot hurt Google rankings?

No — not as a Google ranking lever. GPTBot is OpenAI’s training crawler. Google Search crawl and index is Googlebot. These are different systems, documented on different vendor pages.

What is Googlebot is the Search crawler reference. Blocking GPTBot does not tell Google to demote you. Allowing GPTBot does not boost Google rank. Blocking Google-Extended does not demote you either — Google says the token is not a ranking signal.

ActionGoogle Search effect (per Google / OpenAI docs)ChatGPT search effect (per OpenAI)
Disallow GPTBotNoneNone (training only)
Disallow OAI-SearchBotNoneNot shown in ChatGPT search answers
Disallow Google-ExtendedNone (not a ranking signal)None
Disallow GooglebotYou blocked Search crawlNone
CDN 403 on “AI bots”Depends whether Googlebot is in that ruleOften 403s OAI-SearchBot too

Do not conflate “AI” into one SEO myth. I have been SEO certified since 2021. The AEO version of that work still starts with “which agent, which product.”

How fast do robots.txt changes take effect?

OpenAI documents that for search results, it can take about 24 hours from a site’s robots.txt update for their systems to adjust (bots overview). Other vendors differ. Assume hours to a few days for automated crawlers, then re-verify with log lines and live answer tests.

ChangeExpect
Allow OAI-SearchBot after accidental block~24h for OpenAI search systems to adjust (per OpenAI), then longer for re-crawl of key URLs
Disallow GPTBotFuture training collection should respect the signal; already-trained model memory is a separate, slower clock
Disallow Google-ExtendedFuture Gemini training / listed grounding uses; not a Search recrawl event
CDN unblockImmediate for new fetches; old index residue clears on re-crawl
nosnippet / noindexRequires the page to remain crawlable so the tag can be read (robots meta spec)

Training residue in model weights is not cleared by robots.txt. robots.txt governs future crawl and use signals — not a memory erase.

OpenAI’s publisher FAQ adds a second clock: if they obtain a disallowed URL from a third-party search provider or by crawling other pages, they may still surface just the link and page title. If you do not want that, use noindex — and the crawler must be allowed to read the tag.

How to verify bots can fetch key pages

  1. Confirm robots.txt allows the search agent on the path.
  2. Confirm CDN/WAF allows the vendor’s published IPs / verified bot.
  3. Confirm the page returns 200 without login.
  4. Confirm the page is not noindex if you also care about Google Search / AI Overviews.
  5. Re-run five ChatGPT search prompts that should cite you; log whether your URL returns.
  • robots.txt split training vs search vs Google-Extended
  • WAF exceptions documented
  • About + offer + top 5 answer URLs fetch clean with OAI-SearchBot UA
  • Same URLs fetch clean with a normal browser UA (no accidental auth wall)
  • Prompt panel archived post-change
  • /llms.txt is consistent with what you allow crawlers to read — see llms.txt for brands

Google verification is a different procedure from OpenAI’s JSON lists. Google’s verify crawler requests page gives two methods: reverse-then-forward DNS (googlebot.com / google.com / googleusercontent.com), or match the source IP to the published common-crawler / special-crawler / user-triggered-fetcher ranges.

CheckCommand / sourcePass
robots.txt reachablecurl -sI https://yoursite.com/robots.txt200, text/plain
Search bot fetchcurl -A "OAI-SearchBot" -sI https://yoursite.com/about/200, not 403/401/challenge
OpenAI IPCompare request IP to searchbot.jsonAddress in published ranges
Googlebot IPReverse + forward DNS per Google’s verify docHostname and IP match

Impostor bots and log hygiene

User-agent strings are trivial to spoof. Before you panic about “GPTBot ignoring robots.txt,” match the request IP to the vendor’s published ranges. Impostors wearing the UA are common.

CheckAction
UA says GPTBot, IP not in gptbot.jsonTreat as impostor; do not rewrite policy on fakes
UA + IP match, hits a Disallow pathFile a vendor report if persistent; verify your robots.txt is reachable
robots.txt itself blocked by WAFFix that first — crawlers that cannot read rules cannot honor them
UA says Googlebot, reverse DNS is not googlebot.comImpostor; Google documents this spoof pattern on the Googlebot page

Google common crawlers typically reverse-resolve to crawl-***-***-***-***.googlebot.com or geo-crawl-***-***-***-***.geo.googlebot.com. Special-case crawlers use a different mask (rate-limited-proxy-…google.com) and may or may not honor robots.txt. Do not apply a Googlebot mental model to every Google UA.

Anthropic has warned that IP-blocking their bots can prevent them from reading robots.txt at all. Prefer robots.txt signals over silent IP bans when you want a clean opt-out.

  • Log pipeline stores UA and source IP
  • Weekly job diffs claimed-OpenAI IPs against the three JSON files
  • Googlebot claims go through reverse+forward DNS before a block
  • Security and marketing share one incident channel so a fake UA does not become a policy rewrite

Fill this once. Store it next to the deploy checklist. Re-open it when a vendor changes their crawler page — not when a Slack thread says “AI is stealing our site.”

DecisionOptionsOwnerNotes
OpenAI training (GPTBot)Allow / DisallowLegalFuture collection only
ChatGPT search (OAI-SearchBot)Allow / DisallowMarketingCitation eligibility
ChatGPT user fetch (ChatGPT-User)Allow / monitor / edge-restrictEngrobots.txt may not apply
Anthropic training (ClaudeBot)Allow / DisallowLegalConfirm current help article
Anthropic search (Claude-SearchBot)Allow / DisallowMarketingVisibility in Claude search-style results
Gemini training/grounding (Google-Extended)Allow / DisallowLegalNot Search ranking
Google Search (Googlebot)Allow (default)Marketing + SEODo not mix with Extended
Review cadenceQuarterly or on vendor-doc changeNamed pairDate the last check
  • Worksheet signed off
  • robots.txt matches worksheet
  • CDN rules match worksheet
  • Change log entry with date
  • /llms.txt does not promise access the edge then denies

If marketing wants citations and legal wants training opt-out, that is a normal, supported split — not a conflict that requires blocking everything. OpenAI’s bots page exists to make that split possible. Google’s Google-Extended entry exists for the same reason on the Gemini side.

What robots.txt cannot do

Use this as the “stop asking the file to do that” list.

Askrobots.txt answerActual control
Erase us from a model already trainedNoWeights are not a crawl cache
Remove us from Google AI OverviewsNo (Google-Extended will not)Snippet rules / nosnippet / noindex
Stop a person from pasting our URL into ChatGPTNot reliablyChatGPT-User may ignore robots.txt
Prove a UA is the real vendorNoPublished IP lists + DNS
Survive a WAF 403NoEdge allowlist
Replace a briefing fileNollms.txt is a different artifact

Google’s robots.txt spec also reminds you that user-triggered fetchers and safety crawlers are outside the Robots Exclusion Protocol. A Disallow is a signal to automated crawlers that honor it. It is not a legal lock, a DRM layer, or a ranking boost.

Lane overview for the rest of the visibility stack: /visibility.

FAQ

What is OAI-SearchBot vs GPTBot?

GPTBot crawls for OpenAI foundation-model training. OAI-SearchBot crawls to surface sites in ChatGPT search features. OpenAI treats them as independent robots.txt controls on the official bots page. Blocking training does not equal blocking search — and blocking search is how you disappear from ChatGPT search answers.

What about ClaudeBot and Google-Extended?

ClaudeBot is Anthropic’s training-oriented crawler; use Claude-SearchBot when the question is Claude search indexing. Google-Extended is Google’s product token for certain Gemini training and grounding uses — it does not control Google Search indexing or ranking. Do not block Googlebot thinking you only touched “AI.”

Do WAFs silently block AI bots?

Yes, often. CDN “AI scraper” defaults and bot-fight scores can 403 OAI-SearchBot while your robots.txt looks fine. Verify with user-agent curls and vendor IP lists after every security change. A 200 on the file and a 403 on the page is still a failed citation path.

Will blocking GPTBot hurt Google rankings?

No. Google rankings depend on Google’s crawlers and ranking systems, not on whether OpenAI’s training bot can fetch you. Keep Googlebot decisions separate from GPTBot decisions, and keep Google-Extended separate from both.

How fast do robots.txt changes take effect?

OpenAI notes roughly 24 hours for search systems to adjust after a robots.txt update. Plan for re-crawl lag beyond that. Training opt-outs affect future collection; they do not rewrite model memory overnight.

How do I verify bots can fetch key pages?

Allow the right agents in robots.txt, allow them at the CDN/WAF, confirm HTTP 200 on canonical answer URLs with the bot user-agent, match IPs to OpenAI and Google published lists, then re-test live ChatGPT search prompts and log citations.

CTA

Split training policy from citation eligibility — then prove the fetch with a 200, not a vibes-based robots.txt screenshot.

Lane overview: /visibility. Next step: a visibility audit.

FAQ

What questions does this article answer?

What is OAI-SearchBot vs GPTBot?
`GPTBot` crawls for OpenAI foundation-model training. `OAI-SearchBot` crawls to surface sites in ChatGPT search features. OpenAI treats them as independent robots.txt controls on the [official bots page](https://developers.openai.com/api/docs/bots). Blocking training does not equal blocking search — and blocking search is how you disappear from ChatGPT search answers.
What about ClaudeBot and Google-Extended?
`ClaudeBot` is Anthropic’s training-oriented crawler; use `Claude-SearchBot` when the question is Claude search indexing. `Google-Extended` is Google’s product token for certain Gemini training and grounding uses — it does not control Google Search indexing or ranking. Do not block `Googlebot` thinking you only touched “AI.”
Do WAFs silently block AI bots?
Yes, often. CDN “AI scraper” defaults and bot-fight scores can 403 `OAI-SearchBot` while your robots.txt looks fine. Verify with user-agent curls and vendor IP lists after every security change. A 200 on the file and a 403 on the page is still a failed citation path.
Will blocking GPTBot hurt Google rankings?
No. Google rankings depend on Google’s crawlers and ranking systems, not on whether OpenAI’s training bot can fetch you. Keep `Googlebot` decisions separate from `GPTBot` decisions, and keep `Google-Extended` separate from both.
How fast do robots.txt changes take effect?
OpenAI notes roughly 24 hours for search systems to adjust after a robots.txt update. Plan for re-crawl lag beyond that. Training opt-outs affect future collection; they do not rewrite model memory overnight.
How do I verify bots can fetch key pages?
Allow the right agents in robots.txt, allow them at the CDN/WAF, confirm HTTP 200 on canonical answer URLs with the bot user-agent, match IPs to [OpenAI](https://openai.com/searchbot.json) and [Google](https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests) published lists, then re-test live ChatGPT search prompts and log citations.
Sources

Last reviewed — OpenAI crawler overview, publisher FAQ, Google common crawlers, Google-Extended token, robots.txt spec, Googlebot, and Anthropic crawler docs checked 2026-08-16.

More from this lane

AI Visibility

All →
Book the audit