Spurlock Studios
Contact
Share LinkedIn X
A lime beam hitting a small brass nameplate. Thesis: WHICH AI BOTS SHOULD ALLOW.

Allow citation-useful crawlers if you want answer engines to recommend you. Block scrapers and training crawlers you have no product reason to feed. The seven tokens operators actually argue about are GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Amazonbot, Bytespider, and CCBot — and they do not all do the same job. Googlebot is not Google-Extended. Collapsing them into one “block AI” switch is how brands opt out of ChatGPT search, Perplexity citations, or Google Search while congratulating themselves for a privacy win.

This is the allow-vs-block decision table inside the Answer Engine Optimization playbook. How you write the file, which host it lives on, and how you prove a 200 lives in AI crawlers and robots.txt decisions. Do not treat this spoke as a robots.txt tutorial.

The short answer

  • Want to be recommended? Allow the crawlers that feed citation and retrieval. Block the ones that only scrape or train, unless you have a written reason to donate the corpus.
  • GPTBot and ClaudeBot are training. ChatGPT search and Claude search-style retrieval use different tokens (OAI-SearchBot, Claude-SearchBot). Blocking training does not keep you eligible for citations.
  • PerplexityBot is a citation crawler. Perplexity says it is not used to crawl content for foundation models. Default allow if Perplexity is a surface you care about.
  • Googlebot crawls for Google Search. Google-Extended is a robots.txt product token for listed Gemini training and grounding uses. Blocking one is not blocking the other.
  • Bytespider and CCBot are the usual blocks when you do not need ByteDance products or a public Common Crawl dump. Amazonbot is a policy call unless Alexa / Rufus actually matters.

The allow vs block table

These are the seven bots this post decides. Adjacent twins appear only so you do not block the wrong job.

TokenOperatorJob (vendor wording, checked 2026-09-05)Default if you want citationsDefault if you do not need that product
GPTBotOpenAICrawl that may be used to train generative AI foundation modelsPolicy call (often Disallow)Disallow
ClaudeBotAnthropicCollect web content that could contribute to model trainingPolicy call (often Disallow)Disallow
PerplexityBotPerplexitySurface and link sites in Perplexity search results; not foundation-model crawlAllowDisallow
Google-ExtendedGoogleProduct token: listed Gemini training and grounding usesPolicy callDisallow the token; never Disallow Googlebot as a substitute
AmazonbotAmazonImprove products and services; may be used to train Amazon AI modelsAllow only if Amazon surfaces matterDisallow
BytespiderByteDanceSearch / recommendation / training crawl for ByteDance propertiesDisallow unless you need those productsDisallow
CCBotCommon CrawlPublic web crawl into an open dataset anyone can train on laterDisallow unless you want the dumpDisallow

Citation twins you must not confuse with the rows above:

If you meantDo not useUse instead
ChatGPT search answersGPTBotOAI-SearchBot (OpenAI: opted-out sites are not shown in ChatGPT search answers)
Claude search-style indexingClaudeBotClaude-SearchBot
Google Search crawl / indexGoogle-ExtendedGooglebot
Alexa / Rufus search eligibilityAmazonbotAmzn-SearchBot (Amazon: does not crawl for generative AI training)

Decision list before anyone pastes a CDN preset:

  1. Which answer engines should still be able to cite us?
  2. Which vendors may train on our public pages?
  3. Which scrapers have no product we sell into?
  4. Did a “block AI bots” toggle already answer those three as one?

If legal and marketing cannot sign different answers for (1) and (2), you do not have a policy. You have a mood.

Use the table in a 20-minute meeting, not a Slack poll:

  1. Read each of the seven tokens out loud with the job column. No nicknames.
  2. Marketing marks citation Allow/Deny. Legal marks training Allow/Deny. Security marks scrape Allow/Deny.
  3. Where two owners disagree on one token, that token stays Deny until someone writes a reason to Allow.
  4. Search twins (OAI-SearchBot, Claude-SearchBot, Googlebot) get their own row on the same page even though they are not in the seven. Missing twins are how “we blocked GPTBot” becomes “ChatGPT search never heard of us.”
  5. Date the sheet. The next vendor changelog invalidates folklore, not the jobs.

Citation crawlers, training crawlers, and scrapers

An “AI bot” is not a species. It is a job. Treat three jobs, then assign each token.

JobWhat success looks likeWhat a block actually doesExamples in this table
Citation / retrievalThe engine can fetch or index you and name you with a linkYou become harder to recommend on that surfacePerplexityBot; also OAI-SearchBot, Claude-SearchBot, Googlebot
Training / corpusFuture model weights or listed grounding uses may include youFuture collection should stop; old weights do not rewindGPTBot, ClaudeBot, Google-Extended, often Amazonbot
Public dump / third-party scrapeSomeone else publishes or trains on a copy of the open webYou opt out of that crawl going forwardCCBot; Bytespider when you have no ByteDance surface

User-initiated fetchers (ChatGPT-User, Claude-User, Perplexity-User, Amzn-User) are a fourth bucket. They fire because a person asked. Several vendors say robots.txt may not apply, or generally does not. Do not use those tokens as your citation on/off switch.

BucketHonor robots.txt? (vendor claim)Use it as your allow/block KPI?
Automated search / index crawlerUsually yesYes
Automated training crawlerUsually yesYes, for training policy
Product token with no separate HTTP UAApplied downstream (Google-Extended)Yes, for Gemini listed uses — not for Search
User-triggered fetcherMixed; often noNo

I have been SEO certified since 2021. The AEO version of that work is still “which agent, which product,” not “AI: on or off.”

  • Citation surfaces are named (ChatGPT search, Perplexity, Claude search, Google Search / AI Overviews, Alexa)
  • Training stance is written separately from citation stance
  • Scrapers without a named product are default-deny
  • User-triggered fetchers are not the scoreboard

Should I allow GPTBot?

Allow GPTBot only if you intend OpenAI to use crawled pages in foundation-model training. Do not allow it because you “want to show up in ChatGPT.” That is a different bot.

OpenAI’s crawler overview is explicit: settings are independent. You can allow OAI-SearchBot so ChatGPT search can surface you, and disallow GPTBot so the same pages are not a training signal. If both are allowed, OpenAI may reuse one crawl for both jobs. That is an efficiency note, not a reason to treat the tokens as one.

IntentGPTBotOAI-SearchBotWhat operators get wrong
Citations, no training donationDisallowAllowDisallow both “to be safe”
Citations and trainingAllowAllowFine, if legal signed it
No ChatGPT search, training still OKAllowDisallowRare and usually accidental
Full opt-out of OpenAI automated crawlDisallowDisallowStill not a Google ranking lever

ChatGPT-User is not this decision. OpenAI says it is not used to determine Search appearance, and robots.txt rules may not apply because a user initiated the fetch.

SignalTreat as
Legal wants no OpenAI trainingDisallow GPTBot
Marketing wants ChatGPT search citationsAllow OAI-SearchBot
Someone blocked “GPTBot” in a WAF named “all OpenAI”Audit whether OAI-SearchBot died too
Leadership asks “did we opt out of ChatGPT?”Answer with the search token, not the training token

Version suffixes move. OpenAI’s examples use strings like GPTBot/1.4 and note the version number may change. Match the product token GPTBot, not a pinned UA you copied in 2025.

OpenAI’s publisher FAQ is the citation receipt: for content to be included in summaries and snippets in ChatGPT, do not block OAI-SearchBot. Referral traffic from ChatGPT search can carry utm_source=chatgpt.com. That UTM is evidence the search allow is working. It is not evidence that GPTBot was a good training donation.

You see in analyticsIt meansIt does not mean
utm_source=chatgpt.comChatGPT search (or a ChatGPT-attributed click) found a URLGPTBot is required
No ChatGPT UTM, search prompts still cite youCitations can happen without a clickYour bot table is irrelevant
No citations and no UTM after a “block AI” changeSearch path is probably closedYou need more blog posts first

Default for most citation-seeking brands: Disallow GPTBot. Allow OAI-SearchBot. Write that as two lines of policy, then implement it in the robots.txt spoke.

Should I allow ClaudeBot?

Allow ClaudeBot only as a training-donation decision. Anthropic’s crawler help article splits three robots. Restricting ClaudeBot “signals that the site’s future materials should be excluded from our AI model training datasets.” That sentence is not a citation guarantee and not a citation kill switch.

BotAnthropic’s jobIf you restrict it
ClaudeBotTraining collectionFuture materials should be excluded from training datasets
Claude-SearchBotImprove search result qualityMay reduce visibility and accuracy in user search results
Claude-UserFetch when a person asks ClaudePrevents retrieval in response to that user query

Anthropic states all three honor robots.txt — including the user-initiated agent. That is a documented difference from OpenAI’s ChatGPT-User note and from Perplexity’s Perplexity-User note. Confirm on the vendor page before you brief a lawyer; names and scopes change.

GoalClaude controlDo not substitute
Stay findable in Claude search-style resultsAllow Claude-SearchBotAllowing ClaudeBot instead
Opt out of Anthropic trainingDisallow ClaudeBotDisallowing Claude-SearchBot “to be thorough”
Block live user fetchesDisallow Claude-User (and expect support tickets)Assuming a training Disallow covers it

Anthropic also warns that IP-blocking their bots can prevent them from reading robots.txt, which breaks a clean opt-out. Prefer a token-level Disallow when the goal is a recorded training stance.

Default for most citation-seeking brands: Disallow ClaudeBot. Allow Claude-SearchBot. Same shape as OpenAI. Same refusal to mash them together.

Should I allow PerplexityBot?

Yes — if Perplexity is a surface where you want numbered citations. No — if you have a written decision that Perplexity should not index you.

Perplexity’s crawlers doc says PerplexityBot is designed to surface and link websites in search results on Perplexity, and is not used to crawl content for AI foundation models. That is the opposite of GPTBot. Treating PerplexityBot as “another training scraper” is how you disappear from a retrieval engine while your training policy was already handled elsewhere.

TokenJobFoundation-model crawl?robots.txt
PerplexityBotAutomated search / citation indexPerplexity says noHonor it; this is the allow/block for Perplexity results
Perplexity-UserUser-triggered fetch to answer a questionPerplexity says noGenerally ignores robots.txt because a user requested the fetch
You wantPerplexityBotNotes
Cited in Perplexity answersAllowAlso allow the published IP ranges at the edge, or a WAF will 403 a “yes”
Not in Perplexity’s indexDisallowPerplexity-User can still hit a pasted URL
Training opt-out onlyDo not Disallow this token for that reasonUse GPTBot / ClaudeBot / Google-Extended for training

Perplexity notes settings are independent and may take up to 24 hours to reflect. That is a retrieval clock, not a ranking myth.

  • Perplexity is on the named citation list, or it is explicitly off
  • PerplexityBot is not lumped into a “training scrapers” WAF group
  • Perplexity-User is not your opt-out lever
  • Key answer URLs return 200 to the real bot, not only to Chrome

Default for citation-seeking brands: Allow PerplexityBot.

Why Googlebot is not Google-Extended

Googlebot is the Search crawler. Blocking it is how you leave Google Search — including Discover and other Search features Google lists on the common crawlers page.

Google-Extended is not a crawler you will see as its own HTTP user-agent. Google says crawling is done with existing Google user-agent strings; the robots.txt token is used in a control capacity. What it manages: whether crawled content may be used for training future Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI.

Google also says, in the same entry: Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.

ControlSeparate HTTP UA?Search inclusion / rankingAI Overviews / AI ModeGemini Apps training / listed grounding
GooglebotYesYes — this is the Search crawl pathIndirect: Overviews need an indexed, snippet-eligible pageNot this token
Google-ExtendedNoNo (per Google)No — Google points you to snippet / noindex controlsYes — listed uses
nosnippet / max-snippet / noindexPage-levelnoindex drops SearchSnippet rules apply to AI Overviews and AI Mode per Google’s robots meta specNot a substitute for the Extended token

Google’s AI features page is the other half: robots.txt directives for Googlebot manage crawl for Search. To limit what Search shows, use snippet and index controls. To limit AI training and grounding in some of Google’s other systems, read Google-Extended.

Third-party posts still claim Google-Extended turns off AI Overviews. That is not what Google’s own pages say. If marketing wants “no Overviews” and legal wants “no Gemini training,” those are two tickets.

GoalAllow / keepDo not
Stay in Google SearchGooglebotDisallow Googlebot as an “AI” move
Stay eligible for AI OverviewsIndexed + snippet-eligible (Googlebot path)Expect Google-Extended Disallow to hide the Overview
Opt out of listed Gemini training / groundingDisallow Google-ExtendedAssume that also demotes Search
Ground in Gemini Apps on listed termsAllow Google-ExtendedConfuse that with ChatGPT search

GoogleOther is a third Google name that shows up in logs. Google describes it as a generic crawler for one-off research fetches, not a specific consumer product and not Search. Do not Disallow Googlebot because you saw GoogleOther. Do not assume GoogleOther is Google-Extended. Extended still has no separate HTTP UA.

Log / tokenDecide asCommon false move
GooglebotSearch crawl“That’s the AI bot”
Google-ExtendedGemini training / listed grounding“That’s how we leave AI Overviews”
GoogleOtherUnspecified Google R&D fetchCollapse it into a Googlebot Disallow
Google-InspectionToolSearch Console / Rich Results testsBlocking it “for AI” breaks your own inspections

Default: Never block Googlebot to make an AI statement. Decide Google-Extended as a Gemini training/grounding policy. Most brands that want Search and Overviews leave Googlebot allowed and pick Extended in writing.

Should I allow Amazonbot?

Allow Amazonbot when Amazon product improvement — and possible Amazon model training — is a trade you accept. Block it when Alexa, Rufus, and Amazon’s other surfaces are not a channel you care about.

Amazon’s Amazonbot page now splits three agents. Each setting is independent. Changes may take about 24 hours.

TokenAmazon’s jobTraining?Default for a brand that does not sell the Amazon surface
AmazonbotImprove products and services; may be used to train Amazon AI modelsMaybe (Amazon’s word is “may”)Disallow
Amzn-SearchBotSearch experiences such as Alexa; eligible to appear if permittedAmazon says it does not crawl for generative AI model trainingAllow only if Alexa / Rufus citations matter
Amzn-UserLive fetch for user questions (e.g. Alexa)Amazon says no generative-AI training crawlNot your training switch; may not follow all robots.txt directives

Amazon also notes: if robots.txt does not mention Amzn-SearchBot but allows other search bots, Amzn-SearchBot will crawl in accordance with those other search-bot directives. Silence is not a clean “Amazon off.”

SituationAmazonbotAmzn-SearchBot
DTC brand, no Alexa strategyDisallowDisallow (name it; do not rely on silence)
Local service that wins voice queriesPolicy callAllow if Alexa is a real channel
Marketplace seller who needs RufusPolicy callAllow
Rights-sensitive publisherDisallowDisallow

Amazon documents noarchive as “do not use the page for model training,” plus noindex and none. That is a page-level training hint, not a Google ranking control, and not a ChatGPT control.

Default for most non-Amazon businesses: Disallow Amazonbot. Name Amzn-SearchBot if you also want it off. Do not assume Amazonbot is how ChatGPT or Perplexity find you.

Should I block Bytespider?

Yes for most brands that are not trying to be found inside ByteDance products (TikTok-adjacent search, Doubao, Toutiao). Bytespider is ByteDance’s crawler. It is not a ChatGPT citation path and not Google Search.

ByteDance webmaster material names the token Bytespider. English-language official docs are thinner than OpenAI’s or Google’s. Independent operators report mixed robots.txt honor over the years. Treat a Disallow as a request you should also enforce at the edge if the crawl volume hurts. Do not treat a robots.txt line as a cryptographic guarantee.

Claim you will hearWhat to do instead
“Bytespider is how we get TikTok SEO in the US”Separate TikTok ads / organic from this crawler; do not allow a heavy scrape on a rumor
“Blocking it hides us from ChatGPT”Unrelated tokens
“robots.txt is enough”Log it. If Disallow is ignored, WAF / ASN rules — without blocking Googlebot or OAI-SearchBot in the same group
“The UA is proof”User-agents are cheap to spoof. Confirm operator docs when they exist; Bytespider’s public IP story is weaker than OpenAI’s JSON lists
PolicyBytespiderWhy
Default citation-seeking US/EU brandDisallowNo citation product you are measuring
China-market brand that needs ByteDance searchAllow, then watch bandwidthProduct reason exists
Newsroom under scrape loadDisallow + edgeVolume, not AEO

Hedge: token spelling, UA wrappers, and compliance can change without a nice changelog in English. Re-check logs quarterly. If a CDN “AI bot” list includes Bytespider and OAI-SearchBot, you did not ship a Bytespider decision. You shipped a bundle.

Default: Block Bytespider unless you can name the ByteDance product you are serving.

Should I block CCBot?

Usually yes, if you do not want future snapshots of your site in Common Crawl’s public datasets. CCBot is the crawler for Common Crawl, a nonprofit that publishes open web crawls. Labs, startups, and researchers train and evaluate models on those dumps. Blocking CCBot is opting out of that pipeline going forward. It does not retract crawls already published.

Common Crawl identifies the agent as CCBot (examples historically look like CCBot/2.0). They document robots.txt Disallow and, unlike several commercial vendors, they document Crawl-delay. Version numbers increment. Match CCBot, not a frozen CCBot/1.0 string from a 2018 gist.

IntentCCBotWhat a block does not do
Stop future inclusion in Common Crawl dumpsDisallowErase you from dumps already on disk
Support open research / do not careAllowGet you ChatGPT citations
Reduce third-party training on public mirrorsDisallow, knowing other crawlers still existStop GPTBot or ClaudeBot — those are separate
Cut bandwidth from a polite research crawlDisallow or Crawl-delayStop impersonators using the UA

Common Crawl warns that other crawlers falsely identify as CCBot. They publish IP ranges (including JSON at index.commoncrawl.org/ccbot.json per their bot page). A UA match without the operator’s ranges is not a reason to rewrite policy.

Brand typeDefaultRationale
Most commercial sites measuring ChatGPT / Perplexity / OverviewsDisallowDump is not your citation path
Research org, university, public-data advocateAllowMission fit
Publisher with a rights deskDisallowCorpus donation is a legal question

Default: Block CCBot unless openness is an explicit value, not a leftover Allow.

What should most brands default to?

If you want to be recommended, allow the crawlers that feed recommendations. Block scrapers you do not need. Decide training in writing. That is the whole operating system. The seven-token worksheet for a typical US B2B or local brand:

TokenDefaultWhy
PerplexityBotAllowCitation crawler, not (per Perplexity) foundation-model crawl
GPTBotDisallowTraining; keep OAI-SearchBot allowed
ClaudeBotDisallowTraining; keep Claude-SearchBot allowed
Google-ExtendedWritten policyGemini training / listed grounding — not Search
GooglebotAllowSearch + Overview eligibility path
AmazonbotDisallowNo Amazon surface, possible training
BytespiderDisallowNo ByteDance product to serve
CCBotDisallowPublic dump, not your AEO scoreboard
Company typeChange from the default
Rights-sensitive publisherDisallow training tokens and scrapers; keep search twins if you still want citations
Amazon-heavy catalogConsider Amzn-SearchBot Allow; keep Amazonbot as the training call
China / ByteDance marketBytespider may flip to Allow; watch volume
“We do not want AI at all”You still need Googlebot if you want Google Search. AI Overviews are Search features. Full invisibility is a Search decision, not an Extended token.
Regulated / logged-in appPath-level Denies on /app, /account, /staging for all of the above; do not hide the public About page as collateral

Pages to keep fetchable first — the ones models actually quote:

  1. Canonical About / identity (legal name, what you do, where, since when)
  2. Offer / pricing / service-area pages a buyer would ask an engine about
  3. One comparison or criteria page with a table
  4. FAQ answers that stand alone
URL classCitation crawlersTraining tokensScrapers (Bytespider, CCBot)
/about, /services, /pricingAllowPer signed training stanceDeny unless you have a reason
/blog/ answer postsAllowSame as public marketingDeny is fine for most commercial blogs
/app, /account, /checkoutDenyDenyDeny
Preview / staging hostsDeny all, including GooglebotDenyDeny
/llms.txt, /robots.txtMust remain fetchable or no policy can be readSameSame

Hide apps, carts-in-progress, internal search, and preview hosts. That is a path decision, not an “AI off” decision. Wrong facts that persist after you fix crawl policy are a source problem — avoiding AI-hallucinated brand facts.

  • Marketing signed the citation list
  • Legal signed the training list
  • Scrapers without a product are deny
  • Googlebot was not used as an AI protest
  • Implementation is a separate ticket from this table (robots.txt decisions)

Failure mode: you blocked the bot that cites you

What breaks: a security owner enables “block AI crawlers,” or a plugin Disallows every token that contains “Bot,” or User-agent: * + Disallow: / ships because staging hygiene leaked. GPTBot dies. So does OAI-SearchBot. So does PerplexityBot. Google-Extended may get Disallowed in the same breath as Googlebot if someone matched on “Google.” ChatGPT search stops showing you. Perplexity has nothing to cite. Sometimes Google Search coverage falls. The content team keeps publishing “AEO pages” into a 403.

What it costs: a quarter of work with a citation rate of zero, plus a false story that “AI search doesn’t work in our category.”

What you do instead:

  1. Restore the decision table — citation vs training vs scraper — on one page legal and marketing initial.
  2. Allow the citation twins even if training stays Disallow.
  3. Take Googlebot off every “AI” list.
  4. Confirm the edge did not implement a bundle that ignores the table. The fetch proof belongs in the robots.txt spoke; the policy proof is this matrix with signatures.
  5. Re-run a frozen prompt panel. If you are still unnamed, the next ticket is corroboration and extractability, not another crawler plugin.
SymptomLikely collapsed decisionFirst restore
Perplexity never cites you; ChatGPT search neitherSearch tokens Disallow or WAF 403PerplexityBot + OAI-SearchBot Allow
Google coverage dropped after “AI block”Googlebot caught in the ruleRemove Googlebot from the AI list the same day
Gemini Apps grounding off, Search fineGoogle-Extended Disallow (expected if intended)Leave it if legal wanted that; do not “fix” Search
Bandwidth panic, citations also goneBytespider rule used a shared “AI” groupSplit Bytespider from OpenAI / Perplexity
Brand facts still wrong after a clean AllowNot a bot table problemResidue and corroboration, not more Allows

Bravery is not a restore strategy. The table is.

Tokens and user-agents change

Every table in this post is a worksheet dated by the vendor pages we checked, not a forever spec. OpenAI prints example UAs and says the version number may change. Google prints Chrome/W.X.Y.Z as a placeholder and tells you to wildcard it. Amazon prints Amazonbot/0.1 with a Chrome version wildcard. Common Crawl says they may increment CCBot. ByteDance strings have circulated in more than one wrapper.

Match the product token (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Amazonbot, Bytespider, CCBot). Do not pin GPTBot/1.4 in a WAF and then miss GPTBot/1.5. Do not invent GPT-Bot or GoogleExtended and assume the vendor will fuzzy-match.

DriftWhat breaksCadence
Version suffix incrementsWAF exact-match misses the real bot (or the reverse: only impersonators match)Quarterly token review against vendor docs
New twin appears (search vs training vs user)You keep deciding on the old nameSame review; OpenAI and Anthropic already split
Product token with no HTTP UA (Google-Extended)Log hunters “prove” it does not exist and delete the groupBrief the team: absence in logs is expected
CDN managed list renames “AI bots”Citation crawlers enter a block bundle you did not signRe-read the managed list after every security change
Impersonators reuse the UAYou punish the vendor or ignore a real scrapeIP / ASN verification where the vendor publishes lists

Numbered re-check:

  1. Open the vendor bot page. Copy tokens, not folklore Slack.
  2. Diff against the signed table. Any new twin gets a row: citation, training, or scraper.
  3. Diff against the CDN / WAF managed list. Remove citation crawlers from scrape bundles.
  4. Spot-check logs for the token, not the full historic UA.
  5. Re-run five buyer prompts. Policy without a prompt panel is a file you never measured.

Hedge, out loud: vendors add agents, rename them, and change whether user-triggered fetchers honor robots.txt. If this page disagrees with the vendor page on the day you ship, the vendor page wins.

How do I know the policy is working?

You measure fetch + mention, not a robots.txt screenshot. AEO citation rate is non-deterministic. Log a fixed prompt panel. Treat one lucky ChatGPT run as anecdote.

WeekWhat you recordPass
0Signed table; which tokens are Allow vs Disallow; which CDN bundle existsTwo signatures, not a Slack emoji
1Log lines (or WAF events) for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot vs the deny listCitation bots get 200s on About / offer URLs
2Same 25–40 prompts across ChatGPT (search on), Perplexity, Claude with web, Google AI Overview where it firesNamed / cited / absent / misstated — four buckets
4Repeat prompts; note training tokens still absent if that was the pointNo surprise Allow on Bytespider or CCBot
MetricIs it this post?If it stays red
Citation bots fetch 200YesPolicy vs edge — fix the table or the bundle
Prompt panel names youPartlyFetch is necessary, not sufficient; entity and corroboration next
Training bots absent in logsYes, if you DisallowedImpersonators do not count as vendor compliance
Google impressions collapsedOnly if Googlebot was blockedRestore Googlebot the same day
Hallucinated founding yearNoFact packet and leftover URLs

Prompt panel — freeze the questions, not the engine’s mood that day:

Prompt typeExample shapeWhat a bot-table failure looks like
Brand probe“What does {legal name} do?”Absent or a competitor with a similar string
Category hire“Who should I hire for {category} in {geo}?”Rivals cited; you unnamed
Comparison“{You} vs {rival}”Rival’s table quoted; your offer page never fetched
Local“Best {service} near {city}”Directories win; your GBP-aligned page unused
Correction“What year was {brand} founded?”Stale year — not a robots token issue

Run the same list in ChatGPT with search on, Perplexity, Claude with web tools, and Google (Overview present or not). Four surfaces, one table. If ChatGPT search and Perplexity both go dark in the same week you shipped an “AI” WAF, start with OAI-SearchBot and PerplexityBot, not with a new cluster of blog posts.

What to skip if you only have a week:

  1. Sign the seven-token table (plus search twins).
  2. Allow citation crawlers; deny Bytespider and CCBot unless you have a reason.
  3. Take Googlebot off the AI list.
  4. Freeze 15 buyer prompts and run them twice, a week apart.
  5. Do not rewrite the blog calendar until fetch is proven.

When this is not worth doing yet: you have no public site you want cited, or the only URLs that matter are behind login. Then the table is “deny scrapers, do not invite Googlebot into /app.” Everyone else with a marketing site is already in the game. Ignoring the table is still a decision. It is just an unsigned one.

FAQ

Which AI bots should I allow vs block?

Allow citation-useful crawlers if you want to be recommended — PerplexityBot, plus search twins like OAI-SearchBot and Claude-SearchBot, and Googlebot for Search. Block scrapers you do not need, typically Bytespider and CCBot. Treat GPTBot, ClaudeBot, Google-Extended, and Amazonbot as training or product-policy calls, not as ChatGPT or Google ranking switches.

How do I measure whether my allow vs block policy is working?

Prove citation crawlers can fetch your About and offer URLs, then re-run a frozen prompt panel and score named / cited / absent / misstated. Log lines for PerplexityBot and OAI-SearchBot beat a robots.txt screenshot. If fetch is clean and mentions stay zero, the next problem is corroboration and extractability, not another token.

What usually fails first when teams try this?

They collapse training, citation, and scrapers into one “block AI” CDN or plugin list. OAI-SearchBot and PerplexityBot die with GPTBot and Bytespider. Sometimes Googlebot dies too. The restore is a signed table that splits the three jobs, then an edge rule that matches the table.

How long does this take to show results?

OpenAI and Perplexity document about 24 hours for robots.txt search settings to adjust; Amazon documents a similar window. Re-crawl of key URLs and live citations take longer. Training Disallows affect future collection, not model memory you already dislike. Re-test prompts at one week and one month, not the next morning.

What should I skip if I only have a week?

Skip a full content rewrite. Sign the seven-token table, allow citation crawlers, deny Bytespider and CCBot unless you have a product reason, get Googlebot off the AI list, and run a small prompt panel twice. Implementation details for the file itself belong in the robots.txt spoke, not in another strategy workshop.

When is this not worth doing yet?

If nothing public should be cited — logged-in app, pure staging, or a site you intend to keep out of Search — spend the week on path-level denies and Googlebot hygiene, not Perplexity strategy. If you do have a marketing site and buyer prompts, an unsigned “we’ll deal with AI later” is already a block-or-allow decision. Make it explicit.

CTA

Sign the allow vs block table before you buy another “AI visibility” plugin — citation crawlers on, scrapers off, training in writing.

Lane overview: /visibility. Next step: a visibility audit.

FAQ

What questions does this article answer?

Which AI bots should I allow vs block?
Allow citation-useful crawlers if you want to be recommended — `PerplexityBot`, plus search twins like `OAI-SearchBot` and `Claude-SearchBot`, and `Googlebot` for Search. Block scrapers you do not need, typically `Bytespider` and `CCBot`. Treat `GPTBot`, `ClaudeBot`, `Google-Extended`, and `Amazonbot` as training or product-policy calls, not as ChatGPT or Google ranking switches.
How do I measure whether my allow vs block policy is working?
Prove citation crawlers can fetch your About and offer URLs, then re-run a frozen prompt panel and score named / cited / absent / misstated. Log lines for `PerplexityBot` and `OAI-SearchBot` beat a robots.txt screenshot. If fetch is clean and mentions stay zero, the next problem is corroboration and extractability, not another token.
What usually fails first when teams try this?
They collapse training, citation, and scrapers into one “block AI” CDN or plugin list. `OAI-SearchBot` and `PerplexityBot` die with `GPTBot` and `Bytespider`. Sometimes `Googlebot` dies too. The restore is a signed table that splits the three jobs, then an edge rule that matches the table.
How long does this take to show results?
OpenAI and Perplexity document about 24 hours for robots.txt search settings to adjust; Amazon documents a similar window. Re-crawl of key URLs and live citations take longer. Training Disallows affect future collection, not model memory you already dislike. Re-test prompts at one week and one month, not the next morning.
What should I skip if I only have a week?
Skip a full content rewrite. Sign the seven-token table, allow citation crawlers, deny `Bytespider` and `CCBot` unless you have a product reason, get `Googlebot` off the AI list, and run a small prompt panel twice. Implementation details for the file itself belong in the robots.txt spoke, not in another strategy workshop.
When is this not worth doing yet?
If nothing public should be cited — logged-in app, pure staging, or a site you intend to keep out of Search — spend the week on path-level denies and `Googlebot` hygiene, not Perplexity strategy. If you do have a marketing site and buyer prompts, an unsigned “we’ll deal with AI later” is already a block-or-allow decision. Make it explicit.
Sources

Last reviewed — OpenAI bots overview, Anthropic crawler help, Perplexity crawlers doc, Google common crawlers (Google-Extended), Google AI features, Amazonbot, and Common Crawl CCBot pages checked 2026-09-05. Token names and UA version suffixes change; treat tables as a worksheet.

More from this lane

AI Visibility

All →
Book the audit