Which AI bots should I allow vs block
Allow citation crawlers if you want to be recommended; block Bytespider and CCBot if you do not need them. GPTBot trains. Googlebot is not Google-Extended.
William Spurlock Founder — Spurlock Studios 28 MIN
Allow citation-useful crawlers if you want answer engines to recommend you. Block scrapers and training crawlers you have no product reason to feed. The seven tokens operators actually argue about are GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Amazonbot, Bytespider, and CCBot — and they do not all do the same job. Googlebot is not Google-Extended. Collapsing them into one “block AI” switch is how brands opt out of ChatGPT search, Perplexity citations, or Google Search while congratulating themselves for a privacy win.
This is the allow-vs-block decision table inside the Answer Engine Optimization playbook. How you write the file, which host it lives on, and how you prove a 200 lives in AI crawlers and robots.txt decisions. Do not treat this spoke as a robots.txt tutorial.
The short answer
- Want to be recommended? Allow the crawlers that feed citation and retrieval. Block the ones that only scrape or train, unless you have a written reason to donate the corpus.
GPTBotandClaudeBotare training. ChatGPT search and Claude search-style retrieval use different tokens (OAI-SearchBot,Claude-SearchBot). Blocking training does not keep you eligible for citations.PerplexityBotis a citation crawler. Perplexity says it is not used to crawl content for foundation models. Default allow if Perplexity is a surface you care about.Googlebotcrawls for Google Search.Google-Extendedis a robots.txt product token for listed Gemini training and grounding uses. Blocking one is not blocking the other.BytespiderandCCBotare the usual blocks when you do not need ByteDance products or a public Common Crawl dump.Amazonbotis a policy call unless Alexa / Rufus actually matters.
The allow vs block table
These are the seven bots this post decides. Adjacent twins appear only so you do not block the wrong job.
| Token | Operator | Job (vendor wording, checked 2026-09-05) | Default if you want citations | Default if you do not need that product |
|---|---|---|---|---|
GPTBot | OpenAI | Crawl that may be used to train generative AI foundation models | Policy call (often Disallow) | Disallow |
ClaudeBot | Anthropic | Collect web content that could contribute to model training | Policy call (often Disallow) | Disallow |
PerplexityBot | Perplexity | Surface and link sites in Perplexity search results; not foundation-model crawl | Allow | Disallow |
Google-Extended | Product token: listed Gemini training and grounding uses | Policy call | Disallow the token; never Disallow Googlebot as a substitute | |
Amazonbot | Amazon | Improve products and services; may be used to train Amazon AI models | Allow only if Amazon surfaces matter | Disallow |
Bytespider | ByteDance | Search / recommendation / training crawl for ByteDance properties | Disallow unless you need those products | Disallow |
CCBot | Common Crawl | Public web crawl into an open dataset anyone can train on later | Disallow unless you want the dump | Disallow |
Citation twins you must not confuse with the rows above:
| If you meant | Do not use | Use instead |
|---|---|---|
| ChatGPT search answers | GPTBot | OAI-SearchBot (OpenAI: opted-out sites are not shown in ChatGPT search answers) |
| Claude search-style indexing | ClaudeBot | Claude-SearchBot |
| Google Search crawl / index | Google-Extended | Googlebot |
| Alexa / Rufus search eligibility | Amazonbot | Amzn-SearchBot (Amazon: does not crawl for generative AI training) |
Decision list before anyone pastes a CDN preset:
- Which answer engines should still be able to cite us?
- Which vendors may train on our public pages?
- Which scrapers have no product we sell into?
- Did a “block AI bots” toggle already answer those three as one?
If legal and marketing cannot sign different answers for (1) and (2), you do not have a policy. You have a mood.
Use the table in a 20-minute meeting, not a Slack poll:
- Read each of the seven tokens out loud with the job column. No nicknames.
- Marketing marks citation Allow/Deny. Legal marks training Allow/Deny. Security marks scrape Allow/Deny.
- Where two owners disagree on one token, that token stays Deny until someone writes a reason to Allow.
- Search twins (
OAI-SearchBot,Claude-SearchBot,Googlebot) get their own row on the same page even though they are not in the seven. Missing twins are how “we blocked GPTBot” becomes “ChatGPT search never heard of us.” - Date the sheet. The next vendor changelog invalidates folklore, not the jobs.
Citation crawlers, training crawlers, and scrapers
An “AI bot” is not a species. It is a job. Treat three jobs, then assign each token.
| Job | What success looks like | What a block actually does | Examples in this table |
|---|---|---|---|
| Citation / retrieval | The engine can fetch or index you and name you with a link | You become harder to recommend on that surface | PerplexityBot; also OAI-SearchBot, Claude-SearchBot, Googlebot |
| Training / corpus | Future model weights or listed grounding uses may include you | Future collection should stop; old weights do not rewind | GPTBot, ClaudeBot, Google-Extended, often Amazonbot |
| Public dump / third-party scrape | Someone else publishes or trains on a copy of the open web | You opt out of that crawl going forward | CCBot; Bytespider when you have no ByteDance surface |
User-initiated fetchers (ChatGPT-User, Claude-User, Perplexity-User, Amzn-User) are a fourth bucket. They fire because a person asked. Several vendors say robots.txt may not apply, or generally does not. Do not use those tokens as your citation on/off switch.
| Bucket | Honor robots.txt? (vendor claim) | Use it as your allow/block KPI? |
|---|---|---|
| Automated search / index crawler | Usually yes | Yes |
| Automated training crawler | Usually yes | Yes, for training policy |
| Product token with no separate HTTP UA | Applied downstream (Google-Extended) | Yes, for Gemini listed uses — not for Search |
| User-triggered fetcher | Mixed; often no | No |
I have been SEO certified since 2021. The AEO version of that work is still “which agent, which product,” not “AI: on or off.”
- Citation surfaces are named (ChatGPT search, Perplexity, Claude search, Google Search / AI Overviews, Alexa)
- Training stance is written separately from citation stance
- Scrapers without a named product are default-deny
- User-triggered fetchers are not the scoreboard
Should I allow GPTBot?
Allow GPTBot only if you intend OpenAI to use crawled pages in foundation-model training. Do not allow it because you “want to show up in ChatGPT.” That is a different bot.
OpenAI’s crawler overview is explicit: settings are independent. You can allow OAI-SearchBot so ChatGPT search can surface you, and disallow GPTBot so the same pages are not a training signal. If both are allowed, OpenAI may reuse one crawl for both jobs. That is an efficiency note, not a reason to treat the tokens as one.
| Intent | GPTBot | OAI-SearchBot | What operators get wrong |
|---|---|---|---|
| Citations, no training donation | Disallow | Allow | Disallow both “to be safe” |
| Citations and training | Allow | Allow | Fine, if legal signed it |
| No ChatGPT search, training still OK | Allow | Disallow | Rare and usually accidental |
| Full opt-out of OpenAI automated crawl | Disallow | Disallow | Still not a Google ranking lever |
ChatGPT-User is not this decision. OpenAI says it is not used to determine Search appearance, and robots.txt rules may not apply because a user initiated the fetch.
| Signal | Treat as |
|---|---|
| Legal wants no OpenAI training | Disallow GPTBot |
| Marketing wants ChatGPT search citations | Allow OAI-SearchBot |
| Someone blocked “GPTBot” in a WAF named “all OpenAI” | Audit whether OAI-SearchBot died too |
| Leadership asks “did we opt out of ChatGPT?” | Answer with the search token, not the training token |
Version suffixes move. OpenAI’s examples use strings like GPTBot/1.4 and note the version number may change. Match the product token GPTBot, not a pinned UA you copied in 2025.
OpenAI’s publisher FAQ is the citation receipt: for content to be included in summaries and snippets in ChatGPT, do not block OAI-SearchBot. Referral traffic from ChatGPT search can carry utm_source=chatgpt.com. That UTM is evidence the search allow is working. It is not evidence that GPTBot was a good training donation.
| You see in analytics | It means | It does not mean |
|---|---|---|
utm_source=chatgpt.com | ChatGPT search (or a ChatGPT-attributed click) found a URL | GPTBot is required |
| No ChatGPT UTM, search prompts still cite you | Citations can happen without a click | Your bot table is irrelevant |
| No citations and no UTM after a “block AI” change | Search path is probably closed | You need more blog posts first |
Default for most citation-seeking brands: Disallow GPTBot. Allow OAI-SearchBot. Write that as two lines of policy, then implement it in the robots.txt spoke.
Should I allow ClaudeBot?
Allow ClaudeBot only as a training-donation decision. Anthropic’s crawler help article splits three robots. Restricting ClaudeBot “signals that the site’s future materials should be excluded from our AI model training datasets.” That sentence is not a citation guarantee and not a citation kill switch.
| Bot | Anthropic’s job | If you restrict it |
|---|---|---|
ClaudeBot | Training collection | Future materials should be excluded from training datasets |
Claude-SearchBot | Improve search result quality | May reduce visibility and accuracy in user search results |
Claude-User | Fetch when a person asks Claude | Prevents retrieval in response to that user query |
Anthropic states all three honor robots.txt — including the user-initiated agent. That is a documented difference from OpenAI’s ChatGPT-User note and from Perplexity’s Perplexity-User note. Confirm on the vendor page before you brief a lawyer; names and scopes change.
| Goal | Claude control | Do not substitute |
|---|---|---|
| Stay findable in Claude search-style results | Allow Claude-SearchBot | Allowing ClaudeBot instead |
| Opt out of Anthropic training | Disallow ClaudeBot | Disallowing Claude-SearchBot “to be thorough” |
| Block live user fetches | Disallow Claude-User (and expect support tickets) | Assuming a training Disallow covers it |
Anthropic also warns that IP-blocking their bots can prevent them from reading robots.txt, which breaks a clean opt-out. Prefer a token-level Disallow when the goal is a recorded training stance.
Default for most citation-seeking brands: Disallow ClaudeBot. Allow Claude-SearchBot. Same shape as OpenAI. Same refusal to mash them together.
Should I allow PerplexityBot?
Yes — if Perplexity is a surface where you want numbered citations. No — if you have a written decision that Perplexity should not index you.
Perplexity’s crawlers doc says PerplexityBot is designed to surface and link websites in search results on Perplexity, and is not used to crawl content for AI foundation models. That is the opposite of GPTBot. Treating PerplexityBot as “another training scraper” is how you disappear from a retrieval engine while your training policy was already handled elsewhere.
| Token | Job | Foundation-model crawl? | robots.txt |
|---|---|---|---|
PerplexityBot | Automated search / citation index | Perplexity says no | Honor it; this is the allow/block for Perplexity results |
Perplexity-User | User-triggered fetch to answer a question | Perplexity says no | Generally ignores robots.txt because a user requested the fetch |
| You want | PerplexityBot | Notes |
|---|---|---|
| Cited in Perplexity answers | Allow | Also allow the published IP ranges at the edge, or a WAF will 403 a “yes” |
| Not in Perplexity’s index | Disallow | Perplexity-User can still hit a pasted URL |
| Training opt-out only | Do not Disallow this token for that reason | Use GPTBot / ClaudeBot / Google-Extended for training |
Perplexity notes settings are independent and may take up to 24 hours to reflect. That is a retrieval clock, not a ranking myth.
- Perplexity is on the named citation list, or it is explicitly off
-
PerplexityBotis not lumped into a “training scrapers” WAF group -
Perplexity-Useris not your opt-out lever - Key answer URLs return 200 to the real bot, not only to Chrome
Default for citation-seeking brands: Allow PerplexityBot.
Why Googlebot is not Google-Extended
Googlebot is the Search crawler. Blocking it is how you leave Google Search — including Discover and other Search features Google lists on the common crawlers page.
Google-Extended is not a crawler you will see as its own HTTP user-agent. Google says crawling is done with existing Google user-agent strings; the robots.txt token is used in a control capacity. What it manages: whether crawled content may be used for training future Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI.
Google also says, in the same entry: Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.
| Control | Separate HTTP UA? | Search inclusion / ranking | AI Overviews / AI Mode | Gemini Apps training / listed grounding |
|---|---|---|---|---|
Googlebot | Yes | Yes — this is the Search crawl path | Indirect: Overviews need an indexed, snippet-eligible page | Not this token |
Google-Extended | No | No (per Google) | No — Google points you to snippet / noindex controls | Yes — listed uses |
nosnippet / max-snippet / noindex | Page-level | noindex drops Search | Snippet rules apply to AI Overviews and AI Mode per Google’s robots meta spec | Not a substitute for the Extended token |
Google’s AI features page is the other half: robots.txt directives for Googlebot manage crawl for Search. To limit what Search shows, use snippet and index controls. To limit AI training and grounding in some of Google’s other systems, read Google-Extended.
Third-party posts still claim Google-Extended turns off AI Overviews. That is not what Google’s own pages say. If marketing wants “no Overviews” and legal wants “no Gemini training,” those are two tickets.
| Goal | Allow / keep | Do not |
|---|---|---|
| Stay in Google Search | Googlebot | Disallow Googlebot as an “AI” move |
| Stay eligible for AI Overviews | Indexed + snippet-eligible (Googlebot path) | Expect Google-Extended Disallow to hide the Overview |
| Opt out of listed Gemini training / grounding | Disallow Google-Extended | Assume that also demotes Search |
| Ground in Gemini Apps on listed terms | Allow Google-Extended | Confuse that with ChatGPT search |
GoogleOther is a third Google name that shows up in logs. Google describes it as a generic crawler for one-off research fetches, not a specific consumer product and not Search. Do not Disallow Googlebot because you saw GoogleOther. Do not assume GoogleOther is Google-Extended. Extended still has no separate HTTP UA.
| Log / token | Decide as | Common false move |
|---|---|---|
Googlebot | Search crawl | “That’s the AI bot” |
Google-Extended | Gemini training / listed grounding | “That’s how we leave AI Overviews” |
GoogleOther | Unspecified Google R&D fetch | Collapse it into a Googlebot Disallow |
Google-InspectionTool | Search Console / Rich Results tests | Blocking it “for AI” breaks your own inspections |
Default: Never block Googlebot to make an AI statement. Decide Google-Extended as a Gemini training/grounding policy. Most brands that want Search and Overviews leave Googlebot allowed and pick Extended in writing.
Should I allow Amazonbot?
Allow Amazonbot when Amazon product improvement — and possible Amazon model training — is a trade you accept. Block it when Alexa, Rufus, and Amazon’s other surfaces are not a channel you care about.
Amazon’s Amazonbot page now splits three agents. Each setting is independent. Changes may take about 24 hours.
| Token | Amazon’s job | Training? | Default for a brand that does not sell the Amazon surface |
|---|---|---|---|
Amazonbot | Improve products and services; may be used to train Amazon AI models | Maybe (Amazon’s word is “may”) | Disallow |
Amzn-SearchBot | Search experiences such as Alexa; eligible to appear if permitted | Amazon says it does not crawl for generative AI model training | Allow only if Alexa / Rufus citations matter |
Amzn-User | Live fetch for user questions (e.g. Alexa) | Amazon says no generative-AI training crawl | Not your training switch; may not follow all robots.txt directives |
Amazon also notes: if robots.txt does not mention Amzn-SearchBot but allows other search bots, Amzn-SearchBot will crawl in accordance with those other search-bot directives. Silence is not a clean “Amazon off.”
| Situation | Amazonbot | Amzn-SearchBot |
|---|---|---|
| DTC brand, no Alexa strategy | Disallow | Disallow (name it; do not rely on silence) |
| Local service that wins voice queries | Policy call | Allow if Alexa is a real channel |
| Marketplace seller who needs Rufus | Policy call | Allow |
| Rights-sensitive publisher | Disallow | Disallow |
Amazon documents noarchive as “do not use the page for model training,” plus noindex and none. That is a page-level training hint, not a Google ranking control, and not a ChatGPT control.
Default for most non-Amazon businesses: Disallow Amazonbot. Name Amzn-SearchBot if you also want it off. Do not assume Amazonbot is how ChatGPT or Perplexity find you.
Should I block Bytespider?
Yes for most brands that are not trying to be found inside ByteDance products (TikTok-adjacent search, Doubao, Toutiao). Bytespider is ByteDance’s crawler. It is not a ChatGPT citation path and not Google Search.
ByteDance webmaster material names the token Bytespider. English-language official docs are thinner than OpenAI’s or Google’s. Independent operators report mixed robots.txt honor over the years. Treat a Disallow as a request you should also enforce at the edge if the crawl volume hurts. Do not treat a robots.txt line as a cryptographic guarantee.
| Claim you will hear | What to do instead |
|---|---|
| “Bytespider is how we get TikTok SEO in the US” | Separate TikTok ads / organic from this crawler; do not allow a heavy scrape on a rumor |
| “Blocking it hides us from ChatGPT” | Unrelated tokens |
| “robots.txt is enough” | Log it. If Disallow is ignored, WAF / ASN rules — without blocking Googlebot or OAI-SearchBot in the same group |
| “The UA is proof” | User-agents are cheap to spoof. Confirm operator docs when they exist; Bytespider’s public IP story is weaker than OpenAI’s JSON lists |
| Policy | Bytespider | Why |
|---|---|---|
| Default citation-seeking US/EU brand | Disallow | No citation product you are measuring |
| China-market brand that needs ByteDance search | Allow, then watch bandwidth | Product reason exists |
| Newsroom under scrape load | Disallow + edge | Volume, not AEO |
Hedge: token spelling, UA wrappers, and compliance can change without a nice changelog in English. Re-check logs quarterly. If a CDN “AI bot” list includes Bytespider and OAI-SearchBot, you did not ship a Bytespider decision. You shipped a bundle.
Default: Block Bytespider unless you can name the ByteDance product you are serving.
Should I block CCBot?
Usually yes, if you do not want future snapshots of your site in Common Crawl’s public datasets. CCBot is the crawler for Common Crawl, a nonprofit that publishes open web crawls. Labs, startups, and researchers train and evaluate models on those dumps. Blocking CCBot is opting out of that pipeline going forward. It does not retract crawls already published.
Common Crawl identifies the agent as CCBot (examples historically look like CCBot/2.0). They document robots.txt Disallow and, unlike several commercial vendors, they document Crawl-delay. Version numbers increment. Match CCBot, not a frozen CCBot/1.0 string from a 2018 gist.
| Intent | CCBot | What a block does not do |
|---|---|---|
| Stop future inclusion in Common Crawl dumps | Disallow | Erase you from dumps already on disk |
| Support open research / do not care | Allow | Get you ChatGPT citations |
| Reduce third-party training on public mirrors | Disallow, knowing other crawlers still exist | Stop GPTBot or ClaudeBot — those are separate |
| Cut bandwidth from a polite research crawl | Disallow or Crawl-delay | Stop impersonators using the UA |
Common Crawl warns that other crawlers falsely identify as CCBot. They publish IP ranges (including JSON at index.commoncrawl.org/ccbot.json per their bot page). A UA match without the operator’s ranges is not a reason to rewrite policy.
| Brand type | Default | Rationale |
|---|---|---|
| Most commercial sites measuring ChatGPT / Perplexity / Overviews | Disallow | Dump is not your citation path |
| Research org, university, public-data advocate | Allow | Mission fit |
| Publisher with a rights desk | Disallow | Corpus donation is a legal question |
Default: Block CCBot unless openness is an explicit value, not a leftover Allow.
What should most brands default to?
If you want to be recommended, allow the crawlers that feed recommendations. Block scrapers you do not need. Decide training in writing. That is the whole operating system. The seven-token worksheet for a typical US B2B or local brand:
| Token | Default | Why |
|---|---|---|
PerplexityBot | Allow | Citation crawler, not (per Perplexity) foundation-model crawl |
GPTBot | Disallow | Training; keep OAI-SearchBot allowed |
ClaudeBot | Disallow | Training; keep Claude-SearchBot allowed |
Google-Extended | Written policy | Gemini training / listed grounding — not Search |
Googlebot | Allow | Search + Overview eligibility path |
Amazonbot | Disallow | No Amazon surface, possible training |
Bytespider | Disallow | No ByteDance product to serve |
CCBot | Disallow | Public dump, not your AEO scoreboard |
| Company type | Change from the default |
|---|---|
| Rights-sensitive publisher | Disallow training tokens and scrapers; keep search twins if you still want citations |
| Amazon-heavy catalog | Consider Amzn-SearchBot Allow; keep Amazonbot as the training call |
| China / ByteDance market | Bytespider may flip to Allow; watch volume |
| “We do not want AI at all” | You still need Googlebot if you want Google Search. AI Overviews are Search features. Full invisibility is a Search decision, not an Extended token. |
| Regulated / logged-in app | Path-level Denies on /app, /account, /staging for all of the above; do not hide the public About page as collateral |
Pages to keep fetchable first — the ones models actually quote:
- Canonical About / identity (legal name, what you do, where, since when)
- Offer / pricing / service-area pages a buyer would ask an engine about
- One comparison or criteria page with a table
- FAQ answers that stand alone
| URL class | Citation crawlers | Training tokens | Scrapers (Bytespider, CCBot) |
|---|---|---|---|
/about, /services, /pricing | Allow | Per signed training stance | Deny unless you have a reason |
/blog/ answer posts | Allow | Same as public marketing | Deny is fine for most commercial blogs |
/app, /account, /checkout | Deny | Deny | Deny |
| Preview / staging hosts | Deny all, including Googlebot | Deny | Deny |
/llms.txt, /robots.txt | Must remain fetchable or no policy can be read | Same | Same |
Hide apps, carts-in-progress, internal search, and preview hosts. That is a path decision, not an “AI off” decision. Wrong facts that persist after you fix crawl policy are a source problem — avoiding AI-hallucinated brand facts.
- Marketing signed the citation list
- Legal signed the training list
- Scrapers without a product are deny
-
Googlebotwas not used as an AI protest - Implementation is a separate ticket from this table (robots.txt decisions)
Failure mode: you blocked the bot that cites you
What breaks: a security owner enables “block AI crawlers,” or a plugin Disallows every token that contains “Bot,” or User-agent: * + Disallow: / ships because staging hygiene leaked. GPTBot dies. So does OAI-SearchBot. So does PerplexityBot. Google-Extended may get Disallowed in the same breath as Googlebot if someone matched on “Google.” ChatGPT search stops showing you. Perplexity has nothing to cite. Sometimes Google Search coverage falls. The content team keeps publishing “AEO pages” into a 403.
What it costs: a quarter of work with a citation rate of zero, plus a false story that “AI search doesn’t work in our category.”
What you do instead:
- Restore the decision table — citation vs training vs scraper — on one page legal and marketing initial.
- Allow the citation twins even if training stays Disallow.
- Take
Googlebotoff every “AI” list. - Confirm the edge did not implement a bundle that ignores the table. The fetch proof belongs in the robots.txt spoke; the policy proof is this matrix with signatures.
- Re-run a frozen prompt panel. If you are still unnamed, the next ticket is corroboration and extractability, not another crawler plugin.
| Symptom | Likely collapsed decision | First restore |
|---|---|---|
| Perplexity never cites you; ChatGPT search neither | Search tokens Disallow or WAF 403 | PerplexityBot + OAI-SearchBot Allow |
| Google coverage dropped after “AI block” | Googlebot caught in the rule | Remove Googlebot from the AI list the same day |
| Gemini Apps grounding off, Search fine | Google-Extended Disallow (expected if intended) | Leave it if legal wanted that; do not “fix” Search |
| Bandwidth panic, citations also gone | Bytespider rule used a shared “AI” group | Split Bytespider from OpenAI / Perplexity |
| Brand facts still wrong after a clean Allow | Not a bot table problem | Residue and corroboration, not more Allows |
Bravery is not a restore strategy. The table is.
Tokens and user-agents change
Every table in this post is a worksheet dated by the vendor pages we checked, not a forever spec. OpenAI prints example UAs and says the version number may change. Google prints Chrome/W.X.Y.Z as a placeholder and tells you to wildcard it. Amazon prints Amazonbot/0.1 with a Chrome version wildcard. Common Crawl says they may increment CCBot. ByteDance strings have circulated in more than one wrapper.
Match the product token (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Amazonbot, Bytespider, CCBot). Do not pin GPTBot/1.4 in a WAF and then miss GPTBot/1.5. Do not invent GPT-Bot or GoogleExtended and assume the vendor will fuzzy-match.
| Drift | What breaks | Cadence |
|---|---|---|
| Version suffix increments | WAF exact-match misses the real bot (or the reverse: only impersonators match) | Quarterly token review against vendor docs |
| New twin appears (search vs training vs user) | You keep deciding on the old name | Same review; OpenAI and Anthropic already split |
Product token with no HTTP UA (Google-Extended) | Log hunters “prove” it does not exist and delete the group | Brief the team: absence in logs is expected |
| CDN managed list renames “AI bots” | Citation crawlers enter a block bundle you did not sign | Re-read the managed list after every security change |
| Impersonators reuse the UA | You punish the vendor or ignore a real scrape | IP / ASN verification where the vendor publishes lists |
Numbered re-check:
- Open the vendor bot page. Copy tokens, not folklore Slack.
- Diff against the signed table. Any new twin gets a row: citation, training, or scraper.
- Diff against the CDN / WAF managed list. Remove citation crawlers from scrape bundles.
- Spot-check logs for the token, not the full historic UA.
- Re-run five buyer prompts. Policy without a prompt panel is a file you never measured.
Hedge, out loud: vendors add agents, rename them, and change whether user-triggered fetchers honor robots.txt. If this page disagrees with the vendor page on the day you ship, the vendor page wins.
How do I know the policy is working?
You measure fetch + mention, not a robots.txt screenshot. AEO citation rate is non-deterministic. Log a fixed prompt panel. Treat one lucky ChatGPT run as anecdote.
| Week | What you record | Pass |
|---|---|---|
| 0 | Signed table; which tokens are Allow vs Disallow; which CDN bundle exists | Two signatures, not a Slack emoji |
| 1 | Log lines (or WAF events) for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot vs the deny list | Citation bots get 200s on About / offer URLs |
| 2 | Same 25–40 prompts across ChatGPT (search on), Perplexity, Claude with web, Google AI Overview where it fires | Named / cited / absent / misstated — four buckets |
| 4 | Repeat prompts; note training tokens still absent if that was the point | No surprise Allow on Bytespider or CCBot |
| Metric | Is it this post? | If it stays red |
|---|---|---|
| Citation bots fetch 200 | Yes | Policy vs edge — fix the table or the bundle |
| Prompt panel names you | Partly | Fetch is necessary, not sufficient; entity and corroboration next |
| Training bots absent in logs | Yes, if you Disallowed | Impersonators do not count as vendor compliance |
| Google impressions collapsed | Only if Googlebot was blocked | Restore Googlebot the same day |
| Hallucinated founding year | No | Fact packet and leftover URLs |
Prompt panel — freeze the questions, not the engine’s mood that day:
| Prompt type | Example shape | What a bot-table failure looks like |
|---|---|---|
| Brand probe | “What does {legal name} do?” | Absent or a competitor with a similar string |
| Category hire | “Who should I hire for {category} in {geo}?” | Rivals cited; you unnamed |
| Comparison | “{You} vs {rival}” | Rival’s table quoted; your offer page never fetched |
| Local | “Best {service} near {city}” | Directories win; your GBP-aligned page unused |
| Correction | “What year was {brand} founded?” | Stale year — not a robots token issue |
Run the same list in ChatGPT with search on, Perplexity, Claude with web tools, and Google (Overview present or not). Four surfaces, one table. If ChatGPT search and Perplexity both go dark in the same week you shipped an “AI” WAF, start with OAI-SearchBot and PerplexityBot, not with a new cluster of blog posts.
What to skip if you only have a week:
- Sign the seven-token table (plus search twins).
- Allow citation crawlers; deny
BytespiderandCCBotunless you have a reason. - Take
Googlebotoff the AI list. - Freeze 15 buyer prompts and run them twice, a week apart.
- Do not rewrite the blog calendar until fetch is proven.
When this is not worth doing yet: you have no public site you want cited, or the only URLs that matter are behind login. Then the table is “deny scrapers, do not invite Googlebot into /app.” Everyone else with a marketing site is already in the game. Ignoring the table is still a decision. It is just an unsigned one.
FAQ
Which AI bots should I allow vs block?
Allow citation-useful crawlers if you want to be recommended — PerplexityBot, plus search twins like OAI-SearchBot and Claude-SearchBot, and Googlebot for Search. Block scrapers you do not need, typically Bytespider and CCBot. Treat GPTBot, ClaudeBot, Google-Extended, and Amazonbot as training or product-policy calls, not as ChatGPT or Google ranking switches.
How do I measure whether my allow vs block policy is working?
Prove citation crawlers can fetch your About and offer URLs, then re-run a frozen prompt panel and score named / cited / absent / misstated. Log lines for PerplexityBot and OAI-SearchBot beat a robots.txt screenshot. If fetch is clean and mentions stay zero, the next problem is corroboration and extractability, not another token.
What usually fails first when teams try this?
They collapse training, citation, and scrapers into one “block AI” CDN or plugin list. OAI-SearchBot and PerplexityBot die with GPTBot and Bytespider. Sometimes Googlebot dies too. The restore is a signed table that splits the three jobs, then an edge rule that matches the table.
How long does this take to show results?
OpenAI and Perplexity document about 24 hours for robots.txt search settings to adjust; Amazon documents a similar window. Re-crawl of key URLs and live citations take longer. Training Disallows affect future collection, not model memory you already dislike. Re-test prompts at one week and one month, not the next morning.
What should I skip if I only have a week?
Skip a full content rewrite. Sign the seven-token table, allow citation crawlers, deny Bytespider and CCBot unless you have a product reason, get Googlebot off the AI list, and run a small prompt panel twice. Implementation details for the file itself belong in the robots.txt spoke, not in another strategy workshop.
When is this not worth doing yet?
If nothing public should be cited — logged-in app, pure staging, or a site you intend to keep out of Search — spend the week on path-level denies and Googlebot hygiene, not Perplexity strategy. If you do have a marketing site and buyer prompts, an unsigned “we’ll deal with AI later” is already a block-or-allow decision. Make it explicit.
CTA
Sign the allow vs block table before you buy another “AI visibility” plugin — citation crawlers on, scrapers off, training in writing.
Lane overview: /visibility. Next step: a visibility audit.
What questions does this article answer?
- Which AI bots should I allow vs block?
- Allow citation-useful crawlers if you want to be recommended — `PerplexityBot`, plus search twins like `OAI-SearchBot` and `Claude-SearchBot`, and `Googlebot` for Search. Block scrapers you do not need, typically `Bytespider` and `CCBot`. Treat `GPTBot`, `ClaudeBot`, `Google-Extended`, and `Amazonbot` as training or product-policy calls, not as ChatGPT or Google ranking switches.
- How do I measure whether my allow vs block policy is working?
- Prove citation crawlers can fetch your About and offer URLs, then re-run a frozen prompt panel and score named / cited / absent / misstated. Log lines for `PerplexityBot` and `OAI-SearchBot` beat a robots.txt screenshot. If fetch is clean and mentions stay zero, the next problem is corroboration and extractability, not another token.
- What usually fails first when teams try this?
- They collapse training, citation, and scrapers into one “block AI” CDN or plugin list. `OAI-SearchBot` and `PerplexityBot` die with `GPTBot` and `Bytespider`. Sometimes `Googlebot` dies too. The restore is a signed table that splits the three jobs, then an edge rule that matches the table.
- How long does this take to show results?
- OpenAI and Perplexity document about 24 hours for robots.txt search settings to adjust; Amazon documents a similar window. Re-crawl of key URLs and live citations take longer. Training Disallows affect future collection, not model memory you already dislike. Re-test prompts at one week and one month, not the next morning.
- What should I skip if I only have a week?
- Skip a full content rewrite. Sign the seven-token table, allow citation crawlers, deny `Bytespider` and `CCBot` unless you have a product reason, get `Googlebot` off the AI list, and run a small prompt panel twice. Implementation details for the file itself belong in the robots.txt spoke, not in another strategy workshop.
- When is this not worth doing yet?
- If nothing public should be cited — logged-in app, pure staging, or a site you intend to keep out of Search — spend the week on path-level denies and `Googlebot` hygiene, not Perplexity strategy. If you do have a marketing site and buyer prompts, an unsigned “we’ll deal with AI later” is already a block-or-allow decision. Make it explicit.
Last reviewed — OpenAI bots overview, Anthropic crawler help, Perplexity crawlers doc, Google common crawlers (Google-Extended), Google AI features, Amazonbot, and Common Crawl CCBot pages checked 2026-09-05. Token names and UA version suffixes change; treat tables as a worksheet.
AI Visibility
AI Visibility Cannabis visibility when the ad accounts are banned
Google and Meta will not take the usual spend. The models still answer dispensary, cultivator, and brand questions — if the site can be read and the cart can clear a 21+ order.
AI Visibility When ChatGPT names the franchise, not your shop
Run the best-HVAC-near-me prompt panel. If the model names a national franchise, fix corroboration and entity facts — not another blog calendar.
AI Visibility What belongs in an AI visibility monthly retainer vs a one-time audit
A one-time audit is the baseline plus prioritized fixes. A monthly retainer is prompt-panel tracking, entity hygiene, page jobs, and citation recovery.
AI Visibility Does Wikipedia or Wikidata help AI recommend my brand
Wikipedia is not a paid AI lever. Notability plus independent sources decide the page; a real Wikidata item helps entity consistency, not a promotional stub.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.