Why do ~30% of sites block AI crawlers without knowing it
Top-1,000 GPTBot Disallows hit 35.7% in Originality.AI’s Aug 2024 cut. Accidental blocks come from CDN defaults, plugins, leftover robots, and Googlebot mixups.
William Spurlock Founder — Spurlock Studios 28 MIN
Sites block AI crawlers without knowing it because the block is usually a default, a leftover, or a confused token — not a signed policy. The “~30%” in the question is the order of magnitude of Originality.AI’s top-1,000 robots.txt cut, reported 3 August 2024 by PPC Land: 35.7% of those sites Disallow GPTBot, up from about 5% at GPTBot’s August 2023 launch. That study does not say 30% of the whole web, and it does not measure whether the owner meant to opt out of ChatGPT search.
This spoke is the accidental-block layer of the Answer Engine Optimization playbook. How you write the file lives in AI crawlers and robots.txt decisions. Unblocking origin is necessary, not sufficient: engines still quote third-party threads, which is why Reddit for AI citations stays a separate job.
The short answer
- Cite 35.7% of the top 1,000 Disallowing
GPTBot(Originality.AI via PPC Land, 3 Aug 2024). Do not flatten that into “30% of all sites” or “30% did it by accident.” - “Without knowing it” is a mechanism claim: Cloudflare AI-scrape toggles and new-domain defaults, WordPress plugins and physical leftovers, inherited
Disallowgroups, CDN/WAF 403s, andGoogle-ExtendedvsGooglebotconfusion. - A robots.txt
Allowdoes not prove fetch. Prove a 200 with the real product token. - Training blocks and citation blocks are different tickets. Collapsing them into “block AI” is how marketing ships a quarter of answer pages that ChatGPT search never retrieves.
- Audit the live file and the edge this week. Rewrite the philosophy after the crawler can actually load About.
| Claim in the wild | What a dated source actually measured | What it does not measure |
|---|---|---|
| “~30% of sites block AI crawlers” | 35.7% of the top 1,000 Disallow GPTBot (Originality.AI / PPC Land, 3 Aug 2024) | The long tail; intent; citation crawlers |
| “Everyone on Cloudflare is blocked” | New Cloudflare domains start from a permission default (1 Jul 2025); a one-click scrape toggle exists | Your specific zone’s saved preference |
| “WordPress blocks GPTBot out of the box” | Core virtual robots.txt Disallows /wp-admin/, not AI tokens | Plugin, host, and physical-file overlays |
| “We blocked Google AI so Search is fine” | Google-Extended is a product token, not Googlebot | Whether someone Disallowed Googlebot by mistake |
Where does the ~30% figure actually come from?
It comes from a popular-site robots.txt study, then got rounded in conversation. Treat the question as a pointer to that cut, not as a second unpublished census.
Originality.AI’s AI-bot blocking dashboard tracks which of the top 1,000 sites Disallow named AI tokens. PPC Land’s 3 August 2024 write-up of that program reported:
| Token (that cut) | Share of top 1,000 with a Disallow | What that token is for |
|---|---|---|
GPTBot | 35.7% (was ~5% at Aug 2023 launch) | OpenAI training crawl |
CCBot | 22.1% | Common Crawl dump |
Google-Extended | 13.6% | Gemini training / listed grounding token |
ChatGPT-User | 12.7% | User-initiated ChatGPT fetch |
Newer search twins (OAI-SearchBot, ClaudeBot, anthropic-ai) | Reported in the 1–10% band | Mixed jobs — do not treat the band as one bot |
That is why “~30%” shows up in the title query. 35.7% rounded, on GPTBot, on top 1,000. It is a real dated public number. It is the wrong number if you mean “our SMB site” or “ChatGPT search.”
Cloudflare’s own network sample in the AIndependence post is a different metric: edge block or challenge, not robots.txt text.
| Cloudflare rank band (that post’s June sample) | % of those properties accessed by AI bots | % blocking or challenging those requests |
|---|---|---|
| Top 10 | 80.0% | 40.0% |
| Top 1,000 | 53.2% | 8.8% |
| Top 1,000,000 | 38.73% | 2.98% |
Decision list when someone drops “30%” in a meeting:
- Top 1,000 or the whole web?
GPTBotDisallow, or any AI token, or an edge 403?- Training, citation, or user-fetch?
- Did anyone in the room choose it, or did a default?
If you cannot answer those four, you do not have a statistic. You have a mood.
What does “without knowing it” actually mean?
It means the owner would not sign the same sentence legal would write. The HTTP 403 still happens.
There is no public study I will invent that says “X% of blocks are accidental.” Originality.AI counted files. Cloudflare counted responses. Intent is an audit finding on your zone.
| Layer | What the operator thinks they shipped | What the crawler actually hits | Typical owner |
|---|---|---|---|
| CDN “AI Scrapers and Crawlers” toggle | “Stop scrapers” | Known AI user-agents, including search twins on some presets | Whoever clicked Security → Bots |
| New Cloudflare domain default | “We just put DNS on Cloudflare” | Training (and, depending on the year’s preset, more) blocked until someone opts in | whoever ran signup |
WordPress plugin / physical robots.txt | “Yoast is handling SEO” | Stale Disallow groups the plugin UI no longer shows | whoever migrated the site |
| Inherited gist | “We blocked GPTBot in 2023” | OAI-SearchBot and ChatGPT-User rode along in the paste | an agency, a gist, a security pack |
User-agent: * + Disallow: / | Staging hygiene | Every bot that falls through to *, including Googlebot | a launch checklist that never got reversed |
Google-Extended vs Googlebot | “No Gemini training” or “no Google AI” | Search crawl killed, or training token toggled while Search was the real ask | whoever Googled a snippet |
| WAF / bot-fight / rate-limit | “Block bad bots” | Citation crawlers look like scrapers: bursty, non-browser UA | security, not marketing |
- Marketing can name which engines should still cite us
- Legal can name which vendors may train
- Those two lists are not one checkbox
- Someone has
curl’d the live host this month
If marketing cannot name ChatGPT search vs training, you are already in the accidental bucket — even if the file looks “intentional.”
How does a Cloudflare AI scrape toggle block you?
At the edge, before origin, before robots.txt can help. Cloudflare’s AIndependence post documents a one-click AI Scrapers and Crawlers control under Security → Bots, including on the free tier. The 1 July 2025 press release and Content Independence Day blog then changed the default for new domains: block AI crawlers accessing content without permission or compensation unless the owner opts in. Cloudflare’s own 1 July 2026 recap states that for new domains, AI training crawlers are blocked by default unless the owner chooses otherwise.
That is how a site “blocks AI” without a robots.txt meeting. DNS moved. The toggle was on. Nobody in marketing was in the room.
| Control | Where it lives | What it can do that robots.txt cannot | Failure if you ignore it |
|---|---|---|---|
| AI Scrapers and Crawlers toggle | Cloudflare dashboard, Security → Bots | 403 / challenge known AI fingerprints even when Allow: / | Citation crawlers never see the file |
| New-domain permission default (1 Jul 2025) | Zone onboarding | Starts from deny-until-chosen | A 2025 rebuild inherits a block the 2019 site never had |
| Managed robots / preference sync | Cloudflare-managed robots.txt fragments | Writes Disallow groups you did not type in WordPress | Two sources of truth; the edge still wins |
| Pay Per Crawl / 402 paths | Later Cloudflare crawl-control products | Charge or refuse instead of a quiet 200 | You think you are “open” because the HTML file on origin is open |
Procedure when the zone is on Cloudflare:
- Open the zone → Security → Bots (and any later AI bot policies / Crawl Control screen your plan shows).
- Write down the saved state: allow, block, block-on-ads, or “never touched.”
- Fetch
https://<host>/robots.txtfrom an off-network client. Expecttext/plainand 200, not a challenge HTML page. curl -Awith the citation token you care about (OAI-SearchBot,PerplexityBot,Claude-SearchBot) on About and the money URL. Expect 200.- If robots.txt says Allow and curl is 403, the toggle is the bug. Do not rewrite copy yet.
- Screenshot the Bot / AI policy page with the date
- Diff that screenshot against last quarter — defaults move
- Record whether the zone was created before or after 1 July 2025
- Do not treat “Cloudflare is on” as “we chose a training stance”
The same press release says more than one million customers chose the one-click block after it shipped. That cohort is not “without knowing it.” The accidental cohort is everyone who inherited a default, cloned a zone, or never opened Security → Bots after a redesign. Do not quote the million as your ~30%.
| Situation | Knowing? | What to write in the ticket |
|---|---|---|
| Marketing asked security to “stop AI scraping” | Yes, possibly over-broad | Split training vs citation; re-test search twins |
| New domain after 1 Jul 2025, nobody opened the AI screen | No | Defaults did the work |
| Agency flipped the toggle during a bot incident | Maybe | Screenshot + restore citation tokens |
| UI renamed to AI Crawl Control / bot policies | No if nobody re-read it | Find the current screen; do not search last year’s tutorial |
I have shipped hundreds of production sites. The Cloudflare miss is the one that survives a content sprint because nobody curls with the bot UA.
How do WordPress plugins and host panels do it?
Core WordPress does not ship a GPTBot Disallow. The virtual robots.txt Disallows /wp-admin/ (with Allow for admin-ajax.php) and leaves AI tokens alone. The accidents sit on top of that.
WordPress Reading settings still matter. Discourage search engines from indexing this site is the launch leftover. Since WordPress 5.3 it injects a noindex robots meta (when wp_head runs). Through 5.2, with no physical file, hits to robots.txt could return User-agent: * / Disallow: /. A site that launched on 5.2 and later grew a physical file can still be serving that staging deny.
| Source | What it changes | How you miss it | What to do |
|---|---|---|---|
Virtual WP robots.txt | /wp-admin/ only | You assume a plugin is the file | Fetch the live URL; do not trust the editor screenshot |
Physical robots.txt in web root | Overrides the virtual file and most plugin filters | A migration, an old Yoast “create file,” an agency drop | Diff disk vs dashboard; one file wins |
| SEO plugin file editor (Yoast / Rank Math class) | Edits physical or virtual, depending on product | UI shows rules the crawler never sees, or the reverse | Fetch https://host/robots.txt incognito |
| “Block AI bots” plugin / one-click gist | Writes Disallow: / under GPTBot and often the search twins | Security installed it during a malware scare | Split training vs citation before you delete |
| Host “bot protection” / WAF pack | Edge or origin 403 on non-browser UAs | Panel says “good bots allowed”; AI tokens are not in that list | curl with the token; do not trust the marketing copy on the pack |
| Reading → Discourage indexing | noindex (5.3+) and/or a historic * Disallow | Staging checkbox copied to production | Uncheck; confirm HTML robots and the file |
Wordfence-class firewalls are a maybe, not a documented default AI opt-out. Rate-limit and “block user-agents matching bot” rules can throttle GPTBot and OAI-SearchBot the same way they throttle scrapers. Do not claim Wordfence “blocks AI.” Claim this: if your firewall treats bursty non-browser clients as hostile, citation crawlers look hostile.
- Settings → Reading: Discourage indexing is off on production
-
curl -sI https://host/robots.txtis 200text/plain - Bytes on that URL match the file you think you edit
- No second file on
wwwvs apex - Security plugin UA rules do not match
/bot/iunless you meant to
Plugin UI is not the crawler. The crawler gets bytes.
How do hosting defaults and leftover Disallows stack?
They stack. The crawler sees the strictest door it hits first, not the nicest file in WordPress.
A leftover Disallow on origin plus a CDN “block AI” rule plus a host WAF is three independent nos. Opening one layer does not open the others. That is why teams “fixed robots.txt” and still have zero ChatGPT search citations.
| Layer (outside → in) | Typical default or leftover | What a 200 on this layer still does not prove |
|---|---|---|
| DNS / CDN (Cloudflare, similar) | New-zone AI scrape default; old “block likely bots” | Origin robots.txt policy |
| Host WAF / “bot protection” pack | UA or rate rules aimed at scrapers | That citation IPs are allowlisted |
| App firewall (Wordfence-class) | Burst limits, /bot/ UA rules | That the request was even allowed at the edge |
Physical robots.txt | 2023 gist, staging * deny, agency paste | That the CDN serves this file |
| Virtual WP / plugin editor | Rules the disk file silently overrides | Anything, if a physical file exists |
HTML noindex / robots meta | Reading → Discourage leftover | robots.txt at all — Search still skips the URL |
Order of operations when two people “already checked robots”:
- Name the host the crawler actually requests (
wwwvs apex). - Name the CDN account that answers that host.
- Fetch
/robots.txtthrough that CDN, not from origin SSH. - Fetch a page with the citation UA through that CDN.
- Only then open WordPress.
| Leftover Disallow you still find in 2026 | Usually landed via | Citation risk if left alone |
|---|---|---|
User-agent: * / Disallow: / | Staging, “privacy” checkbox era, launch panic | Catastrophic — Search and AI retrieval |
GPTBot + ChatGPT-User + OAI-SearchBot as one block | 2023–2024 “block ChatGPT” gist | Training and search twins |
Google-Extended copied as Googlebot | Sloppy snippet | Google Search |
CCBot / Bytespider only | Reasonable scrape hygiene | Low, if search twins are open |
Empty or HTML robots.txt | Plugin 404 page with status 200 | Cooperating bots cannot read rules |
- One spreadsheet row per public host: CDN, origin, file SHA, last curl date
- “We use Cloudflare” is not a row — the zone setting is the row
- Leftover Disallows get a keep/drop owner, not a vibe
- After any host migration, re-run the stack; defaults reset
Hosting did not secretly publish a 30% statistic. Hosting is how a site joins whatever statistic already exists without a policy meeting.
How do inherited robots.txt files silently Disallow?
A Disallow does not expire. 2023 gists still win in 2026 if nobody deleted the group.
The usual inheritance paths:
| Inheritance | How it lands | What it usually Disallows | Why nobody “knows” |
|---|---|---|---|
| Agency 2023 “block GPTBot” paste | Copied into Yoast / a physical file | GPTBot, often ChatGPT-User, sometimes OAI-SearchBot the week SearchGPT shipped | The SOW said “protect content” |
| Theme or boilerplate repo | Shipped in public/robots.txt | A catch-all * or a long AI list | Developers never opened it after launch |
| Staging → production copy | Disallow: / under * | Everything that honors * | Launch day never reversed the deny |
Apex vs www | Only one host got the cleanup | The other host still denies | Each host has its own file |
| CDN “managed robots” plus origin file | Two documents | The stricter edge policy | Marketing edited origin; Cloudflare serves something else |
Leftover patterns to grep for — then decide, do not panic-delete:
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: OAI-SearchBot
Disallow: /
User-agent: *
Disallow: /
| Pattern | Citation consequence | Training consequence | First fix |
|---|---|---|---|
GPTBot Disallow only | ChatGPT search can still be allowed if OAI-SearchBot is open | Future OpenAI training collection should stop | Keep or drop as policy, not as an accident |
OAI-SearchBot Disallow | ChatGPT search answers should drop you | Training is a different tag | This is the accidental citation kill if you meant “no training” |
ChatGPT-User Disallow | Live user fetches may fail; OpenAI notes robots.txt may not apply | Not your training control | Do not use this as the Search opt-out |
* + Disallow: / | Search engines that fall through to * are out | Same | Reverse staging; add explicit Allows if you must keep a tight * |
Googlebot Disallow | Google Search crawl/index path is wounded | Gemini token is unrelated | Restore Googlebot unless the site is meant to be private |
Numbered cleanup:
- Save the live file with a date stamp.
- Highlight every
User-agentgroup that names an AI or*deny. - Label each group training, citation, user-fetch, or Search.
- Delete or split groups that were never a signed decision.
- Re-fetch. Then — and only then — write the policy file the other spoke describes.
Do not “simplify” by putting every token under *. That is how Google Search dies in the same commit as Bytespider.
How does Google-Extended vs Googlebot confusion opt you out?
People hear “Google AI” and Disallow the wrong token.
Google’s common crawlers list is explicit: Google-Extended has no separate HTTP user-agent. Crawling uses existing Google UAs; the robots.txt token is a control for whether crawled content may be used for listed Gemini training and grounding uses. Google states Google-Extended does not impact inclusion in Google Search and is not a ranking signal.
Googlebot is the Search crawler. Disallow Googlebot and you are not “opting out of AI Overviews.” You are opting out of the Search crawl/index path that Overviews still sit on.
| What someone typed | What they thought it did | What Google documents | Typical wreckage |
|---|---|---|---|
User-agent: Google-Extended + Disallow: / | “No Google AI” / “no AI Overviews” | Gemini training / listed grounding opt-out; not Search ranking | Overviews can still use Search-indexed snippets unless you also change snippet rules |
User-agent: Googlebot + Disallow: / | “Just the AI crawler” | Search crawl/index for those URLs | Rankings and Overview eligibility both starve |
User-agent: Google (not a real group token you meant) | “Catch all Google” | Groups match tokens, not vibes | Miss or over-match; test, do not guess |
| CDN “block AI” that includes Google user-agents | “Scrapers only” | Depends on the rule pack | Googlebot 403 looks like a ranking collapse |
-
Googlebotis allowed on public URLs you want in Search -
Google-Extendedis a legal ticket, written separately - Nobody used
Google-Extendedas an AI Overviews off switch - Snippet controls (
nosnippet,max-snippet) are a different ticket if the actual goal was Overview text
Originality.AI’s 13.6% Google-Extended Disallow rate (same Aug 2024 top-1,000 cut) is a token rate. It is not “13.6% left Google Search.” Do not quote it that way.
How do CDN WAF and bot-fight rules 403 citation crawlers?
robots.txt is a suggestion to cooperating crawlers. A WAF is a door.
Cloudflare’s AIndependence post is blunt: well-behaved bots honor robots.txt; many scrapers spoof a browser. Their answer is fingerprinting and a Bot Score. The same machinery that catches a liar can catch OAI-SearchBot if you pointed “likely bot” at challenge-or-block and never allowlisted the published ranges.
| Layer | Symptom in logs | What it is not | Fix |
|---|---|---|---|
| Managed AI scrape rule | 403 with a Cloudflare or WAF challenge body | A robots.txt Disallow | Change the AI bot policy; re-curl |
| Super Bot Fight / “block likely bots” | JS challenge, 403, or empty | “Google is fine so AI is fine” | Allowlist verified AI ranges or lower the action on known good tokens |
| Rate limit (plugin or host) | 429 after the first burst of paths | A policy stance | Raise crawler limits; do not require a browser cookie |
| Country or ASN block | Timeouts from US model-vendor IPs | A content decision | If OpenAI or Anthropic egress is blocked, citations die quietly |
robots.txt itself challenged | HTML login wall with status 200 | A valid rules file | Allow anonymous GET of /robots.txt |
Checklist that actually moves the ticket:
-
curl -sI https://host/robots.txt— 200,text/plain, no JS challenge -
curl -A "OAI-SearchBot" -sI https://host/about/— 200 - Repeat for
PerplexityBotandClaude-SearchBotif those surfaces matter - Repeat from a second network so you are not testing your own allowlist
- Confirm the UA IP against the vendor’s published JSON when you start blocking impostors
If the file Allows and the edge 403s, you do not have a content problem. You have a door problem. The robots.txt decisions spoke will not save a fetch that never reaches origin.
What failure mode burns a quarter of AEO work?
What breaks: the team ships answer-first pages, FAQ blocks, and entity facts. ChatGPT search and Perplexity never retrieve them. The deck still says “we did AEO.” The cause is a 403 or a leftover OAI-SearchBot Disallow from a 2023 gist.
What it costs: a quarter of writing against a closed door. Brand search does not move. The prompt panel stays “absent.” Leadership concludes “AEO does not work.”
What you do instead:
- Freeze 20 prompts you actually want to win.
- Before rewriting a word, pass the fetch checklist on homepage, About, and one money URL.
- Only then score cited / mentioned / absent / hallucinated.
- If fetch fails, stop the content calendar. Fix the door.
- After a 200, wait for the vendor’s recrawl window before you declare the copy a failure.
| False diagnosis | Evidence that it is false | Real diagnosis |
|---|---|---|
| “Our writing is not AEO enough” | curl 403 with OAI-SearchBot | Edge toggle |
| “We need FAQ schema” | File Disallows the search bot | Leftover robots group |
| “Google hates us” | Googlebot Disallow or noindex leftover | Search crawl, not Gemini |
| “Perplexity never cites SMBs” | PerplexityBot never got a 200 | WAF |
| “Reddit is stealing our citations” | Your URLs 403; reddit.com 200s | You are unfetchable; Reddit is not — see Reddit for AI citations after you open the door |
Bravery is not a 403 strategy.
How do you audit accidental blocks in one afternoon?
You do not need a new vendor. You need the live host, four curls, and a named owner.
| Step | Command or place | Pass | Fail |
|---|---|---|---|
| 1. Live file | curl -sI https://host/robots.txt | 200 text/plain | Challenge, 403, HTML |
| 2. Bytes | Save body; grep Disallow and User-agent | Groups match the worksheet | Mystery * deny or search-twin Disallow |
| 3. Citation fetch | curl -A search tokens on About + money URL | 200 | 401/403/429/challenge |
| 4. Search crawl | Confirm Googlebot is not Disallowed; HTML is not noindex on public URLs | Indexed path still open | “We blocked Google AI” actually hit Search |
| 5. Edge UI | Cloudflare / host WAF / WP firewall | Policy matches legal + marketing | Toggle on, nobody owns it |
| 6. Host split | Repeat 1–3 on apex and www | Same policy | One host still staging-denied |
Numbered afternoon (90 minutes if the logins exist):
- Inventory public hosts (apex,
www, docs, blog). - Dump every
/robots.txt. - Screenshot CDN AI / bot screens.
- Run the curl table.
- Write a three-line verdict: fetch OK / file leftover / edge block.
- Assign one owner to change one layer. Do not change file and WAF in the same hour unless you can re-test both.
- Date-stamped robots dump in the ticket
- Date-stamped curl headers in the ticket
- Named human for the next change
- Re-test scheduled after the vendor’s search recrawl window (OpenAI documents about 24 hours for ChatGPT search systems to adjust after a robots.txt change on the bots overview)
If you cannot log into Cloudflare, you cannot finish the audit. Get the login before you hire more words.
Curl recipes that belong in the ticket (replace host and path):
| Job | Command | You are looking for |
|---|---|---|
| File headers | curl -sI https://www.example.com/robots.txt | 200, text/plain, no cf-mitigated challenge |
| File body | curl -s https://www.example.com/robots.txt | Real groups, not an HTML error document |
| ChatGPT search twin | curl -sI -A "OAI-SearchBot" https://www.example.com/about/ | 200 (copy the current UA from OpenAI’s bots page if a vendor requires the full string) |
| Perplexity | curl -sI -A "PerplexityBot" https://www.example.com/about/ | 200 if that surface matters; confirm the string on Perplexity’s crawler doc |
| Google Search crawler | curl -sI -A "Googlebot" https://www.example.com/about/ | 200; then confirm HTML is not noindex |
| Apex vs www | Repeat the four above on the other host | Same policy, or a documented 301 before the fetch |
UA version suffixes change. Match the product token, not a pinned GPTBot/1.4 string. If a 403 appears only on the bot UA and a browser 200s, you found the WAF, not a content bug.
What should you fix first if you want citations?
Fetch, then tokens, then copy. Not the other way around.
| Priority | Fix | Skip until this passes |
|---|---|---|
| 0 | Anonymous 200 on /robots.txt and on About / money URL for citation UAs | New blog calendar |
| 1 | Remove OAI-SearchBot / PerplexityBot / Claude-SearchBot Disallows you did not mean | Training philosophy workshops |
| 2 | Turn off or split the Cloudflare AI scrape preset so search twins can 200 | Homepage theater |
| 3 | Restore Googlebot if someone Disallowed it as “Google AI” | Gemini-token debates |
| 4 | Then write the split policy file | Gist pastes from 2023 |
| 5 | Then answer-first HTML on the pages models quote | llms.txt as a substitute for fetch |
Decision list:
- Do we want to be retrieved and named in ChatGPT search, Perplexity, Claude search-style answers, Google Search? If yes, those crawlers need a 200.
- Do we want those vendors to train on the public pages? Separate Allow/Disallow. Do not reuse the citation answer.
- Is any remaining block a WAF? Fix the WAF, not the H1.
- After 200s land, measure the prompt panel. If you are still absent, the problem moved to extractable proof — that is the playbook, not this spoke.
Default for most brands that want recommendations: open the citation door; decide training like an adult. The accident is closing both doors with one “block AI” click.
How do you measure whether the accidental block is gone?
You measure fetch, then inclusion, not a feeling that robots.txt looks cleaner.
| Signal | Tool | Cadence | Pass |
|---|---|---|---|
| robots.txt 200 + intended groups | curl + diff vs worksheet | Every security change | Bytes match policy |
| Citation UA 200 on money URLs | curl -A | Same | No challenge body |
| ChatGPT search / Perplexity / Claude inclusion | Frozen prompt panel, screenshots | Weekly for 4 weeks after the 200 | Cited or at least retrieved, not “I cannot access that site” |
| Google Search | Search Console URL inspection + index coverage | After Googlebot restores | URL is indexable |
| Residual mentions | Same panel | Monthly | Hallucinations drop once the engine can fetch you; Reddit and PR still do their own job |
Do not report “we unblocked AI.” Report: OAI-SearchBot 200 on /about/ on 2026-08-14; prompt 7 cited us on 2026-08-21. That is a measurement. A green toggle is not.
Weekly loop after the first clean 200:
- Monday: re-curl the citation UAs. If they 403 again, security shipped a rule. Stop scoring copy.
- Wednesday: run the frozen prompt panel. Log cited / mentioned / absent / hallucinated. Save the source list.
- Friday: URL-inspect the money URL in Search Console if
Googlebotwas in the mess. - Week 4: decide whether you are still in this spoke (fetch) or in the playbook (selection).
| Week after the 200 | Fetch still the story? | Selection now the story? |
|---|---|---|
| 0–1 | Yes — recrawl windows | No — do not rewrite the homepage yet |
| 2–3 | Only if 403s return | Start watching source panels |
| 4+ | Only if a new WAF shipped | Thin About, no table, or Reddit already answers it |
If inclusion stays zero after fetch is clean, you have a different post’s problem: thin entity, no citeable table, or engines preferring Reddit threads that already answer the query. Opening the crawler does not mint a citation. It only stops you from donating the slot.
What should you skip if you only have a week?
Skip a new allow/block philosophy, a 40-bot gist, and a homepage redesign. You have one job: prove or kill the accidental door.
- Curl
robots.txton every public host - Curl citation UAs on homepage, About, one money URL
- Open Cloudflare / host WAF / WP firewall and screenshot AI / bot rules
- Delete or split only the groups that are obviously leftovers (
*sitewide deny,OAI-SearchBotDisallow next to a “we want ChatGPT citations” OKR) - Leave
GPTBot/ClaudeBot/Google-Extendeduntouched if legal has not signed — those are policy, not accidents, until someone says otherwise - Freeze 10 prompts and run them once after the first clean 200, then again a week later
Do not install a new “AI visibility” plugin that writes more Disallows. Do not flip Cloudflare to “allow all AI” if legal forbade training — split the categories if the dashboard can, or allowlist search tokens only.
A week is enough to stop the self-own. It is not enough to win the category. Say that out loud so nobody expects a citation spike from a 403 fix on Friday.
When is this not worth doing yet?
When the site is meant to be private, when legal has already signed “no retrieval, no training, no user fetch,” or when you cannot fetch your own /robots.txt because basic hosting is down.
| Situation | Do this spoke | Do something else first |
|---|---|---|
| Public brand that wants ChatGPT / Perplexity / Google citations | Yes — audit accidental blocks now | — |
| Staging, passworded, or pre-launch | Keep the deny; do not copy it to production | Launch checklist |
| Legal signed a total retrieval opt-out | Confirm the file and the edge match that letter | Do not “open the door” on a Slack vibe |
Money URLs noindex, canonicalized away, or 404 | Fetch still matters, but indexability is the floor | Fix the five pages that should exist |
| No Search Console, no Cloudflare login, no WP admin | You cannot finish the audit | Get access; a screenshot of origin is not the edge |
| You already 200 with citation UAs | Stop rereading robots.txt | Write extractable answers; work the playbook |
If legal wants the door closed, this post is still useful: it tells you to close it on purpose (file + edge + hosts) instead of discovering a 2023 gist and calling it strategy.
Accidental blocks are cheap to create and expensive to notice. The ~30% headline is a top-1,000 GPTBot rate. Your job is the curl.
FAQ
Why do ~30% of sites block AI crawlers without knowing it?
The ~30% figure tracks Originality.AI’s August 2024 top-1,000 robots.txt cut — 35.7% Disallow GPTBot — not a measured “without knowing” share of the whole web. Accidental blocks still happen because Cloudflare AI-scrape toggles and new-domain defaults, WordPress leftovers, inherited Disallow groups, WAF 403s, and Google-Extended vs Googlebot mixups sit outside the meeting where someone would have chosen a citation policy. Count files and HTTP status codes on your hosts; do not treat 35.7% as your SMB base rate.
How do I measure whether why do ~30% of sites block AI crawlers without knowing it is working?
Measure whether your accidental door is closed: robots.txt bytes match the worksheet, citation user-agents get 200s on About and the money URL, and a frozen prompt panel stops returning fetch failures. Then watch cited / mentioned / absent weekly for a month. A prettier robots.txt screenshot is not the KPI.
What usually fails first when teams try this?
They edit origin robots.txt and never open the CDN. Allow: / for OAI-SearchBot plus a Cloudflare AI scrape 403 is the classic miss. Second place is deleting GPTBot Disallow when the real citation kill was OAI-SearchBot or a * staging deny copied to production.
How long does this take to show results?
The fetch fix is same day once someone has dashboard access. ChatGPT search systems are documented at about 24 hours to adjust after a robots.txt change; other vendors differ, so re-curl and re-prompt for a week. Citation selection after you are fetchable still takes weeks of extractable pages, not one Friday toggle.
What should I skip if I only have a week?
Skip a new training-vs-citation philosophy offsite, a 40-token gist, and a content calendar. Curl the file, curl the citation UAs, screenshot the WAF, and reverse leftovers. Leave signed training Disallows alone until legal speaks.
When is this not worth doing yet?
When the property is supposed to be unindexed, when legal already signed a total retrieval opt-out, or when you cannot log into the CDN that actually answers the bot. Fix access and intent first. Opening a door you meant to lock is not an AEO win.
CTA
If citation crawlers never got a 200, the content sprint is theater. Audit the file, the plugin, and the edge, then write for the engines that can actually fetch you.
Lane: /visibility · Book a visibility audit.
What questions does this article answer?
- Why do ~30% of sites block AI crawlers without knowing it?
- The ~30% figure tracks Originality.AI’s August 2024 top-1,000 robots.txt cut — 35.7% Disallow `GPTBot` — not a measured “without knowing” share of the whole web. Accidental blocks still happen because Cloudflare AI-scrape toggles and new-domain defaults, WordPress leftovers, inherited `Disallow` groups, WAF 403s, and `Google-Extended` vs `Googlebot` mixups sit outside the meeting where someone would have chosen a citation policy. Count files and HTTP status codes on *your* hosts; do not treat 35.7% as your SMB base rate.
- How do I measure whether why do ~30% of sites block AI crawlers without knowing it is working?
- Measure whether **your** accidental door is closed: `robots.txt` bytes match the worksheet, citation user-agents get 200s on About and the money URL, and a frozen prompt panel stops returning fetch failures. Then watch cited / mentioned / absent weekly for a month. A prettier robots.txt screenshot is not the KPI.
- What usually fails first when teams try this?
- They edit origin `robots.txt` and never open the CDN. `Allow: /` for `OAI-SearchBot` plus a Cloudflare AI scrape 403 is the classic miss. Second place is deleting `GPTBot` Disallow when the real citation kill was `OAI-SearchBot` or a `*` staging deny copied to production.
- How long does this take to show results?
- The fetch fix is same day once someone has dashboard access. ChatGPT search systems are documented at about 24 hours to adjust after a robots.txt change; other vendors differ, so re-curl and re-prompt for a week. Citation *selection* after you are fetchable still takes weeks of extractable pages, not one Friday toggle.
- What should I skip if I only have a week?
- Skip a new training-vs-citation philosophy offsite, a 40-token gist, and a content calendar. Curl the file, curl the citation UAs, screenshot the WAF, and reverse leftovers. Leave signed training Disallows alone until legal speaks.
- When is this not worth doing yet?
- When the property is supposed to be unindexed, when legal already signed a total retrieval opt-out, or when you cannot log into the CDN that actually answers the bot. Fix access and intent first. Opening a door you meant to lock is not an AEO win.
Last reviewed — Originality.AI top-1,000 dashboard, PPC Land 3 Aug 2024 report of that cut, Cloudflare AIndependence toggle post, Cloudflare 1 Jul 2025 Content Independence Day / press default, Google common-crawlers Google-Extended entry, WordPress Reading settings, OpenAI bots overview checked 2026-09-05.
AI Visibility
AI Visibility Cannabis visibility when the ad accounts are banned
Google and Meta will not take the usual spend. The models still answer dispensary, cultivator, and brand questions — if the site can be read and the cart can clear a 21+ order.
AI Visibility When ChatGPT names the franchise, not your shop
Run the best-HVAC-near-me prompt panel. If the model names a national franchise, fix corroboration and entity facts — not another blog calendar.
AI Visibility How do I get cited by Perplexity specifically
Allow PerplexityBot, put a liftable answer and unique numbers in HTML, then log numbered sources on a frozen prompt panel. There is no bought citation rate.
AI Visibility Does Wikipedia or Wikidata help AI recommend my brand
Wikipedia is not a paid AI lever. Notability plus independent sources decide the page; a real Wikidata item helps entity consistency, not a promotional stub.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.