Spurlock Studios
Contact
Share LinkedIn X
A small text-file card with no glyphs. Thesis: 30 SITES BLOCK AI CRAWLERS.

Sites block AI crawlers without knowing it because the block is usually a default, a leftover, or a confused token — not a signed policy. The “~30%” in the question is the order of magnitude of Originality.AI’s top-1,000 robots.txt cut, reported 3 August 2024 by PPC Land: 35.7% of those sites Disallow GPTBot, up from about 5% at GPTBot’s August 2023 launch. That study does not say 30% of the whole web, and it does not measure whether the owner meant to opt out of ChatGPT search.

This spoke is the accidental-block layer of the Answer Engine Optimization playbook. How you write the file lives in AI crawlers and robots.txt decisions. Unblocking origin is necessary, not sufficient: engines still quote third-party threads, which is why Reddit for AI citations stays a separate job.

The short answer

  • Cite 35.7% of the top 1,000 Disallowing GPTBot (Originality.AI via PPC Land, 3 Aug 2024). Do not flatten that into “30% of all sites” or “30% did it by accident.”
  • “Without knowing it” is a mechanism claim: Cloudflare AI-scrape toggles and new-domain defaults, WordPress plugins and physical leftovers, inherited Disallow groups, CDN/WAF 403s, and Google-Extended vs Googlebot confusion.
  • A robots.txt Allow does not prove fetch. Prove a 200 with the real product token.
  • Training blocks and citation blocks are different tickets. Collapsing them into “block AI” is how marketing ships a quarter of answer pages that ChatGPT search never retrieves.
  • Audit the live file and the edge this week. Rewrite the philosophy after the crawler can actually load About.
Claim in the wildWhat a dated source actually measuredWhat it does not measure
“~30% of sites block AI crawlers”35.7% of the top 1,000 Disallow GPTBot (Originality.AI / PPC Land, 3 Aug 2024)The long tail; intent; citation crawlers
“Everyone on Cloudflare is blocked”New Cloudflare domains start from a permission default (1 Jul 2025); a one-click scrape toggle existsYour specific zone’s saved preference
“WordPress blocks GPTBot out of the box”Core virtual robots.txt Disallows /wp-admin/, not AI tokensPlugin, host, and physical-file overlays
“We blocked Google AI so Search is fine”Google-Extended is a product token, not GooglebotWhether someone Disallowed Googlebot by mistake

Where does the ~30% figure actually come from?

It comes from a popular-site robots.txt study, then got rounded in conversation. Treat the question as a pointer to that cut, not as a second unpublished census.

Originality.AI’s AI-bot blocking dashboard tracks which of the top 1,000 sites Disallow named AI tokens. PPC Land’s 3 August 2024 write-up of that program reported:

Token (that cut)Share of top 1,000 with a DisallowWhat that token is for
GPTBot35.7% (was ~5% at Aug 2023 launch)OpenAI training crawl
CCBot22.1%Common Crawl dump
Google-Extended13.6%Gemini training / listed grounding token
ChatGPT-User12.7%User-initiated ChatGPT fetch
Newer search twins (OAI-SearchBot, ClaudeBot, anthropic-ai)Reported in the 1–10% bandMixed jobs — do not treat the band as one bot

That is why “~30%” shows up in the title query. 35.7% rounded, on GPTBot, on top 1,000. It is a real dated public number. It is the wrong number if you mean “our SMB site” or “ChatGPT search.”

Cloudflare’s own network sample in the AIndependence post is a different metric: edge block or challenge, not robots.txt text.

Cloudflare rank band (that post’s June sample)% of those properties accessed by AI bots% blocking or challenging those requests
Top 1080.0%40.0%
Top 1,00053.2%8.8%
Top 1,000,00038.73%2.98%

Decision list when someone drops “30%” in a meeting:

  1. Top 1,000 or the whole web?
  2. GPTBot Disallow, or any AI token, or an edge 403?
  3. Training, citation, or user-fetch?
  4. Did anyone in the room choose it, or did a default?

If you cannot answer those four, you do not have a statistic. You have a mood.

What does “without knowing it” actually mean?

It means the owner would not sign the same sentence legal would write. The HTTP 403 still happens.

There is no public study I will invent that says “X% of blocks are accidental.” Originality.AI counted files. Cloudflare counted responses. Intent is an audit finding on your zone.

LayerWhat the operator thinks they shippedWhat the crawler actually hitsTypical owner
CDN “AI Scrapers and Crawlers” toggle“Stop scrapers”Known AI user-agents, including search twins on some presetsWhoever clicked Security → Bots
New Cloudflare domain default“We just put DNS on Cloudflare”Training (and, depending on the year’s preset, more) blocked until someone opts inwhoever ran signup
WordPress plugin / physical robots.txt“Yoast is handling SEO”Stale Disallow groups the plugin UI no longer showswhoever migrated the site
Inherited gist“We blocked GPTBot in 2023”OAI-SearchBot and ChatGPT-User rode along in the pastean agency, a gist, a security pack
User-agent: * + Disallow: /Staging hygieneEvery bot that falls through to *, including Googlebota launch checklist that never got reversed
Google-Extended vs Googlebot“No Gemini training” or “no Google AI”Search crawl killed, or training token toggled while Search was the real askwhoever Googled a snippet
WAF / bot-fight / rate-limit“Block bad bots”Citation crawlers look like scrapers: bursty, non-browser UAsecurity, not marketing
  • Marketing can name which engines should still cite us
  • Legal can name which vendors may train
  • Those two lists are not one checkbox
  • Someone has curl’d the live host this month

If marketing cannot name ChatGPT search vs training, you are already in the accidental bucket — even if the file looks “intentional.”

How does a Cloudflare AI scrape toggle block you?

At the edge, before origin, before robots.txt can help. Cloudflare’s AIndependence post documents a one-click AI Scrapers and Crawlers control under Security → Bots, including on the free tier. The 1 July 2025 press release and Content Independence Day blog then changed the default for new domains: block AI crawlers accessing content without permission or compensation unless the owner opts in. Cloudflare’s own 1 July 2026 recap states that for new domains, AI training crawlers are blocked by default unless the owner chooses otherwise.

That is how a site “blocks AI” without a robots.txt meeting. DNS moved. The toggle was on. Nobody in marketing was in the room.

ControlWhere it livesWhat it can do that robots.txt cannotFailure if you ignore it
AI Scrapers and Crawlers toggleCloudflare dashboard, Security → Bots403 / challenge known AI fingerprints even when Allow: /Citation crawlers never see the file
New-domain permission default (1 Jul 2025)Zone onboardingStarts from deny-until-chosenA 2025 rebuild inherits a block the 2019 site never had
Managed robots / preference syncCloudflare-managed robots.txt fragmentsWrites Disallow groups you did not type in WordPressTwo sources of truth; the edge still wins
Pay Per Crawl / 402 pathsLater Cloudflare crawl-control productsCharge or refuse instead of a quiet 200You think you are “open” because the HTML file on origin is open

Procedure when the zone is on Cloudflare:

  1. Open the zone → Security → Bots (and any later AI bot policies / Crawl Control screen your plan shows).
  2. Write down the saved state: allow, block, block-on-ads, or “never touched.”
  3. Fetch https://<host>/robots.txt from an off-network client. Expect text/plain and 200, not a challenge HTML page.
  4. curl -A with the citation token you care about (OAI-SearchBot, PerplexityBot, Claude-SearchBot) on About and the money URL. Expect 200.
  5. If robots.txt says Allow and curl is 403, the toggle is the bug. Do not rewrite copy yet.
  • Screenshot the Bot / AI policy page with the date
  • Diff that screenshot against last quarter — defaults move
  • Record whether the zone was created before or after 1 July 2025
  • Do not treat “Cloudflare is on” as “we chose a training stance”

The same press release says more than one million customers chose the one-click block after it shipped. That cohort is not “without knowing it.” The accidental cohort is everyone who inherited a default, cloned a zone, or never opened Security → Bots after a redesign. Do not quote the million as your ~30%.

SituationKnowing?What to write in the ticket
Marketing asked security to “stop AI scraping”Yes, possibly over-broadSplit training vs citation; re-test search twins
New domain after 1 Jul 2025, nobody opened the AI screenNoDefaults did the work
Agency flipped the toggle during a bot incidentMaybeScreenshot + restore citation tokens
UI renamed to AI Crawl Control / bot policiesNo if nobody re-read itFind the current screen; do not search last year’s tutorial

I have shipped hundreds of production sites. The Cloudflare miss is the one that survives a content sprint because nobody curls with the bot UA.

How do WordPress plugins and host panels do it?

Core WordPress does not ship a GPTBot Disallow. The virtual robots.txt Disallows /wp-admin/ (with Allow for admin-ajax.php) and leaves AI tokens alone. The accidents sit on top of that.

WordPress Reading settings still matter. Discourage search engines from indexing this site is the launch leftover. Since WordPress 5.3 it injects a noindex robots meta (when wp_head runs). Through 5.2, with no physical file, hits to robots.txt could return User-agent: * / Disallow: /. A site that launched on 5.2 and later grew a physical file can still be serving that staging deny.

SourceWhat it changesHow you miss itWhat to do
Virtual WP robots.txt/wp-admin/ onlyYou assume a plugin is the fileFetch the live URL; do not trust the editor screenshot
Physical robots.txt in web rootOverrides the virtual file and most plugin filtersA migration, an old Yoast “create file,” an agency dropDiff disk vs dashboard; one file wins
SEO plugin file editor (Yoast / Rank Math class)Edits physical or virtual, depending on productUI shows rules the crawler never sees, or the reverseFetch https://host/robots.txt incognito
“Block AI bots” plugin / one-click gistWrites Disallow: / under GPTBot and often the search twinsSecurity installed it during a malware scareSplit training vs citation before you delete
Host “bot protection” / WAF packEdge or origin 403 on non-browser UAsPanel says “good bots allowed”; AI tokens are not in that listcurl with the token; do not trust the marketing copy on the pack
Reading → Discourage indexingnoindex (5.3+) and/or a historic * DisallowStaging checkbox copied to productionUncheck; confirm HTML robots and the file

Wordfence-class firewalls are a maybe, not a documented default AI opt-out. Rate-limit and “block user-agents matching bot” rules can throttle GPTBot and OAI-SearchBot the same way they throttle scrapers. Do not claim Wordfence “blocks AI.” Claim this: if your firewall treats bursty non-browser clients as hostile, citation crawlers look hostile.

  • Settings → Reading: Discourage indexing is off on production
  • curl -sI https://host/robots.txt is 200 text/plain
  • Bytes on that URL match the file you think you edit
  • No second file on www vs apex
  • Security plugin UA rules do not match /bot/i unless you meant to

Plugin UI is not the crawler. The crawler gets bytes.

How do hosting defaults and leftover Disallows stack?

They stack. The crawler sees the strictest door it hits first, not the nicest file in WordPress.

A leftover Disallow on origin plus a CDN “block AI” rule plus a host WAF is three independent nos. Opening one layer does not open the others. That is why teams “fixed robots.txt” and still have zero ChatGPT search citations.

Layer (outside → in)Typical default or leftoverWhat a 200 on this layer still does not prove
DNS / CDN (Cloudflare, similar)New-zone AI scrape default; old “block likely bots”Origin robots.txt policy
Host WAF / “bot protection” packUA or rate rules aimed at scrapersThat citation IPs are allowlisted
App firewall (Wordfence-class)Burst limits, /bot/ UA rulesThat the request was even allowed at the edge
Physical robots.txt2023 gist, staging * deny, agency pasteThat the CDN serves this file
Virtual WP / plugin editorRules the disk file silently overridesAnything, if a physical file exists
HTML noindex / robots metaReading → Discourage leftoverrobots.txt at all — Search still skips the URL

Order of operations when two people “already checked robots”:

  1. Name the host the crawler actually requests (www vs apex).
  2. Name the CDN account that answers that host.
  3. Fetch /robots.txt through that CDN, not from origin SSH.
  4. Fetch a page with the citation UA through that CDN.
  5. Only then open WordPress.
Leftover Disallow you still find in 2026Usually landed viaCitation risk if left alone
User-agent: * / Disallow: /Staging, “privacy” checkbox era, launch panicCatastrophic — Search and AI retrieval
GPTBot + ChatGPT-User + OAI-SearchBot as one block2023–2024 “block ChatGPT” gistTraining and search twins
Google-Extended copied as GooglebotSloppy snippetGoogle Search
CCBot / Bytespider onlyReasonable scrape hygieneLow, if search twins are open
Empty or HTML robots.txtPlugin 404 page with status 200Cooperating bots cannot read rules
  • One spreadsheet row per public host: CDN, origin, file SHA, last curl date
  • “We use Cloudflare” is not a row — the zone setting is the row
  • Leftover Disallows get a keep/drop owner, not a vibe
  • After any host migration, re-run the stack; defaults reset

Hosting did not secretly publish a 30% statistic. Hosting is how a site joins whatever statistic already exists without a policy meeting.

How do inherited robots.txt files silently Disallow?

A Disallow does not expire. 2023 gists still win in 2026 if nobody deleted the group.

The usual inheritance paths:

InheritanceHow it landsWhat it usually DisallowsWhy nobody “knows”
Agency 2023 “block GPTBot” pasteCopied into Yoast / a physical fileGPTBot, often ChatGPT-User, sometimes OAI-SearchBot the week SearchGPT shippedThe SOW said “protect content”
Theme or boilerplate repoShipped in public/robots.txtA catch-all * or a long AI listDevelopers never opened it after launch
Staging → production copyDisallow: / under *Everything that honors *Launch day never reversed the deny
Apex vs wwwOnly one host got the cleanupThe other host still deniesEach host has its own file
CDN “managed robots” plus origin fileTwo documentsThe stricter edge policyMarketing edited origin; Cloudflare serves something else

Leftover patterns to grep for — then decide, do not panic-delete:

User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: *
Disallow: /
PatternCitation consequenceTraining consequenceFirst fix
GPTBot Disallow onlyChatGPT search can still be allowed if OAI-SearchBot is openFuture OpenAI training collection should stopKeep or drop as policy, not as an accident
OAI-SearchBot DisallowChatGPT search answers should drop youTraining is a different tagThis is the accidental citation kill if you meant “no training”
ChatGPT-User DisallowLive user fetches may fail; OpenAI notes robots.txt may not applyNot your training controlDo not use this as the Search opt-out
* + Disallow: /Search engines that fall through to * are outSameReverse staging; add explicit Allows if you must keep a tight *
Googlebot DisallowGoogle Search crawl/index path is woundedGemini token is unrelatedRestore Googlebot unless the site is meant to be private

Numbered cleanup:

  1. Save the live file with a date stamp.
  2. Highlight every User-agent group that names an AI or * deny.
  3. Label each group training, citation, user-fetch, or Search.
  4. Delete or split groups that were never a signed decision.
  5. Re-fetch. Then — and only then — write the policy file the other spoke describes.

Do not “simplify” by putting every token under *. That is how Google Search dies in the same commit as Bytespider.

How does Google-Extended vs Googlebot confusion opt you out?

People hear “Google AI” and Disallow the wrong token.

Google’s common crawlers list is explicit: Google-Extended has no separate HTTP user-agent. Crawling uses existing Google UAs; the robots.txt token is a control for whether crawled content may be used for listed Gemini training and grounding uses. Google states Google-Extended does not impact inclusion in Google Search and is not a ranking signal.

Googlebot is the Search crawler. Disallow Googlebot and you are not “opting out of AI Overviews.” You are opting out of the Search crawl/index path that Overviews still sit on.

What someone typedWhat they thought it didWhat Google documentsTypical wreckage
User-agent: Google-Extended + Disallow: /“No Google AI” / “no AI Overviews”Gemini training / listed grounding opt-out; not Search rankingOverviews can still use Search-indexed snippets unless you also change snippet rules
User-agent: Googlebot + Disallow: /“Just the AI crawler”Search crawl/index for those URLsRankings and Overview eligibility both starve
User-agent: Google (not a real group token you meant)“Catch all Google”Groups match tokens, not vibesMiss or over-match; test, do not guess
CDN “block AI” that includes Google user-agents“Scrapers only”Depends on the rule packGooglebot 403 looks like a ranking collapse
  • Googlebot is allowed on public URLs you want in Search
  • Google-Extended is a legal ticket, written separately
  • Nobody used Google-Extended as an AI Overviews off switch
  • Snippet controls (nosnippet, max-snippet) are a different ticket if the actual goal was Overview text

Originality.AI’s 13.6% Google-Extended Disallow rate (same Aug 2024 top-1,000 cut) is a token rate. It is not “13.6% left Google Search.” Do not quote it that way.

How do CDN WAF and bot-fight rules 403 citation crawlers?

robots.txt is a suggestion to cooperating crawlers. A WAF is a door.

Cloudflare’s AIndependence post is blunt: well-behaved bots honor robots.txt; many scrapers spoof a browser. Their answer is fingerprinting and a Bot Score. The same machinery that catches a liar can catch OAI-SearchBot if you pointed “likely bot” at challenge-or-block and never allowlisted the published ranges.

LayerSymptom in logsWhat it is notFix
Managed AI scrape rule403 with a Cloudflare or WAF challenge bodyA robots.txt DisallowChange the AI bot policy; re-curl
Super Bot Fight / “block likely bots”JS challenge, 403, or empty“Google is fine so AI is fine”Allowlist verified AI ranges or lower the action on known good tokens
Rate limit (plugin or host)429 after the first burst of pathsA policy stanceRaise crawler limits; do not require a browser cookie
Country or ASN blockTimeouts from US model-vendor IPsA content decisionIf OpenAI or Anthropic egress is blocked, citations die quietly
robots.txt itself challengedHTML login wall with status 200A valid rules fileAllow anonymous GET of /robots.txt

Checklist that actually moves the ticket:

  • curl -sI https://host/robots.txt — 200, text/plain, no JS challenge
  • curl -A "OAI-SearchBot" -sI https://host/about/ — 200
  • Repeat for PerplexityBot and Claude-SearchBot if those surfaces matter
  • Repeat from a second network so you are not testing your own allowlist
  • Confirm the UA IP against the vendor’s published JSON when you start blocking impostors

If the file Allows and the edge 403s, you do not have a content problem. You have a door problem. The robots.txt decisions spoke will not save a fetch that never reaches origin.

What failure mode burns a quarter of AEO work?

What breaks: the team ships answer-first pages, FAQ blocks, and entity facts. ChatGPT search and Perplexity never retrieve them. The deck still says “we did AEO.” The cause is a 403 or a leftover OAI-SearchBot Disallow from a 2023 gist.

What it costs: a quarter of writing against a closed door. Brand search does not move. The prompt panel stays “absent.” Leadership concludes “AEO does not work.”

What you do instead:

  1. Freeze 20 prompts you actually want to win.
  2. Before rewriting a word, pass the fetch checklist on homepage, About, and one money URL.
  3. Only then score cited / mentioned / absent / hallucinated.
  4. If fetch fails, stop the content calendar. Fix the door.
  5. After a 200, wait for the vendor’s recrawl window before you declare the copy a failure.
False diagnosisEvidence that it is falseReal diagnosis
“Our writing is not AEO enough”curl 403 with OAI-SearchBotEdge toggle
“We need FAQ schema”File Disallows the search botLeftover robots group
“Google hates us”Googlebot Disallow or noindex leftoverSearch crawl, not Gemini
“Perplexity never cites SMBs”PerplexityBot never got a 200WAF
“Reddit is stealing our citations”Your URLs 403; reddit.com 200sYou are unfetchable; Reddit is not — see Reddit for AI citations after you open the door

Bravery is not a 403 strategy.

How do you audit accidental blocks in one afternoon?

You do not need a new vendor. You need the live host, four curls, and a named owner.

StepCommand or placePassFail
1. Live filecurl -sI https://host/robots.txt200 text/plainChallenge, 403, HTML
2. BytesSave body; grep Disallow and User-agentGroups match the worksheetMystery * deny or search-twin Disallow
3. Citation fetchcurl -A search tokens on About + money URL200401/403/429/challenge
4. Search crawlConfirm Googlebot is not Disallowed; HTML is not noindex on public URLsIndexed path still open“We blocked Google AI” actually hit Search
5. Edge UICloudflare / host WAF / WP firewallPolicy matches legal + marketingToggle on, nobody owns it
6. Host splitRepeat 1–3 on apex and wwwSame policyOne host still staging-denied

Numbered afternoon (90 minutes if the logins exist):

  1. Inventory public hosts (apex, www, docs, blog).
  2. Dump every /robots.txt.
  3. Screenshot CDN AI / bot screens.
  4. Run the curl table.
  5. Write a three-line verdict: fetch OK / file leftover / edge block.
  6. Assign one owner to change one layer. Do not change file and WAF in the same hour unless you can re-test both.
  • Date-stamped robots dump in the ticket
  • Date-stamped curl headers in the ticket
  • Named human for the next change
  • Re-test scheduled after the vendor’s search recrawl window (OpenAI documents about 24 hours for ChatGPT search systems to adjust after a robots.txt change on the bots overview)

If you cannot log into Cloudflare, you cannot finish the audit. Get the login before you hire more words.

Curl recipes that belong in the ticket (replace host and path):

JobCommandYou are looking for
File headerscurl -sI https://www.example.com/robots.txt200, text/plain, no cf-mitigated challenge
File bodycurl -s https://www.example.com/robots.txtReal groups, not an HTML error document
ChatGPT search twincurl -sI -A "OAI-SearchBot" https://www.example.com/about/200 (copy the current UA from OpenAI’s bots page if a vendor requires the full string)
Perplexitycurl -sI -A "PerplexityBot" https://www.example.com/about/200 if that surface matters; confirm the string on Perplexity’s crawler doc
Google Search crawlercurl -sI -A "Googlebot" https://www.example.com/about/200; then confirm HTML is not noindex
Apex vs wwwRepeat the four above on the other hostSame policy, or a documented 301 before the fetch

UA version suffixes change. Match the product token, not a pinned GPTBot/1.4 string. If a 403 appears only on the bot UA and a browser 200s, you found the WAF, not a content bug.

What should you fix first if you want citations?

Fetch, then tokens, then copy. Not the other way around.

PriorityFixSkip until this passes
0Anonymous 200 on /robots.txt and on About / money URL for citation UAsNew blog calendar
1Remove OAI-SearchBot / PerplexityBot / Claude-SearchBot Disallows you did not meanTraining philosophy workshops
2Turn off or split the Cloudflare AI scrape preset so search twins can 200Homepage theater
3Restore Googlebot if someone Disallowed it as “Google AI”Gemini-token debates
4Then write the split policy fileGist pastes from 2023
5Then answer-first HTML on the pages models quotellms.txt as a substitute for fetch

Decision list:

  1. Do we want to be retrieved and named in ChatGPT search, Perplexity, Claude search-style answers, Google Search? If yes, those crawlers need a 200.
  2. Do we want those vendors to train on the public pages? Separate Allow/Disallow. Do not reuse the citation answer.
  3. Is any remaining block a WAF? Fix the WAF, not the H1.
  4. After 200s land, measure the prompt panel. If you are still absent, the problem moved to extractable proof — that is the playbook, not this spoke.

Default for most brands that want recommendations: open the citation door; decide training like an adult. The accident is closing both doors with one “block AI” click.

How do you measure whether the accidental block is gone?

You measure fetch, then inclusion, not a feeling that robots.txt looks cleaner.

SignalToolCadencePass
robots.txt 200 + intended groupscurl + diff vs worksheetEvery security changeBytes match policy
Citation UA 200 on money URLscurl -ASameNo challenge body
ChatGPT search / Perplexity / Claude inclusionFrozen prompt panel, screenshotsWeekly for 4 weeks after the 200Cited or at least retrieved, not “I cannot access that site”
Google SearchSearch Console URL inspection + index coverageAfter Googlebot restoresURL is indexable
Residual mentionsSame panelMonthlyHallucinations drop once the engine can fetch you; Reddit and PR still do their own job

Do not report “we unblocked AI.” Report: OAI-SearchBot 200 on /about/ on 2026-08-14; prompt 7 cited us on 2026-08-21. That is a measurement. A green toggle is not.

Weekly loop after the first clean 200:

  1. Monday: re-curl the citation UAs. If they 403 again, security shipped a rule. Stop scoring copy.
  2. Wednesday: run the frozen prompt panel. Log cited / mentioned / absent / hallucinated. Save the source list.
  3. Friday: URL-inspect the money URL in Search Console if Googlebot was in the mess.
  4. Week 4: decide whether you are still in this spoke (fetch) or in the playbook (selection).
Week after the 200Fetch still the story?Selection now the story?
0–1Yes — recrawl windowsNo — do not rewrite the homepage yet
2–3Only if 403s returnStart watching source panels
4+Only if a new WAF shippedThin About, no table, or Reddit already answers it

If inclusion stays zero after fetch is clean, you have a different post’s problem: thin entity, no citeable table, or engines preferring Reddit threads that already answer the query. Opening the crawler does not mint a citation. It only stops you from donating the slot.

What should you skip if you only have a week?

Skip a new allow/block philosophy, a 40-bot gist, and a homepage redesign. You have one job: prove or kill the accidental door.

  • Curl robots.txt on every public host
  • Curl citation UAs on homepage, About, one money URL
  • Open Cloudflare / host WAF / WP firewall and screenshot AI / bot rules
  • Delete or split only the groups that are obviously leftovers (* sitewide deny, OAI-SearchBot Disallow next to a “we want ChatGPT citations” OKR)
  • Leave GPTBot / ClaudeBot / Google-Extended untouched if legal has not signed — those are policy, not accidents, until someone says otherwise
  • Freeze 10 prompts and run them once after the first clean 200, then again a week later

Do not install a new “AI visibility” plugin that writes more Disallows. Do not flip Cloudflare to “allow all AI” if legal forbade training — split the categories if the dashboard can, or allowlist search tokens only.

A week is enough to stop the self-own. It is not enough to win the category. Say that out loud so nobody expects a citation spike from a 403 fix on Friday.

When is this not worth doing yet?

When the site is meant to be private, when legal has already signed “no retrieval, no training, no user fetch,” or when you cannot fetch your own /robots.txt because basic hosting is down.

SituationDo this spokeDo something else first
Public brand that wants ChatGPT / Perplexity / Google citationsYes — audit accidental blocks now—
Staging, passworded, or pre-launchKeep the deny; do not copy it to productionLaunch checklist
Legal signed a total retrieval opt-outConfirm the file and the edge match that letterDo not “open the door” on a Slack vibe
Money URLs noindex, canonicalized away, or 404Fetch still matters, but indexability is the floorFix the five pages that should exist
No Search Console, no Cloudflare login, no WP adminYou cannot finish the auditGet access; a screenshot of origin is not the edge
You already 200 with citation UAsStop rereading robots.txtWrite extractable answers; work the playbook

If legal wants the door closed, this post is still useful: it tells you to close it on purpose (file + edge + hosts) instead of discovering a 2023 gist and calling it strategy.

Accidental blocks are cheap to create and expensive to notice. The ~30% headline is a top-1,000 GPTBot rate. Your job is the curl.

FAQ

Why do ~30% of sites block AI crawlers without knowing it?

The ~30% figure tracks Originality.AI’s August 2024 top-1,000 robots.txt cut — 35.7% Disallow GPTBot — not a measured “without knowing” share of the whole web. Accidental blocks still happen because Cloudflare AI-scrape toggles and new-domain defaults, WordPress leftovers, inherited Disallow groups, WAF 403s, and Google-Extended vs Googlebot mixups sit outside the meeting where someone would have chosen a citation policy. Count files and HTTP status codes on your hosts; do not treat 35.7% as your SMB base rate.

How do I measure whether why do ~30% of sites block AI crawlers without knowing it is working?

Measure whether your accidental door is closed: robots.txt bytes match the worksheet, citation user-agents get 200s on About and the money URL, and a frozen prompt panel stops returning fetch failures. Then watch cited / mentioned / absent weekly for a month. A prettier robots.txt screenshot is not the KPI.

What usually fails first when teams try this?

They edit origin robots.txt and never open the CDN. Allow: / for OAI-SearchBot plus a Cloudflare AI scrape 403 is the classic miss. Second place is deleting GPTBot Disallow when the real citation kill was OAI-SearchBot or a * staging deny copied to production.

How long does this take to show results?

The fetch fix is same day once someone has dashboard access. ChatGPT search systems are documented at about 24 hours to adjust after a robots.txt change; other vendors differ, so re-curl and re-prompt for a week. Citation selection after you are fetchable still takes weeks of extractable pages, not one Friday toggle.

What should I skip if I only have a week?

Skip a new training-vs-citation philosophy offsite, a 40-token gist, and a content calendar. Curl the file, curl the citation UAs, screenshot the WAF, and reverse leftovers. Leave signed training Disallows alone until legal speaks.

When is this not worth doing yet?

When the property is supposed to be unindexed, when legal already signed a total retrieval opt-out, or when you cannot log into the CDN that actually answers the bot. Fix access and intent first. Opening a door you meant to lock is not an AEO win.

CTA

If citation crawlers never got a 200, the content sprint is theater. Audit the file, the plugin, and the edge, then write for the engines that can actually fetch you.

Lane: /visibility · Book a visibility audit.

FAQ

What questions does this article answer?

Why do ~30% of sites block AI crawlers without knowing it?
The ~30% figure tracks Originality.AI’s August 2024 top-1,000 robots.txt cut — 35.7% Disallow `GPTBot` — not a measured “without knowing” share of the whole web. Accidental blocks still happen because Cloudflare AI-scrape toggles and new-domain defaults, WordPress leftovers, inherited `Disallow` groups, WAF 403s, and `Google-Extended` vs `Googlebot` mixups sit outside the meeting where someone would have chosen a citation policy. Count files and HTTP status codes on *your* hosts; do not treat 35.7% as your SMB base rate.
How do I measure whether why do ~30% of sites block AI crawlers without knowing it is working?
Measure whether **your** accidental door is closed: `robots.txt` bytes match the worksheet, citation user-agents get 200s on About and the money URL, and a frozen prompt panel stops returning fetch failures. Then watch cited / mentioned / absent weekly for a month. A prettier robots.txt screenshot is not the KPI.
What usually fails first when teams try this?
They edit origin `robots.txt` and never open the CDN. `Allow: /` for `OAI-SearchBot` plus a Cloudflare AI scrape 403 is the classic miss. Second place is deleting `GPTBot` Disallow when the real citation kill was `OAI-SearchBot` or a `*` staging deny copied to production.
How long does this take to show results?
The fetch fix is same day once someone has dashboard access. ChatGPT search systems are documented at about 24 hours to adjust after a robots.txt change; other vendors differ, so re-curl and re-prompt for a week. Citation *selection* after you are fetchable still takes weeks of extractable pages, not one Friday toggle.
What should I skip if I only have a week?
Skip a new training-vs-citation philosophy offsite, a 40-token gist, and a content calendar. Curl the file, curl the citation UAs, screenshot the WAF, and reverse leftovers. Leave signed training Disallows alone until legal speaks.
When is this not worth doing yet?
When the property is supposed to be unindexed, when legal already signed a total retrieval opt-out, or when you cannot log into the CDN that actually answers the bot. Fix access and intent first. Opening a door you meant to lock is not an AEO win.
Sources

Last reviewed — Originality.AI top-1,000 dashboard, PPC Land 3 Aug 2024 report of that cut, Cloudflare AIndependence toggle post, Cloudflare 1 Jul 2025 Content Independence Day / press default, Google common-crawlers Google-Extended entry, WordPress Reading settings, OpenAI bots overview checked 2026-09-05.

More from this lane

AI Visibility

All →
Book the audit