Can blocking AI crawlers remove me from ChatGPT answers
No. Blocking GPTBot stops future training crawls. ChatGPT search uses OAI-SearchBot. Trained weights do not rewind, and other sources can still name you.
William Spurlock Founder — Spurlock Studios 34 MIN
No. Blocking AI crawlers does not delete you from ChatGPT the way a noindex plus a Search Console removal request deletes a URL from Google. OpenAI’s crawler overview splits the work: GPTBot crawls content that may be used to train generative AI foundation models; OAI-SearchBot surfaces sites in ChatGPT search features. Those tags are independent. A robots.txt Disallow is a signal about future collection. It does not rewind weights already trained, and ChatGPT can still mention you from other sources, user-triggered fetches, or navigational links.
This spoke is the ChatGPT-specific consequence layer of the Answer Engine Optimization playbook. How you write the file lives in AI crawlers and robots.txt decisions. Why citation is a different scoreboard than rank lives in AEO vs SEO. Do not treat this post as a multi-vendor allow/block table.
The short answer
- Blocking
GPTBottells OpenAI your content should not be used in foundation-model training. It is not a ChatGPT search off switch. - Blocking
OAI-SearchBotis the Search opt-out. OpenAI says opted-out sites will not be shown in ChatGPT search answers, though they can still appear as navigational links. - OpenAI documents about 24 hours for search systems to adjust after a robots.txt change. That clock is for Search, not for erasing trained parameters.
ChatGPT-Userfetches a page when a person asks. OpenAI says robots.txt rules may not apply, and this agent is not used to decide Search inclusion.- ChatGPT can still name you from a third-party search provider, from other pages that mention you, from Atlas link-and-title surfacing, or from whatever the model already knew.
| If you meant | The control OpenAI documents | What you still might see |
|---|---|---|
| Stop future training collection | User-agent: GPTBot + Disallow | ChatGPT search cites, browse fetches, training residue |
| Stop ChatGPT search answers | User-agent: OAI-SearchBot + Disallow | Navigational links; no-search mentions; other-source mentions |
| Stop a live “open this URL” fetch | Not a reliable robots.txt job | ChatGPT-User may still fetch; OpenAI says rules may not apply |
| Vanish from ChatGPT as a brand | No such control on the bots page | Wikipedia, news, directories, reviews, competitors |
Decide which of those four you actually wanted. Then measure that surface. “We blocked AI” is not a surface.
Does blocking GPTBot remove you from ChatGPT answers?
No. GPTBot is the training crawler. ChatGPT answers are not a single pipe, and the search pipe is a different user agent.
OpenAI’s wording, checked 2026-09-05 on the bots overview: GPTBot is used to crawl content that may be used in training generative AI foundation models. Disallowing GPTBot indicates a site’s content should not be used in that training. The same page says each setting is independent — you can allow OAI-SearchBot to appear in search results while disallowing GPTBot.
| Claim people make | What OpenAI actually wrote | ChatGPT consequence |
|---|---|---|
| “We blocked GPTBot so ChatGPT cannot see us” | GPTBot is training crawl | Search and user fetch are other agents |
| “One AI robots tag covers ChatGPT” | OAI-SearchBot and GPTBot tags are independent | Blocking one leaves the other on |
| “A training block also blocks search because they share a crawl” | If both are allowed, OpenAI may reuse one crawl for both jobs | That efficiency note is not “Disallow GPTBot = Disallow search” |
| “ChatGPT will forget us this week” | Not documented on the bots page | Weights already trained are not a robots.txt field |
If both bots are allowed, OpenAI may use results from just one crawl for both use cases to avoid duplicate crawling. Read that as an efficiency note. It is not a reason to treat the tokens as one bot, and it is not a reason to expect a GPTBot Disallow to starve Search.
Run this before you tell legal the brand is “out of ChatGPT”:
- Name the ChatGPT surface you care about: no-search chat, Search answers, live URL fetch, or ads landing-page checks.
- Map that surface to one token:
GPTBot,OAI-SearchBot,ChatGPT-User, orOAI-AdsBot. - Confirm the robots.txt group uses that exact token. Version suffixes change; OpenAI’s examples currently show
GPTBot/1.4andOAI-SearchBot/1.4and say the version number may change. - Prove a fetch against the published IP list for that agent, not against a spoofable UA string.
- Re-run a frozen ChatGPT prompt panel with Search on and Search off. Log cited / mentioned / link-only / absent.
If step 1 was “ChatGPT answers” and step 2 was GPTBot, you filled the wrong ticket.
Three channels, three tickets. Do not merge them in the change request:
| Channel | Fires when | OpenAI agent | robots.txt as the off switch? |
|---|---|---|---|
| Training residue | The model already knows a public fact | None this week — past GPTBot collection plus everything else in the corpus | No. GPTBot Disallow is future-only |
| ChatGPT Search | Search is on and retrieval hits the web | OAI-SearchBot | Yes, for answers. Navigational links are a documented leftover |
| Live browse | A person points ChatGPT at a URL, or a Custom GPT / GPT Action fetches | ChatGPT-User | Unreliable. OpenAI says rules may not apply |
Independence is not a slogan. OpenAI’s example on the bots page is the whole point: allow Search, disallow training, or the reverse. A snippet that only fits this post:
# ChatGPT Search answers — keep if you want to be quoted
User-agent: OAI-SearchBot
Allow: /
# Foundation-model training — policy call
User-agent: GPTBot
Disallow: /
That pair does not hide you from ChatGPT. It stops future training collection while leaving Search eligible. Invert both lines if you want the opposite. How you host the file, how you prove a 200, and the syntax traps live in the robots.txt decisions spoke. This page only cares what ChatGPT still does after you ship.
What does OAI-SearchBot actually control in ChatGPT?
OAI-SearchBot is the Search crawler. It is the token OpenAI tells you to use if you want to appear in — or opt out of — ChatGPT search features.
From the same crawler overview: OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features. Sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from the published IP ranges. For search results, it can take about 24 hours from a robots.txt update for systems to adjust.
| Search outcome | OpenAI’s documented meaning | Your next check |
|---|---|---|
| Shown in ChatGPT search answers | Needs OAI-SearchBot allowed (and a fetchable page) | Cited snippet or quoted passage, not just a URL |
| Navigational link only | Still possible after Search opt-out | Title + URL, no summary from your HTML |
| Absent from Search | Opt-out honored or you were never retrieved | Confirm Search actually ran on that prompt |
Referral with utm_source=chatgpt.com | Publisher FAQ: ChatGPT adds that UTM on referrals | Analytics, not a robots screenshot |
The Publishers and Developers FAQ is blunt: for content to be included in summaries and snippets in ChatGPT, do not block OAI-SearchBot. If you want to be in the answer, this is the allow. If you want to be out of search answers, this is the Disallow. Neither move is GPTBot.
- robots.txt has an explicit
User-agent: OAI-SearchBotgroup, not a nickname - You did not match on
OAI-SearchBot/1.4as if the1.4were a forever contract - Published ranges at openai.com/searchbot.json are allowed at the edge
- Money URLs return 200 to that bot, not a JS challenge
- You waited the documented ~24 hours before declaring Search “unchanged”
Search answers are the one ChatGPT surface OpenAI does give you a crawler opt-out for. Treat navigational links as a remaining hole, not as proof the opt-out failed.
If the leftover is a title-and-URL in Atlas and you actually want that gone too, OpenAI’s publisher FAQ points at noindex — with a crawl requirement that fights a total Disallow.
| Goal | Move that matches OpenAI’s FAQ | Move that fights itself |
|---|---|---|
| Stay in Search answers and snippets | Allow OAI-SearchBot; keep pages crawlable | Disallow Search, then wonder why referrals died |
| Leave Search answers, accept leftover links | Disallow OAI-SearchBot; live with navigational links | Call leftover links a crawler bug |
| Leave Search answers and suppress Atlas link+title | Allow crawl long enough for noindex to be read, then decide on Search | Disallow Search so the crawler cannot read noindex |
| Stop future training only | Disallow GPTBot; leave Search as a separate line | One group named “AI” |
Procedure for the Atlas leftover, in order:
- Confirm the leftover is your URL as a link+title, not a quoted passage from your HTML. Quoted HTML is a Search-answer problem.
- Confirm Search is the surface (Atlas / Search), not a no-search chat mention.
- If you need the tag read,
OAI-SearchBotmust be allowed to fetch that URL. A sitewide Disallow blocks the read. - Ship
noindexon the URLs you do not want surfaced as links. Wait for a fetch. Then decide whether Search answers are still a goal on other URLs. - Do not use
GPTBotfor any step in this list.
Path-level Search policy is allowed by robots.txt generally, but do not invent OpenAI-specific path semantics they did not write. If only the article directory should be out of Search answers, say that in the OAI-SearchBot group and re-check the official page when tokens change.
Why blocking does not rewind already-trained weights
A robots.txt line is not a model-edit API. OpenAI documents what GPTBot may collect going forward. The bots page does not document a way to erase a domain from foundation-model parameters that already exist.
That is the gap operators skip. They ship User-agent: GPTBot / Disallow: /, wait a day, ask ChatGPT “who is [brand],” and treat a remaining mention as a crawler bug. It is usually training residue, a Search hit, or a third-party page. OpenAI has not published a deletion SLA for already-trained weights. I will not invent one.
| Layer | What a Disallow can do | What a Disallow cannot do | Hedge |
|---|---|---|---|
Future GPTBot crawl | Signal: do not use this site’s crawled content for foundation-model training | Wipe parameters from models already shipped | Forward-looking language only |
| ChatGPT no-search answer | Nothing by itself | Stop the model from using prior knowledge of a public brand | No OpenAI “forget this entity” control on the bots page |
| ChatGPT Search | Only if you Disallow OAI-SearchBot | Hide navigational links OpenAI still allows | ~24 hour Search adjustment, not a weight rewind |
| User browse | Unreliable via robots.txt | Guarantee ChatGPT-User never fetches | OpenAI: robots.txt may not apply |
| Other people’s pages | Nothing | Stop Wikipedia, press, reviews, or competitors from naming you | Those hosts have their own robots.txt |
Decision list when someone asks “are we out of the model yet?”:
- If the mention appears with Search off, you are not looking at
OAI-SearchBot. You are looking at prior knowledge, user context, or something else that is not this week’s crawl. - If the mention appears with Search on and a citation to your URL, Search still sees you — check
OAI-SearchBot, notGPTBot. - If the mention cites a different URL, you blocked the wrong host.
- If legal needs “never used in training, including past runs,” that is a contract and a counsel problem. It is not a robots group.
Bravery is not a restore strategy. Neither is a Disallow.
OpenAI’s Atlas note in the publisher FAQ actually reinforces the split: publishers should disallow GPTBot on sites they want excluded from potential training, and OpenAI says it respects that signal for content acquired via users’ interactions in Atlas. That is still a going-forward training signal. It is not a rewind of GPT-family weights already sitting in production. The same FAQ adds: if users opt in to training, webpages that opt out of GPTBot will not be trained on. That sentence is about opt-in user training vs your Disallow. It is still not a delete of weights already shipped.
Worked example — same brand, four questions, Search off vs on. Fake name on purpose; copy the columns.
| Prompt | Search | What you see | What the robots line did |
|---|---|---|---|
| “Who is Example HVAC in Traverse City?” | Off | A paragraph with city, owner, old phone | Nothing this week. Residue or other-source memory |
| Same prompt | On | Citation to examplehvac.com/about | OAI-SearchBot still allowed (or <24h) |
| Same prompt after Search Disallow + 24h | On | Competitor + a directory; your URL as a blue link | Answers opt-out; navigational leftover |
| “Summarize https://examplehvac.com/pricing” | n/a | A live fetch of the page | ChatGPT-User; training Disallow irrelevant |
If row 1 still happens after you blocked every OpenAI token you can name, you did not fail robots.txt. You asked a foundation model a question about a public business. There is no crawler that un-knows a NAP that was already in the weights.
- Legal has been told “future collection” in writing, not “ChatGPT amnesia”
- The branded Search-off prompt is on the baseline sheet
- Nobody is using row 1 as evidence that
OAI-SearchBotis broken - Atlas opt-in training language is not being quoted as a Search control
Can ChatGPT still mention you from other sources?
Yes. Even a clean Search opt-out does not make you unnameable. ChatGPT can still talk about a brand that exists on the rest of the web.
OpenAI’s publisher FAQ is the primary source for one of those paths: if they obtain the URL of a disallowed page from a third-party search provider or by crawling other pages, and they have signals the page is relevant, they may surface just the link and page title in ChatGPT Atlas. If you do not want that, use a noindex meta tag. Catch: for their crawler to read that tag, it must be allowed to crawl the page. A total OAI-SearchBot Disallow can block the tag-read you needed to suppress the leftover link.
| Mention path | Needs your HTML? | Needs OAI-SearchBot allowed on your host? | Typical leftover after a block |
|---|---|---|---|
| ChatGPT search answer quoting your page | Yes | Yes | Goes away after opt-out + ~24h, if it was really Search |
| Atlas / Search navigational link to your URL | Title/URL only | Not necessarily | Can remain via third-party URL discovery |
| Citation to Wikipedia, Crunchbase, a newspaper, a directory | No | No | Survives your robots.txt entirely |
| Competitor comparison page that names you | No | No | You are a string on someone else’s site |
| No-search chat using prior training | No | No | Weights do not rewind |
| User pastes your URL into the thread | The paste | No | ChatGPT-User may fetch; robots.txt may not apply |
Other-source mentions are why “we blocked crawlers” and “ChatGPT still said our name” can both be true. Your robots.txt does not govern nytimes.com, a Reddit thread, or a partner bio.
- Prompt panel logs the cited domain, not just “they mentioned us”
- You separately scored Search-on vs Search-off
- You checked whether the leftover URL is yours or a third party
- If the leftover is a navigational link to you, you read the Atlas
noindexnote before tightening crawl further - You did not promise leadership “unmentionable” as a robots.txt outcome
If the leftover citation is a trade pub, the fix is a correction request to that pub — or better on-site facts so the next retrieval prefers you. Blocking GPTBot will not unpublish someone else’s article.
Classify leftovers before you open a crawler ticket. Fake domains on purpose:
| What ChatGPT showed | Cited host | Your robots.txt relevant? | Next owner |
|---|---|---|---|
Quoted your /pricing table | you.com | Yes — Search | Eng: OAI-SearchBot group + 200s |
| Named you, linked wikipedia.org | wikipedia.org | No | Entity / PR, not robots |
| Named you, linked a G2 or Angi profile | review site | No | Profile accuracy |
| Named you, no link, Search off | none | No | Expectation management; weights |
| Linked you.com with title only | you.com | Partial — Atlas leftover | noindex path, not GPTBot |
| Summarized a URL the user pasted | you.com | ChatGPT-User maybe | WAF vs accept live reads |
Pages to inspect first when the leftover is yours: /, /about, /pricing or /services, the money location page, and the URL currently ranking for the buyer question. A homepage-only robots check is how Search keeps quoting an old article while you congratulate yourselves on /robots.txt.
What ChatGPT-User fetches that a Disallow may not stop
ChatGPT-User is not a crawler in the search-index sense. It is a user-triggered fetch.
OpenAI’s bots page: when users ask ChatGPT or a Custom GPT a question, it may visit a web page with a ChatGPT-User agent. Users may also hit external apps via GPT Actions. ChatGPT-User is not used for crawling the web in an automatic fashion. Because the actions are user-initiated, robots.txt rules may not apply. ChatGPT-User is not used to determine whether content may appear in Search. Use OAI-SearchBot for Search opt-outs and automatic crawl.
| Behavior | ChatGPT-User | OAI-SearchBot | GPTBot |
|---|---|---|---|
| Automatic sitewide crawl | No (OpenAI: not automatic) | Yes — Search features | Yes — training collection |
| Honors robots.txt | May not | Yes, as the Search tag | Yes, as the training tag |
| Controls Search answers | No | Yes | No |
| Fires because a person asked | Yes | No | No |
| Example token in UA (version may change) | ChatGPT-User/1.0 | OAI-SearchBot/1.4 | GPTBot/1.4 |
| IP list | chatgpt-user.json | searchbot.json | gptbot.json |
If a customer pastes https://yoursite.com/pricing into ChatGPT and asks it to read the page, that fetch is this agent. A Search Disallow does not mean the summary never happens. A training Disallow does not mean the summary never happens. OpenAI told you robots.txt may not apply here.
Procedure when logs show ChatGPT-User:
- Confirm the IP against chatgpt-user.json. UA spoofing is cheap.
- Do not rewrite the Search policy because of this line. Search is
OAI-SearchBot. - If you need a hard block on live fetches, that is an edge/WAF conversation against published ranges — and it will break the “user asked ChatGPT to read us” path on purpose. Know that cost.
- Re-check the bots page before you pin
ChatGPT-User/1.0in a WAF rule as if the1.0were eternal. OpenAI already flags version drift on the sibling agents.
User-triggered traffic is not proof your Search opt-out failed. It is proof a human pointed ChatGPT at a URL.
Custom GPTs and GPT Actions sit in this bucket because OpenAI listed them on the same ChatGPT-User row. A prospect running a “research this vendor” GPT that fetches your pricing page is not OAI-SearchBot indexing you. A GPT Action hitting your API is not foundation-model training. Collapsing those log lines into “ChatGPT crawled us” is how teams undo a Search Allow they still needed.
| User-triggered shape | Likely agent | Block with robots.txt? | Product cost if you hard-block at the edge |
|---|---|---|---|
| Paste URL, “summarize this” | ChatGPT-User | OpenAI: may not apply | User cannot get a live read of your page |
| Custom GPT configured to browse your docs | ChatGPT-User | Same hedge | Support / sales GPTs go blind |
| GPT Action calling your origin | ChatGPT-User (OpenAI groups Actions here) | Same hedge | The Action fails; that may be intended |
| ChatGPT Search citation with no paste | OAI-SearchBot | Yes — Search tag | You leave Search answers |
- Log lines with
ChatGPT-Userare labeled “user fetch,” not “Search” - Custom GPT traffic is not a training KPI
- WAF rules match the current token and the JSON ranges, not a year-old UA paste
- Sales is warned if you intend to break live “read this page” behavior
The four OpenAI agents mapped to ChatGPT behavior
OpenAI currently documents four user agents on the crawler overview. Only two are answer-surface controls. One is ads. One is live fetch. Tokens and UA strings change; copy from developers.openai.com/api/docs/bots, not from a 2025 gist.
| User agent | Job (OpenAI wording, checked 2026-09-05) | ChatGPT answer lever? | If you Disallow |
|---|---|---|---|
GPTBot | Crawl that may be used in training generative AI foundation models | No — training | Future training collection should stop; Search stays unless you also change Search |
OAI-SearchBot | Surface sites in ChatGPT search features | Yes — Search answers | Not shown in search answers; navigational links still possible |
ChatGPT-User | User actions in ChatGPT / Custom GPTs may visit a page | Indirect — live read of a URL | Unreliable; OpenAI says robots.txt may not apply |
OAI-AdsBot | Validate landing pages submitted as ChatGPT ads; not used to train foundation models | No — ads safety/relevance | Only relevant if you run ChatGPT ads |
When OpenAI fetches robots.txt itself, the UA may include an extra robots.txt marker so you can tell a rules fetch from a content fetch in logs that omit paths. That marker is not a fifth product. Do not write a group named robots.txt.
| Log you have | Treat as | Do not treat as |
|---|---|---|
GPTBot + IP in gptbot.json | Training crawl | ChatGPT Search |
OAI-SearchBot + IP in searchbot.json | Search crawl | Training opt-out proof |
ChatGPT-User + IP in chatgpt-user.json | User-triggered fetch | Search inclusion control |
OAI-AdsBot + IP in adsbot.json | Ad landing-page check | An answers crawler |
| Matching UA, IP not in the JSON | Impostor | Policy evidence |
OAI-AdsBot exists so you do not stuff ads validation into the training or Search argument. OpenAI says it only visits pages submitted as ads, and that collected data is not used to train foundation models. If you do not run ChatGPT ads, leave it alone.
Match the product token (GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot), not a pinned full string. OpenAI’s own examples say the version number may change, and OAI-SearchBot rides a full Chrome-looking UA with the token appended. A naive “starts with Mozilla/5.0” filter will misfile it as a human.
What still appears after you think you opted out
You need a test matrix, not a feeling. Same prompts. Same ChatGPT plan class. Search on and Search off. Write down the cited URL.
| Test | What “removed” looks like | What a leftover usually is |
|---|---|---|
| Search on, “best [category] for [job]” | Your domain is not in the answer body | Competitor, directory, or you still allowed Search |
| Search on, branded query | No summary scraped from your HTML | Navigational link; Atlas third-party URL path |
| Search off, branded query | You were never counting on crawl for this | Training residue or prompt context |
| User pastes your URL | Page fetch fails or is empty | ChatGPT-User; robots.txt may not apply |
| Referral analytics | utm_source=chatgpt.com dries up on your URLs | Search answers were the click path; mentions can continue with zero clicks |
Do this as a numbered lab, not a Slack poll:
- Freeze 20 prompts. Ten branded, ten category. Save them in a sheet. Do not improvise live.
- Run each with Search enabled. Screenshot citations. Record domain, not vibes.
- Run the branded ten with Search disabled. Label the column
no-search. - Change only one robots group (
GPTBotorOAI-SearchBot). Wait the Search clock if you touched Search (~24 hours per OpenAI). - Re-run the same 20. Diff cited / mentioned / link-only / absent.
- If Search-off mentions did not move, stop telling people the training Disallow “removed us from ChatGPT.” It did not claim that job.
| Panel result after a Search Disallow | Likely interpretation | Wrong interpretation |
|---|---|---|
| Search-on citations to you gone; Search-off mentions remain | Search opt-out working; weights/other sources remain | “OpenAI ignored robots.txt” |
| Search-on still cites your HTML | Bot still allowed, WAF 403, wrong host, or <24h | “GPTBot is the Search bot” |
| Search-on cites a newspaper that quotes you | Other-source mention | “Our Disallow failed” |
| Navigational link, no snippet | Documented leftover | “Answers opt-out is broken” |
utm_source=chatgpt.com still hits a URL you Disallowed | Cached client, different host, or ad/path you did not cover | Ignore it without checking the host |
A mention without a link is not a Search citation. A Search citation without your HTML is not a GPTBot event. Label the columns or you will “fix crawlers” forever.
The failure mode: celebrating a block that did not hide you
The expensive failure is not “ChatGPT can still say our name.” The expensive failure is shipping a block, declaring the brand invisible, and then either (a) losing Search citations you still wanted or (b) telling the board the privacy work is done when ChatGPT still recommends you from Wikipedia.
I have watched this movie on visibility audits: legal pastes a “block AI” snippet, marketing hears “we’re out of ChatGPT,” sales keeps getting “ChatGPT told me to call you,” and nobody can reconcile the three stories because they were never the same surface.
| Failure | What it costs | What you do instead |
|---|---|---|
Disallow GPTBot, expect Search to go dark | Zero Search change; false confidence | Disallow OAI-SearchBot if Search answers are the goal |
Disallow OAI-SearchBot, still wanted citations | You opted out of the one ChatGPT surface that quotes your pages | Allow Search; decide training separately |
User-agent: * Disallow: / as “AI block” | You may have opted out of far more than ChatGPT | Split tokens; read the robots spoke before you paste |
Block ChatGPT-User in robots.txt and call it Search | Search policy unchanged; live fetches may continue | Use OAI-SearchBot for Search; treat User as a WAF decision |
| Ignore other-source mentions | Leadership thinks the work failed | Log cited domain; fix or accept third-party pages |
| Skip the ~24 hour Search clock | You revert the file because “nothing happened overnight” | Timestamp the change; re-run the panel the next day |
Cost is not theoretical: you either donate the citation slot to a competitor who left OAI-SearchBot allowed, or you burn a quarter arguing about a training Disallow that never claimed to hide a no-search mention. Spurlock Studios visibility work starts by naming the surface. I will not invent a “percent of ChatGPT answers from training vs Search” figure. OpenAI does not publish that mix. Your panel is the measurement.
- Written stance: training allow/deny (
GPTBot) - Written stance: Search answers allow/deny (
OAI-SearchBot) - Those two stances are signed by different owners if they disagree
- Prompt panel exists before the robots change, so you have a baseline
- Nobody used “ChatGPT” as a synonym for
GPTBotin the change ticket
If legal and marketing cannot sign different answers for training vs Search, you do not have a ChatGPT policy. You have a mood.
How to measure whether a ChatGPT block is actually working
Measure the surface you changed. Do not use GPTBot hit volume as a ChatGPT-answers KPI.
| KPI | Proves | Does not prove |
|---|---|---|
GPTBot 200s drop after Disallow | Training crawler is honoring the file (if IPs match gptbot.json) | ChatGPT stopped naming you |
OAI-SearchBot 200s drop after Disallow | Search crawler stopped fetching | No-search mentions stopped |
| Search-on panel: your URL absent from answers | Search answers opt-out is doing its documented job | You are unmentionable |
| Search-off panel: branded mention gone | Maybe; also maybe the prompt or the model week changed | The robots file did it |
utm_source=chatgpt.com sessions | Search referral clicks | Mentions with no click |
| Navigational-link count in Atlas / Search | Leftover path OpenAI still documents | Opt-out “failed” |
Build the sheet once:
- Date, prompt, ChatGPT plan, Search on/off, cited domains, your URL role (
cited/mentioned/link-only/absent), robots group last changed, hours since change. - Freeze the account class. Free vs paid vs workspace can change whether Search even runs. Do not mix them in one column and call it science.
- Re-run weekly for four weeks after a Search change. One overnight pass is how people revert a working opt-out.
- Keep training logs in a second sheet. Mixing them is how
GPTBotvolume becomes a fake answers dashboard.
| Cadence | Artifact | Owner |
|---|---|---|
| Before any robots change | 20-prompt baseline, Search on and off | SEO / visibility |
+24 hours after OAI-SearchBot change | Same 20, Search on | SEO |
| Weekly for a month | Same 20; note model-week drift as a hedge | SEO |
| After every WAF change | IP-verified 200/403 for the four OpenAI JSON lists | Eng |
| Quarterly | Re-read the official bots page; tokens and UA versions change | Eng + legal |
I will cite receipts I can defend: SEO certified since 2021, 500+ automations built, 20,000+ hours architecting agentic systems, 35,000+ hours saved for clients, hundreds of production sites. I will not cite a fabricated “ChatGPT forget curve.” If a vendor dashboard claims one, treat it as marketing until it maps to OpenAI’s published agents.
Example log (fake domains — copy the columns):
| Date | Prompt | Search | Cited domains | Your URL | Hours since robots change | Token last changed |
|---|---|---|---|---|---|---|
| 2026-08-02 | “best HVAC company Traverse City” | On | directory.com, rival.com | Absent | 0 (baseline) | none |
| 2026-08-02 | “who is Example HVAC” | Off | none | Mention, no link | 0 | none |
| 2026-08-04 | Category prompt | On | rival.com | Link-only | 30 | OAI-SearchBot Disallow |
| 2026-08-04 | Branded, Search off | Off | none | Mention, no link | 30 | OAI-SearchBot Disallow |
| 2026-08-04 | “summarize [pricing URL]” | n/a | you.com (fetch) | Cited via paste | 30 | GPTBot unchanged |
Row 3 moving while row 4 does not is a successful Search opt-out, not a failed training block. Row 5 is ChatGPT-User. If your sheet cannot tell those rows apart, you are not measuring ChatGPT answers. You are measuring folklore.
How long until a robots change shows up in ChatGPT?
OpenAI publishes one clock, and it is for Search: about 24 hours from a robots.txt update for search systems to adjust. That is not a training-rewind clock. That is not a ChatGPT-User clock. That is not a promise that third-party pages update.
| Change you made | Documented wait | What you should see first | What you should not expect |
|---|---|---|---|
OAI-SearchBot Disallow | ~24 hours (OpenAI, search systems) | Search answers stop quoting your HTML | Immediate unmentionability; link leftovers gone |
OAI-SearchBot Allow | ~24 hours | Eligibility to be in Search answers, not a guaranteed cite | Rank-style “we’re #1 in ChatGPT” |
GPTBot Disallow | Not published as a ChatGPT-answers SLA | Training crawler 404/empty if it honors the file | No-search answers going blank this week |
ChatGPT-User robots Disallow | OpenAI: rules may not apply | Maybe nothing | A Search opt-out |
noindex while still allowing crawl | Not given as hours on the bots page | Atlas leftover links may drop if the crawler can read the tag | Working if you also Disallow the crawler that needs to read it |
CDN cache of /robots.txt | Your cache TTL, not OpenAI’s | Bots still seeing yesterday’s file | “OpenAI is slow” while your edge is stale |
Procedure when the panel has not moved:
curlthe livehttps://<host>/robots.txtbytes. Diff against the worksheet. Apex andwwware different files.- Check cache headers. A 24-hour OpenAI clock plus a 24-hour CDN cache is two days, not one.
- Confirm the bot that matters can fetch a money URL from an IP in the matching JSON file.
- Confirm you waited 24 hours after the bytes OpenAI would actually fetch had changed.
- Only then open a “they ignored us” ticket.
OpenAI also notes that when it fetches robots.txt, the UA may include a robots.txt marker. Use that to debug “we never saw the bot” arguments. A rules fetch is not a content fetch. Missing content logs are not proof Search is off if you never allowed the IPs.
Host mixups look like “OpenAI ignored us” and are usually you:
| Host fact | ChatGPT leftover it causes | Fix that is not a new token |
|---|---|---|
Apex has the Disallow; www still Allows Search | Search still quotes www URLs | Same groups on every live host |
CDN cached /robots.txt for 24h | Search clock starts when OpenAI sees new bytes | Purge; confirm curl bytes |
| Staging Disallow merged to production | You opted out harder than the ticket said | Diff production robots against the worksheet |
Money site is app.example.com; robots live on marketing apex | Search still cites the app host | Per-host files. robots.txt is not global |
-
curl -sI https://example.com/robots.txtand thewwwtwin - Cache-Control / CDN TTL recorded next to the OpenAI 24-hour note
- The host in the ChatGPT citation matches the host you edited
- Staging
User-agent: * Disallow: /did not ride a merge to prod
What a one-week ChatGPT-consequence pass looks like
One week is enough to stop lying to yourselves. It is not enough to make a famous brand unnameable.
| Day | Action | Done looks like |
|---|---|---|
| 1 | Write the two stances: training vs Search answers | Signed lines, not a Slack emoji |
| 2 | Baseline 20-prompt panel, Search on and off | Sheet with cited domains |
| 3 | Inventory which OpenAI tokens you currently Allow/Disallow | Exact tokens; no GPT-Bot typos |
| 4 | Change only the token that matches the written stance | One PR, one host at a time |
| 5 | Allowlist published IPs for the agents you still want | 200s, not a JS wall |
| 6 | Wait the Search clock if you touched OAI-SearchBot | Timestamp in the sheet |
| 7 | Re-run the panel; classify leftovers (Search / residue / other source / User) | Three-column postmortem, not “it failed” |
Skip if you only have a week:
- A multi-vendor crawler religion argument
- Rewriting the entire robots.txt tutorial (that is the other spoke)
- Promising the board ChatGPT will “forget” the brand
- Blocking
ChatGPT-Userin robots.txt and calling the ticket done - Collapsing training and Search into one “AI” group
- Measuring success as
GPTBotlog volume
After week one, the operating rhythm is boring: re-read the official bots page when you touch the file, keep the panel frozen, and treat leftover mentions as a classified object. If you want citations, spend the next week on quoteable pages — not on a second Disallow.
What “quoteable” means for this spoke is narrow: the HTML of the cited URL still contains the sentence you are trying to suppress or the sentence you are trying to earn. A Search Disallow on a JS shell changes nothing ChatGPT could quote. A Search Allow on a JS shell also changes nothing. Fetch the money URL as a bot, save the HTML, and grep the claim. If the claim is not in the bytes, you do not have a crawler problem. You have a rendering problem, and AEO vs SEO is the scoreboard shift — citation, not rank.
| Page | Why it shows up in ChatGPT leftovers | First check |
|---|---|---|
/ | Branded queries | NAP vs what no-search ChatGPT recites |
/about | “Who is” prompts | Entity facts vs Wikipedia |
/pricing or /services | Buyer prompts | Is the number in HTML? |
| Location / service area URL | Local “best X” | Address consistency |
| One high-traffic blog URL | Old claims Search still retrieves | Date, facts, noindex if you truly want it out |
When is blocking ChatGPT surfaces not worth doing yet?
Skip the block when you still want to be the cited source, when you have no baseline, or when the thing you hate is a Wikipedia sentence robots.txt cannot touch.
| Condition | Block this week? | Do this first |
|---|---|---|
| You want ChatGPT search to recommend you | No — allow OAI-SearchBot | Make the page quoteable; see the playbook |
| Legal wants no future OpenAI training | Yes — GPTBot only | Do not touch Search in the same PR |
| You have never run a Search-on vs Search-off panel | Not yet | Baseline, then change one token |
| Money URLs are a JS shell or a bot wall | Pointless theater | Server-render the answer; then decide crawl policy |
| The leftover mention is a press hit | Your robots.txt is the wrong tool | Correction, entity facts, or accept the cite |
| You need “unmentionable in no-search ChatGPT” | Not available via these crawlers | Counsel, contracts, and expectations management |
You run ChatGPT ads and blocked OAI-AdsBot by accident | Undo | Ads validation is not training and not Search |
| Nobody owns the two stances | Do not ship a mood | Split owners, then ship |
A visibility audit is the right next step when ChatGPT still names you and three teams each think a different robots line was supposed to stop it. DIY the one-week pass if someone can edit /robots.txt and run 20 prompts. Hire when the argument is political or the leftovers are mixed (Search cites + press cites + no-search residue) and nobody can read the sheet. The lane page is /visibility.
Blocking crawlers is a collection-policy tool. ChatGPT answers are a retrieval-plus-memory product. Those are adjacent. They are not the same switch.
FAQ
Can blocking AI crawlers remove me from ChatGPT answers?
No — not as one switch, and not as a memory wipe. Blocking GPTBot signals that crawled content should not be used in OpenAI foundation-model training. ChatGPT search answers are controlled with OAI-SearchBot; opted-out sites are not shown in those answers but can still appear as navigational links. Already-trained weights do not rewind, and ChatGPT can still mention you from other sources or a user-triggered ChatGPT-User fetch.
How do I measure whether blocking AI crawlers is removing me from ChatGPT answers?
Run a frozen 20-prompt panel with Search on and Search off, and log cited / mentioned / link-only / absent plus the cited domain. A GPTBot hit drop only proves the training crawler changed. A Search-on citation drop after an OAI-SearchBot Disallow is the actual Search-answers test. Wait OpenAI’s documented ~24 hours for search systems before you call the change a no-op.
What usually fails first when teams try this?
They Disallow GPTBot and expect ChatGPT search to go dark, or they paste a User-agent: * deny and call it an AI policy. The other early fail is treating a leftover Wikipedia or press mention as proof OpenAI ignored robots.txt. Classify the leftover: your HTML, a navigational link, a third-party URL, no-search residue, or a user paste.
How long does this take to show results?
OpenAI says search systems can take about 24 hours to adjust after a robots.txt update. That clock is for ChatGPT search features, not for erasing foundation-model weights. Training residue, other-source mentions, and ChatGPT-User fetches have no published “gone from answers” SLA on the crawler overview. Timestamp the live bytes you served, then re-run the same panel.
What should I skip if I only have a week?
Skip promising unmentionability, skip a multi-vendor crawler rewrite, and skip blocking ChatGPT-User in robots.txt as a Search stand-in. Spend the week on two written stances (training vs Search), a 20-prompt baseline, one token change, IP allowlisting for the agents you still want, and a classified leftover list. That is enough to stop guessing.
When is this not worth doing yet?
If you still want ChatGPT search to quote you, do not Disallow OAI-SearchBot. If you have no Search-on vs Search-off baseline, do not ship a block you cannot measure. If the leftover is a third-party page or a no-search mention of a public brand, robots.txt on your host will not finish the job. Fix quoteability and classification first; then decide collection policy.
CTA
If ChatGPT still names you after a crawler block, stop treating GPTBot as a delete key — split training from Search, then measure the leftover.
Lane: /visibility · Book a visibility audit.
What questions does this article answer?
- Can blocking AI crawlers remove me from ChatGPT answers?
- No — not as one switch, and not as a memory wipe. Blocking `GPTBot` signals that crawled content should not be used in OpenAI foundation-model training. ChatGPT search answers are controlled with `OAI-SearchBot`; opted-out sites are not shown in those answers but can still appear as navigational links. Already-trained weights do not rewind, and ChatGPT can still mention you from other sources or a user-triggered `ChatGPT-User` fetch.
- How do I measure whether blocking AI crawlers is removing me from ChatGPT answers?
- Run a frozen 20-prompt panel with Search on and Search off, and log cited / mentioned / link-only / absent plus the cited domain. A `GPTBot` hit drop only proves the training crawler changed. A Search-on citation drop after an `OAI-SearchBot` Disallow is the actual Search-answers test. Wait OpenAI’s documented ~24 hours for search systems before you call the change a no-op.
- What usually fails first when teams try this?
- They Disallow `GPTBot` and expect ChatGPT search to go dark, or they paste a `User-agent: *` deny and call it an AI policy. The other early fail is treating a leftover Wikipedia or press mention as proof OpenAI ignored robots.txt. Classify the leftover: your HTML, a navigational link, a third-party URL, no-search residue, or a user paste.
- How long does this take to show results?
- OpenAI says search systems can take about 24 hours to adjust after a robots.txt update. That clock is for ChatGPT search features, not for erasing foundation-model weights. Training residue, other-source mentions, and `ChatGPT-User` fetches have no published “gone from answers” SLA on the crawler overview. Timestamp the live bytes you served, then re-run the same panel.
- What should I skip if I only have a week?
- Skip promising unmentionability, skip a multi-vendor crawler rewrite, and skip blocking `ChatGPT-User` in robots.txt as a Search stand-in. Spend the week on two written stances (training vs Search), a 20-prompt baseline, one token change, IP allowlisting for the agents you still want, and a classified leftover list. That is enough to stop guessing.
- When is this not worth doing yet?
- If you still want ChatGPT search to quote you, do not Disallow `OAI-SearchBot`. If you have no Search-on vs Search-off baseline, do not ship a block you cannot measure. If the leftover is a third-party page or a no-search mention of a public brand, robots.txt on your host will not finish the job. Fix quoteability and classification first; then decide collection policy.
Last reviewed — OpenAI crawler overview and Publishers and Developers FAQ checked 2026-09-05. Product tokens and user-agent version suffixes change; match the token, not a pinned GPTBot/1.4 string.
AI Visibility
AI Visibility Cannabis visibility when the ad accounts are banned
Google and Meta will not take the usual spend. The models still answer dispensary, cultivator, and brand questions — if the site can be read and the cart can clear a 21+ order.
AI Visibility How do I get cited by Perplexity specifically
Allow PerplexityBot, put a liftable answer and unique numbers in HTML, then log numbered sources on a frozen prompt panel. There is no bought citation rate.
AI Visibility What belongs in an AI visibility monthly retainer vs a one-time audit
A one-time audit is the baseline plus prioritized fixes. A monthly retainer is prompt-panel tracking, entity hygiene, page jobs, and citation recovery.
AI Visibility Does Wikipedia or Wikidata help AI recommend my brand
Wikipedia is not a paid AI lever. Notability plus independent sources decide the page; a real Wikidata item helps entity consistency, not a promotional stub.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.