Google uses the Search foundation
Google states that AI features rely on existing Search systems. There is no special AI schema requirement; index eligibility, snippets and normal Search controls remain relevant.
OpenAI separates search from training
OAI-SearchBot supports search discovery, while GPTBot is a separate control associated with potential model training. Publishers should decide on each purpose independently.
Anthropic and Perplexity publish distinct controls
Anthropic and Perplexity document crawler identities and retrieval behaviour. Robots rules should be paired with live CDN or WAF checks because a policy file alone cannot prove successful access.
DIRECT ANSWER / CRAWLER POLICY
AI crawlers should be controlled by named agent and purpose, not with one blanket rule.
Search discovery, a fetch made because a person asked a question and collection for model development are different activities. Several platforms publish separate agent names precisely so a site owner can make separate decisions. Allowing an automatic search crawler while declining a training crawler can be a coherent policy; treating every agent containing the letters AI as the same visitor is not.
The decision begins in robots.txt but does not end there. The origin, CDN, firewall, bot-management product, rate limiter and application must all deliver the intended public response. A user-agent string can be spoofed, so a security allow-list should combine the declared agent with the platform's current IP publication or verification method. A robots rule expresses publisher preference; it is not proof that the genuine crawler visited, indexed or cited a page.
For a publisher seeking discovery, the practical default is to keep public search and user-requested retrieval available, decide model-development access independently, protect non-public routes with authentication and verify representative URLs from outside the deployment environment. Keep a dated policy register because platform names, purposes and network ranges can change.
THREE PURPOSES / THREE DECISIONS
Name the activity before writing the rule.
The same platform may publish more than one agent. A useful policy records the purpose, desired outcome and acceptable operational cost for each one.- 01
Search discovery
An automatic crawler discovers and refreshes documents for a search or answer index. Blocking it may reduce or remove eligibility for that platform's search experience, but permission still does not guarantee indexing or selection.
Decide: do we want this public material discoverable in this platform's search results? - 02
User-requested retrieval
A fetcher visits because a person asked a product to open or use a page. Some providers state that ordinary robots.txt rules may not apply to this user-directed activity. Authentication, authorization and application security remain the real boundary for private information.
Decide: may a user ask this product to retrieve the same page that any anonymous visitor can open? - 03
Model development
A separate crawler or control token expresses whether public content may be collected for improving or training foundation models. This decision can be independent of search discovery and should be reviewed with the publisher's legal, licensing and commercial policy.
Decide: does the publisher permit this documented development use under its rights and risk policy?
PLATFORM / AGENT MATRIX
Use the current official name, then verify the delivery path.
The table records documented roles as of the review date. Recheck the linked primary source before a policy change; do not copy a static allow-list into a permanent security rule without network verification.POLICY / VERIFICATION WORKFLOW
Move from publishing intent to evidence in five passes.
01 / CLASSIFY
Inventory content and decide what is genuinely public.
Separate public editorial and commercial pages from account areas, intake payloads, private files, preview hosts, internal search results and administrative routes. robots.txt is not access control: a route containing confidential material belongs behind authentication even when every crawler is disallowed.
Record the rights position for books, datasets, images and client material. A crawler policy cannot grant rights the publisher does not hold, and a permissive site-wide rule should not silently expose material whose publication status is uncertain.
02 / CHOOSE
Make an independent decision for each documented purpose.
Write a small decision table: agent, provider, purpose, desired policy, owner, date, source URL and review date. Search access, user retrieval and model development should not inherit one another merely because they come from the same company.
If the objective is discoverability, preserve the automatic search agents needed for that objective. A publisher may still choose a different model-development policy. State the trade-off plainly and obtain the appropriate legal or rights decision where the content warrants it.
03 / IMPLEMENT
Align robots.txt, edge security and the application.
Publish exact agent groups without accidental precedence conflicts. Then inspect CDN bot rules, managed challenges, country blocks, rate limits, origin allow-lists and application middleware. An Allow line cannot override a firewall that returns 403, a challenge page or an empty 200 response.
Use bounded rate controls. A harsh blanket rule can make a platform abandon retrieval, while no protection at all can expose the origin to spoofed traffic. Where a provider publishes IP ranges, combine those ranges with the declared user-agent and refresh them from the official endpoint.
04 / TEST
Verify policy parsing, network delivery and rendered meaning separately.
First test whether the target URL is allowed by the relevant robots group. Then send a named-agent request from outside the origin environment and retain status, final URL, headers, body hash and challenge result. Finally render representative pages to ensure the answer, evidence and links survive.
A synthetic user-agent test proves only how the stack responds to that string from the test network. It does not authenticate a platform visit. Genuine log classification requires the provider's current IP publication or its prescribed DNS verification method.
05 / OBSERVE
Keep access, visit, index and citation as separate records.
A green delivery test establishes technical eligibility at one time. Origin logs may later show a verified visit; webmaster tools may show crawl or indexing; answer observations or platform reports may show a citation. None of these stages should be backfilled from another.
Review the policy when official documentation, crawler names, IP feeds, firewall products, hosting topology or publication rights change. Save the former version and retest before reporting the new state as live.
MINIMUM POLICY REGISTER
A usable crawler policy has six records beyond robots.txt.
The public file is only one expression of a wider operational decision.Purpose register
Map every named agent to search discovery, user-requested retrieval, model development or another documented use. Do not infer purpose from the brand name alone.
Content boundary
List the public route families included in the policy and the private systems protected by authentication. Include subdomains and alternate hosts explicitly.
Primary-source record
Keep the official documentation URL, date reviewed, agent token, verification endpoint and material caveats. Set a review date for volatile details.
Edge implementation
Record CDN, WAF, rate-limit, geographic and origin rules that can override apparent robots permission, along with the responsible owner.
Acceptance evidence
Retain parser result, named-agent HTTP response, final body signature, resource failures and a representative rendered capture for each important route class.
Outcome ledger
Log verified visits, indexing records, citations and referrals separately. Never convert a successful access test into an unstated claim about downstream use.
POLICY RED FLAGS
Eight mistakes that break either discovery or security.
Most crawler failures come from collapsing unlike controls or treating documentation as permanent.Every AI-labelled agent is blocked together
This can disable search discovery or user-requested retrieval when the intended decision concerned only model development.
Every agent is allowed to every route
Public discoverability does not justify exposing account, administration, private file or intake surfaces. Authentication is the boundary for secrets.
The user-agent string is treated as identity
Agent names are easy to spoof. Security exceptions should use the provider's documented network-verification method as well.
robots.txt is the only test
The CDN, firewall, rate limiter or origin may still return a challenge, error, different body or blocked resource.
Google-Extended is presented as Google Search control
Google explicitly states that this token does not affect Search inclusion or ranking.
ChatGPT-User is used to manage ChatGPT Search
OpenAI identifies OAI-SearchBot as the Search control and says ChatGPT-User does not determine Search inclusion.
A bot hit is reported as training or citation
A log request proves neither how content was used nor whether it appeared in an answer. Purpose and outcomes need their own evidence.
A copied IP list is never refreshed
Published network ranges and agent versions can change. Pinning an old list can block genuine traffic or trust addresses no longer assigned to the provider.
CONNECTED EVIDENCE
Connect crawler policy to the rest of the discovery system.
Use the policy with technical delivery, a bounded audit, a discovery model and outcome measurement rather than treating it as a standalone ranking tactic.QUESTIONS / PRECISE ANSWERS
What crawler controls can and cannot decide
01Can I allow ChatGPT Search but block OpenAI model training?+
OpenAI documents OAI-SearchBot and GPTBot as independent controls. A publisher can allow OAI-SearchBot for Search while disallowing GPTBot to express a different training preference. The live network and page response still need verification.
02Does blocking Google-Extended remove my site from Google Search?+
Google says no. Google-Extended is a separate control token for specified Gemini training and grounding uses and does not affect inclusion or ranking in Google Search. Googlebot remains the relevant search crawler.
03Will an Allow rule guarantee an AI citation?+
No. It can remove one access barrier. A genuine visit, indexing or retrieval, selection as grounding material, citation and recommendation are later and separate events.
04Is robots.txt enough to protect private content?+
No. It is a crawler preference file, not an authorization system. Private pages, files and APIs require authentication, authorization and appropriate application or network controls.
05How do I know that a request came from a real crawler?+
Do not trust the user-agent string alone. Use the provider's current published IP ranges, verification endpoint or prescribed reverse-and-forward DNS method, then retain the date and result with the request evidence.
Practical checklist
- Document the purpose of each crawler rule
- Separate search retrieval from training preferences
- Test live responses with the published user agents
- Review official documentation after platform changes
Official sources
- Google Search CentralAI features and your website↗
- Google Search CentralGoogle's common crawlers↗
- Google Search CentralIntroduction to robots.txt↗
- OpenAIOverview of OpenAI crawlers↗
- Anthropic Help CenterAnthropic web crawlers↗
- Perplexity DocsPerplexity crawlers↗
- Microsoft Bing Webmaster ToolsWhich crawlers does Bing use?↗
- Microsoft Bing Webmaster BlogIntroducing AI Performance in Bing Webmaster Tools↗
- Yandex WebmasterYandex robot user agents↗

