← Back to GEO LibraryTECHNICAL / CRAWLER POLICY

AI crawler controls by platform

AI access policy should be explicit by crawler and purpose. A blanket allow or block can create consequences the publisher did not intend.

Updated
02/09/2026
Publisher
NobleJackal
Version
1.1
Source review
Primary sources reviewed
01

Google uses the Search foundation

Google states that AI features rely on existing Search systems. There is no special AI schema requirement; index eligibility, snippets and normal Search controls remain relevant.

02

OpenAI separates search from training

OAI-SearchBot supports search discovery, while GPTBot is a separate control associated with potential model training. Publishers should decide on each purpose independently.

03

Anthropic and Perplexity publish distinct controls

Anthropic and Perplexity document crawler identities and retrieval behaviour. Robots rules should be paired with live CDN or WAF checks because a policy file alone cannot prove successful access.

DIRECT ANSWER / CRAWLER POLICY

AI crawlers should be controlled by named agent and purpose, not with one blanket rule.

Search discovery, a fetch made because a person asked a question and collection for model development are different activities. Several platforms publish separate agent names precisely so a site owner can make separate decisions. Allowing an automatic search crawler while declining a training crawler can be a coherent policy; treating every agent containing the letters AI as the same visitor is not.

The decision begins in robots.txt but does not end there. The origin, CDN, firewall, bot-management product, rate limiter and application must all deliver the intended public response. A user-agent string can be spoofed, so a security allow-list should combine the declared agent with the platform's current IP publication or verification method. A robots rule expresses publisher preference; it is not proof that the genuine crawler visited, indexed or cited a page.

For a publisher seeking discovery, the practical default is to keep public search and user-requested retrieval available, decide model-development access independently, protect non-public routes with authentication and verify representative URLs from outside the deployment environment. Keep a dated policy register because platform names, purposes and network ranges can change.

THREE PURPOSES / THREE DECISIONS

Name the activity before writing the rule.

The same platform may publish more than one agent. A useful policy records the purpose, desired outcome and acceptable operational cost for each one.
  1. 01

    Search discovery

    An automatic crawler discovers and refreshes documents for a search or answer index. Blocking it may reduce or remove eligibility for that platform's search experience, but permission still does not guarantee indexing or selection.

    Decide: do we want this public material discoverable in this platform's search results?
  2. 02

    User-requested retrieval

    A fetcher visits because a person asked a product to open or use a page. Some providers state that ordinary robots.txt rules may not apply to this user-directed activity. Authentication, authorization and application security remain the real boundary for private information.

    Decide: may a user ask this product to retrieve the same page that any anonymous visitor can open?
  3. 03

    Model development

    A separate crawler or control token expresses whether public content may be collected for improving or training foundation models. This decision can be independent of search discovery and should be reviewed with the publisher's legal, licensing and commercial policy.

    Decide: does the publisher permit this documented development use under its rights and risk policy?

PLATFORM / AGENT MATRIX

Use the current official name, then verify the delivery path.

The table records documented roles as of the review date. Recheck the linked primary source before a policy change; do not copy a static allow-list into a permanent security rule without network verification.
Platform / agentDocumented rolePublisher controlWhat to verify
Google / GooglebotBuilds the Google Search index. Google says rules for Googlebot affect Search, including its Search features.robots.txt token: Googlebot. Index exclusion requires a readable noindex or access control, not a crawl block alone.Use Google's published crawler IP ranges or reverse and forward DNS; inspect Search Console separately.
Google / Google-ExtendedA control token for Gemini model training and grounding in specified Gemini and Vertex AI experiences; it has no separate request user-agent.robots.txt token: Google-Extended. Google states it does not affect Google Search inclusion or ranking.Verify the policy text itself; do not search logs for a Google-Extended request that the documentation says does not exist.
Microsoft / BingbotBing's standard crawler for its search index. Bing Webmaster Tools separately reports search health and supported AI citation activity.robots.txt token: Bingbot; crawl-rate controls are also available in Bing Webmaster Tools.Use Bing's verification tool or published verification methods because the user-agent can be spoofed.
Yandex / YandexBotYandex's main indexing crawler. Yandex may also use purpose-specific agents with their own tokens.robots.txt token: YandexBot for the main robot, or Yandex for the documented broader group where appropriate.Use Yandex Webmaster response and robots analysers; verify suspicious traffic using Yandex's IP or hostname method.
OpenAI / OAI-SearchBotAutomatic search crawler used to surface websites in ChatGPT search features.robots.txt token: OAI-SearchBot. OpenAI documents this independently from GPTBot.Match the current published OAI-SearchBot IP ranges as well as the user-agent; test public HTML and required resources.
OpenAI / ChatGPT-UserFetcher used for certain actions initiated by ChatGPT or Custom GPT users; it is not the automatic search crawler.OpenAI says robots.txt rules may not apply to these user-initiated actions and that this agent does not determine Search inclusion.Use OpenAI's published ChatGPT-User IP list and enforce genuine private boundaries with authentication.
OpenAI / GPTBotCrawler for content that may be used to improve and train OpenAI generative foundation models.robots.txt token: GPTBot. It can be allowed or disallowed independently of OAI-SearchBot.Use the current GPTBot IP publication; record the development-use decision separately from search access.
Anthropic / Claude-SearchBotNavigates the web to improve search-result relevance and accuracy for Claude users.robots.txt token: Claude-SearchBot; Anthropic also documents Crawl-delay support where appropriate.Check the current Anthropic IP publication, robots response, rate behaviour and the final page body.
Anthropic / Claude-UserRetrieves web content at a Claude user's direction.robots.txt token: Claude-User under Anthropic's published policy; blocking it may reduce user-directed web retrieval.Separate user retrieval from automatic search and model-development traffic in logs and policy records.
Anthropic / ClaudeBotCollects public web content that may contribute to model training and improvement.robots.txt token: ClaudeBot; it is independent of Claude-SearchBot and Claude-User.Confirm the current network source, apply a measured crawl delay if required and retain the policy date.
Perplexity / PerplexityBotAutomatic crawler designed to surface and link websites in Perplexity search; its documentation says it is not for foundation-model training.robots.txt token: PerplexityBot; Perplexity recommends allowing its current published IP ranges for search visibility.Combine user-agent and the current PerplexityBot IP endpoint; monitor WAF challenges and response integrity.
Perplexity / Perplexity-UserFetcher that may visit a page in response to a person's question and link it in the answer.Perplexity says this user-requested fetcher generally ignores robots.txt and is not a training crawler.Use the separate Perplexity-User IP endpoint; protect non-public data with authentication rather than robots.txt.

POLICY / VERIFICATION WORKFLOW

Move from publishing intent to evidence in five passes.

01 / CLASSIFY

Inventory content and decide what is genuinely public.

Separate public editorial and commercial pages from account areas, intake payloads, private files, preview hosts, internal search results and administrative routes. robots.txt is not access control: a route containing confidential material belongs behind authentication even when every crawler is disallowed.

Record the rights position for books, datasets, images and client material. A crawler policy cannot grant rights the publisher does not hold, and a permissive site-wide rule should not silently expose material whose publication status is uncertain.

02 / CHOOSE

Make an independent decision for each documented purpose.

Write a small decision table: agent, provider, purpose, desired policy, owner, date, source URL and review date. Search access, user retrieval and model development should not inherit one another merely because they come from the same company.

If the objective is discoverability, preserve the automatic search agents needed for that objective. A publisher may still choose a different model-development policy. State the trade-off plainly and obtain the appropriate legal or rights decision where the content warrants it.

03 / IMPLEMENT

Align robots.txt, edge security and the application.

Publish exact agent groups without accidental precedence conflicts. Then inspect CDN bot rules, managed challenges, country blocks, rate limits, origin allow-lists and application middleware. An Allow line cannot override a firewall that returns 403, a challenge page or an empty 200 response.

Use bounded rate controls. A harsh blanket rule can make a platform abandon retrieval, while no protection at all can expose the origin to spoofed traffic. Where a provider publishes IP ranges, combine those ranges with the declared user-agent and refresh them from the official endpoint.

04 / TEST

Verify policy parsing, network delivery and rendered meaning separately.

First test whether the target URL is allowed by the relevant robots group. Then send a named-agent request from outside the origin environment and retain status, final URL, headers, body hash and challenge result. Finally render representative pages to ensure the answer, evidence and links survive.

A synthetic user-agent test proves only how the stack responds to that string from the test network. It does not authenticate a platform visit. Genuine log classification requires the provider's current IP publication or its prescribed DNS verification method.

05 / OBSERVE

Keep access, visit, index and citation as separate records.

A green delivery test establishes technical eligibility at one time. Origin logs may later show a verified visit; webmaster tools may show crawl or indexing; answer observations or platform reports may show a citation. None of these stages should be backfilled from another.

Review the policy when official documentation, crawler names, IP feeds, firewall products, hosting topology or publication rights change. Save the former version and retest before reporting the new state as live.

MINIMUM POLICY REGISTER

A usable crawler policy has six records beyond robots.txt.

The public file is only one expression of a wider operational decision.
01

Purpose register

Map every named agent to search discovery, user-requested retrieval, model development or another documented use. Do not infer purpose from the brand name alone.

02

Content boundary

List the public route families included in the policy and the private systems protected by authentication. Include subdomains and alternate hosts explicitly.

03

Primary-source record

Keep the official documentation URL, date reviewed, agent token, verification endpoint and material caveats. Set a review date for volatile details.

04

Edge implementation

Record CDN, WAF, rate-limit, geographic and origin rules that can override apparent robots permission, along with the responsible owner.

05

Acceptance evidence

Retain parser result, named-agent HTTP response, final body signature, resource failures and a representative rendered capture for each important route class.

06

Outcome ledger

Log verified visits, indexing records, citations and referrals separately. Never convert a successful access test into an unstated claim about downstream use.

POLICY RED FLAGS

Eight mistakes that break either discovery or security.

Most crawler failures come from collapsing unlike controls or treating documentation as permanent.
01

Every AI-labelled agent is blocked together

This can disable search discovery or user-requested retrieval when the intended decision concerned only model development.

02

Every agent is allowed to every route

Public discoverability does not justify exposing account, administration, private file or intake surfaces. Authentication is the boundary for secrets.

03

The user-agent string is treated as identity

Agent names are easy to spoof. Security exceptions should use the provider's documented network-verification method as well.

04

robots.txt is the only test

The CDN, firewall, rate limiter or origin may still return a challenge, error, different body or blocked resource.

05

Google-Extended is presented as Google Search control

Google explicitly states that this token does not affect Search inclusion or ranking.

06

ChatGPT-User is used to manage ChatGPT Search

OpenAI identifies OAI-SearchBot as the Search control and says ChatGPT-User does not determine Search inclusion.

07

A bot hit is reported as training or citation

A log request proves neither how content was used nor whether it appeared in an answer. Purpose and outcomes need their own evidence.

08

A copied IP list is never refreshed

Published network ranges and agent versions can change. Pinning an old list can block genuine traffic or trust addresses no longer assigned to the provider.

CONNECTED EVIDENCE

Connect crawler policy to the rest of the discovery system.

Use the policy with technical delivery, a bounded audit, a discovery model and outcome measurement rather than treating it as a standalone ranking tactic.

QUESTIONS / PRECISE ANSWERS

What crawler controls can and cannot decide

01Can I allow ChatGPT Search but block OpenAI model training?

OpenAI documents OAI-SearchBot and GPTBot as independent controls. A publisher can allow OAI-SearchBot for Search while disallowing GPTBot to express a different training preference. The live network and page response still need verification.

02Does blocking Google-Extended remove my site from Google Search?

Google says no. Google-Extended is a separate control token for specified Gemini training and grounding uses and does not affect inclusion or ranking in Google Search. Googlebot remains the relevant search crawler.

03Will an Allow rule guarantee an AI citation?

No. It can remove one access barrier. A genuine visit, indexing or retrieval, selection as grounding material, citation and recommendation are later and separate events.

04Is robots.txt enough to protect private content?

No. It is a crawler preference file, not an authorization system. Private pages, files and APIs require authentication, authorization and appropriate application or network controls.

05How do I know that a request came from a real crawler?

Do not trust the user-agent string alone. Use the provider's current published IP ranges, verification endpoint or prescribed reverse-and-forward DNS method, then retain the date and result with the request evidence.

Practical checklist

  • Document the purpose of each crawler rule
  • Separate search retrieval from training preferences
  • Test live responses with the published user agents
  • Review official documentation after platform changes

Official sources

  1. Google Search CentralAI features and your website
  2. Google Search CentralGoogle's common crawlers
  3. Google Search CentralIntroduction to robots.txt
  4. OpenAIOverview of OpenAI crawlers
  5. Anthropic Help CenterAnthropic web crawlers
  6. Perplexity DocsPerplexity crawlers
  7. Microsoft Bing Webmaster ToolsWhich crawlers does Bing use?
  8. Microsoft Bing Webmaster BlogIntroducing AI Performance in Bing Webmaster Tools
  9. Yandex WebmasterYandex robot user agents