SEO, GEO & AI discovery

Should Your Business Block AI Crawlers?

The short answer: do not make one blanket decision for every AI crawler. Separate search-discovery crawlers, model-training crawlers and user-triggered retrieval. Most service and commerce businesses should allow verified search crawlers on public marketing pages, decide training access from content value and policy, and protect private, paid or operational material with authentication—not robots.txt alone.

Editorial illustration of one business knowledge archive with an open discovery arch, a controlled retrieval gate and a closed training vault
One archive, three access decisions · Original illustration by ThomPerformance

“Allow or block AI” is the wrong business question

A website platform presents one switch: block AI crawlers. The choice looks simple. The commercial consequences are not.

Some automated visitors help an AI product discover and cite public pages. Others collect content that may be used to train a model. A third group retrieves a page because a user asked an assistant to visit it. Blocking every category can remove useful discovery. Allowing every category can expose valuable content to a use the business never intended.

My verdict is direct: set policy by purpose and page value, not by the letters “AI” in a crawler name. Public service, product and evidence pages usually have a different job from paid research, customer data or an internal knowledge base.

A robots.txt file is a public instruction file that compliant crawlers read before requesting pages. It is useful for expressing preferences, but the Robots Exclusion Protocol is not authentication. It does not turn a public URL into private content, prove that a requester is genuine or enforce a licence.

This guide owns the access-policy decision. Use How to Get Your Business Mentioned in AI Search for evidence and citation readiness, or the wider Business Discoverability Stack when indexing, identity and authority are the primary problem.

The Three-Door AI Access Policy

I separate AI-related access into three doors. The categories are more durable than copying a long crawler list because each door begins with the business purpose.

OpenAI documents these purposes separately. OAI-SearchBot supports visibility in ChatGPT search. GPTBot covers content that may be used to train OpenAI's generative AI foundation models. ChatGPT-User supports certain user-triggered actions; OpenAI notes that robots.txt rules may not apply to those requests.

This distinction lets a business allow search discovery while declining model-training access. OpenAI also notes that a robots.txt change may take about 24 hours to affect its search systems, so a policy test needs a reasonable evidence window.

Google uses a different structure. Google-Extended is a control token rather than a separate requesting crawler. Google says it can govern certain Gemini training and grounding uses without affecting inclusion or ranking in Google Search. Blocking Googlebot is a different decision and can remove pages from conventional and AI-powered Google search experiences.

Apply the policy to content, not only to crawler names

Public growthDesigned to be found

Service pages, product information, useful guides, public case evidence and business facts usually benefit from verified discovery access.

Public but valuableUseful with limits

Original research, detailed comparisons or licensable archives may justify search access while training access is restricted.

Customer-onlyRequires authentication

Paid content, portals, account data and contracted deliverables need enforceable access—not a crawler preference file.

OperationalShould not be public

Personal data, internal documents, unpublished pricing, staging environments and system endpoints belong behind real controls.

For most service and ecommerce businesses, the website is a demand asset. Preventing verified search crawlers from accessing public pages can work against the reason those pages exist. A publisher whose content itself is the paid product faces a different economic trade-off and may need legal, licensing and infrastructure input.

Cloudflare's AI Crawl Control documentation reflects that need for granularity: it supports monitoring AI service requests, individual crawler policies, robots.txt compliance and enforcement rules. Its current guidance also distinguishes publisher use cases from ecommerce and general business sites.

Choose the access decision that matches the commercial risk

Business situationDiscovery accessTraining accessRequired safeguard
Service business seeking qualified enquiriesUsually allow verified search crawlers on public expertise and service pagesOwner policy decisionKeep forms, CRM data and unpublished materials outside public URLs
Ecommerce site with public productsUsually allow on product, category and guidance pagesReview by content and platformProtect customer data, private pricing and operational endpoints
Publisher or research businessAllow selectively where discovery supports subscriptions or licensingOften restrict or negotiateAuthentication, metering, licensing advice and enforceable controls
Customer portal or member areaDo not rely on crawler exclusionsDo not rely on crawler exclusionsAuthentication, authorisation and no public indexing
Bot traffic creates material server costPreserve verified beneficial crawlers where possibleRate-limit or block by policyLogs, verified identities, firewall rules and rollback monitoring

Allowing a crawler does not guarantee a mention, citation or customer. It only preserves technical eligibility. The page still needs a direct answer, clear business identity, original evidence and credible corroboration. Measure those signals with the AI Visibility Evidence Ladder.

Likewise, blocking model-training access does not make public content confidential or prevent every form of retrieval. Treat legal rights, platform terms, security, indexing and commercial discovery as related but separate decisions. This article provides a marketing-governance framework, not legal advice.

A 30-day crawler-access review

Days 1–5

Inventory

List the public pages, protected areas, crawler rules, firewall settings, host controls and named decision owner.

Days 6–12

Classify

Separate discovery, retrieval and training purposes. Verify current crawler documentation rather than trusting a copied list.

Days 13–20

Apply

Set page-tier policy, protect non-public material and make the smallest reversible change needed.

Days 21–30

Verify

Check logs, important URLs, search eligibility, AI answer retrieval and unintended blocks before making the rule permanent.

Record the date, rule, reason, expected benefit, risk and rollback condition. Review the policy when a provider changes its crawler purpose or when a hosting platform changes default AI-bot controls. Do not assume yesterday's user-agent list still describes today's products.

If discoverability is the goal, connect the technical policy to AI growth support, growth partnership services, documented case evidence, evidence standards and Thomas's operating model. Access is useful only when the public page helps a buyer make a better decision.

Practitioner note: I would never approve a blanket “block AI” switch from its label alone. I first ask which crawler is affected, which pages it can reach, which commercial outcome may be lost and whether the control is advisory or enforceable.

Sources and evidence notes

Sources and search results were checked on 5 September 2026. Crawler purposes and platform controls can change; verify current provider documentation before implementation. The Three-Door AI Access Policy, content tiers, decision matrix and 30-day review are original ThomPerformance analysis. No search volume, citation outcome or legal conclusion is claimed.

  1. OpenAI: Overview of OpenAI Crawlers
  2. Google: Google-Extended crawler control
  3. Google Search Central: AI features and your website
  4. Cloudflare: AI Crawl Control
  5. RFC 9309: Robots Exclusion Protocol

Frequently asked questions

Will blocking GPTBot remove my business from ChatGPT search?

Not by itself. OpenAI documents GPTBot as the control for content that may be used to train its generative AI foundation models. OAI-SearchBot is the separate crawler used to surface websites in ChatGPT search. A business can allow OAI-SearchBot while disallowing GPTBot, although visibility and citation are never guaranteed.

Does blocking Google-Extended hurt Google rankings or AI Overviews?

Google states that Google-Extended does not affect inclusion or ranking in Google Search. It controls certain uses for Gemini model training and grounding in Gemini products. Google’s AI search features rely on Google Search systems, so blocking Googlebot or making pages ineligible for Search is a separate and potentially damaging decision.

Is robots.txt enough to protect private or paid content?

No. A robots.txt file is a public set of crawler instructions, not a privacy or security control. Sensitive, customer-only, paid, confidential or operational content should require authentication or another enforceable access layer. Server or firewall controls may be appropriate for unwanted automated traffic, but they must be designed carefully to avoid blocking legitimate services.

Should a service or ecommerce business normally allow AI search crawlers?

Usually on public pages intended to attract customers, provided the crawler is verified and the business accepts the platform’s terms. Service pages, product information, useful guides and public proof can support discovery. Customer portals, private pricing, unpublished inventory, licensed research and internal documents need a separate policy.

How often should an AI crawler policy be reviewed?

Review it at least quarterly and whenever the website platform, firewall, crawler documentation or commercial model changes. Keep an owner, a dated decision record and a small set of test pages. Crawler names and product uses change, so a one-time copied robots.txt list is not durable governance.

Keep discovery open without treating every use as the same

Decide separately who may discover public pages, who may retrieve them for a user and who may use them for model development. Then apply that choice by content tier, protect genuinely private material with real access controls and verify the commercial effect. A precise policy preserves both control and qualified discoverability.

Which door is currently unclear for your business: discovery, user retrieval or training?

Request an AI access diagnostic

About the author: Thomas Ho is a Paid Digital Marketing & AI Growth Partner helping business leaders connect paid acquisition, organic discovery, customer evidence and practical AI to qualified pipeline and revenue.

Free 48-hour audit

Protect valuable content without hiding the business from future buyers.

Separate discovery, retrieval and training access—then document the smallest safe policy for each content tier.

Request an AI access diagnostic Review AI growth support

Free operating template

Stop reviewing paid ads with screenshots and green arrows.

Use the same weekly review structure I use to connect spend with qualified leads, opportunities, pipeline and decisions.

  • Commercial scorecard
  • Creative test log
  • Decision ownership
Get a free audit