“Allow or block AI” is the wrong business question
A website platform presents one switch: block AI crawlers. The choice looks simple. The commercial consequences are not.
Some automated visitors help an AI product discover and cite public pages. Others collect content that may be used to train a model. A third group retrieves a page because a user asked an assistant to visit it. Blocking every category can remove useful discovery. Allowing every category can expose valuable content to a use the business never intended.
My verdict is direct: set policy by purpose and page value, not by the letters “AI” in a crawler name. Public service, product and evidence pages usually have a different job from paid research, customer data or an internal knowledge base.
A robots.txt file is a public instruction file that compliant crawlers read before requesting pages. It is useful for expressing preferences, but the Robots Exclusion Protocol is not authentication. It does not turn a public URL into private content, prove that a requester is genuine or enforce a licence.
This guide owns the access-policy decision. Use How to Get Your Business Mentioned in AI Search for evidence and citation readiness, or the wider Business Discoverability Stack when indexing, identity and authority are the primary problem.
The Three-Door AI Access Policy
I separate AI-related access into three doors. The categories are more durable than copying a long crawler list because each door begins with the business purpose.
Search and citation
A crawler builds or refreshes a search index so public pages can appear in answers, summaries or source links.
User-requested access
An assistant fetches a page because a user asked it to read, compare or act on that specific public information.
Model development
A crawler collects content that may be used to improve or train foundation models, without promising a search visit or citation.
OpenAI documents these purposes separately. OAI-SearchBot supports visibility in ChatGPT search. GPTBot covers content that may be used to train OpenAI's generative AI foundation models. ChatGPT-User supports certain user-triggered actions; OpenAI notes that robots.txt rules may not apply to those requests.
This distinction lets a business allow search discovery while declining model-training access. OpenAI also notes that a robots.txt change may take about 24 hours to affect its search systems, so a policy test needs a reasonable evidence window.
Google uses a different structure. Google-Extended is a control token rather than a separate requesting crawler. Google says it can govern certain Gemini training and grounding uses without affecting inclusion or ranking in Google Search. Blocking Googlebot is a different decision and can remove pages from conventional and AI-powered Google search experiences.
Apply the policy to content, not only to crawler names
Service pages, product information, useful guides, public case evidence and business facts usually benefit from verified discovery access.
Original research, detailed comparisons or licensable archives may justify search access while training access is restricted.
Paid content, portals, account data and contracted deliverables need enforceable access—not a crawler preference file.
Personal data, internal documents, unpublished pricing, staging environments and system endpoints belong behind real controls.
For most service and ecommerce businesses, the website is a demand asset. Preventing verified search crawlers from accessing public pages can work against the reason those pages exist. A publisher whose content itself is the paid product faces a different economic trade-off and may need legal, licensing and infrastructure input.
Cloudflare's AI Crawl Control documentation reflects that need for granularity: it supports monitoring AI service requests, individual crawler policies, robots.txt compliance and enforcement rules. Its current guidance also distinguishes publisher use cases from ecommerce and general business sites.
Choose the access decision that matches the commercial risk
| Business situation | Discovery access | Training access | Required safeguard |
|---|---|---|---|
| Service business seeking qualified enquiries | Usually allow verified search crawlers on public expertise and service pages | Owner policy decision | Keep forms, CRM data and unpublished materials outside public URLs |
| Ecommerce site with public products | Usually allow on product, category and guidance pages | Review by content and platform | Protect customer data, private pricing and operational endpoints |
| Publisher or research business | Allow selectively where discovery supports subscriptions or licensing | Often restrict or negotiate | Authentication, metering, licensing advice and enforceable controls |
| Customer portal or member area | Do not rely on crawler exclusions | Do not rely on crawler exclusions | Authentication, authorisation and no public indexing |
| Bot traffic creates material server cost | Preserve verified beneficial crawlers where possible | Rate-limit or block by policy | Logs, verified identities, firewall rules and rollback monitoring |
Allowing a crawler does not guarantee a mention, citation or customer. It only preserves technical eligibility. The page still needs a direct answer, clear business identity, original evidence and credible corroboration. Measure those signals with the AI Visibility Evidence Ladder.
Likewise, blocking model-training access does not make public content confidential or prevent every form of retrieval. Treat legal rights, platform terms, security, indexing and commercial discovery as related but separate decisions. This article provides a marketing-governance framework, not legal advice.
A 30-day crawler-access review
Inventory
List the public pages, protected areas, crawler rules, firewall settings, host controls and named decision owner.
Classify
Separate discovery, retrieval and training purposes. Verify current crawler documentation rather than trusting a copied list.
Apply
Set page-tier policy, protect non-public material and make the smallest reversible change needed.
Verify
Check logs, important URLs, search eligibility, AI answer retrieval and unintended blocks before making the rule permanent.
Record the date, rule, reason, expected benefit, risk and rollback condition. Review the policy when a provider changes its crawler purpose or when a hosting platform changes default AI-bot controls. Do not assume yesterday's user-agent list still describes today's products.
If discoverability is the goal, connect the technical policy to AI growth support, growth partnership services, documented case evidence, evidence standards and Thomas's operating model. Access is useful only when the public page helps a buyer make a better decision.
Practitioner note: I would never approve a blanket “block AI” switch from its label alone. I first ask which crawler is affected, which pages it can reach, which commercial outcome may be lost and whether the control is advisory or enforceable.
Sources and evidence notes
Sources and search results were checked on 5 September 2026. Crawler purposes and platform controls can change; verify current provider documentation before implementation. The Three-Door AI Access Policy, content tiers, decision matrix and 30-day review are original ThomPerformance analysis. No search volume, citation outcome or legal conclusion is claimed.
Frequently asked questions
Will blocking GPTBot remove my business from ChatGPT search?
Not by itself. OpenAI documents GPTBot as the control for content that may be used to train its generative AI foundation models. OAI-SearchBot is the separate crawler used to surface websites in ChatGPT search. A business can allow OAI-SearchBot while disallowing GPTBot, although visibility and citation are never guaranteed.
Does blocking Google-Extended hurt Google rankings or AI Overviews?
Google states that Google-Extended does not affect inclusion or ranking in Google Search. It controls certain uses for Gemini model training and grounding in Gemini products. Google’s AI search features rely on Google Search systems, so blocking Googlebot or making pages ineligible for Search is a separate and potentially damaging decision.
Is robots.txt enough to protect private or paid content?
No. A robots.txt file is a public set of crawler instructions, not a privacy or security control. Sensitive, customer-only, paid, confidential or operational content should require authentication or another enforceable access layer. Server or firewall controls may be appropriate for unwanted automated traffic, but they must be designed carefully to avoid blocking legitimate services.
Should a service or ecommerce business normally allow AI search crawlers?
Usually on public pages intended to attract customers, provided the crawler is verified and the business accepts the platform’s terms. Service pages, product information, useful guides and public proof can support discovery. Customer portals, private pricing, unpublished inventory, licensed research and internal documents need a separate policy.
How often should an AI crawler policy be reviewed?
Review it at least quarterly and whenever the website platform, firewall, crawler documentation or commercial model changes. Keep an owner, a dated decision record and a small set of test pages. Crawler names and product uses change, so a one-time copied robots.txt list is not durable governance.
Keep discovery open without treating every use as the same
Decide separately who may discover public pages, who may retrieve them for a user and who may use them for model development. Then apply that choice by content tier, protect genuinely private material with real access controls and verify the commercial effect. A precise policy preserves both control and qualified discoverability.
Which door is currently unclear for your business: discovery, user retrieval or training?
