Digital Insights Blog > AI Website Crawling: Should You Let AI Read Your Website?
AI Website Crawling: Should You Let AI Read Your Website?
- 7 min read
Five years ago, most marketing teams had little reason to care which bots were reading their websites. Googlebot was welcome, obvious troublemakers weren’t and much of what happened in between was an IT problem. AI has made that comfortable division obsolete.
AI website crawling can now mean several very different things. A company may be reading your pages so its AI search engine can recommend you. Another may be collecting your articles for model training. An AI agent might be visiting because an actual customer has asked it to research your products. And now somebody may even want to pay you for access to your content.
That last possibility became more real on September 30, when Cloudflare announced Pay Per Use, an experimental program that allows participating AI companies to compensate publishers when their content is used in specified ways.
It is another sign that the old choice between “let the bots in” and “keep the bots out” isn’t going to be sufficient much longer.
AI Website Crawling Doesn’t Mean Just One Thing
Cloudflare now divides AI-related automated traffic into three broad categories: Search, Agent and Training. The names are fairly self-explanatory, but the differences matter.
Search crawlers collect and index information so it can later be retrieved or referenced. Training crawlers collect material that may be used to train or fine-tune AI models. Agents are different again: they can visit a site in real time while doing something on behalf of a user.
Imagine an association that has spent 20 years building a valuable collection of industry information. It probably wants someone asking ChatGPT about professional certification in its field to learn about its certification program. That’s the sort of visibility the association has spent years trying to achieve through Google.
The same association might also be happy for an AI agent to find its annual conference, check the registration deadline and report the price to a member who asked.
But what if a crawler wants to ingest 5,000 articles from the association’s archive to improve an AI model? Or it starts retrieving an expensive annual compensation survey that exists primarily as a benefit for dues-paying members?
Those aren’t really the same proposition.
Cloudflare’s own numbers suggest website owners have begun to recognize that. In September 2026, the company reported that fewer than 1% of websites on its network choose to block Search bots, compared with 17% that block Training bots.
The message isn’t difficult to decipher: we want to be found, but that doesn’t automatically mean everything we publish is available for every purpose.
You Can Already Make Some of These Choices
This isn’t entirely about technology that’s coming someday. OpenAI already provides separate controls for different uses. OAI-SearchBot helps surface websites in ChatGPT search results. GPTBot crawls content that may be used to train OpenAI’s generative AI foundation models. A website can allow the former and disallow the latter.
Google handles the distinction somewhat differently. Google-Extended isn’t a separate crawler. It is a robots.txt control that allows publishers to limit whether content Google crawls can be used for certain Gemini training and grounding purposes. Google says the setting does not affect inclusion or ranking in Google Search.
For the person responsible for a company’s website, that’s a fairly significant development.
The question no longer has to be, “Do we let AI read our website?” It can be, “What are we comfortable letting AI do with it?”
As AI website crawling becomes more common, that second question is going to matter much more.
What If the AI Company Paid You?
Cloudflare added another wrinkle on September 30 with Pay Per Use.
The program is still in beta, so it would be premature for publishers to start adding AI licensing revenue to next year’s budget. But the idea behind it is interesting.
Cloudflare had already been experimenting with Pay Per Crawl, in which payment is associated with accessing content. Pay Per Use goes a step further by attempting to associate payment with what happens to that content afterward.
Take a publisher that pays reporters and analysts to produce original financial research.
One AI company might crawl an article and never use it. Another might use information from that article when answering a subscriber’s question about a company or market. Pay Per Use is an attempt to distinguish between those situations and put a value on the latter.
Participating AI companies report qualifying uses, while Cloudflare manages usage records, billing and payments. Cloudflare says participating companies are required to report qualifying usage and that those reports are checked against participating publishers.
Whether this particular model takes off is anyone’s guess. The more important development is that the market is starting to treat access, indexing, training and use as different things.
A website owner may increasingly be able to say yes to one and no to another.
Why Not Just Block All of Them?
For some organizations, restricting AI access will make sense. Blocking everything, however, comes with a potential cost.
Imagine a nonprofit working on Chesapeake Bay restoration. A prospective donor who once might have typed “Chesapeake Bay environmental nonprofits” into Google could now ask an AI assistant, “Which organizations are restoring waterways around the Chesapeake Bay, and what does each one do?”
The nonprofit would probably like to be part of that answer.
A procurement officer could ask an AI system to find digital agencies with experience building accessible Drupal websites for federal organizations. An engineer could ask it to identify manufacturers producing a particular component. A prospective student could ask which universities offer a specialized master’s program within 100 miles of Washington, D.C.
These are search queries even if nobody sees a conventional search-results page.
Blocking a particular AI crawler doesn’t necessarily make an organization invisible to every AI service. Different systems discover information in different ways. OpenAI, for example, notes that a URL could still appear as a navigational link even when OAI-SearchBot has been blocked if OpenAI discovers that URL elsewhere.
Still, organizations have spent decades trying to make themselves discoverable online. It would be strange to abandon that goal just as the definition of “discoverable” is changing.
But Letting Everybody in Isn’t Much of a Strategy Either
Suppose a professional association publishes a short article summarizing its annual salary research. The association wants that article found. It demonstrates expertise, attracts professionals in the industry and gives prospective members another reason to investigate the organization.
Behind the member login, however, sits the complete 80-page compensation report. Perhaps the organization spent tens of thousands of dollars collecting and analyzing the data, and access to the report is one reason people pay annual dues.
The association doesn’t necessarily have to think about those two pieces of content in the same way.
The same is true of a consulting company that has accumulated 15 years of original articles, a university with public admissions information alongside licensed research, or a manufacturer that wants product specifications widely distributed but doesn’t want automated systems poking around areas intended for customers or distributors.
This is where AI website crawling stops being merely a technical subject. Someone has to decide what’s valuable to expose, what’s valuable to protect and where the organization benefits from machine access.
robots.txt Isn’t a Locked Door
There is another point worth understanding because it is easy to miss.
Putting instructions in robots.txt doesn’t physically prevent a crawler from entering your website. It communicates what you want crawlers to do, and compliant crawlers honor those instructions.
Think of it as a sign on the door. Well-behaved visitors read the sign and follow the rules. The sign itself doesn’t lock anything.
Cloudflare’s AI Crawl Control and security tools can provide stronger enforcement when an organization wants to actually prevent particular automated systems from accessing its site.
That distinction probably hasn’t mattered much to the average communications director in the past. It may matter considerably more when the organization’s research, publications and other intellectual property are involved.
It May Be Time for an AI Access Policy
Most organizations already have rules about digital access, even if they don’t call them an access policy.
The homepage is public. Staff information requires authentication. An association keeps certain research inside its member portal. An ecommerce site wants Google crawling product pages but not a customer’s shopping cart. APIs have permissions. Firewalls keep unwanted traffic out.
AI belongs in that conversation now.
The result doesn’t need to be a 30-page policy document. For many organizations, the first version might be quite simple:
Traditional search engines are welcome. AI search systems are welcome where they improve discovery. Training crawlers are evaluated separately. User-directed agents can access appropriate public information. Proprietary research and authenticated content remain protected.
Another organization might reach completely different conclusions.
An ecommerce company may eventually be delighted to see shopping agents crawling product information because those agents bring customers.
A university may want an AI assistant to easily retrieve admissions deadlines, tuition, program requirements and campus information.
A publisher whose primary product is original reporting may be much more restrictive.
That’s the point. There isn’t one correct configuration for every website.
Some of Your Website Visitors Won’t Be People
This change is already happening faster than many website owners realize. Cloudflare reported in September that daily requests from AI agents across its network had grown by more than 1,700% in a year. The company also says more than half of the traffic reaching sites on its network is now automated. Some of those machines are simply gathering information. Others increasingly represent actual people.
Picture someone planning a conference. Rather than opening six hotel websites and building a spreadsheet, she asks an AI assistant to compare hotels by meeting-room capacity, airport distance, room count and amenities.
Those hotels may compete for her business without her personally visiting all six websites.
Now take it one step further. A member tells an assistant, “Renew my membership and register me for the annual meeting.”
If the website and its integrations permit it, the eventual visitor may be software carrying out the member’s instructions.
At that point, knowing that the visitor is a “bot” isn’t terribly informative. You need to know why it is there.
Find Out What’s Happening Before You Decide
Five years ago, nobody in the communications department was sitting around debating whether GPTBot should read the annual salary survey. Now somebody probably should be.
But the first move shouldn’t necessarily be blocking anything. Start by finding out what is happening.
Which AI crawlers are visiting the website? Where are they going? Are AI systems already sending visitors to the site? Is the material you most want AI systems to understand readily accessible, or is it buried in PDFs and poorly structured pages? Are automated systems repeatedly accessing content you consider particularly valuable?
The answers may surprise you.
They also give marketing, IT and leadership something much more useful to discuss than whether the organization is simply “for” or “against” AI crawlers.
New Target Can Help You Make That Decision
The AI website crawling question touches several parts of a modern website at once: hosting, security, Cloudflare configuration, analytics, SEO and AEO, content strategy and increasingly AI itself.
Those areas shouldn’t be considered separately.
New Target works across that entire environment, including managed hosting and security, WordPress and Drupal development, analytics, SEO and AEO, content strategy and AI services. We can help organizations understand which automated systems are accessing their websites, determine where AI discovery creates value, protect content where appropriate and put the technical controls in place to support those decisions.
For years, the arrangement was easy to understand: search engines read your website, and in return they helped people find it. The next version of the web won’t be quite so simple.
The useful question isn’t whether you should let AI crawl your website. It’s which AI you want to let in, what you want it to see, and what you’re getting out of the deal. Contact us. We can help.
A global team of digerati with offices in Washington, D.C. and Southern California, we provide digital marketing, web design, and creative for brands you know and nonprofits you love.
Follow us to receive the latest digital insights:
- 7 min read
Five years ago, most marketing teams had little reason to care which bots were reading their websites. Googlebot was welcome, obvious troublemakers weren’t and much of what happened in between...
- 8 min read
Artificial intelligence is rapidly becoming another place where consumers research products, evaluate services, and make purchasing decisions. That naturally raises a question for marketers: Can we advertise there? As of...
- 4 min read
Website Design Trends to Watch For several years, website design seemed to be converging on a familiar formula: a large hero image, a short headline, a row of cards, plenty...
- 10 min read
For the past couple of years, marketers have been told that search is changing. Google AI Overviews, AI Mode and other generative AI tools increasingly answer questions directly, changing both...