AI crawlers and what you can control

You will be able to read your robots.txt for AI crawlers and make an informed choice about allowing or blocking them.

Priya read a post in a small business group that said every website should block AI bots immediately, because AI companies were taking content for free. Someone replied that blocking AI bots would make you invisible in ChatGPT. A third person said the two were the same bots anyway. Nobody linked to a source, and Priya ended the thread more confused than when she started.

The confusion is understandable, because the companies run several crawlers each, for different purposes, under different names. This lesson explains how to read your robots.txt for AI crawlers and make a decision you can explain, crawler by crawler.

AI companies run named crawlers

Lesson 4.1, Robots.txt, noindex and canonicals do different jobs, explained that robots.txt gives instructions to crawlers by name, using a line that starts "User-agent:". AI companies publish the names of their crawlers so site owners can do exactly that.

The important point is that one company may run several crawlers for different jobs. OpenAI is a clear example. Its documentation describes OAI-SearchBot, which it uses to find and surface websites in ChatGPT's search features, and GPTBot, which it uses to collect content that may be used to train its models, and each one is controlled separately. A site can allow OAI-SearchBot, so that its pages can appear and be cited in ChatGPT search answers, while blocking GPTBot, so that its content is not used for training. Or it can do the reverse, or allow both, or block both.

Other providers publish their own crawler names and purposes, Perplexity and Anthropic among them. Read each provider's current documentation to see what its crawlers are called and what each one does, because one name may be for search, another for training and another for fetching a page when a user asks.

Google's training control leaves Search alone

Google handles this differently. Googlebot crawls for Google Search, and that includes AI Overviews and AI Mode, which lesson 8.1, How AI Overviews and AI Mode choose their links, explained are part of Search. Separately, Google offers a product token called Google-Extended. Adding a robots.txt rule for Google-Extended lets you control whether content Google crawls from your site may be used for training and grounding its Gemini models.

Google says Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal. So blocking Google-Extended does not remove you from Google Search or from AI Overviews. If you want to limit how your content appears in AI Overviews, Google points to the same controls it uses for ordinary results, such as nosnippet and noindex, with the same effects on how your pages appear in Search.

Decide per crawler, not with a blanket block

The tempting move, after a post like the one Priya read, is to block everything with "AI" in its name. The trouble is that a search crawler and a training crawler do different things for you.

A search crawler fetches pages so an assistant can cite them in answers, usually with a link. Blocking it means that assistant is less able to retrieve and cite your pages, which for a small business that wants customers to find it is often a loss. A training crawler collects content that may be used to train future models. Blocking it is a choice about how your content is used, and it does not, according to the providers that separate them, stop your pages being cited by their search features.

So decide crawler by crawler. Many small businesses choose to allow the search crawlers, because they want to be cited, and make a separate decision about training crawlers based on how they feel about their content being used. Reasonable owners land in different places here, and what matters is that you choose deliberately, rather than inheriting a rule from a plugin or a forum thread.

Never use a blanket rule that blocks all crawlers, such as "User-agent: *" followed by "Disallow: /", to deal with AI. That also blocks Googlebot and every other search engine.

Check other places where crawlers can be blocked as well. Some hosting, security and content delivery services offer settings that block AI crawlers, sometimes switched on by default. If your robots.txt allows a crawler but your hosting blocks it, the hosting setting wins.

Names and rules change

AI crawlers are new, and the companies add, rename and redefine them. The names and purposes in this lesson are those described in the providers' documentation at the time of writing. Before you edit robots.txt, check each provider's current documentation for the exact user-agent names and what each controls. A typo in a crawler name means the rule does nothing.

After you change robots.txt, check it still allows Googlebot to crawl the pages you want found. Search Console has a robots.txt report that shows the file Google last fetched and any problems it found.

Reading your own file

Open your domain followed by /robots.txt. Look for lines naming AI crawlers, and for any blanket rules. Priya's file, generated by a plugin, blocked GPTBot and nothing else, which nobody remembered choosing. She decided to keep that rule, allow the search crawlers explicitly, and write down why. The activity below asks you to list every rule affecting AI crawlers in your own file and record your decision for each.

Read your robots.txt, list every rule that affects AI crawlers, and write which you will allow or block and why.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).