Definition

What is a retrieval crawler?

A retrieval crawler is an agent that fetches a page at the moment a person asks an AI system a question, in order to answer from up-to-date content and cite the source. It differs from the training crawler, which collects content to feed a model with no link to a specific question. Blocking it removes you from the answers your customers read.

How to recognize one

Every crawler identifies itself with a name (its user agent) that its publisher documents. Anthropic describes Claude-User as the crawler that fetches a page when a person asks Claude a question. OpenAI describes ChatGPT-User as the agent for certain actions triggered by a person in ChatGPT. The test is always the same: the crawler's visit follows a specific question.

These visits do not show up in your visit reports, because these crawlers do not run the tracking code. They leave a line in your server logs.

Three crawler families, three decisions

AI publishers now separate their crawlers by function. At Anthropic, ClaudeBot collects for training, Claude-SearchBot indexes content for search and Claude-User fetches a page on demand. At OpenAI, GPTBot serves training and OAI-SearchBot surfaces sites in ChatGPT search. Blocking one does not block the others: each line in the robots.txt file is a separate decision.

Its limits

The robots.txt file expresses a preference, not a barrier. The major publishers document their crawlers and say they follow the protocol; others do not. OpenAI even states that the file's rules may not apply to ChatGPT-User, since the action comes from a person. To really prevent access to a section, you need server-side blocking or authentication.

Example

A purchasing director asks Claude which packaging manufacturers serve Quebec. Claude runs a search, keeps your services page among the results, then Claude-User fetches it to read your lead times and certifications. The answer cites your page. If your robots.txt blocks Claude-User, Claude can no longer retrieve the page in response to the question, even if it is indexed: the answer could then be written from your competitors' pages.

How we read it

We read the retrieval crawler as a visibility channel, with a link and a possible visitor at the end. For almost every business, whose content serves to sell something else, it stays open. The training decision is separate: it depends on the market value of your content, not on your visibility.

The most common risk is a block nobody decided on. A robots.txt template copied online, or edited years ago by a former vendor, shuts out training crawlers and citing crawlers alike. Nobody measures an absence: the effect is discovered months later.

Not to be confused with

Training crawler
It collects content to feed a future model, with no link to a question and no link back to you. Blocking it limits the future use of your content with no immediate effect on your visibility.
Search crawler
Claude-SearchBot or OAI-SearchBot crawl the web ahead of time to build the index the assistant will search. The retrieval crawler comes later, at the moment of the question. Blocking the search crawler can remove you from the assistant entirely.
Google-Extended
It is not a crawler but a token (the name you allow or block in robots.txt). It covers Gemini training and the grounding of its answers in the Gemini apps, not Google Search.

Related concepts

Further reading

Related services

Frequently asked questions

How do I know if my site blocks retrieval crawlers?

Open your robots.txt file (at yourdomain.com/robots.txt) and read every line that names an AI crawler. Then check your server logs to see which crawlers actually visit and on which pages.

← The full glossary