DOCUMENTED
Directly supported by the platform provider’s own documentation.
RESEARCH · REFERENCE V1.0 · SOURCES READ 2026-09-05
Which crawlers and fetchers each AI search provider documents, what the provider says each one is for, and how to check that a request claiming to be one of them is genuine. Every row links the provider’s own document, carries one of five classes, and shows the date that document was last read.
This page describes controls. It does not recommend allowing or blocking anything — what to do with these controls depends on your content and your business — and no provider on this page guarantees that an allowed crawler leads to a citation.
Five classes · a verified date on every row · a re-verification date
Every row on this page carries three things: a link to the provider’s own document, one of five classes, and the date that document was last read. The classes are the same five we use on every claim we make — defined once, on the methodology page.
Directly supported by the platform provider’s own documentation.
Seen in our testing, but not established as a platform rule.
A reasonable reading of the evidence. Not a mechanism.
Insufficient evidence. Says so, and stays unused.
A deliberate test of something unverified, registered before it runs.
On this page a row is DOCUMENTED only when the provider states it in its own documentation, linked and dated. A crawler that appears in server logs but in no provider document is OBSERVED, never DOCUMENTED. Where we draw a conclusion from documented facts, that conclusion is marked INFERRED, separately from the facts. Nothing is promoted a class because it is convenient.
Published endpoints · reverse DNS · the layer in front of the origin
A user-agent string is a claim, not an identity. Anyone can send one. What a provider publishes to settle the question is an IP list, a reverse-DNS pattern or a tool — and every endpoint below answered on the date in its row.
| Provider | Method | Published endpoints | Checked | Class | Verified | Source |
|---|---|---|---|---|---|---|
| OpenAI | Published IP ranges, one list per agent. | HTTP 200 on | DOCUMENTED | OpenAI · crawler documentation | ||
| Anthropic | If a crawler has a source IP address on this list, it indicates that the crawler is coming from Anthropic. One published IP list for all Anthropic agents. | HTTP 200 on | DOCUMENTED | Anthropic · does Anthropic crawl data from the web | ||
| Perplexity | Published IP ranges, one list per agent. | HTTP 200 on | DOCUMENTED | Perplexity · crawlers | ||
The common crawlers generally crawl from the IP ranges published in the common-crawlers.json object, and the reverse DNS mask of their hostname matches crawl-***-***-***-***.googlebot.com or geo-crawl-***-***-***-***.geo.googlebot.com. Published IP ranges per crawler group, plus a reverse-DNS mask for the common crawlers. | HTTP 200 on | DOCUMENTED | Google · common crawlers | |||
| Apple | Traffic coming from Applebot is generally identified by using reverse DNS in the *.applebot.apple.com domain. Another way is to match the IP address with a CIDR prefix contained in the following JSON file Reverse DNS in *.applebot.apple.com, or a published IP list. | HTTP 200 on | DOCUMENTED | Apple · about Applebot | ||
| Mistral | Published IP lists for the index and user agents. | HTTP 200 on | DOCUMENTED | Mistral · robots | ||
| Amazon | Published IP addresses, one page per agent. | HTTP 200 on | DOCUMENTED | Amazon · Amazonbot | ||
| Microsoft | The Verify Bingbot tool inside Bing Webmaster Tools. | | a tool behind a login — nothing to fetch | DOCUMENTED | Bing · which crawlers does Bing use transcribed from the rendered page — the table is JavaScript-rendered |
Alternate methods like blocking IP address(es) from which Anthropic Bots operates may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file.Anthropic · DOCUMENTED · 2026-09-05
The IP list is published for recognising Anthropic’s crawlers. Anthropic states that blocking those addresses is not a reliable way to opt out — robots.txt is the documented control.
If you’re using a Web Application Firewall (WAF) to protect your site, you may need to explicitly whitelist Perplexity’s bots to ensure they can access your content.Perplexity · DOCUMENTED · 2026-09-05
The check may have to happen in front of the origin. A robots.txt that permits a crawler does not help if a firewall or bot-protection layer answers that crawler with a challenge — a failure that is invisible in robots.txt and invisible in analytics. Perplexity’s documentation says so in as many words and lists firewall settings for it.
Search · user retrieval · training
The tables on this page are sorted by what an agent does, because that is what the controls attach to. The same provider often runs one of each. This three-way split is our framing — INFERRED from the documented purposes, not a taxonomy any provider publishes — and the rows below say, per provider, where the documentation actually separates the three and where it does not.
| Job | What it does | What blocking it costs |
|---|---|---|
| Search | Builds an index so the provider can surface and link your pages in its search product. | Your pages stop being eligible for that product’s answers. |
| User retrieval | Fetches one page because a person asked a question that needs it, right now. | The assistant cannot open your page for a user who asked about it. |
| Training | Collects content that may be used to train foundation models. | Your content is excluded from future training sets — with the providers that document the separation. |
The agent to allow if you want to appear in that provider’s answers
| Provider | User agent | Provider’s own words | Class | Verified | Source |
|---|---|---|---|---|---|
| OpenAI | OAI-SearchBot | OAI-SearchBot is for search. OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features. Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links. | DOCUMENTED | OpenAI · crawler documentation | |
| Googlebot | Crawling preferences addressed to the Googlebot user agent affect Google Search (including Discover and all Google Search features), as well as other products such as Google Images, Google Video, Google News, and Discover. | DOCUMENTED | Google · common crawlers | ||
| Microsoft | bingbot | our standard crawler … handles most of our crawling needs each day transcribed from the rendered page — the table is JavaScript-rendered | DOCUMENTED | Bing · which crawlers does Bing use transcribed from the rendered page — the table is JavaScript-rendered | |
| Perplexity | PerplexityBot | PerplexityBot is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models. | DOCUMENTED | Perplexity · crawlers | |
| Anthropic | Claude-SearchBot | Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses. Disabling Claude-SearchBot on your site prevents our system from indexing your content for search optimization, which may reduce your site’s visibility and accuracy in user search results. | DOCUMENTED | Anthropic · does Anthropic crawl data from the web |
OpenAI also documents a fourth agent, OAI-AdsBot, which
belongs to none of the three classes: OAI-AdsBot only visits pages submitted as ads, and the data collected by OAI-AdsBot is not used to train generative AI foundation models.
— DOCUMENTED · 2026-09-05 ·
OpenAI · crawler documentation.
Google documents no separate crawler for AI Overviews or AI Mode.
To be eligible to be shown in generative AI features on Google Search, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements.Just because a page meets all requirements, best practices, and complies with the policies, doesn’t mean that Google will crawl, index, or serve its content. Indexing and serving aren’t guaranteed.Google · DOCUMENTED · 2026-09-05
Google’s list of common crawlers contains no agent specific to its generative AI features, and its guide ties eligibility for those features to Search eligibility. Our reading: there is nothing extra to allow for Google’s AI answers — Googlebot is the control, and the documented extra condition sits in Search Console, not in robots.txt (see the page-level controls below). INFERRED · Google · AI features and your website
Microsoft documents no separate Copilot crawler.
Bing’s documented crawlers are bingbot, AdIdxBot, BingPreview, MicrosoftPreview and BingVideoPreview — read from the rendered page, because the table is JavaScript-rendered. None is specific to Copilot. Our reading: the controls Microsoft documents for AI answers are page-level meta values, not a user agent (see the page-level controls below). INFERRED · Bing · which crawlers does Bing use · transcribed from the rendered page
One page, because a person asked — and robots.txt may not apply
These agents fetch a page because someone asked a question that needs it. Of the providers on this page that document one, five state a user-initiated exception to robots.txt — OpenAI, Perplexity, Google, Meta and Amazon, the last two in the five-provider table further down. Two document such an agent without stating an exception: Anthropic, which states blanket compliance, and Mistral. Whether that difference shows in practice is not something a table of documentation can answer — that would be OBSERVED, and it would need server logs. It is recorded as an asymmetry in the documentation, nothing more.
| Provider | User agent | Provider’s own words | robots.txt | Class | Verified | Source |
|---|---|---|---|---|---|---|
| OpenAI | ChatGPT-User | When users ask ChatGPT or a CustomGPT a question, it may visit a web page with a ChatGPT-User agent. […] ChatGPT-User is not used for crawling the web in an automatic fashion. | Because these actions are initiated by a user, robots.txt rules may not apply. | DOCUMENTED | OpenAI · crawler documentation | |
| Perplexity | Perplexity-User | Perplexity-User supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response. […] It is not used for web crawling or to collect content for training AI foundation models. | Since a user requested the fetch, this fetcher generally ignores robots.txt rules. | DOCUMENTED | Perplexity · crawlers | |
| Anthropic | Claude-User | Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent. Claude-User allows site owners to control which sites can be accessed through these user-initiated requests. | Anthropic’s Bots respect “do not crawl” signals by honoring industry standard directives in robots.txt. No user-initiated exception is stated. | DOCUMENTED | Anthropic · does Anthropic crawl data from the web | |
| Google-Agent · Google-GeminiNotebook · Google-Read-Aloud · Google-Pinpoint · … | Google-Agent is used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request. The Gemini Notebook fetcher requests individual URLs that Gemini Notebook users have provided as sources for their projects. | Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules. | DOCUMENTED | Google · user-triggered fetchers |
ChatGPT-User is not used to determine whether content may appear in Search. Please use OAI-SearchBot in robots.txt for managing Search opt outs and automatic crawl.
Allowing ChatGPT-User does not get a page into ChatGPT search.
A data-licensing decision — separate from search where the provider documents it as separate
Blocking these is a decision about your content. With the three providers below, the documentation separates that decision from search visibility — twice in the provider’s own words, once as our inference — and the column says which is which.
| Provider | User agent | Provider’s own words | Blocking it costs search visibility? | Class | Verified | Source |
|---|---|---|---|---|---|---|
| OpenAI | GPTBot | GPTBot is used to make our generative AI foundation models more useful and safe. It is used to crawl content that may be used in training our generative AI foundation models. | a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI’s generative AI foundation models. DOCUMENTED | DOCUMENTED | OpenAI · crawler documentation | |
| Anthropic | ClaudeBot | ClaudeBot helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training. When a site restricts ClaudeBot access, it signals that the site’s future materials should be excluded from our AI model training datasets. | A separate agent from Claude-SearchBot, each with its own documented effect. That blocking one leaves the other untouched is not stated in those words. INFERRED | DOCUMENTED | Anthropic · does Anthropic crawl data from the web | |
| Google-Extended robots.txt token — no user-agent string of its own | Google-Extended is a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding […] in Gemini Apps and Grounding with Google Search on Vertex AI. Google-Extended doesn’t have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity. | Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search. DOCUMENTED | DOCUMENTED | Google · common crawlers |
Google-Extended does not govern AI Overviews or AI Mode. Those are
Search, and Search is governed by Googlebot. Google-Extended governs Gemini
model training and grounding in Gemini Apps and Vertex AI, and Google states
that it does not affect a site’s inclusion in Search. That a site which
disallows Google-Extended can still appear in Google’s AI answers follows from
those two documented statements — INFERRED,
not quoted.
Separate controls do not always mean separate crawls. OpenAI:
If your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling.
— DOCUMENTED · 2026-09-05.
Page level: a meta value, a Search Console setting
Not every control is a line in robots.txt. These work at page level, without any user agent, and four providers document one. Microsoft’s two tags separate three states: in the AI answer with content (no tag), in the answer as URL, title and snippet only (NOCACHE), out of the answer entirely (NOARCHIVE) — and Microsoft states that both tags keep the page in ordinary search results.
| Control | Provider | Provider’s own words | Class | Verified | Source |
|---|---|---|---|---|---|
| NOCACHE (robots meta) | Microsoft | Content with the NOCACHE tag may be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer; Going forward, for content in our Bing Index that is labeled NOCACHE, only URLs, Titles and Snippets may be used in training Microsoft’s generative AI foundation models. Blog post of September 2023; the product is called “Bing Chat” there. | DOCUMENTED | Bing Webmaster Blog · NOCACHE and NOARCHIVE | |
| NOARCHIVE (robots meta) | Microsoft | Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers. Going forward, for content in our Bing Index that is labeled NOARCHIVE, we will not use the content for training Microsoft’s generative AI foundation models. We can assure publishers that content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results.Blog post of September 2023; the product is called “Bing Chat” there. | DOCUMENTED | Bing Webmaster Blog · NOCACHE and NOARCHIVE | |
| Search Console · “Search generative AI features” | In addition to the technical requirements for Search, a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search. | DOCUMENTED | Google · AI features and your website | ||
| nosnippet (robots meta) | Apple | Apple will not use data tagged nosnippet as additional context and up-to-date content when AI models are used to generate output for display in Apple products and services. Even if you disallow Applebot-Extended and tag website content with the nosnippet meta tag, your website instructions may still allow Applebot to crawl your webpages. | DOCUMENTED | Apple · about Applebot | |
| noarchive (robots meta) | Amazon | When these user agents access web pages they respect the link-level rel=nofollow directive, and page level robots meta tags of noarchive (do not use the page for model training), noindex (do not index the page) and none (do not index the page). | DOCUMENTED | Amazon · Amazonbot |
Mistral · Meta · Apple · Amazon · Common Crawl
Five providers that document their agents and are not in the tables above. Mistral documents a clean triple. Meta documents a search-class agent, a user fetcher and a training crawler — read from the rendered page, because the raw fetch was refused, and marked as such in every row. Apple’s search crawler also feeds its model training unless a second token is disallowed. Amazon documents three agents with independent settings. Common Crawl is neither a search engine nor an assistant.
| Provider | User agent | Job | Provider’s own words | robots.txt | Class | Verified | Source |
|---|---|---|---|---|---|---|---|
| Mistral | MistralAI-Index | Search | MistralAI-Index is for automated crawling of the web for indexing purposes only. It indexes content for Mistral search, which helps answer user questions in Vibe. Content crawled by MistralAI-Index is not used for generative AI training of any kind. | — | DOCUMENTED | Mistral · robots | |
| Mistral | MistralAI-User | User retrieval | MistralAI-User is for user actions in Vibe. When users ask Vibe a question, it may visit a web page to help answer and include a link to the source in its response. […] It is not used for crawling the web in any automatic fashion, nor to crawl content for generative AI training. | — | DOCUMENTED | Mistral · robots | |
| Mistral | MistralAI-Training | Training | MistralAI-Training crawls web content to help build datasets for training Mistral generative AI models. […] This crawler is not used for search indexing or to answer live user queries in Vibe. | — | DOCUMENTED | Mistral · robots | |
| Meta | Meta-WebIndexer | Search | navigates the web to improve Meta AI search result quality cite and link to your content in Meta AI’s responsestranscribed from the rendered page — the raw fetch was refused | — | DOCUMENTED | Meta · web crawlers transcribed from the rendered page — the raw fetch was refused | |
| Meta | Meta-ExternalFetcher | User retrieval | fetches individual links at a user’s request transcribed from the rendered page — the raw fetch was refused | might bypass robots.txt transcribed from the rendered page — the raw fetch was refused | DOCUMENTED | Meta · web crawlers transcribed from the rendered page — the raw fetch was refused | |
| Meta | Meta-ExternalAgent | Training | crawls the web for use cases such as training foundation AI models or improving products transcribed from the rendered page — the raw fetch was refused | — | DOCUMENTED | Meta · web crawlers transcribed from the rendered page — the raw fetch was refused | |
| Apple | Applebot | Search | The data crawled by Applebot is used to power various features, such as the search technology integrated into many user experiences in Apple’s ecosystem including Spotlight, Siri, and Safari. The data crawled by Applebot may also be used to help train Apple foundation models powering generative AI features across Apple products […]. Web publishers can opt-out from having their content used to train generative foundation models by disallowing Applebot-Extended in the robots.txt file. | — | DOCUMENTED | Apple · about Applebot | |
| Apple | Applebot-Extended robots.txt token — no user-agent string of its own | Training | Applebot-Extended does not crawl webpages. Webpages that disallow Applebot-Extended can still be included in search results. Applebot-Extended is only used to determine how to use the data crawled by the Applebot user agent. | — | DOCUMENTED | Apple · about Applebot | |
| Amazon | Amzn-SearchBot | Search | Amzn-SearchBot is used to improve search experiences in Amazon products and services. By permitting Amzn-SearchBot access to your website, your content is eligible to appear in search experiences such as Alexa. If robots.txt files don’t mention Amzn-SearchBot but allow other search bots, Amzn-SearchBot will crawl in accordance with the robots.txt directives given to other search bots. Amzn-SearchBot does not crawl content for generative AI model training. | — | DOCUMENTED | Amazon · Amazonbot | |
| Amazon | Amzn-User | User retrieval | Amzn-User supports user actions, such as responding to Alexa queries that require up-to-date information. […] Amzn-User does not crawl content for generative AI model training. | Because actions taken by Amzn-User can be initiated by a user, it may not follow all robots.txt directives. | DOCUMENTED | Amazon · Amazonbot | |
| Amazon | Amazonbot | Training | Amazonbot is used to improve our products and services. This helps us provide more accurate information to customers and may be used to train Amazon AI models. This page describes how webmasters can control Amazonbot, Amzn-SearchBot, and Amzn-User interactions with their site. Each user agent setting is independent of the othersOne agent for product improvement and possible training. Amazon documents no finer split inside Amazonbot — the training opt-out is Amazonbot itself; search and user retrieval have their own agents. | — | DOCUMENTED | Amazon · Amazonbot | |
| Common Crawl | CCBot | Open archive | Common Crawl is a non-profit foundation founded with the goal of democratizing access to web information by producing and maintaining an open repository of web crawl data that is universally accessible and analyzable by anyone. Not a search engine and not an assistant. Blocking it affects an archive many parties use, model builders among them — not a product you can appear in. | — | DOCUMENTED | Common Crawl · CCBot |
Two names that circulate, and no provider source for either
Two crawler names come up constantly and appear in no document by their provider. They are UNVERIFIED here, and they stay that way until the provider documents them. Third-party directories were found and refused as sources — several exist, and none is a provider source.
| Provider | What we found | Class | Verified | Source |
|---|---|---|---|---|
| DeepSeek | No crawler documentation by the provider was located. deepseek.com/robots.txt names no agent of its own. Third-party directories list a “DeepSeekBot”; none is a provider source. It stays unverified until DeepSeek documents it. User-Agent: * Allow: / | UNVERIFIED | No provider source located. | |
| ByteDance (“Bytespider”) | No crawler documentation by the provider was located. The agent name circulates in third-party lists and server logs; neither is a provider statement. Same treatment. | UNVERIFIED | No provider source located. |
Google’s words, twice — and the file we maintain anyway
Google Search ignores it. In Google’s words:
You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them.Google · DOCUMENTED · 2026-09-05
It’s completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files. Doing so will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.Google · DOCUMENTED · 2026-09-05
That is the whole Google answer, and this page does not soften it. Which tools do read the file? Some do. This page does not list them, because we have not verified a single one from a provider source — any list here would be UNVERIFIED dressed as a reference. If we verify one, it gets a row, a link and a date like everything else.
We publish an llms.txt on this site and maintain it by hand, for the tools that read it. It is not part of any claim about Google.
| Common claim | Google’s words | Class | Verified | Source |
|---|---|---|---|---|
| Content must be “chunked” into small pieces for AI. | There’s no requirement to break your content into tiny pieces for AI to better understand it. | DOCUMENTED | Google · AI features and your website | |
| Special structured data is needed for AI features. | Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add. | DOCUMENTED | Google · AI features and your website | |
| A separate writing style is needed for AI search. | The best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems. | DOCUMENTED | Google · AI features and your website |
Disclosure · OBSERVED · checkable against the live file
blackoarstudio.com/robots.txt · read
User-agent: *
Allow: /
Sitemap: https://blackoarstudio.com/sitemap-index.xml Every agent on this page is currently allowed on this site, including the training bots. That is a deliberate position for a studio with no licensed archive to protect, and it is not advice. A publisher with content worth licensing might reasonably decide the opposite — the training table above is the table that lets them do so without losing search visibility, where the provider documents the separation.
OBSERVED — a live capture of our own file, 2026-09-05. Not a provider statement.
Next: 2026-12-05
This page describes other people’s systems and will go out of date without warning — in a way a broken-link check will not catch. Our own source list was compiled on 26 August 2026 and re-read on 2026-09-05. What the re-read found — OBSERVED:
Content drifts under stable URLs. That is why every row carries a date rather than just a link.
Provider documents only · read 2026-09-05 · preserved with checksums
Two documents could not be captured raw and are marked above and in their rows. Both were read from the rendered page. Recorded so a reader knows which rows rest on a preserved document and which on a transcription. The IP endpoints in the verification table were fetched on the same date and preserved alongside the documents.