{"schemaVersion":"2026-09-26.2","title":"Lasting Content AI crawler directory","description":"Verified crawler controls and trade-offs from official operator documentation.","canonicalUrl":"https://lastingcontent.com/bots","methodology":"Each record uses the crawler operator's own documentation. Unclear claims stay unclear. Missing official evidence is not replaced with third-party claims.","totalEntries":40,"lastVerifiedAt":"2026-09-26","nextScheduledReviewAt":"2026-12-26","fields":{"slug":"Stable page and record identifier.","name":"Crawler name used by its operator.","userAgentToken":"Exact token documented for robots.txt or requests.","operator":"Organization that operates the crawler.","purposeSlug":"Directory purpose group.","whatItDoes":"Plain-language summary of the operator's documented purpose.","respectsRobots":"Operator claim: true, false, partial, or unclear (the operator does not say).","robotsBlock":"Copy-ready robots.txt rule for the crawler.","verificationMethod":"Official request verification method, or null when none is published.","blockingCost":"Product or discovery access lost when the crawler is blocked.","sourceUrl":"Official operator documentation used as evidence.","sourceTitle":"Label for the official source.","accessedAt":"Date the official source was checked, in YYYY-MM-DD format.","recheckAt":"Scheduled source review date, in YYYY-MM-DD format.","fullUserAgent":"Optional. Full user-agent string as the operator documents it.","ipListUrl":"Optional. Operator's published IP-range file.","reverseDns":"Optional. Hostname suffixes a verified reverse-DNS lookup ends in.","publishedRanges":"Optional. IP ranges the operator lists on its bot page rather than in a separate file.","notes":"Optional. Operator-specific facts that bear on the block decision, from the operator's documentation.","decision":"Optional. Lasting Content's editorial reading of who should block the crawler; not an operator claim.","extraSources":"Optional. Further operator pages used for the notes.","titleJob":"Optional. Short job label used in the page title.","noindexReason":"Optional. Why the card is kept out of search: no verifiable operator documentation.","titleOperator":"Optional. Short operator name used in the page title."},"changeHistory":[{"date":"2026-09-21","summary":"Published the first verified directory with official operator sources, robots.txt controls, verification methods, and scheduled review dates."},{"date":"2026-09-21","summary":"Omitted Bytespider because no current official operator documentation was available to verify its purpose and control behavior."},{"date":"2026-09-25","summary":"Spot-checked 15 entries (OpenAI, Anthropic, Google, Bing, Perplexity, the three partial/unclear Meta and Amazon rows) against current operator documentation. Corrected Claude-User's robots.txt field from false to true: Anthropic's own bot documentation states plainly that opting out of any of its three bots, including Claude-User, is done through the same robots.txt Disallow mechanism, unlike OpenAI's ChatGPT-User and Perplexity's Perplexity-User, which explicitly say robots.txt rules may not apply to them. No other checked field was wrong."},{"date":"2026-09-26","summary":"Re-read the DuckDuckGo, Meta, Semrush, Ahrefs and Majestic crawler pages and added per-bot full user agents, IP-list links, reverse-DNS suffixes, documented notes, and a block-or-allow reading. Corrected three fields: SemrushBot's token is the bare SemrushBot (Semrush documents no version suffix); meta-externalfetcher's robots.txt bypass is for user-requested fetches, not security checks (that reason belongs to FacebookExternalHit); DuckDuckBot's blocking cost now says DuckDuckGo sources most traditional links from Bing."},{"date":"2026-09-26","summary":"Re-read the OpenAI, Anthropic, Google and Bing crawler pages and added the same per-bot detail to their ten cards. Corrected Bingbot's reverse-DNS check (Bing documents only search.msn.com, not bing.com) and changed ChatGPT-User's robots.txt field from false to partial, since OpenAI says the rules may not apply, the same wording mapped to partial for Meta's fetcher."},{"date":"2026-09-26","summary":"Re-read the Apple, Perplexity, Common Crawl, Amazon, You.com and Diffbot crawler pages and added per-bot detail to their ten cards. Corrected Applebot-Extended, which Apple says does not crawl and only controls training use; Diffbot's token, which Diffbot documents as the bare Diffbot; PerplexityBot's blocking cost, which no longer claims a block removes you from Perplexity answers, since Perplexity-User fetches separately; YouBot's purpose and cost, which You.com describes only in terms of its search results; and Diffbot's robots.txt field, now partial, since Diffbot says robots.txt can be overridden under a customer's agreement with a site."},{"date":"2026-09-26","summary":"Re-read the Ai2, ImageSift, Huawei Petal, DataForSEO, Mistral and Liner crawler pages and added per-bot detail. Corrected AI2Bot's robots.txt field to unclear, now labelled not documented (Ai2's notice does not mention robots.txt) and dropped an unsourced dataset name; moved ImagesiftBot to other product crawlers (ImageSift describes a similar-image index, not a public dataset) and fixed its verification field; moved LinerBot to search indexers, as Liner describes it. Timpi documents Collector crawling only in general terms and names no user agent, so the Timpibot card now says so and is kept out of search."},{"date":"2026-09-26","summary":"Added three crawlers from their operators' pages: MistralAI-Index and MistralAI-Training, so Mistral's search and training controls each have a card, and OpenAI's OAI-AdsBot, which reviews ChatGPT ad landing pages. OAI-AdsBot's robots.txt field is unclear, because OpenAI does not say whether it reads robots.txt."},{"date":"2026-09-26","summary":"Added a publishedRanges field for ranges an operator lists on its bot page rather than in a file (YouBot's /24, DataForSEO's three IPv4 and three IPv6 subnets), and PetalBot's two reverse-DNS domains as a field, so the AI crawler IP verifier can check them."}],"entries":[{"slug":"gptbot","name":"GPTBot","userAgentToken":"GPTBot/1.4","operator":"OpenAI","purposeSlug":"training","whatItDoes":"It collects public web content to train OpenAI's future models. Blocking it keeps your words out of the next training set; it does not remove you from today's AI answers.","respectsRobots":true,"robotsBlock":"User-agent: GPTBot\nDisallow: /","verificationMethod":"IP match against OpenAI's published gptbot.json list","blockingCost":"OpenAI treats the block as a signal that your content should not be used to train its generative AI foundation models. ChatGPT search and user-requested visits are controlled by other tokens.","sourceUrl":"https://developers.openai.com/api/docs/bots","sourceTitle":"OpenAI bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot","ipListUrl":"https://openai.com/gptbot.json","notes":["OpenAI says each robots.txt setting is independent: a site can allow OAI-SearchBot to appear in ChatGPT search while disallowing GPTBot to keep content out of training.","If a site allows both GPTBot and OAI-SearchBot, OpenAI may use one crawl for both purposes to avoid crawling twice.","When it fetches robots.txt, GPTBot may add a robots.txt marker to its user-agent string, so those requests are easy to spot in logs that omit the path."],"decision":"Block it if you do not want your writing used to train OpenAI's models. The block costs nothing in ChatGPT search, which runs on a separate crawler, OAI-SearchBot, under a separate setting. That makes GPTBot one of the cheapest blocks in this directory. The case for leaving it open is indirect: you may want future models to know your brand, product names, or definitions without a search step. OpenAI documents no traffic benefit to allowing it."},{"slug":"oai-searchbot","name":"OAI-SearchBot","userAgentToken":"OAI-SearchBot/1.4","operator":"OpenAI","purposeSlug":"search-index","whatItDoes":"It crawls websites so OpenAI can surface them in ChatGPT search results and citations. It is separate from GPTBot training controls.","respectsRobots":true,"robotsBlock":"User-agent: OAI-SearchBot\nDisallow: /","verificationMethod":"IP match against OpenAI's published searchbot.json list","blockingCost":"Your pages will not be shown in ChatGPT search answers, though OpenAI says they can still appear as navigational links.","sourceUrl":"https://developers.openai.com/api/docs/bots","sourceTitle":"OpenAI bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot","ipListUrl":"https://openai.com/searchbot.json","notes":["OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, but can still appear as navigational links.","A robots.txt change can take about 24 hours to reach ChatGPT search, according to OpenAI.","OpenAI recommends allowing both the token in robots.txt and requests from its published IP ranges."],"decision":"Leave it open if you want ChatGPT search to cite and link to you; it is the OpenAI crawler tied to those answers. Blocking it removes you from ChatGPT search answers, apart from navigational links, and does nothing about training, which GPTBot governs. One mistake to check for: a CDN or firewall bot rule that blocks OpenAI's addresses wholesale, so robots.txt says yes while the network says no. Compare your bot-management settings with the published IP list before assuming you are open."},{"slug":"chatgpt-user","name":"ChatGPT-User","userAgentToken":"ChatGPT-User/1.0","operator":"OpenAI","purposeSlug":"retrieval","whatItDoes":"It fetches pages on demand when a user's request needs them, so what it reads feeds live answers and citations rather than a training set.","respectsRobots":"partial","robotsBlock":"User-agent: ChatGPT-User\nDisallow: /","verificationMethod":"IP match against OpenAI's published chatgpt-user.json list","blockingCost":"ChatGPT users cannot have ChatGPT open your pages, if the block holds. OpenAI says robots.txt may not apply to these user-initiated visits.","sourceUrl":"https://developers.openai.com/api/docs/bots","sourceTitle":"OpenAI bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot","ipListUrl":"https://openai.com/chatgpt-user.json","notes":["It visits a page when a ChatGPT or Custom GPT user's question needs it, and for GPT Actions. OpenAI says it is not used for automatic crawling.","Because a user started the action, OpenAI says robots.txt rules may not apply.","OpenAI says it is not used to decide whether content may appear in Search; OAI-SearchBot handles search opt-outs."],"decision":"A robots.txt block is a weak lever here: OpenAI says the rule may not apply, and each visit comes from a person's question. Leave it open for public pages. If a section must stay out of ChatGPT sessions, put it behind a login. If you must refuse the traffic, match the published IP list at the CDN, knowing you are turning away a reader's request, not a crawler. To leave ChatGPT search, block OAI-SearchBot instead."},{"slug":"oai-adsbot","name":"OAI-AdsBot","userAgentToken":"OAI-AdsBot/1.0","operator":"OpenAI","purposeSlug":"other-product","titleJob":"ad-review crawler","whatItDoes":"It visits the landing pages of ads submitted to ChatGPT, to check them against OpenAI's ad policies and to judge when each ad is relevant to show. It visits no other pages.","respectsRobots":"unclear","robotsBlock":"User-agent: OAI-AdsBot\nDisallow: /","verificationMethod":"IP match against OpenAI's published adsbot.json list","blockingCost":"OpenAI does not say. The crawler only reviews pages submitted as ChatGPT ads, so a block can only affect ads that point at your pages.","sourceUrl":"https://developers.openai.com/api/docs/bots","sourceTitle":"OpenAI bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot","ipListUrl":"https://openai.com/adsbot.json","notes":["OpenAI says OAI-AdsBot only visits pages submitted as ads on ChatGPT. When an ad is submitted, OpenAI may visit its landing page to check that it complies with OpenAI's policies.","OpenAI may also use content from the landing page to decide when the ad is most relevant to show to users.","OpenAI says the data OAI-AdsBot collects is not used to train generative AI foundation models.","OpenAI's page names OAI-SearchBot and GPTBot as its robots.txt controls and does not say whether OAI-AdsBot reads robots.txt.","On 2026-09-26 adsbot.json listed two IPv4 ranges, both /25, in a file dated 2026-05-12."],"decision":"If you advertise in ChatGPT, keep it open on your landing pages, at the CDN and firewall as well as in robots.txt: it is how OpenAI reviews the page your ad sends people to. If you do not advertise, it should only arrive when someone else submits an ad that links to your page. A request that names OAI-AdsBot from an address outside adsbot.json did not come from OpenAI. Because OpenAI does not say the bot reads robots.txt, a firewall rule on the user agent and the IP list is the only way to refuse it that does not depend on OpenAI."},{"slug":"claudebot","name":"ClaudeBot","userAgentToken":"ClaudeBot","operator":"Anthropic","purposeSlug":"training","whatItDoes":"It collects public web content to train Anthropic's future models. Blocking it keeps your words out of the next training set; it does not remove you from today's AI answers.","respectsRobots":true,"robotsBlock":"User-agent: ClaudeBot\nDisallow: /","verificationMethod":"IP match against Anthropic's published bots.json list","blockingCost":"Anthropic says the block signals that the site's future materials should be excluded from its model training datasets. Claude search indexing and user-requested fetches are controlled separately.","sourceUrl":"https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler","sourceTitle":"Anthropic bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","ipListUrl":"https://claude.com/crawling/bots.json","notes":["Anthropic says blocking ClaudeBot signals that the site's future materials should be excluded from its training datasets. The wording covers material going forward.","It supports the non-standard Crawl-delay extension, for example Crawl-delay: 1, as a softer alternative to a block.","Anthropic asks for the robots.txt rule on every subdomain you want to opt out, and says its bots will not try to bypass CAPTCHAs."],"decision":"Block it if you do not want future material in Anthropic's training data. The block does not remove you from Claude's search or from user-requested fetches, which have their own tokens. If the issue is load rather than training, Crawl-delay is the documented middle ground. Use robots.txt rather than an IP block: Anthropic says an IP block may not hold as an opt-out, because the bot can no longer read your rules."},{"slug":"claude-user","name":"Claude-User","userAgentToken":"Claude-User","operator":"Anthropic","purposeSlug":"retrieval","whatItDoes":"It fetches pages on demand when a user's request needs them, so what it reads feeds live answers and citations rather than a training set.","respectsRobots":true,"robotsBlock":"User-agent: Claude-User\nDisallow: /","verificationMethod":"IP match against Anthropic's published bots.json list","blockingCost":"Anthropic says disabling Claude-User stops Claude retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search.","sourceUrl":"https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler","sourceTitle":"Anthropic bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","ipListUrl":"https://claude.com/crawling/bots.json","notes":["When someone asks Claude a question, Claude may visit websites as Claude-User. Anthropic says the token lets site owners control which sites these user-initiated requests can reach.","Anthropic documents no robots.txt exception for Claude-User. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it."],"decision":"Leave it open for public pages you want Claude users to be able to read on request. Because Anthropic says Claude-User honors robots.txt, a block here does work, which makes it the right tool for sections you want out of live Claude answers while Claude-SearchBot indexes the rest. Blocking it for training reasons achieves nothing: training is ClaudeBot's job."},{"slug":"claude-searchbot","name":"Claude-SearchBot","userAgentToken":"Claude-SearchBot","operator":"Anthropic","purposeSlug":"search-index","whatItDoes":"It crawls and indexes web content to improve the relevance and accuracy of Claude search results.","respectsRobots":true,"robotsBlock":"User-agent: Claude-SearchBot\nDisallow: /","verificationMethod":"IP match against Anthropic's published bots.json list","blockingCost":"Anthropic says disabling Claude-SearchBot stops it indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results.","sourceUrl":"https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler","sourceTitle":"Anthropic bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","ipListUrl":"https://claude.com/crawling/bots.json","notes":["Anthropic says it navigates the web to improve search result quality for users, analyzing content to make search responses more relevant and accurate."],"decision":"Allow it if you want Claude's search to know your pages; the documented cost of a block is lower visibility and accuracy when Claude users search. A site that wants out of training but in search can block ClaudeBot and leave Claude-SearchBot and Claude-User open, the split Anthropic's three separate tokens allow. Block it only for sections you want kept out of Claude's search entirely.","titleJob":"search bot"},{"slug":"googlebot","name":"Googlebot","userAgentToken":"Googlebot","operator":"Google","purposeSlug":"search-index","whatItDoes":"It crawls pages for Google's search index and search products, including Images, Video, News, and Discover.","respectsRobots":true,"robotsBlock":"User-agent: Googlebot\nDisallow: /","verificationMethod":"Reverse DNS ending in googlebot.com, confirmed by a forward lookup, or an IP match against common-crawlers.json","blockingCost":"Content excluded from Google Search index and Google products (Images, Video, News, Discover).","sourceUrl":"https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers","sourceTitle":"Google bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)","ipListUrl":"https://developers.google.com/static/crawling/ipranges/common-crawlers.json","reverseDns":["googlebot.com"],"notes":["A Googlebot rule covers Google Search, including Discover and all Search features, plus Google Images, Video, and News. Google says AI is built into Search and that robots.txt rules for Googlebot are the control for how Search crawls your site.","To limit what Search shows from a page, including in AI features, Google points to nosnippet, data-nosnippet, max-snippet, and noindex, not to robots.txt.","Googlebot-Image, Googlebot-Video, and Googlebot-News have their own tokens, and Google lists Googlebot as a second token for each, so a Googlebot group applies to them when they have no group of their own.","The Chrome version in the user agent changes over time. Google says to match it with a wildcard, not an exact version.","The string shown above is Googlebot Smartphone. Googlebot Desktop sends Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36."],"decision":"Do not block it on any page you want in Google. It is the crawler behind Search, and Google says it is also the control for AI features in Search, so no Googlebot setting keeps you in Search but out of AI Overviews. To limit what AI features can show, Google points to snippet controls: nosnippet for a whole page, data-nosnippet for passages, max-snippet for length. They limit your ordinary snippet too, and Google does not describe them as removing a page from AI features. Blocking Google-Extended does not keep you out of AI features either: Google says that token does not affect Search.","extraSources":[{"url":"https://developers.google.com/search/docs/appearance/ai-features","title":"Google: AI features and your website"},{"url":"https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests","title":"Google: verify requests from Google crawlers"}]},{"slug":"google-extended","name":"Google-Extended","userAgentToken":"Google-Extended","operator":"Google","purposeSlug":"training-grounding","whatItDoes":"It is a robots.txt control token for Gemini model training and for grounding Gemini responses with content from Google's search index. It is not an HTTP user agent.","respectsRobots":true,"robotsBlock":"User-agent: Google-Extended\nDisallow: /","verificationMethod":"Google-Extended is a robots.txt control token, not an HTTP user-agent string","blockingCost":"Content excluded from Gemini model training and grounding uses.","sourceUrl":"https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers","sourceTitle":"Google bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","notes":["Google-Extended has no user-agent string of its own. Google crawls with its existing crawlers and uses the token only to control how the fetched content may be used.","It covers training future Gemini models (Gemini Apps and the Vertex AI API for Gemini) and grounding in Gemini Apps and Grounding with Google Search on Vertex AI.","Google says it does not affect a site's inclusion in Google Search and is not used as a ranking signal."],"decision":"Block it if you do not want Google using your content to train Gemini or to ground Gemini app answers; Google says the block costs nothing in Search. It does not keep you out of AI Overviews or AI Mode, which are Search features governed by Googlebot and snippet controls. You will never see it in your logs, because it never visits: the rule in your robots.txt is the only evidence the block exists.","extraSources":[{"url":"https://developers.google.com/search/docs/appearance/ai-features","title":"Google: AI features and your website"}],"titleJob":"Gemini control"},{"slug":"googleother","name":"GoogleOther","userAgentToken":"GoogleOther","operator":"Google","purposeSlug":"other-product","whatItDoes":"It is Google's generic crawler for public content used by various product teams, including one-off internal research and development crawls.","respectsRobots":true,"robotsBlock":"User-agent: GoogleOther\nDisallow: /","verificationMethod":"Reverse DNS ending in googlebot.com, confirmed by a forward lookup, or an IP match against common-crawlers.json","blockingCost":"Content excluded from various Google product uses for fetching publicly accessible content.","sourceUrl":"https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers","sourceTitle":"Google bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GoogleOther) Chrome/W.X.Y.Z Safari/537.36","ipListUrl":"https://developers.google.com/static/crawling/ipranges/common-crawlers.json","reverseDns":["googlebot.com"],"notes":["Google says crawling preferences for GoogleOther do not affect any specific product. Product teams use it to fetch public content, for example in one-off crawls for internal research and development.","GoogleOther-Image and GoogleOther-Video fetch images and video. Google lists GoogleOther as a second token for both, so a GoogleOther group applies to them when they have no group of their own.","It is one of Google's common crawlers, which Google says always respect robots.txt for automatic crawls."],"decision":"Blocking it is low-risk and low-reward. Google says crawl preferences for GoogleOther don't affect any specific product, so a block should not touch Search; it removes you only from product and research crawls. Block it if Google's crawl load exceeds what Search needs, or if you object to research use; otherwise it is not worth a line in robots.txt.","titleJob":"research crawler"},{"slug":"bingbot","name":"Bingbot","userAgentToken":"bingbot","operator":"Microsoft","purposeSlug":"search-index","whatItDoes":"It crawls pages for the Bing search index.","respectsRobots":true,"robotsBlock":"User-agent: bingbot\nDisallow: /","verificationMethod":"Reverse DNS ending in search.msn.com, confirmed by a forward lookup, or an IP match against bingbot.json","blockingCost":"Content excluded from the Bing Search index, and from search products that draw on Bing's results, such as the traditional links DuckDuckGo says it largely sources from Bing.","sourceUrl":"https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0","sourceTitle":"Microsoft bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36","ipListUrl":"https://www.bing.com/toolbox/bingbot.json","reverseDns":["search.msn.com"],"notes":["Once Bingbot finds a robots.txt section for itself, Bing says it ignores the generic * section, so repeat the general rules in its section.","Bing supports Crawl-delay in a Bingbot section, and Bing Webmaster Tools has a Crawl control tool that sets the crawl rate by the hour.","Bing says search engines cache robots.txt for at least a few hours, so a change can take a few hours to show in its crawling.","BingPreview, which makes page snapshots, sends the same bingbot user-agent string. Bing says to refresh its IP list daily rather than hardcode ranges."],"decision":"Do not block it on pages you want found. Bing's index reaches beyond Bing: DuckDuckGo, for one, says it sources most of its traditional links from Bing. If Bingbot's load is a problem, use Crawl-delay or the hourly Crawl control in Bing Webmaster Tools rather than a block. Whenever you add a Bingbot section, copy every generic rule into it, because Bing stops reading the * section once it finds its own.","extraSources":[{"url":"https://www.bing.com/webmasters/help/how-to-verify-bingbot-3905dc26","title":"Bing: how to verify Bingbot"},{"url":"https://www.bing.com/webmasters/help/how-to-create-a-robots-txt-file-cb7c31ec","title":"Bing: how to create a robots.txt file"},{"url":"https://duckduckgo.com/duckduckgo-help-pages/results/sources","title":"DuckDuckGo: where search results come from"}]},{"slug":"applebot","name":"Applebot","userAgentToken":"Applebot","operator":"Apple","purposeSlug":"search-index","whatItDoes":"It crawls pages for search in Spotlight, Siri, and Safari. Apple says the data may also help train its foundation models and add context to AI-generated answers in its products.","respectsRobots":true,"robotsBlock":"User-agent: Applebot\nDisallow: /","verificationMethod":"Reverse DNS in applebot.apple.com, or an IP match against Apple's published applebot.json CIDR list","blockingCost":"Your pages drop out of Apple's search results in Spotlight, Siri, and Safari. Training and AI-answer use have their own controls, so a block is not needed for those.","sourceUrl":"https://support.apple.com/en-us/119829","sourceTitle":"Apple bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)","ipListUrl":"https://search.developer.apple.com/applebot.json","reverseDns":["applebot.apple.com"],"notes":["If robots.txt does not mention Applebot but does mention Googlebot, Apple says Applebot follows the Googlebot rules.","Applebot does not follow Crawl-delay. Apple says it slows down on its own when a site slows or returns errors.","The nosnippet tag, in a meta tag or an X-Robots-Tag: applebot header, stops Apple generating a description or web answer for a page and stops it using that content as context for AI-generated output.","Pages marked isAccessibleForFree: false in structured data stay eligible for search, but Apple will not use them as context for AI-generated output."],"decision":"Leave it open on pages you want found in Siri, Spotlight, and Safari; there is rarely a reason to block the crawler itself. Apple separates the concerns publishers usually have: disallow Applebot-Extended to keep content out of foundation-model training, and use nosnippet, or isAccessibleForFree set to false on paywalled pages, to keep it out of AI-generated answers. The cost of nosnippet: Apple says suggestions for that page then show only its title. One trap: a site with Googlebot rules and no Applebot section gets the Googlebot rules applied to Apple, so check that those rules suit Apple too."},{"slug":"applebot-extended","name":"Applebot-Extended","userAgentToken":"Applebot-Extended","operator":"Apple","purposeSlug":"training","whatItDoes":"It is a robots.txt control token, not a crawler. Apple uses it to decide whether content Applebot crawls may train Apple's foundation models for Apple Intelligence, Services, and Developer Tools.","respectsRobots":true,"robotsBlock":"User-agent: Applebot-Extended\nDisallow: /","verificationMethod":"Applebot-Extended is a robots.txt control token, not an HTTP user-agent string","blockingCost":"Apple will not use your content to train its general-purpose foundation models. Search appearance in Spotlight, Siri, and Safari is unaffected.","sourceUrl":"https://support.apple.com/en-us/119829","sourceTitle":"Apple bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","notes":["Apple says Applebot-Extended does not crawl webpages and is used only to decide how data crawled by Applebot may be used.","Pages that disallow Applebot-Extended can still appear in search results, and Apple says its rules are not considered in search ranking.","It does not cover AI-generated answers in Siri and Search. Apple's opt-out for those is the nosnippet tag on specific content."],"decision":"Block it if you do not want Apple training its foundation models on your content; Apple says the block costs nothing in search ranking or appearance. It is the cheapest training opt-out in this directory, next to Google-Extended. It does not stop Apple using your pages as context for AI answers in Siri; for that, add nosnippet to the content you want withheld. Because it never visits, your logs will not show it: the robots.txt rule is the only record.","titleJob":"training control"},{"slug":"perplexitybot","name":"PerplexityBot","userAgentToken":"PerplexityBot/1.0","operator":"Perplexity","purposeSlug":"search-index","whatItDoes":"It gathers and indexes public web content so Perplexity can surface and link websites in search results.","respectsRobots":true,"robotsBlock":"User-agent: PerplexityBot\nDisallow: /","verificationMethod":"IP match against Perplexity's published perplexitybot.json list","blockingCost":"Perplexity will not surface or link your pages in its search results. Perplexity-User can still visit a page when a user asks about it.","sourceUrl":"https://docs.perplexity.ai/docs/resources/perplexity-crawlers","sourceTitle":"Perplexity bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)","ipListUrl":"https://www.perplexity.com/perplexitybot.json","notes":["Perplexity says PerplexityBot surfaces and links websites in its search results and is not used to crawl content for AI foundation models.","Perplexity recommends allowing it in robots.txt and permitting requests from its published IP ranges; for a web application firewall it suggests a rule that matches both the user agent and the IP list.","A robots.txt change may take up to 24 hours to apply."],"decision":"Leave it open if you want Perplexity to cite and link you: it is the agent behind Perplexity's search results, and Perplexity says it is not used to crawl content for AI foundation models. Blocking it removes you from those results but does not stop Perplexity-User fetching a page a user asks about. If you use a WAF, follow Perplexity's pattern: an allow rule that matches both the user agent and the published IP list. It keeps impostors out only if you also block other requests that use the same user agent."},{"slug":"perplexity-user","name":"Perplexity-User","userAgentToken":"Perplexity-User/1.0","operator":"Perplexity","purposeSlug":"retrieval","whatItDoes":"It fetches pages on demand when a user's request needs them, so what it reads feeds live answers and citations rather than a training set.","respectsRobots":false,"robotsBlock":"User-agent: Perplexity-User\nDisallow: /","verificationMethod":"IP match against Perplexity's published perplexity-user.json list","blockingCost":"Where the block is enforced, Perplexity cannot visit your pages to answer a user's question. Perplexity says this fetcher generally ignores robots.txt.","sourceUrl":"https://docs.perplexity.ai/docs/resources/perplexity-crawlers","sourceTitle":"Perplexity bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)","ipListUrl":"https://www.perplexity.com/perplexity-user.json","notes":["When a user asks Perplexity a question, it may visit a page to answer and include a link to that page in its response.","Because a user requested the fetch, Perplexity says this fetcher generally ignores robots.txt rules.","Perplexity says it is not used for web crawling or to collect content for training AI foundation models."],"decision":"Treat it as a reader, not a crawler: it arrives because of someone's question, and Perplexity says the answer can include a link to your page. Leave it open for public pages. A robots.txt rule will generally be ignored, so if a section must stay out of Perplexity answers, the controls that work are a login or a CDN rule on the published IP list. Blocking it does nothing about indexing, which is PerplexityBot's job.","titleJob":"live fetcher"},{"slug":"duckassistbot","name":"DuckAssistBot","userAgentToken":"DuckAssistBot/1.2","operator":"DuckDuckGo","purposeSlug":"search-index","whatItDoes":"It crawls pages in real time for DuckDuckGo's AI-assisted answers, which cite their sources. DuckDuckGo says the data is not used to train AI models.","respectsRobots":true,"robotsBlock":"User-agent: DuckAssistBot\nDisallow: /","verificationMethod":"IP match against DuckDuckGo's published duckassistbot.json list","blockingCost":"Content excluded from DuckDuckGo AI-assisted answers; does not affect organic search ranking.","sourceUrl":"https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot","sourceTitle":"DuckDuckGo bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)","ipListUrl":"https://duckduckgo.com/duckassistbot.json","notes":["A robots.txt block takes effect after 72 hours, according to DuckDuckGo. Opt-out problems go to crawling@duckduckgo.com.","DuckDuckGo says opting out does not affect organic rankings or whether a site appears in its search results.","It crawls from the same addresses as DuckDuckBot: on 2026-09-26 the two published IP files were identical. Separate the two by user-agent token, not by IP."],"decision":"Block it if you do not want DuckDuckGo's AI answers to summarize your text, and you accept losing the source link those answers carry. DuckDuckGo's help page makes this a clean trade: the opt-out leaves organic rankings alone and the data is not used for training, so a block gives up AI-answer visibility. Sites that want DuckDuckGo traffic should usually leave it open, because the answers cite their sources. Wait three days before checking whether a block worked.","titleJob":"AI-answer bot"},{"slug":"duckduckbot","name":"DuckDuckBot","userAgentToken":"DuckDuckBot","operator":"DuckDuckGo","purposeSlug":"search-index","whatItDoes":"It crawls pages for DuckDuckGo search results.","respectsRobots":true,"robotsBlock":"User-agent: DuckDuckBot\nDisallow: /","verificationMethod":"IP match against DuckDuckGo's published duckduckbot.json list","blockingCost":"Your pages drop out of the indexes DuckDuckGo builds with its own crawler. DuckDuckGo says it sources most of the traditional links in its results from Bing. It does not say whether a DuckDuckBot block touches those; expect them to follow your Bingbot rules.","sourceUrl":"https://duckduckgo.com/duckduckgo-help-pages/results/duckduckbot","sourceTitle":"DuckDuckGo bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)","ipListUrl":"https://duckduckgo.com/duckduckbot.json","notes":["DuckDuckGo says it maintains DuckDuckBot and its own indexes to support its results, but sources the traditional links and images in those results largely from Bing.","Its IP list is the same as DuckAssistBot's: on 2026-09-26 duckduckbot.json and duckassistbot.json were identical files of 486 addresses each, so an IP match proves a request came from DuckDuckGo, not which of the two bots sent it.","DuckDuckGo states only that it respects WWW::RobotRules, the name of a Perl robots.txt parser. It documents no crawl-delay support and no time for a robots.txt change to take effect."],"decision":"Almost no public site should block it. It identifies itself with a checkable IP list and feeds DuckDuckGo's own indexes. If the goal is to keep your text out of DuckDuckGo's AI-assisted answers, block DuckAssistBot instead: DuckDuckGo says that opt-out does not change organic rankings. Block DuckDuckBot only for sections no search engine should crawl, and give Bingbot the same rule, because Bing supplies most of DuckDuckGo's ordinary links.","extraSources":[{"url":"https://duckduckgo.com/duckduckgo-help-pages/results/sources","title":"DuckDuckGo: where search results come from"}]},{"slug":"ccbot","name":"CCBot","userAgentToken":"CCBot/2.0","operator":"Common Crawl","purposeSlug":"archive-data","whatItDoes":"It builds a public dataset that other systems, including AI trainers, consume downstream.","respectsRobots":true,"robotsBlock":"User-agent: CCBot\nDisallow: /","verificationMethod":"Reverse DNS in crawl.commoncrawl.org (IPv4 only), or an IP match against ccbot.json","blockingCost":"Content excluded from Common Crawl archives and public web datasets.","sourceUrl":"https://commoncrawl.org/ccbot","sourceTitle":"Common Crawl bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"CCBot/2.0 (https://commoncrawl.org/faq/)","ipListUrl":"https://index.commoncrawl.org/ccbot.json","reverseDns":["crawl.commoncrawl.org"],"notes":["Common Crawl says it knows of crawlers falsely identifying themselves as CCBot, and recommends verifying requests. Reverse DNS works over IPv4; it is not yet supported over IPv6, where the IP file is the check.","It obeys Crawl-delay, slows down on HTTP 429 or 5xx responses, and by default waits a few seconds between requests to the same site.","Common Crawl says its dataset is a random sample of the web, not a full archive of any site. It does not execute JavaScript or use cookies.","A robots.txt block stops future crawls. Common Crawl does not document removing pages from past crawls; for material already collected, its route is a legal opt-out request to info@commoncrawl.org, which it lists in its public Opt-Out Ledger to alert downstream users."],"decision":"Block it if you do not want future pages in open web datasets. Common Crawl publishes its archive for anyone to use, including AI developers who never crawl you themselves, so this one rule reaches further than a block on any single AI company's crawler. The block covers future crawls; for pages already collected, Common Crawl's route is a legal opt-out request. The cost to readers is nil; the loss is research and archive presence. Verify requests before judging its behaviour from logs, because Common Crawl says impostors use its name.","extraSources":[{"url":"https://commoncrawl.org/faq","title":"Common Crawl FAQ"}]},{"slug":"meta-externalagent","name":"meta-externalagent","userAgentToken":"meta-externalagent/1.1","operator":"Meta","purposeSlug":"training-grounding","whatItDoes":"It collects public web content to train Meta's future models and to improve Meta products with indexed content.","respectsRobots":true,"robotsBlock":"User-agent: meta-externalagent\nDisallow: /","verificationMethod":"No official IP ranges published","blockingCost":"Content excluded from Meta AI training datasets and from product features that use its web index.","sourceUrl":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","sourceTitle":"Meta Web Crawlers (Meta for Developers)","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)","notes":["Meta gives two example uses: training foundation AI models, and improving products by indexing content directly. One token covers both, so you cannot allow the indexing use and refuse the training.","Meta's documentation points site owners to robots.txt rather than non-standard formats like NoAI tags. It does not document support for a noai meta tag."],"decision":"Block it if you do not want your content in Meta's model training. The documented cost is small and vague: Meta says the same crawler improves products by indexing content, but names no product that sends readers back. Meta ties citations and links in Meta AI answers to a different crawler, Meta-WebIndexer, so a site can refuse training here and still allow Meta-WebIndexer. Publishers who license their content, or whose archive is the product, have the clearest case for a block.","titleJob":"AI training crawler"},{"slug":"meta-externalfetcher","name":"meta-externalfetcher","userAgentToken":"meta-externalfetcher/1.1","operator":"Meta","purposeSlug":"retrieval","whatItDoes":"It fetches individual links at a user's request, including when Meta's AI navigates a site to complete a task for someone.","respectsRobots":"partial","robotsBlock":"User-agent: meta-externalfetcher\nDisallow: /","verificationMethod":"No official IP ranges published","blockingCost":"Meta users may be unable to have Meta's AI open your pages for them, to the extent the fetcher honors the block. Meta says it may bypass robots.txt because the fetch was requested by a user.","sourceUrl":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","sourceTitle":"Meta Web Crawlers (Meta for Developers)","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers)","notes":["Meta says this crawler may bypass robots.txt because it performs fetches a user requested. A robots.txt rule records your preference; it does not enforce it.","Meta also lists evaluating and improving agentic AI capabilities as a use, so these fetches are not only one-off page reads for a person."],"decision":"A robots.txt block here is mostly a statement of preference, since Meta says this fetcher may ignore it. If the concern is load or access to private pages, fix it where it can be enforced: put private content behind a login and rate-limit at the CDN. Meta publishes no IP list, so a firewall rule usually has to match the user-agent string, which any client can send or omit. For public pages you want people to reach through Meta's assistant, leave it open."},{"slug":"facebookexternalhit","name":"FacebookExternalHit","userAgentToken":"facebookexternalhit/1.1","operator":"Meta","purposeSlug":"other-product","whatItDoes":"It fetches and caches the title, description, and thumbnail for links shared on Facebook, Instagram, and Messenger.","respectsRobots":"partial","robotsBlock":"User-agent: facebookexternalhit\nDisallow: /","verificationMethod":"No official IP ranges published","blockingCost":"Shared links may not show an up-to-date title, description, or thumbnail in Meta apps.","sourceUrl":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","sourceTitle":"Meta Web Crawlers (Meta for Developers)","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)","notes":["It must be able to fetch the page within a few seconds, or Facebook cannot display the preview.","Open Graph tags must appear within the first 1 MB of the page, and the server must support gzip and deflate encoding.","Meta says it might bypass robots.txt when it runs security or integrity checks, such as scanning for malware.","After you fix a page, Meta's Sharing Debugger tool or the Sharing API forces a fresh crawl, so an old preview does not linger."],"decision":"Leave it open on every page you want people to share. Blocking it does not stop sharing; it breaks the preview, so links to your site appear on Facebook, Instagram, and Messenger without a proper title, description, or image. The common reason to block it is a private or staging area, and there robots.txt is the wrong tool, because Meta says this crawler may bypass robots.txt for security checks. Put those pages behind authentication instead.","titleJob":"link previewer"},{"slug":"meta-webindexer","name":"Meta-WebIndexer","userAgentToken":"meta-webindexer/1.1","operator":"Meta","purposeSlug":"search-index","whatItDoes":"It crawls and indexes web content to improve Meta AI search results, citations, and links.","respectsRobots":true,"robotsBlock":"User-agent: meta-webindexer\nDisallow: /","verificationMethod":"No official IP ranges published","blockingCost":"Content may not be cited or linked in Meta AI search responses.","sourceUrl":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","sourceTitle":"Meta Web Crawlers (Meta for Developers)","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"meta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers)","notes":["Meta says allowing it helps Meta cite and link to your content in Meta AI's responses. It is the Meta crawler tied to citations.","Meta describes its purpose as improving Meta AI search result quality. Unlike meta-externalagent, its description does not mention model training."],"decision":"Allow it if you want Meta AI answers to link to you; it is the one Meta crawler whose documented job includes citation. A site that wants to stay out of Meta's model training and keep those citations can block meta-externalagent and leave Meta-WebIndexer open. Block it if you do not want Meta AI to index your pages for answers, knowing the block also gives up the links. It does not stop meta-externalfetcher, which fetches pages at a user's request and may bypass robots.txt.","titleJob":"AI search crawler"},{"slug":"meta-externalads","name":"Meta-ExternalAds","userAgentToken":"meta-externalads/1.1","operator":"Meta","purposeSlug":"other-product","whatItDoes":"It crawls public web content to improve Meta advertising and other business products and services.","respectsRobots":true,"robotsBlock":"User-agent: meta-externalads\nDisallow: /","verificationMethod":"No official IP ranges published","blockingCost":"Meta loses your pages for whatever advertising and business-product uses the crawler serves; Meta does not say more, and no referral, preview, or citation feature depends on it in Meta's documentation.","sourceUrl":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","sourceTitle":"Meta Web Crawlers (Meta for Developers)","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"meta-externalads/1.1 (+/documentation/sharing/webmasters/web-crawlers)","notes":["Meta's whole description is one sentence: it crawls for use cases such as improving advertising and other business-related products and services. Meta does not say which pages it reads, how often, or where the data goes after that.","It has no documented role in link previews. Those come from FacebookExternalHit, so blocking this crawler does not change how shared links look."],"decision":"Meta documents no benefit to a publisher in allowing it: the stated use is Meta's own advertising and business products, not referrals, previews, or citations. For most sites a block is low-risk. If you advertise on Meta, note that the crawler page does not say whether ad landing pages depend on this crawler; treat that as unknown rather than assume either way.","titleJob":"ads crawler"},{"slug":"amazonbot","name":"Amazonbot","userAgentToken":"Amazonbot/0.1","operator":"Amazon","purposeSlug":"other-product","whatItDoes":"It crawls public content to improve Amazon products and services, and Amazon says the content may also train its AI models.","respectsRobots":true,"robotsBlock":"User-agent: Amazonbot\nDisallow: /","verificationMethod":"IP match against Amazon's published Amazonbot address list","blockingCost":"Content excluded from Amazon product-improvement crawls and AI model training.","sourceUrl":"https://developer.amazon.com/amazonbot","sourceTitle":"Amazon bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36","ipListUrl":"https://developer.amazon.com/amazonbot/ip-addresses/","notes":["Amazon says Amazonbot improves its products and services and may be used to train Amazon AI models.","Amazon honors the noarchive robots meta tag as do not use the page for model training, so a page can stay crawlable but out of training.","Amazon says a change may take about 24 hours to apply, but the bot may use a robots.txt copy cached for up to 30 days, and treats an unreachable robots.txt as if none exists. It does not support Crawl-delay.","Amazon says sites that allow Amazonbot may be eligible for its Content Partners program, which lists a +1% affiliate commission boost on eligible sales, free hosting credits, and AI traffic management tools."],"decision":"Amazon is the one operator here that ties allowing its crawler to a benefit: sites that allow Amazonbot may be eligible for Amazon Content Partners, which lists a +1% affiliate commission boost on eligible sales, hosting credits, and traffic tools (terms at contentpartners.amazon.com). Beyond that, allowing it mostly feeds Amazon's products and possibly its model training. A middle path Amazon documents: allow crawling and add noarchive to pages you do not want used for training. Blocking it does not affect Alexa search, which uses Amzn-SearchBot. Amazon says a settings change may take about 24 hours to apply, and the bot may work from a robots.txt copy up to 30 days old.","titleJob":"product crawler"},{"slug":"amzn-searchbot","name":"Amzn-SearchBot","userAgentToken":"Amzn-SearchBot/0.1","operator":"Amazon","purposeSlug":"search-index","whatItDoes":"It crawls and indexes web content for search experiences in Amazon products and services, including Alexa.","respectsRobots":true,"robotsBlock":"User-agent: Amzn-SearchBot\nDisallow: /","verificationMethod":"IP match against Amazon's published Amzn-SearchBot address list","blockingCost":"Content is not eligible to appear in Amazon search experiences such as Alexa.","sourceUrl":"https://developer.amazon.com/amazonbot","sourceTitle":"Amazon bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36","ipListUrl":"https://developer.amazon.com/amazonbot/searchbot-ip-addresses/","notes":["If robots.txt does not mention Amzn-SearchBot but allows other search bots, Amazon says it crawls under the rules given to those bots."],"decision":"Allow it if you want to be eligible for Amazon search experiences such as Alexa answers; Amazon says it does not collect training data, so the usual training objection does not apply. Note the fallback: with no Amzn-SearchBot group, it follows whatever you set for other search bots, so a site that blocks everything except Googlebot and Bingbot may still admit it. Name it explicitly if you want it out."},{"slug":"amzn-user","name":"Amzn-User","userAgentToken":"Amzn-User/0.1","operator":"Amazon","purposeSlug":"retrieval","whatItDoes":"It fetches current web information for user actions, such as an Alexa question that needs a live answer.","respectsRobots":"partial","robotsBlock":"User-agent: Amzn-User\nDisallow: /","verificationMethod":"IP match against Amazon's published Amzn-User address list","blockingCost":"Amazon users may not be able to retrieve the site's current information for a live request.","sourceUrl":"https://developer.amazon.com/amazonbot","sourceTitle":"Amazon bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36","ipListUrl":"https://developer.amazon.com/amazonbot/live-ip-addresses/","notes":["It fetches live information when a customer asks something that needs it, such as an Alexa question about current details.","Because a user can start its actions, Amazon says it may not follow all robots.txt directives."],"decision":"Leave it open for public pages whose current details people ask Alexa about, such as opening hours, prices, or availability; blocking it risks an answer built from stale or third-party data. Amazon says it may not follow every robots.txt rule, so enforce any hard block at the CDN with the published IP list. It does not train models, so blocking it for training reasons achieves nothing."},{"slug":"youbot","name":"YouBot","userAgentToken":"YouBot/1.0","operator":"You.com","purposeSlug":"search-index","whatItDoes":"It powers the You.com search engine, discovering and indexing pages for its search results.","respectsRobots":true,"robotsBlock":"User-agent: YouBot\nDisallow: /","verificationMethod":"Cloudflare Web Bot Auth signatures, reverse DNS in search.you.com, or the 68.67.112.0/24 range","blockingCost":"Your pages drop out of You.com search results.","sourceUrl":"https://you.com/docs/youbot","sourceTitle":"You.com bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; YouBot/1.0; +https://docs.you.com/youbot; env:prod) Chrome/X.X.X.X Safari/537.36","reverseDns":["search.you.com"],"publishedRanges":["68.67.112.0/24"],"notes":["You.com recommends cryptographic verification: YouBot signs requests under Cloudflare's Web Bot Auth standard, with public keys at you.com/.well-known/http-message-signatures-directory.","Genuine requests come from 68.67.112.0/24, and reverse DNS names take the form youbot-68-67-112-106.search.you.com.","It honors Crawl-delay, slows down on HTTP 429, and caches robots.txt for 30 minutes, so changes apply quickly."],"decision":"YouBot is the easiest crawler in this directory to verify and to throttle: a single published /24, reverse DNS, signed requests, Crawl-delay support, and a 30-minute robots.txt cache. That makes a block or a slowdown low-effort and reliable. Whether to allow it depends on whether You.com's search users are an audience for you; the crawl feeds its search results. Most sites lose little either way, so throttle before you block."},{"slug":"diffbot","name":"Diffbot","userAgentToken":"Diffbot","operator":"Diffbot","purposeSlug":"search-index","whatItDoes":"It proactively crawls the web for Diffbot's general search engine and Knowledge Graph. Diffbot says it is not used for AI training.","respectsRobots":"partial","robotsBlock":"User-agent: Diffbot\nDisallow: /","verificationMethod":"No official IP ranges published","blockingCost":"Your pages drop out of the web index behind the Diffbot Knowledge Graph and its search services.","sourceUrl":"https://www.diffbot.com/docs/crawl/faq/robots-txt","sourceTitle":"Diffbot bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","notes":["Diffbot says its crawler builds a general search engine and Knowledge Graph and is not used for AI training.","It adheres to robots.txt, including Crawl-delay, by default. Diffbot says robots.txt can be overridden in specific cases, typically when its customer has an agreement with the site being crawled.","A second agent, Diffbot-User, handles requests a person makes through Diffbot software. It has its own robots.txt token.","Diffbot's customers can also run their own crawls through its Crawlbot and Extract products, where Diffbot recommends they set their own user agent and robots.txt adherence is a setting that is on by default and can be switched off. Those requests may not say Diffbot at all."],"decision":"Blocking the Diffbot token is low-cost for most sites: it removes you from a knowledge graph and search index that people reach through Diffbot's APIs rather than by browsing. Two limits to know. The block does not cover Diffbot-User, which has its own token. And crawls that Diffbot's customers run through its software may use their own user agent, so a Diffbot rule will not catch them. Block both tokens if you want Diffbot's own traffic out.","titleJob":"knowledge-graph crawler"},{"slug":"ai2bot","name":"AI2Bot","userAgentToken":"AI2Bot","operator":"Allen Institute for AI","purposeSlug":"training","whatItDoes":"It explores certain domains to find web content, which Ai2 says is used to train open language models.","respectsRobots":"unclear","robotsBlock":"User-agent: AI2Bot\nDisallow: /","verificationMethod":"No IP list or reverse-DNS method published","blockingCost":"Where the block holds, your pages stay out of the content Ai2 collects to train open language models. Ai2's notice does not say whether the bot reads robots.txt.","sourceUrl":"https://allenai.org/crawler","sourceTitle":"Ai2 crawling notice","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler)","notes":["Ai2's crawling notice gives the bot's purpose and its user-agent string, and says the string can be used to filter or reject its traffic. It does not mention robots.txt.","Ai2 publishes no IP list or reverse-DNS method, so the user-agent string is the only documented identifier."],"decision":"Block it if you do not want your pages in the data Ai2 uses to train open language models. Because the notice does not say the bot reads robots.txt, add the robots.txt rule and also reject the user agent at your server or CDN, which is the method Ai2's own notice points to. Some publishers allow it because the models it trains are open; Ai2 documents no benefit to your site beyond that.","titleOperator":"Ai2"},{"slug":"imagesiftbot","name":"ImagesiftBot","userAgentToken":"ImagesiftBot","operator":"ImageSift (Hive)","purposeSlug":"other-product","whatItDoes":"It collects publicly available images, with their page text and alt text, for ImageSift's web intelligence products, which search for similar images.","respectsRobots":true,"robotsBlock":"User-agent: ImagesiftBot\nDisallow: /","verificationMethod":"No IP list or reverse-DNS method published","blockingCost":"Your images and their page text drop out of ImageSift's similar-image index.","sourceUrl":"https://imagesift.com/about","sourceTitle":"ImageSift (Hive) bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (compatible; ImagesiftBot; +imagesift.com)","notes":["Along with images, it saves the host URL, the text on the page, and the image alt text, and indexes them so ImageSift's products can find similar images.","It honors Crawl-delay as the minimum gap between the start of requests: with Crawl-delay: 5 it makes at most one request in each 5-second slot.","If no rule names ImagesiftBot but one names Googlebot, it follows the Googlebot rules. ImageSift takes opt-out requests at support@imagesift.com."],"decision":"The case for blocking it is image reuse: its index exists so ImageSift's customers can find where similar images appear, and ImageSift documents no way it sends visitors back. Check the Googlebot inheritance before you decide: with a Googlebot group and no ImagesiftBot group, it takes Googlebot's permissions, which are usually generous. Name it explicitly to block it, or set a Crawl-delay if load is the only concern.","titleJob":"image crawler","titleOperator":"ImageSift"},{"slug":"petalbot","name":"PetalBot","userAgentToken":"PetalBot","operator":"Huawei","purposeSlug":"search-index","whatItDoes":"It crawls pages for Petal Search indexing and related Huawei search services.","respectsRobots":true,"robotsBlock":"User-agent: PetalBot\nDisallow: /","verificationMethod":"Reverse DNS then forward DNS. Huawei's steps name aspiegel.com, but its worked example resolves to petalsearch.com","blockingCost":"Petal says your pages become unsearchable in Petal search and in the Huawei Assistant and AI Search services it powers.","sourceUrl":"https://webmaster.petalsearch.com/site/petalbot","sourceTitle":"Huawei bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (compatible;PetalBot;+https://webmaster.petalsearch.com/site/petalbot)","reverseDns":["aspiegel.com","petalsearch.com"],"notes":["Petal says PetalBot builds the index behind Petal search and the content recommendations in Huawei Assistant and AI Search, both powered by Petal Search.","Huawei's verification steps say the reverse-DNS name should be in aspiegel.com, but the worked example on the same page resolves to petalbot-114-119-128-10.petalsearch.com. Accept a request only when the name ends in aspiegel.com or petalsearch.com and a forward lookup of that name returns the same IP.","After a block, Petal says it may take several months to clear pages already in its index. Urgent removal requests go to petalbot@huawei.com.","Petal documents no Crawl-delay support; it says it adjusts crawling to server capacity, site quality, and how often the site updates."],"decision":"Allow it if people using Huawei devices are part of your audience, since PetalBot feeds both Petal search and the recommendations in Huawei Assistant and AI Search. If they are not, a block costs little. Two practical points: an existing index entry can take months to clear after a block, so email petalbot@huawei.com if removal is urgent; and verify requests by checking that the reverse-DNS name ends in aspiegel.com or petalsearch.com and that its forward lookup returns the same IP."},{"slug":"semrushbot","name":"SemrushBot","userAgentToken":"SemrushBot","operator":"Semrush","purposeSlug":"seo-tools","whatItDoes":"It crawls for Semrush's SEO and marketing database; readers never see its fetches.","respectsRobots":true,"robotsBlock":"User-agent: SemrushBot\nDisallow: /","verificationMethod":"No IP list published; Semrush asks sites not to block it by IP because it uses no consecutive IP blocks","blockingCost":"Semrush stops reading your pages, so their content and outbound links drop out of its backlink index. Links that other sites point at you are still found on those sites.","sourceUrl":"https://www.semrush.com/bot/","sourceTitle":"Semrush bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","notes":["The SemrushBot token covers the backlink index. Other Semrush tools crawl under their own tokens, each blocked separately: SiteAuditBot, SemrushBot-BA, SemrushBot-SI, SemrushBot-SWA, SplitSignalBot, SemrushBot-OCOB, SemrushBot-FT, RyteBot, and SemrushBot-ESI.","It honors Crawl-delay up to 10 seconds; higher values are cut to 10. With no crawl-delay set, it adjusts its rate to your server load.","Your robots.txt must return HTTP 200. On a 4xx response SemrushBot assumes there are no restrictions; on a 5xx it stops crawling the whole site.","Semrush says it can take up to one hour or 100 requests to notice a robots.txt change, and each subdomain needs its own robots.txt file."],"decision":"Blocking it costs readers nothing, but it does not do what many site owners expect. Your backlink profile stays visible in Semrush, because links to you are found on the linking sites; what disappears is the content and outbound links of your own pages. Most sites should throttle instead: a Crawl-delay of up to 10 seconds cuts the load and keeps your pages current in the tool. Block it outright when crawl load is a real cost, and add the tool-specific tokens if you want every Semrush crawler out."},{"slug":"ahrefsbot","name":"AhrefsBot","userAgentToken":"AhrefsBot/7.0","operator":"Ahrefs","purposeSlug":"seo-tools","whatItDoes":"It crawls for Ahrefs's SEO and marketing database; readers never see its fetches.","respectsRobots":true,"robotsBlock":"User-agent: AhrefsBot\nDisallow: /","verificationMethod":"Published IP ranges, and reverse DNS ending in ahrefs.com or ahrefs.net","blockingCost":"Content excluded from Ahrefs backlink database and Yep search engine index.","sourceUrl":"https://ahrefs.com/robot","sourceTitle":"Ahrefs bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)","ipListUrl":"https://api.ahrefs.com/v3/public/crawler-ip-ranges","reverseDns":["ahrefs.com","ahrefs.net"],"notes":["AhrefsBot feeds two products: the Ahrefs marketing-intelligence database and Yep, a search engine Ahrefs runs. Blocking it removes you from both.","Ahrefs' Site Audit tool crawls under a separate token, AhrefsSiteAudit, limited to 30 URLs a minute. Owners who verify a site in Ahrefs can let that audit crawler ignore robots.txt on their own site.","Crawl-delay is followed for HTML pages but not while rendering JavaScript, when the bot may fetch several page assets at once.","Returning 4xx or 5xx status codes makes it slow down automatically, which Ahrefs suggests for outages or maintenance windows."],"decision":"Unlike the other SEO crawlers in this directory, AhrefsBot also feeds a search engine, Yep, so a block costs a little search visibility as well as tool coverage. For most sites the better lever is Crawl-delay, which Ahrefs honors for HTML fetches. Block it when crawl load is a real cost, or when you do not want your pages' content in a competitor-research database. A block does not hide links pointing to you: AhrefsBot records those when it crawls the linking sites."},{"slug":"mj12bot","name":"MJ12bot","userAgentToken":"MJ12bot/v1.4.8","operator":"Majestic","purposeSlug":"seo-tools","whatItDoes":"It crawls for Majestic's SEO and marketing database; readers never see its fetches.","respectsRobots":true,"robotsBlock":"User-agent: MJ12bot\nDisallow: /","verificationMethod":"No IP list: a distributed crawler. Majestic adds a pre-arranged ident string to its requests when asked by email","blockingCost":"Majestic stops mapping the links on your pages. Links that other sites point at you are still recorded when it crawls those sites.","sourceUrl":"https://mj12bot.com/","sourceTitle":"Majestic bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","notes":["MJ12bot runs as a distributed community crawler with no fixed IP range, and Majestic asks sites not to block it by IP.","To verify it, email bot@majestic12.co.uk. Majestic will send a private ident string with its requests to your site, in a CRAWLER-IDENT header or the user-agent.","It honors Crawl-delay up to 20 seconds. An MJ12bot section in robots.txt replaces the * section rather than adding to it, so copy any crawl-delay into it.","It stores the link graph, not page content. If it cannot fetch robots.txt it assumes crawling is allowed, but some failures, such as 403 Forbidden, count as a full disallow."],"decision":"MJ12bot is the hardest SEO crawler here to verify: no IP list and no documented reverse-DNS check, only an ident string you request by email. Anyone can claim its user-agent, so a firewall rule on that string blocks impostors and the real bot alike, and robots.txt stays the practical control. Throttle it with Crawl-delay if its load is acceptable; block it if a place in Majestic's link index is not worth the requests. The block removes the links on your pages from Majestic, not the links other sites point at you."},{"slug":"dataforseobot","name":"DataForSeoBot","userAgentToken":"DataForSeoBot","operator":"DataForSEO","purposeSlug":"seo-tools","whatItDoes":"It crawls for DataForSEO's SEO and marketing database; readers never see its fetches.","respectsRobots":true,"robotsBlock":"User-agent: DataForSeoBot\nDisallow: /","verificationMethod":"Reverse DNS crawling-gateway-*.dataforseo.com, within published subnets 136.243.220.208/29, 136.243.228.176/29 and 136.243.228.192/29 (IPv4) and three IPv6 /64s","blockingCost":"DataForSEO stops adding the links on your pages to its backlink database. Links other sites point at you are still found on those sites.","sourceUrl":"https://dataforseo.com/dataforseo-bot","sourceTitle":"DataForSEO bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 (compatible; DataForSeoBot; +https://dataforseo.com/dataforseo-bot)","reverseDns":["dataforseo.com"],"publishedRanges":["136.243.220.208/29","136.243.228.176/29","136.243.228.192/29","2a01:4f8:2b03:38b::/64","2a01:4f8:2b03:38c::/64","2a01:4f8:2b03:38d::/64"],"notes":["DataForSEO says the bot adds the links it finds on your pages to its backlink database and collects no other data, triggers no ads, and does not add to your Google Analytics traffic.","Its default gap between requests is 5 seconds; a Crawl-delay line in its robots.txt group changes that.","It publishes three IPv4 subnets and three IPv6 subnets, and reverse DNS names of the form crawling-gateway-136-243-228-176.dataforseo.com.","DataForSEO lists a Backlinks API and a Backlink Database among its products, the commercial outlets for this data."],"decision":"DataForSEO mainly sells data through APIs and databases: the links its bot records feed a Backlinks API and Backlink Database, so they can end up in products built on that data. Blocking it costs readers nothing and removes the links on your own pages from that supply; it does not hide links pointing to you. With a 5-second default gap it is gentle, so block it for data reasons rather than load. Its small, published subnets make it one of the easiest bots to verify."},{"slug":"timpibot","name":"Timpibot","userAgentToken":"Timpibot","operator":"Timpi","purposeSlug":"search-index","whatItDoes":"Timpi's knowledge base says community-run Collector nodes crawl the internet to discover public pages for its search index. It does not name a user agent.","respectsRobots":"unclear","robotsBlock":"User-agent: Timpibot\nDisallow: /","verificationMethod":"None possible: Timpi says Collectors run on any system and need no static IP address","blockingCost":"Unknown: Timpi does not document what the crawler feeds.","sourceUrl":"https://timpi.gitbook.io/timpis-knowledge-base/timpis-knowledge-base/timpi-nodes/what-nodes-run-on-timpis-decentralised-network","sourceTitle":"Timpi knowledge base: what nodes run on Timpi's network","accessedAt":"2026-09-26","recheckAt":"2026-12-26","notes":["Timpi says Collectors are run by the community, work on any system, and need no static IP address or open ports. That is likely why Timpi publishes no IP list or reverse-DNS check."],"decision":"With no operator documentation of the bot itself, there is nothing to weigh a block against and no way to verify a request. If Timpibot traffic is a problem, add a robots.txt rule and block the user agent at your CDN, accepting that Timpi does not say whether its Collectors read robots.txt. If it is not a problem, leave it. Recheck when Timpi publishes a crawler page.","noindexReason":"Timpi documents its crawling only in general terms: its knowledge base describes community-run Collector nodes but names no user agent and says nothing about robots.txt. The Timpibot token comes from server logs and third-party bot lists, not from Timpi. The card stays for readers who see Timpibot in their logs, but it is kept out of search results until Timpi documents the bot."},{"slug":"mistralai-user","name":"MistralAI-User","userAgentToken":"MistralAI-User/1.0","operator":"Mistral","purposeSlug":"retrieval","whatItDoes":"It fetches pages on demand when a user's request needs them, so what it reads feeds live answers and citations rather than a training set.","respectsRobots":true,"robotsBlock":"User-agent: MistralAI-User\nDisallow: /","verificationMethod":"IP match against Mistral's published MistralAI-User address list","blockingCost":"Vibe users cannot have Mistral's assistant visit your pages to answer a question.","sourceUrl":"https://docs.mistral.ai/robots","sourceTitle":"Mistral bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)","ipListUrl":"https://mistral.ai/mistralai-user-ips.json","notes":["Mistral says MistralAI-User is for user actions in Vibe. It is not used to crawl the web automatically, nor to crawl content for generative AI training.","Mistral documents two more agents with their own tokens: MistralAI-Index crawls automatically to index content for Mistral search, and MistralAI-Training builds datasets for training Mistral models.","When it visits a page to help answer, Mistral says the response may include a link to the source."],"decision":"Leave it open for public pages: it visits because a Vibe user asked something, and Mistral says the answer can link to the source. Blocking it does not keep you out of Mistral's search index or training data, which belong to MistralAI-Index and MistralAI-Training. To opt out of Mistral training, disallow MistralAI-Training by name; to leave Mistral search, disallow MistralAI-Index."},{"slug":"mistralai-index","name":"MistralAI-Index","userAgentToken":"MistralAI-Index/1.0","operator":"Mistral","purposeSlug":"search-index","whatItDoes":"It crawls the web automatically to build the index behind Mistral search, which Vibe uses to answer questions. Mistral says nothing it collects is used for AI training.","respectsRobots":true,"robotsBlock":"User-agent: MistralAI-Index\nDisallow: /","verificationMethod":"IP match against Mistral's published MistralAI-Index address list","blockingCost":"Your pages leave the Mistral search index that Vibe draws on to answer questions.","sourceUrl":"https://docs.mistral.ai/robots","sourceTitle":"Mistral bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)","ipListUrl":"https://mistral.ai/mistralai-index-ips.json","notes":["Mistral says MistralAI-Index crawls for indexing purposes only, and that content it crawls is not used for generative AI training of any kind.","On 2026-09-26 its IP file listed two single IPv4 addresses (/32) in a file dated 2026-04-19. None of them appears in the MistralAI-User file, so an address tells you which Mistral agent sent a request."],"decision":"Leave it open if you want Vibe to find your pages through Mistral search; disallow it if you would rather stay out of that index. It is not a training opt-out, because Mistral says this crawler feeds no training; that job belongs to MistralAI-Training. It also does not stop MistralAI-User, which fetches a page when a Vibe user's question needs it. The IP file is small enough to check by hand: two published addresses, so a request outside them that claims this name is not Mistral's."},{"slug":"mistralai-training","name":"MistralAI-Training","userAgentToken":"MistralAI-Training/1.0","operator":"Mistral","purposeSlug":"training","whatItDoes":"It crawls web content to build datasets for training Mistral's generative AI models. Blocking it keeps your pages out of those datasets; it does not remove you from Mistral search or Vibe answers.","respectsRobots":true,"robotsBlock":"User-agent: MistralAI-Training\nDisallow: /","verificationMethod":"No IP list published: Mistral publishes address files for MistralAI-User and MistralAI-Index, not for this crawler","blockingCost":"Your content is kept out of the datasets Mistral builds to train its models. Mistral says this crawler serves no search index and no live answers.","sourceUrl":"https://docs.mistral.ai/robots","sourceTitle":"Mistral bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)","notes":["Mistral says webmasters can disallow MistralAI-Training in robots.txt, and that it is not used for search indexing or to answer live user queries in Vibe.","Its section of Mistral's page gives a user-agent string but, unlike the two other Mistral agents, no IP file."],"decision":"Block it if you do not want Mistral training on your writing. Mistral separates training from search and live fetches, so the block costs nothing in Vibe answers or Mistral search. That makes it the cheapest Mistral rule, like GPTBot for OpenAI. The limit: with no published addresses, a request that claims to be MistralAI-Training cannot be confirmed as Mistral's, so the robots.txt rule is the lever, not a firewall list."},{"slug":"linerbot","name":"LinerBot","userAgentToken":"LinerBot/1.0","operator":"Liner","purposeSlug":"search-index","whatItDoes":"It gathers information from the web and indexes it for Liner's search engine.","respectsRobots":true,"robotsBlock":"User-agent: LinerBot\nDisallow: /","verificationMethod":"IP match against Liner's published linerbot.json list","blockingCost":"Your pages drop out of Liner's search index.","sourceUrl":"https://docs.getliner.com/docs/linerbot","sourceTitle":"Liner bot documentation","accessedAt":"2026-09-26","recheckAt":"2026-12-26","fullUserAgent":"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; LinerBot/1.0; +https://docs.getliner.com/docs/linerbot)","ipListUrl":"https://docs.getliner.com/linerbot.json","notes":["Liner describes LinerBot as a crawler that gathers information from the internet and indexes it for its search engine.","It accepts path-level rules, so you can allow /public/ and disallow /private/ rather than blocking the whole site.","Liner says its IP pool may change over time and points to linerbot.json for the current list."],"decision":"The stakes are small either way: LinerBot feeds one company's search engine. Allow it if you want your pages findable in Liner; block it, or only its private paths, if you do not. Liner's page does not say whether the index is used for model training, so treat that as unknown. Verify against the published IP file before blaming it for load, since the user agent is easy to copy."}]}