Free interactive tool

AI Crawler IP Verifier

Paste an IP address and the user-agent string from one log line, and the verifier checks the request the way the crawler’s operator says to: it fetches the operator’s current IP file and looks for the address, and it runs a reverse DNS lookup and confirms the name with a forward lookup. It covers the 26 crawlers in our AI crawler directory whose operators publish something to check, and it says so plainly for the 12 that publish nothing. The answer is one of verified, not verified, inconclusive, or cannot be verified.

How the verifier decides

1  Crawler: the one you pick, or the token the user-agent string
     names (longest match wins, so Applebot-Extended is not Applebot).
2  IP file: fetch the operator file linked from the crawler's card
     (cached up to 1 hour). Pass if the IP is inside any listed range.
3  Ranges on the operator's page: same test, for YouBot and
     DataForSeoBot, which list ranges on their bot page, not in a file.
4  Reverse DNS: look up the IP's name. Pass only if the name ends in
     the operator's documented domain AND a forward lookup of that
     name returns the same IP.
5  Verdict: any pass = verified. No pass and a check that could not
     run = inconclusive. Every check failed = not verified. No method
     published = cannot be verified.

Sources: the reverse-then-forward DNS check is the one Google documents in Verify requests from Google crawlers and fetchers (page last updated 2026-03-20, read 2026-09-26); Bing, Apple, Common Crawl, Ahrefs, Huawei, DataForSEO and You.com document the same kind of check for their own domains. Each IP file and domain comes from the crawler’s card, where the operator’s page is linked and dated. The verifier fetches only those 23 files: it takes a crawler name, never a URL. 9 crawlers have a reverse-DNS domain.

Verify a request from your log

The client IP, not your CDN's. One address per check.

This string names GPTBot (OpenAI).

Worked example: three log lines

Three hypothetical log lines, checked on 2026-09-26:

66.249.66.1    "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
74.7.175.130   "... compatible; GPTBot/1.4; +https://openai.com/gptbot"
203.0.113.50   "... ClaudeBot ..."
  • 66.249.66.1 as Googlebot: verified twice. It is inside 66.249.66.0/27 in Google’s common-crawlers.json, and its reverse DNS name, crawl-66-249-66-1.googlebot.com, resolves back to the same address. This is the example on Google’s own verification page.
  • 74.7.175.130 as GPTBot: verified, with a catch. It is inside 74.7.175.128/25 in gptbot.json. That range is also in OpenAI’s searchbot.json, one of six ranges the two files share, so the verifier adds that the address proves OpenAI, and only the user-agent string says it was GPTBot rather than OAI-SearchBot. If you block GPTBot in robots.txt and see OAI-SearchBot requests from this range, OpenAI is not ignoring your rule.
  • 203.0.113.50 as ClaudeBot: not verified. The address is in none of the entries in Anthropic’s bots.json, and Anthropic publishes no reverse-DNS method, so there is nothing else it could pass. Whoever sent it, it was not Anthropic. (This address is from a block reserved for documentation, so it can never be a real crawler.)

The next step differs for each line. The first two follow your robots.txt policy, since the operator honours it. For the third, block that address, or block requests that say ClaudeBot and come from outside Anthropic’s file. Do not block the ClaudeBot name itself: that also stops the real crawler.

What does a verdict not tell you?

  • Which bot, when an operator shares addresses. Anthropic publishes one file for ClaudeBot, Claude-User and Claude-SearchBot; Googlebot and GoogleOther share a file; DuckDuckGo’s two files were identical when we compared them. A pass there proves the operator, and the user-agent string names the bot.
  • Anything about the 12 crawlers with nothing to check. Meta, Diffbot, Ai2, ImageSift, Semrush, Timpi and MistralAI-Training publish no IP list and no reverse-DNS domain, and Majestic verifies MJ12bot only with a private string it sends after you email it. For these, “cannot be verified” is the honest answer: rate-limit by behaviour rather than trusting or refusing the name.
  • Signed requests. YouBot signs its requests under Web Bot Auth, and Google is testing signatures with some agents. Checking a signature needs the request headers, not an IP, so the verifier does not do it. Your CDN may.
  • Whether your log shows the real client. Behind a CDN or proxy, the IP in a default log line is often the proxy’s, and every check fails. Use the header your CDN documents for the original client IP.
  • What the operator does with the page. A verified request tells you who fetched it, not whether your robots.txt allows it. The robots.txt tester answers that.

Privacy: the IP and user agent you enter are used for this one check and are not stored. The lookups run from our server, so the operator’s file host and the DNS servers see our server’s request, not yours.