How to Verify AI Crawlers in Your Server Logs

A request that says it is GPTBot, ClaudeBot or Googlebot proves nothing until you check it the way the operator documents: an IP match against the published file, a reverse DNS lookup confirmed by a forward lookup, or a signature check. Here is the procedure, which of 40 crawlers support each method, and what an IP match can and cannot tell you.

To verify a request that claims to be an AI or search crawler, check it the way its operator documents, because the user-agent string is free text that anyone can send. There are three methods. Match the request's IP address against the operator's published IP file. Or run a reverse DNS lookup on the IP and confirm the hostname with a forward lookup. Or, for the few crawlers that sign their requests, verify the signature. On 2026-09-26, of the 40 crawlers in our AI crawler directory, 23 had a published IP file, 9 had a documented reverse-DNS name and 1 signed its requests (some support more than one). MJ12bot offers an ident string on request, 2 are control tokens that never visit, and 11 published nothing you can check.

The check has one more limit. Some operators send several bots from the same addresses, so an IP match can prove the request came from OpenAI, Anthropic, Google or DuckDuckGo without proving which of their bots it was. The last sections cover that case and what to do with each result. To run the IP-file and reverse-DNS checks on one request without a script, paste it into our AI crawler IP verifier.

Which verification method does each crawler support?

Every row below comes from the crawler's card in the directory, where the operator's page is linked and dated.

Method the operator publishes Crawlers (2026-09-26)
IP file and reverse DNS Googlebot, GoogleOther, Bingbot, Applebot, CCBot, AhrefsBot
IP file only GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, DuckDuckBot, DuckAssistBot, Amazonbot, Amzn-SearchBot, Amzn-User, MistralAI-User, MistralAI-Index, LinerBot
Reverse DNS, no IP file YouBot (which also signs requests and lists one range on its page), DataForSeoBot (which lists its subnets on its page), PetalBot
An ident string sent on request MJ12bot: Majestic adds a private string to its requests after you email it
Nothing you can check meta-externalagent, meta-externalfetcher, facebookexternalhit, meta-webindexer, meta-externalads, Diffbot, AI2Bot, ImagesiftBot, SemrushBot, Timpibot, MistralAI-Training
Nothing to check: a control token, not a visitor Google-Extended, Applebot-Extended

Two rows need a word. Google-Extended and Applebot-Extended never appear in logs: Google and Apple crawl with Googlebot and Applebot and use those tokens only to decide how the content may be used. If you see either name in a user-agent string, it is not from Google or Apple. PetalBot's reverse-DNS step is inconsistent in Huawei's own documentation: the steps name aspiegel.com, but the worked example resolves to petalsearch.com. Accept either until Huawei corrects it.

Step 1: pull the requests you want to check

You need four fields per request: the client IP, the user-agent string, the path, and the time. Filter by the name the bot claims:

grep -i "GPTBot" access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head -20

That lists the top 20 IPs claiming to be GPTBot, with request counts. Check at least the busiest ten. If your site sits behind a CDN or reverse proxy, make sure the IP in your log is the visitor's, not the proxy's. Otherwise every check below tests your CDN's address, and every request fails. Your CDN's documentation names the header that carries the original client IP.

Step 2: match the IP against the operator's file

Every JSON IP file linked from the directory uses the same shape, a prefixes list of ipv4Prefix or ipv6Prefix ranges; we ran the script below against each of them on 2026-09-26. It checks one IP against one file with Python's standard library. Save it as check_ip.py:

import ipaddress, json, sys, urllib.request

url, ip = sys.argv[1], ipaddress.ip_address(sys.argv[2])
# Some operators refuse Python's default user agent, so send a plain one.
req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0 (ip-check)"})
data = json.load(urllib.request.urlopen(req))
ranges = [p.get("ipv4Prefix") or p.get("ipv6Prefix") for p in data["prefixes"]]
hit = next((r for r in ranges if ip in ipaddress.ip_network(r)), None)
print(f"{ip} is in {hit}" if hit else f"{ip} is NOT in {url}")

A worked example, with a hypothetical log IP of 132.196.86.10:

python3 check_ip.py https://openai.com/gptbot.json 132.196.86.10
# 132.196.86.10 is in 132.196.86.0/24

On 2026-09-26, 132.196.86.0/24 was the first range in OpenAI's gptbot.json, so a request from that address saying GPTBot passes. The same user-agent string from an address outside the file fails, whatever it claims.

Three habits keep this check honest. Fetch the file fresh rather than copying ranges into a config file: Bing says to refresh its list daily, and OpenAI's chatgpt-user.json was regenerated on 2026-09-25, the day before we read it. Check the file the crawler's card links, not a similar-looking one. And note the exception: Amazon publishes its three lists as web pages of single addresses (1,292, 816 and 1,023 of them on 2026-09-26), not as JSON, so the script needs a small change to read them.

Step 3: confirm reverse DNS with a forward lookup

Reverse DNS turns the IP into a hostname. The owner of an IP range controls that answer, so on its own a hostname proves little: anyone can make their own IP claim to be crawl-1-2-3-4.googlebot.com. The forward lookup closes the gap. It asks the domain in the hostname, which only the operator controls, for the IP, and the two must agree.

Google's documentation gives this example:

host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The request passes when both checks hold: the hostname ends in the operator's domain, and the forward lookup returns the IP from your log. For Google, the accepted domains are googlebot.com, google.com and googleusercontent.com, depending on the crawler type. The other suffixes in the directory are search.msn.com for Bingbot, applebot.apple.com for Applebot, crawl.commoncrawl.org for CCBot (IPv4 only), ahrefs.com or ahrefs.net for AhrefsBot, dataforseo.com for DataForSeoBot, and search.you.com for YouBot.

Step 4: check signed requests, where they exist

A signed request carries three headers defined by HTTP Message Signatures (RFC 9421) and the Web Bot Auth drafts: Signature-Agent names where the bot's public keys live, and Signature-Input and Signature carry the signature. Your server fetches the keys from /.well-known/http-message-signatures-directory on that host and checks the signature. A valid signature proves the request came from the key holder, whatever its IP.

In the directory, YouBot is the only crawler whose operator documents signing: You.com says it signs under Web Bot Auth and publishes keys at you.com/.well-known/http-message-signatures-directory. Google is testing the protocol with some AI agents on its infrastructure, but its page says it does not sign every request and tells sites to fall back to IP and reverse-DNS checks. That is the practical rule for signatures today: treat a valid signature as proof, and a missing one as nothing. Google also notes that major CDNs, WAFs and bot-detection services can verify these signatures for you, which is easier than writing the check yourself.

What an IP match proves, and what it does not

We fetched every IP file the directory links on 2026-09-26 and compared the files of each operator that runs more than one bot.

Operator What the files showed What an IP match tells you
Anthropic One file for ClaudeBot, Claude-User and Claude-SearchBot Anthropic, not which bot
Google Googlebot and GoogleOther share one file Google, not which crawler
DuckDuckGo DuckDuckBot's and DuckAssistBot's files were identical, 486 addresses each DuckDuckGo, not which bot
OpenAI Four separate files; GPTBot and OAI-SearchBot share six identical /25 ranges; ChatGPT-User and OAI-AdsBot share nothing with the others The bot, except in 6 of GPTBot's 18 ranges, which searchbot.json also lists
Perplexity Two files, no shared addresses The bot
Mistral Two files for its user and index agents, no shared addresses; none for MistralAI-Training The bot, for the two with files
Amazon Three web-page lists, no shared addresses The bot

The OpenAI overlap matches what OpenAI says: when a site allows both GPTBot and OAI-SearchBot, it may use one crawl for both purposes. So the user-agent string, not the IP, tells you which bot you are looking at, once the IP has shown the request is genuine.

Here is where that matters. Say you disallowed GPTBot and allowed OAI-SearchBot. Your log shows requests from a shared OpenAI range. If they say OAI-SearchBot, OpenAI is doing what your robots.txt allows. They are not GPTBot ignoring your rule. Only a request that says GPTBot and fetches a disallowed path is a real problem. Our robots.txt tester shows which rule applies to each token, so you can check what your file asks of both.

What should you do with each result?

  • Verified. The request is the operator's. Treat it by your robots.txt policy. If you want it gone, change robots.txt and wait for the operator's stated delay, rather than blocking it at the network. The exception is the user-requested fetchers whose operators say robots.txt may not apply (ChatGPT-User, Perplexity-User, Amzn-User): for those, a CDN rule on the published IP list is the network-level way to refuse them; a login also works.
  • Failed. The request claims a name its operator did not send. Block it at the CDN or firewall with a rule that matches both parts: the claimed user agent and an IP outside the operator's list. Do not block the user-agent string alone, because that also refuses the real crawler.
  • Unverifiable. For the 11 crawlers with nothing to check, the user agent is only a claim. Do not allowlist traffic because of it. If the load is a problem, rate-limit by behaviour rather than trusting or refusing the name.

Two operator warnings apply before you block anything by IP. Anthropic says an IP block may not work as an opt-out, because it also stops its bots from reading your robots.txt. Semrush asks sites not to block it by IP, because it uses no fixed blocks of addresses.

Common mistakes

  • Verifying the proxy. A log that records your CDN's IP makes every crawler fail. Fix the log format first.
  • Stopping at reverse DNS. Without the forward lookup, a spoofer who controls their own reverse DNS passes.
  • Hardcoding ranges. Files change. A copied list goes stale and starts failing real crawlers.
  • Reading an IP match as the bot's name. At Anthropic, Google, DuckDuckGo and inside OpenAI's shared ranges, the IP proves the operator only.
  • Expecting to see Google-Extended. It is a robots.txt token. A request that uses it as a user agent is not from Google.

FAQ

Can a request with a Googlebot user agent be fake?

Yes. The user-agent header is set by whoever sends the request. Google publishes its verification methods because people claim to be Google. Run the reverse and forward DNS check or match the IP against Google's published ranges.

How do I verify GPTBot?

Match the IP against https://openai.com/gptbot.json. OpenAI publishes no reverse-DNS method for its bots. If the IP is in a range that also appears in searchbot.json, the IP proves OpenAI; the user-agent string tells you which bot.

Does Anthropic publish IP ranges for ClaudeBot?

Yes. Anthropic publishes one file covering ClaudeBot, Claude-User and Claude-SearchBot, so a match proves the request is Anthropic's and the user agent tells you which of the three it is.

Why can't I verify Meta's crawlers?

On 2026-09-26, Meta published user-agent strings for its five crawlers but no IP list or reverse-DNS method. Nothing Meta documents lets you prove a request came from Meta, so treat the name as a claim.

Do I need Web Bot Auth?

Not yet. Only one crawler in our directory documents signing, and Google says its own signing is experimental and partial. Check whether your CDN already verifies signatures, and keep the IP and DNS checks as the main method.

Claim ledger

Claim Source Checked
Google's reverse-then-forward DNS check, its accepted domains and the 66.249.66.1 example Google, Verify requests from Google crawlers and fetchers (last updated 2026-03-20) 2026-09-26
Google tests Web Bot Auth with some AI agents, does not sign every request, and says to keep IP and reverse-DNS checks Google, Authenticate requests with Web Bot Auth (last updated 2026-05-04) 2026-09-26
OpenAI publishes separate IP files per bot and may use one crawl for GPTBot and OAI-SearchBot OpenAI bot documentation 2026-09-26
Anthropic's three bots share one IP file; an IP block may stop its bots reading robots.txt Anthropic bot documentation 2026-09-26
YouBot signs requests under Web Bot Auth You.com YouBot documentation 2026-09-26
Method per crawler, suffixes, and Semrush's and Bing's statements Each crawler's card in the directory, with its operator source 2026-09-26
Overlaps and counts in the IP files Our fetch of every linked IP file 2026-09-26

Review this page when an operator changes its verification method, or by 2026-12-26 when the directory's scheduled recheck is due.

Sources

  1. https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests
  2. https://developers.google.com/crawling/docs/crawlers-fetchers/web-bot-auth
  3. https://developers.openai.com/api/docs/bots
  4. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
  5. https://you.com/docs/youbot
  6. https://www.rfc-editor.org/rfc/rfc9421.html
  7. https://lastingcontent.com/bots/data.json
  8. Lasting Content's comparison of every crawler IP file linked from its crawler directory, fetched 2026-09-26

Reviewed

Scope: Post-AI SEO and blog growth. We update this guide as the underlying search behaviour changes.