Should you block Googlebot? Google's search crawler

Googlebot is operated by Google. It crawls pages for Google's search index and search products, including Images, Video, News, and Discover.

Who should block Googlebot, and who should not?

Do not block it on any page you want in Google. It is the crawler behind Search, and Google says it is also the control for AI features in Search, so no Googlebot setting keeps you in Search but out of AI Overviews. To limit what AI features can show, Google points to snippet controls: nosnippet for a whole page, data-nosnippet for passages, max-snippet for length. They limit your ordinary snippet too, and Google does not describe them as removing a page from AI features. Blocking Google-Extended does not keep you out of AI features either: Google says that token does not affect Search.

The verified facts

User-agent tokenGooglebot
Full user agent in logsMozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
OperatorGoogle
PurposeSearch and AI-search indexers
Honors robots.txtYes, per the operator's documentation.
Published IP rangesdevelopers.google.com/static/crawling/ipranges/common-crawlers.json
Reverse DNS ends ingooglebot.com

What does Google’s documentation add?

  • A Googlebot rule covers Google Search, including Discover and all Search features, plus Google Images, Video, and News. Google says AI is built into Search and that robots.txt rules for Googlebot are the control for how Search crawls your site.
  • To limit what Search shows from a page, including in AI features, Google points to nosnippet, data-nosnippet, max-snippet, and noindex, not to robots.txt.
  • Googlebot-Image, Googlebot-Video, and Googlebot-News have their own tokens, and Google lists Googlebot as a second token for each, so a Googlebot group applies to them when they have no group of their own.
  • The Chrome version in the user agent changes over time. Google says to match it with a wildcard, not an exact version.
  • The string shown above is Googlebot Smartphone. Googlebot Desktop sends Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36.

How do you block Googlebot?

Add this to your robots.txt:

User-agent: Googlebot
Disallow: /

Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.

How do you tell a real Googlebot request from a fake one?

A user-agent match proves nothing, since any client can send Googlebot. Check the IP against the published range file above, re-fetched on a schedule. Or reverse-resolve the IP, check the name ends in googlebot.com, and forward-resolve that name back to the same IP. A request that fails is not Google. The log verification guide has a script for the IP check and a worked reverse-DNS example, or paste the IP into the AI crawler IP verifier, which runs these checks against Google’s current data.

What does blocking Googlebot cost you?

Content excluded from Google Search index and Google products (Images, Video, News, Discover).

What else does Google document?

Google also documents Google-Extended and GoogleOther, each with its own robots.txt token. How Google’s crawlers fit together.

Which crawlers in the same group should you decide on at the same time?

A robots.txt group for Googlebot does nothing to crawlers with other tokens. In the same group, the directory also covers Bingbot (Microsoft), Applebot (Apple), PerplexityBot (Perplexity), DuckAssistBot (DuckDuckGo), DuckDuckBot (DuckDuckGo) and Meta-WebIndexer (Meta), each with its own token and its own documented cost of blocking.