
Shocking the SEO World: Google Admits to Hundreds of Undisclosed Crawlers — Who's Actually Scraping Your Site?
Hi everyone, this is Neo.
If you do SEO for overseas independent sites, “Googlebot” is a name you know well. Whether you’re analyzing server logs or configuring robots.txt, Googlebot seems to be the “one true god” that decides your website’s fate.
But recently, Gary Illyes, a Google Search Relations engineer, dropped a bombshell on the Search Off The Record podcast: Googlebot is essentially a relic of the past — Google actually runs hundreds (or more) of undisclosed crawlers and fetchers internally!
What’s going on? What are all these mysterious crawlers doing? Today, let’s dig into Google’s massive internal crawling infrastructure and see what signals hide behind it for independent site sellers.
1. “Googlebot” Is Just a Historical Codename
A lot of people picture Googlebot as one giant, independently operating super-program. In reality, the name dates back to the early 2000s. Back then, Google’s product line was simple — basically just web search — so having a single crawler called Googlebot made sense.
But as products like AdWords (now Google Ads), Google Images, and Google News launched one after another, each product line needed its own crawler to fetch data from the web. Over time, “Googlebot” became a habit — arguably a misnomer.
Today’s Googlebot is no longer a single crawling system. It’s just one of many “clients” interacting with Google’s sprawling underlying crawl infrastructure.
Neo’s take: Stop treating Googlebot as a single entity. When you see a Google User-Agent in your server logs, it could represent Google Search, Google Ads, or even some AI project being tested. Our SEO strategies need to be more granular — you can’t just watch one name.
2. Inside Google’s Crawl Infrastructure: The SaaS-Style “Jack”
If Googlebot isn’t the core system, what is?
Gary revealed that Google has a unified crawling infrastructure (let’s call it “Jack” for now). “Jack” works a lot like a familiar SaaS platform.
Any internal Google product team — search, shopping, or AI — that needs data from the internet doesn’t write its own crawler. They just make an API call to “Jack.”
When they call it, teams can customize parameters like:
- Which User-Agent to use?
- Which product token in robots.txt to honor?
- How long to wait for data (timeout)?
If there are no special requirements, “Jack” provides a set of defaults. Essentially, it’s a service deployed in the cloud or a data center: you give it a command — “go fetch that webpage, but don’t crash their server” — and it does the job.
Neo’s take: This SaaS-style architecture means Google’s crawling capability is highly modular, scalable, and extremely efficient. It also explains how Google can re-parse and crawl the entire web so quickly whenever it ships a new AI feature (like AI Overviews). For independent sites, keeping your server stable and fast (server response time) matters more than ever — because you’re facing a machine army that can be called in parallel at any moment.
3. Why So Many Crawlers, and Why Won’t They Tell You?
If there are so many crawlers, why does Google’s official documentation (the Google Crawlers page) list only a few dozen?
The answer is brutally practical: there’s no room to document them all — and no need.
Google is a giant company. Internally, countless teams have all kinds of odd crawling needs (data research, algorithm testing, etc.). If every tiny crawl action were documented, the page would explode.
So Google takes a “keep the big, drop the small” approach. Only the “major crawlers” — the ones with huge crawl volumes and obvious impact on the broader internet — are publicly documented. The small-scale, low-frequency internal crawlers become “invisible” to the SEO world. That said, Gary mentioned he has a monitoring tool: if an internal crawler’s volume crosses a certain threshold, he goes and has a chat with the team behind it, then decides whether it should be documented publicly.
Neo’s take: When technical folks scan server logs, they often see requests from Google IP ranges with unfamiliar User-Agents. In the past, people would get nervous — is it a fake bot? Now we know the truth: it’s probably some obscure internal Google team doing its job. As long as these crawls aren’t tanking your server and the IP really is Google’s, don’t rush to block them. Blocking blindly could hurt your own exposure in new Google features.
4. The Real Difference Between Crawlers and Fetchers
While discussing these mysterious crawls, Gary also explained a fundamental concept: crawlers and fetchers are not the same thing.
- Crawler:
- Working mode: Batch work.
- Characteristics: They continuously chew through massive URL lists like a tireless assembly line. They don’t care what time it is — if there are resources, they crawl.
- Fetcher:
- Working mode: Single-URL processing.
- Characteristics: Usually triggered by a “user” or a specific action. For example, you click “Test Live URL” in Search Console, or use the Rich Results Test. On the other end of that chain, a real user is anxiously waiting for a result.
Neo’s take: Knowing the difference is a huge help for technical SEO debugging. If you notice your site getting hit with high-frequency requests for the same page at one moment, that’s likely fetcher behavior (maybe your own team testing frantically, or a live app calling in). But if you see steady, wide-ranging page access, that’s a real crawler building an index. Understanding these behavior patterns helps you allocate your crawl budget better.
Summary
This post packs a lot of information — here are the key takeaways:
- Googlebot is just a codename. In reality it’s one of many clients interacting with Google’s unified, SaaS-like crawl infrastructure.
- Google runs hundreds of undisclosed crawlers. Internal teams have endless needs and the docs have finite space, so only high-volume crawlers are published.
- Don’t blindly block unknown crawlers. As long as the IP is genuinely Google’s, unfamiliar crawl requests are likely internal product lines fetching data.
- Understand the crawler/fetcher distinction. The former crawls continuously in bulk in the background; the latter is a one-off, real-time fetch triggered by a command.
Embrace change, see through the underlying logic, and you’ll stay ahead in the independent site SEO game.