Google Engineer Leaks the Truth About Googlebot Crawl Limits — Is 15MB Just a Myth?
Hi everyone, this is Neo.
If you do SEO for independent sites, you probably know Googlebot (Google’s crawler) pretty well. Every day we try every trick to get Googlebot to crawl more, crawl deeper, so our pages get indexed faster and rank higher.
Recently, two big names from the Google Search team — Gary Illyes and Martin Splitt — dropped some seriously hardcore knowledge on their podcast Search Off The Record, and it completely upends a lot of what the outside world believes about Googlebot’s crawl limits.
Especially that legendary “15MB crawl limit” — the truth may be nothing like what you imagine. Today I’m going to break it all down and explain what this intel actually means for people running independent sites.
1. Googlebot Isn’t Really a “Program”
First, Martin and Gary dropped a bombshell: Googlebot is not actually a standalone program.
Inside Google, they run a massive, cloud-based crawling infrastructure (a SaaS platform) with the internal codename “Jack” (they say the name doesn’t matter, but let’s call it that anyway).
The Googlebot we all know is really just a client of this “Jack” platform.
What does that mean? Think of it like a ride-hailing app: you’re the passenger (Google Search), you post a request to the platform (Jack), and the platform dispatches a car to pick someone up (crawl the webpage). Beyond Google Search, other Google products — Google Ads, Google Shopping, Google News, even the latest Gemini AI — are all clients of this same platform. They all share this one crawling infrastructure.
Neo’s take: This means Google’s crawling power is massive and highly standardized. Whether you’re doing SEO or running ads, your site faces the same underlying crawl machinery. It also explains why the experience on your ad landing pages can sometimes affect SEO — the underlying data is probably shared.
2. The 15MB Limit: A Default, Not a Law
There’s long been a rumor floating around the SEO world: Googlebot only crawls the first 15MB of a webpage. If your HTML exceeds 15MB, the rest gets truncated and Google never sees it.
Gary Illyes confirmed it: yes, there is a 15MB default limit at the infrastructure level. It exists to keep Google’s servers from getting overwhelmed.
But! Here’s the kicker:
This limit can be overridden — and it frequently is.
Gary revealed a stunning detail: For Google Search, in the pursuit of efficiency and speed, they’ve actually lowered the limit to 2MB!
Yes, you read that right. For a typical HTML page, Google Search may only crawl the first 2MB. If your page’s code is bloated past 2MB and your important content (body text, core keywords) sits at the end — it may genuinely be “wasted effort.”
3. Why Different Limits?
So the default is 15MB — why does search only use 2MB, while other situations allow more?
Think of it like a buffet.
- Google Search is like fast food — it wants “fast” and “accurate.” It needs to process a page in milliseconds and extract the title, body, and links. If every webpage downloaded 15MB, the internet would clog up and Google’s servers couldn’t handle it.
- PDF files: Gary mentioned that when crawling PDFs, the limit can stretch to 64MB or more. PDFs are just inherently large, and they usually contain rich information worth the download time.
- Images: Same logic — high-res images easily exceed 2MB, so the image search crawler config surely allows much larger file sizes.
Gary gave an example: one version of the official HTML spec document is a single page weighing 14MB. Google simply won’t crawl that single-page version — it’s too resource-intensive to process. Instead, it picks the split-up, individual small section pages.
Neo’s take: The lesson for us DTC sellers is direct: putting your pages on a diet is critical! A lot of Shopify themes or WordPress plugins stuff in huge amounts of inline CSS and JS — or even embed images as Base64 right in the HTML — and the HTML file size balloons. Please check your page’s raw source size (View Source) and make sure the core content (product description, price, reviews) sits as early as possible, with total HTML size ideally kept under 2MB.
4. Google’s Crawler Is a “SaaS” Service
Martin Splitt summed it up: Google’s crawling system isn’t a “monolithic” block — it’s a flexible service.
That means crawl configurations can be dynamically adjusted. For example, if Google notices your site updates frequently or has extremely high content quality, it may internally tune up the crawl parameters for your site. Conversely, if your site is all junk code, it might treat you with the most resource-efficient settings (like stricter size limits).
Summary: What Should DTC Sellers Do?
After hearing the Google engineers’ intel, here are a few practical tips from Neo:
- Watch your HTML size: The “hard limit” is 15MB, but for SEO results, keep your page’s HTML under 2MB.
- Put important content first: Keep your SEO title, H1 tags, and core product descriptions near the front of the HTML. Don’t let mountains of CSS/JS push core content thousands of lines down.
- Optimize resource loading: Don’t embed big chunks of Base64 image data directly in HTML — it wastes crawl budget massively.
- Understand the differences: Googlebot crawls HTML, images, and PDFs with different standards. Your product manual (PDF) can be bigger, but product pages (HTML) must stay lightweight.
Google’s technology keeps evolving, but the core logic never changes: efficiently retrieve the most valuable information. Our job is to make things easy for Google — and then Google gives us the traffic.
Hope this post helps you optimize your independent site!