Common Crawl Just Published an AI Visibility Manual — Is Your Site Even in the Training Data?
Hi everyone, this is Neo.
Today’s topic sits one layer deeper than “AI search citations” — and it’s a lot more brutal. Yet I’d bet 99% of independent site owners have never even heard of it.
Your website may not be in AI training data at all. That means ChatGPT, Claude, and Gemini never knew your brand existed before they shipped — no matter how well you’ve optimized for search.
The root of all this is a nonprofit called Common Crawl. Every month it fetches over 2 billion web pages and gives the copy away to anyone who wants it. Almost every major model (GPT, Claude, Llama, Mistral…) trains on data that traces back to this archive. When Mozilla audited 47 LLMs released between 2019 and 2023, 64% had trained on Common Crawl data in some form, and roughly 60% of GPT-3’s training weight was filtered Common Crawl.
In other words: if an AI model already knows your brand without searching the web, odds are this is where it learned it.
Last week something big happened: Common Crawl published The AI Visibility Audit — an official field guide teaching site owners how to check whether they exist in the training data. SEJ contributor Suganthan Mohanadasan read it, shrugged, said “this is all manual work, so I automated it,” and shipped a free tool.
Today I’m breaking the whole thing down: what CCBot is, whether your site is in the archive, why sites built after July 2025 are almost certainly blocked by default, and the 5-step checklist.
1. First, understand how AI models “meet” your website
Traditional SEO is: Index → Rank. Googlebot crawls your page, it goes into the index, and it ranks by relevance.
The AI era is: Train → Retrieve. Models don’t search the web in real time — they read a huge slice of the internet into their weights before they ship.
And that “slice of the internet” is mostly Common Crawl.
Here’s the pipeline:
- CCBot (Common Crawl’s crawler) fetches roughly 2 billion pages a month — the July 2026 crawl holds 2.14 billion.
- The raw snapshots are published for free.
- Labs don’t train on the raw crawl. They take filtered derivatives: C4, RefinedWeb, RedPajama, Dolma, FineWeb — every training set you can name descends from Common Crawl.
- Models train on that filtered data.
So the path your content takes into a model is: CCBot crawl → Common Crawl archive → filtered training set → model weights. Each step filters out sites. If your page gets stopped at step one, everything after it is wasted.
One key detail: every step has capacity limits. Since March 2025, Common Crawl raised its per-page storage cap from 1 MiB to 5 MiB. That’s more forgiving than before — but it’s still not unlimited.
2. Can CCBot actually read your site? The reality is harsh
The manual’s most brutal line:
“Your CDN may be blocking CCBot without telling you, keeping your pages out of the archive that trains most models.”
This isn’t fear-mongering. The data is right there:
- BBC’s robots.txt blocks 13 of the 14 AI crawlers the checker tracks. CCBot asked for permission, read the refusal, and left. One of the biggest news sites on earth doesn’t exist as far as the training-data pipeline is concerned.
- BuzzStream measured that 75% of top US and UK publishers are deliberately blocking training crawlers.
- Chris Green’s HTTP Archive numbers: roughly 492,000 sites name CCBot in robots.txt — and about 95% of those mentions exist to block it.
- The kicker: since July 1, 2025, Cloudflare blocks AI crawlers by default on every new domain it serves. You don’t have to do anything — if your domain lives on Cloudflare, AI crawlers can’t get in by default. Cloudflare’s managed robots.txt reached 3.8 million more domains that never wrote their own file.
- Ad plugins add their own blocklists to robots.txt.
In 2026, a CCBot block is more likely a platform default than a decision you made.
What does this mean for independent site owners? You spent real money on content, links, and on-page optimization — and the AI models never met you before they shipped. When a user asks ChatGPT “which brand’s product is best,” the model answers from whatever was in its training data — which is usually your competitors. Because you were never in there.
3. The official manual + a free tool that automates it
Common Crawl knows this is a problem. Their team spent the past year at SEO conferences worldwide, and everyone kept asking the same question:
“Why does a page that ranks well in Google stay invisible to ChatGPT, Gemini, Claude, and Perplexity?”
So they wrote The AI Visibility Audit — an 18-page free guide covering how AI systems actually discover content, why training-data inclusion behaves like a ranking factor, and how to run a repeatable five-check audit using only free tools in about 90 minutes.
The manual’s logic is clean. The five checks run from most decisive to most strategic:
- CDN/firewall check: curl your homepage with CCBot’s user-agent and compare the response to a browser’s. If your CDN or WAF challenges CCBot at the edge (instead of honoring robots.txt), no robots.txt edit will help. The manual’s own advice: check your CDN dashboard before you touch anything on your server.
- Index API check: query Common Crawl’s index server for your domain and see how many pages were captured each month.
- Web Graph rank: Common Crawl publishes a monthly web graph with Harmonic Centrality and PageRank — and these metrics directly drive CCBot’s crawl priority. The lower your domain sits, the less often it gets crawled.
- Sitemap diff: compare your sitemap against the URLs actually captured, and turn the gap into a work queue of pages that should be in but aren’t.
- Stored-copy check: pull the actual bytes CCBot archived for your homepage, extract the text, and compare it to your live HTML. Both sides are read before JavaScript runs — because CCBot never runs it.
Sounds professional? Right — it’s all manual work. It needs scripts, API calls, and data reading. That’s exactly why Suganthan automated it into the Common Crawl Visibility Checker (free, no signup, checks one domain at a time, straight from Common Crawl’s own archives).
The tool’s standout features:
- Robots.txt history: Common Crawl stores the robots.txt it read before every crawl. The checker diffs those stored copies month by month — so a block gets a start date. It even recognizes blocking templates: Cloudflare’s managed robots.txt has a distinctive license comment, Squarespace’s default has its own signature. You’ll instantly know “this was the platform, not me.”
- Live probe: besides history, it fetches your homepage as CCBot/2.0 next to a normal browser and a control bot — catching the sneaky case where robots.txt reads clean but a firewall challenges the crawler at the edge. If it’s an edge challenge, the card shows the firewall vendor and links CCBot’s published IP ranges.
- Next-crawl date: it tells you when the next crawl happens — because a fix made today only shows up in next month’s data.
- Sitemap panel: it fetches your sitemap and diffs it against captured URLs. The author’s own site: 42 of 97 pages in the July crawl, with the other 55 downloading as a CSV work queue.
- Per-page history: paste a specific page URL and see its capture history across the year.
The author’s own site went from 2 captured pages in last August’s crawl to 55 in July’s — “the most flattering chart anyone has ever drawn of it.”
4. A detail you must not miss: CCBot doesn’t run JavaScript
I want to hammer this one separately, because way too many independent sites die on it.
CCBot fetches pages without executing any JavaScript. If your site renders content client-side (SPA, React/Vue hydration), CCBot sees an empty shell. And empty shells get filtered out of training sets.
The manual’s advice is blunt: content must render on the server (SSR). This is why I keep telling site owners: skip the fancy front-end frameworks — SEO-friendliness comes first. Googlebot can execute JS now, but CCBot can’t. The training-data “standard” is more conservative than the search index.
Two more counterintuitive points:
- CCBot doesn’t browse your site the way you imagine. It doesn’t “explore” — it follows whatever links the web actually contains. The author found one of his own pages was discovered through a newsletter link, archived with long-dead ConvertKit tracking parameters still attached. Don’t just build internal links; the links pointing at you from the outside world are CCBot’s entry points.
- Crawl priority follows “harmonic centrality.” Common Crawl doesn’t crawl every site equally — it prioritizes domains by how central they are in the web’s link graph. A link from a site deeply embedded in the web’s core is worth far more than dozens of links from isolated sites. Link topology matters more than link volume — I’ve said this before, and here’s another confirmation.
5. The 90-minute self-check: is your site in AI training data?
Here’s the manual + tool condensed into a checklist you can actually run:
| Step | What to do | How to judge |
|---|---|---|
| 1 | Search your domain on the Common Crawl index server | Look at monthly capture counts; 0 pages = serious problem |
| 2 | curl your homepage with the CCBot/2.0 user-agent | Compare with a browser response; 403/challenge = CDN is blocking |
| 3 | Check whether your robots.txt explicitly allows CCBot | No rule = allowed by default; Disallow = blocked |
| 4 | Pull an archived copy of a page and check the text | Empty shell = client-side rendering problem, go SSR |
| 5 | Check robots history with the free tool | Find the month the block started; check if you switched CDNs or installed a plugin then |
Neo’s take:
First, a lot of sites are getting blocked from AI crawlers by accident. Domains created on Cloudflare after July 1, 2025 block AI crawlers by default. You did nothing — and the models just never met you. I wrote about the AI-crawler debate last year when everyone was agonizing over whether to block OpenAI’s crawler. Now the story has flipped: you’re not blocking them — your platform is doing it for you, and you don’t even know. Go check, especially if you’re on Cloudflare.
Second, training-data visibility is becoming a “factory setting” ranking factor. The old formula was index + rank; the new one is train + retrieve. If the model never trained on you, you can’t even enter the retrieval race. This isn’t GEO mysticism — it’s physics. And note: this is more fundamental than AI-search citations. You can’t monitor it with a prompt panel, and you can’t fix it overnight — a fix today shows up in next month’s crawl, and the model has to retrain after that. The earlier you act, the better.
Third, on the “should I let AI crawlers in” debate, my position hasn’t changed: letting your content into training data costs you nothing. Content is your résumé. When AI crawlers read your page, you’re giving away information but gaining presence in the AI’s mind. What you should actually guard against is plagiarism and abuse — that’s a copyright problem, not a robots.txt problem. BBC blocks training crawlers because it has a moat. If you don’t have BBC’s moat, don’t copy BBC.
Fourth, this free tool is a snapshot of where SEO is in 2026. The official body publishes methodology, the community builds automation, and the data is all open. Whoever closes the “check → fix → verify” loop first wins the AI-visibility race. Information gaps are closing; execution gaps are widening.
One last note: the checker covers CCBot specifically — AI visibility also involves GPTBot, ClaudeBot, and the rest (the author has a separate free AI Crawler Access Checker for that). But Common Crawl is special: GPTBot fetches your page for OpenAI alone. A CCBot visit ends up with everyone who’ll ever train on the open web. That’s why it deserves your attention first.