Machine Traffic Will Be 1,000x Human? Behind Cloudflare's Scary Prediction, a Fake AI Crawler Was Stealing SSH Keys
Hi everyone, this is Neo.
The SEO world is buzzing over a number: in five years, non-human traffic on the internet will be 1,000 times human traffic. That line didn’t come from some influencer — it came straight from Cloudflare’s CFO Thomas Seifert on the company’s Q2 2026 earnings call. And the full quote is even more dramatic: “humans will be a rounding error on the internet.”
Scary stuff. But almost immediately, a veteran SEO consultant used his own website’s raw logs to pour cold water on that number: his site’s biggest “AI crawler” over a 24-hour window wasn’t there to read his articles. It was a credential scanner wearing Common Crawl’s name, hunting for SSH keys.
Today I want to put these two stories together, because the machine-traffic surge is real — but what’s hiding behind the “1,000x” headline is far more interesting. And for anyone running an indie ecommerce site, it directly affects your server security, your bandwidth bills, and your AI-era traffic strategy.
Where the “1,000x” number comes from
Quick background. Cloudflare’s Q2 2026 results (released August 6) were genuinely strong: revenue of $696.1 million, up 36% year over year. CEO Matthew Prince leaned hard into the “Agentic Internet” narrative, saying the web is being rewritten for machine-to-machine traffic.
But the headline-grabber was the CFO’s line on the call:
“If the current trends continue, we think in five years, non-human traffic will be as much as 1,000 times as much as human traffic. Humans will be a rounding error on the internet — not because human traffic goes down, but that’s just how fast we’re seeing non-human traffic grow.”
Here’s the kicker: Seifert added his own caveat — “with the big caveat that I have called it wrong at every point along the way.” Cloudflare previously predicted machine traffic would pass human traffic in 2027. It happened in May 2026. His errors keep running toward underestimating, which is the strongest argument for taking the projection seriously.
Machines being the majority is a fact, not a prediction
Before you laugh at “1,000x,” let’s be clear: machine traffic overtaking humans is already here.
Cloudflare’s official blog post Content Independence Day, one year on (July 2026) laid out the numbers:
- In 2026, traffic crossed a historic threshold: more than 50% of all internet traffic is now non-human
- 52% of crawler requests are now for AI training (as of June 2026), up from 22% in spring 2025
- Mixed-use crawlers (blending search, agent use, and training) represent over 36% of activity
- Pure search crawling is a small and declining share of crawler traffic
That same week, another Cloudflare post, The Agentic Internet, added a punchline: a lot of well-behaved bots keep re-fetching pages that haven’t changed — billions of requests. In their words: “an enormous amount of machine effort, attached to no outcome at all.” Site owners pay to serve it, bots pay to make it, and nobody gets anything.
So yes, machines are the majority of requests. That’s measured, and it’s true. The debate is what those machines actually are, where they come from, and what they want.
The log dump that changes the picture
This is the most valuable part of the story. Instead of chanting along with the prediction, SEO consultant Slobodan Manic (founder of No Hacks) pulled his own site’s AI crawler view and read the actual paths.
Over 24 hours: roughly 3,000 requests, about a third unsuccessful — up more than 1,000% from the previous period.
By crawler: CCBot (Common Crawl) 1,510, ChatGPT-User 375, ClaudeBot 296, Googlebot 245, PetalBot 107… Thirteen others sharing 353 between them.
Looks normal at first. CCBot is the long-running non-profit crawler whose corpus trained a good share of the models everyone argues about — having it as your top visitor seems unremarkable. Then he exported the paths:
/id_rsa,/id_ecdsa,/private-key— SSH private keys/ssl/localhost.key— SSL private keys/.aws/config— AWS credentials/key.json,/serviceAccountKey.json— Firebase/Google service account keys/actuator/configprops— Spring Boot config leakage/api/v1/env— environment variables/.env— environment config/@fs/proc/self/environ— a known path-traversal bug probe/blog/wp-login.php— WordPress login brute-force (on a site that has never run WordPress)
A hundred paths, 1,028 requests, 6.7 MB transferred, zero referrals. The number of requests for anything he’d actually written rounds to nothing.
This is not a crawler. This is a credential scanner — working through a checklist of sensitive files, the same checklist it runs everywhere, and his website is just a row in the loop.
And here’s the detail that should unsettle you: this scanner calls itself CCBot. Real Common Crawl traffic is verifiable — genuine CCBot comes from documented IP blocks and reverse-resolves to hostnames ending in crawl.commoncrawl.org. Manic couldn’t confirm the impersonation from his plan’s logs, but on Cloudflare’s AI crawler dashboard, this traffic was attributed to Common Crawl as the operator, and counted toward his AI crawler totals.
In other words: a credential thief, wearing a research nonprofit’s name, sitting legitimately in your “AI crawler” stats — while completely invisible in your security logs. Because if you’re not blocking it, it passes through, gets served, and leaves no mark. It shows up in exactly one place on the dashboard: the AI crawler view, next to ChatGPT-User and Googlebot.
The new faces on the scanning wordlist
Buried in that log were two paths that didn’t exist on scanner wordlists a year ago:
/.mcp.json— requested 30 times/.continue/config.json— requested 24 times
What are these? /.mcp.json is the config file for MCP (Model Context Protocol) servers — the protocol that lets AI agents call your tools. /.continue/config.json is the settings file for the Continue coding assistant. Both routinely hold API keys and access tokens, because that’s what you put in them to let an agent reach your services.
Translation: someone has added agent tooling configs to the standard secret-scanning wordlist. The same automated sweep that’s been asking every website for /.env since forever now also asks for the file that lists which tools your agents can call and what they authenticate with. Nobody announced it, and it happened fast.
For indie site owners, the takeaway is uncomfortable: every configuration file you create for an AI tool is a new attack surface. Don’t assume Cloudflare makes you safe — Cloudflare deflects traffic attacks, but it can’t stop someone who already has your leaked credentials from logging straight into your server.
Cloudflare’s play: owns the meter and the valve
Manic also flagged something subtle that I think deserves its own section — Cloudflare’s commercial choreography is textbook. In the first week of August alone:
- Earnings call projection about “1,000x machine traffic” (frames the problem)
- Blog post quantifying how much of the web is no longer human (amplifies the problem)
- An agent-readiness scanner that tells you you’re not ready (sells the problem)
- An AI-visibility product that scores you (sells the fix)
- A “bridge” to expose your website’s tools to agents (new revenue)
- A default that starts blocking some agents in September unless you opt out (forces a decision)
Every product is reasonable on its own. But together: the company measuring the problem, framing the problem, and selling the fix is one company. They now own both the meter and the valve.
I’m not calling Cloudflare the villain — Pay-per-crawl was the right idea, Content Independence Day (letting owners decide which machines get in) was the right idea. But as users, we should separate their numbers from their narrative: trust the traffic data, discount the sales story.
Five things indie site owners should do right now
Enough analysis — here’s what you can actually do, starting today:
1. Read paths, not totals
Don’t just stare at “how much AI traffic you got.” Export the logs and look at which paths were requested. That’s exactly how Manic found the problem. If you only watch totals, your biggest AI crawler looks like a diligent Common Crawl — a perfect illusion.
2. Cross-reference crawler logs with security logs
The most ironic part of this story: the scanner was fully visible in the AI crawler view and completely invisible in security events. Security logs only record requests that trip a rule, and if you’re not blocking it, it leaves no trace.
Compare your AI crawler traffic against your WAF/security events. If any “AI crawler” is hitting sensitive paths (/.env, /.ssh, /wp-login.php), whatever it calls itself, that’s a red flag.
3. Verify the identity of known crawlers
Common Crawl publishes its verification method: genuine CCBot traffic comes from official IP blocks and reverse-resolves to *.crawl.commoncrawl.org. Google has official verification tools for Googlebot too. Add an identity check layer in Cloudflare or at the server level — this beats UA-string blocking, because UA strings are trivially spoofed.
4. Audit your sensitive file exposure
Curl these paths from outside and see what’s publicly reachable:
/.env/.ssh/id_rsa/.aws/config/.mcp.json/.continue/config.json/key.json,/serviceAccountKey.json/actuator/configprops,/api/v1/env
Any sensitive file reachable from the public internet is a ticking bomb. Pay special attention to MCP config files — the newest additions to scanner wordlists, and they usually contain API keys.
5. Don’t be scared by “1,000x,” don’t be numbed by “50% non-human”
Read the two numbers together: machine traffic is indeed the mainstream — that’s reality. But machine traffic is a mixed bag: training, retrieval, noise, and theft. The machine traffic that actually matters to you (AI search citing you, agents reading your content) is only a fraction of it. Optimize for the valuable machine traffic, not for all of it.
Neo’s take
Some honest opinions to wrap up.
The “1,000x machine traffic” prediction will probably come true — but its meaning is different from what Cloudflare wants you to believe. Their narrative is “the Agentic Internet is here, you need our full toolkit.” Manic’s logs tell us that the surge is stuffed with worthless and even hostile traffic.
My read for indie site owners:
-
Machine traffic is the inflation of the AI era. Just like currency printing dilutes money, exploding machine traffic dilutes the meaning of “traffic.” Going forward, don’t report traffic numbers — report effective machine traffic: how many AIs actually read and cited your content.
-
Security is now part of SEO. We used to care about site security because getting hacked hurts rankings. Now security directly affects your AI visibility — if a scanner steals your credentials and your server gets compromised, forget AI citations, Google may remove you from the index entirely. A monthly routine of cross-referencing security logs with crawler logs should be in every indie site owner’s checklist.
-
A costume is not an identity. UA strings and self-declared crawler names are just costumes. Real identity comes from IPs, reverse DNS, and behavior patterns. Spending ten extra minutes on this could save you from a server compromise.
-
Cloudflare’s problem is also its opportunity. It sells the anxiety and the antidote in one package — but flip it around: when machine traffic really hits 1,000x, the platform that gives site owners genuine choice over which machines get in is basically Cloudflare. So the conclusion isn’t “don’t use Cloudflare.” It’s “use it — but with your eyes open.”
The machine traffic tide is here, and you can’t stop it. But you can do this: let the valuable machines in, keep the thieves out, and make Cloudflare’s numbers work for you instead of letting them sell you fear. That’s the survival posture for indie site owners in 2026.
That’s all for today. If you spot something weird in your logs, come talk to me — maybe your case is the next article.