DOJ Files Finally Reveal: Can Click Fraud Boost Your Rankings? The Truth Isn't What You Think.
Hi everyone, this is Neo.
If you’ve been doing SEO for a while, you’ve heard this old question: is click volume actually a Google ranking factor?
The SEO world has been arguing about it for over twenty years. Some people are convinced “fake clicks = higher rankings,” so they go hunting for click farms. Others think “clicks influence rankings” is a legend cooked up by the rumor mill. And plenty of people pay for “CTR optimization” tools, burning cash every month without knowing if any of it works.
A huge news story finally gave this question an official answer — the US Department of Justice’s September 2025 antitrust memorandum opinion peeled back the curtain on Google’s internal mechanisms. Combined with a click-related patent Google filed way back in 2006, we can now tell you very clearly:
Clicks are not a direct ranking factor. But clicks do influence rankings. Those two statements don’t contradict each other.
Sounds contradictory? Don’t worry — I’m going to explain the whole thing in plain English today.
Twenty years of debate, settled by the DOJ in one word
Start with the most critical passage in that September DOJ document. Google’s expert witness, Professor James Allan of UMass Amherst, testified using a core concept:
“Signals have different levels of complexity. There are ‘raw signals,’ such as click counts, page content, and the words in a query.”
The key term here is raw signal.
A “raw signal” is the most basic data point a search system can observe directly — before it’s been interpreted or used for training data.
The DOJ document classifies all of the following as “raw signals”:
- Click data
- Page content
- Human rater scores
- Search queries
Notice the classification here — you might not have realized this, but the scores from Google’s “quality raters” and real user clicks sit at the same level inside Google’s system. Both are raw data that need further processing before they can influence rankings.
The document has an even sharper line:
“At the other end of the spectrum are innovative deep learning models. Deep models discover and exploit patterns in massive data sets. They provide unique capabilities at high cost.”
In other words: clicks are one end (the most basic raw data), deep learning models are the other end (the highly complex final system). Between them lie countless layers of data processing.
Neo’s take: A lot of independent site owners have been brainwashed by various SEO tools into thinking clicks are a ranking signal. Where’s the flaw? They’re ignoring the processing pipeline in between. Clicks are like crude oil — what ultimately determines how a car performs isn’t the oil itself, it’s the refinery, the engine, and a thousand engineering systems. Staring at crude oil won’t tell you how fast the car goes.
How exactly are “raw signals” processed?
The DOJ document mentions a name over and over: Navboost.
Plenty of self-media accounts love to say “Google uses Navboost to rank sites based on user clicks.” But in the DOJ document, Navboost is described as a system for measuring popularity:
“…popularity measured through user intent and feedback systems (including Navboost/Glue)…”
Note the wording: “measure popularity”, not “rank directly.” Those are two completely different concepts.
The full data pipeline looks roughly like this (simplified from the DOJ document):
raw click data
↓ noise removal (spam traffic, bots)
↓ aggregation (accumulated by query-URL pair)
↓ AI model training (RankEmbed, RankEmbedBERT, etc.)
↓ quality/relevance signal generation
↓ fusion with other signals (links, freshness, user location, etc.)
↓ final ranking
In other words, from the moment you click a mouse to the moment that action actually affects some site’s ranking, there are at least 5–6 stages of data processing in between. Every stage filters for spam and manipulation.
Neo’s take: This pipeline matters enormously for independent site operators. It means —
- If you fake 100 clicks, they never even reach the level that influences rankings, because they get filtered out during noise removal
- But real users clicking your site consistently over time become part of AI training data, making the AI “think more highly” of your content
- So “manufacturing clicks” and “building a brand” are two different things. The former is useless; the latter works.
The “70-day search log” has been badly misread
The DOJ document has one line that gets quoted everywhere:
“70 days of search logs plus scores generated by human raters”
Lots of SEO bloggers ran with it: “Whoa, Google ranks with 70 days of click data!”
But read the full context and you’ll see that’s not what it says at all:
“RankEmbed and its successor RankEmbedBERT are ranking models that rely on two primary data sources: [redacted]% of 70 days of search logs, plus scores generated by human raters that Google uses to measure the quality of organic search results.”
See that? Those 70 days of data are used to train the RankEmbedBERT AI model — not fed directly into the ranking engine.
Even more important, that 70-day dataset:
- is aggregated data (not individual one-off clicks)
- is training data (not a direct ranking input)
- after RankEmbedBERT processes it, the model ranks pages based on natural language analysis
Think of it this way: you send a chef 70 days of ingredients, the chef turns them into a menu, and the menu decides what the restaurant sells. Ingredients don’t equal restaurant revenue — the menu does.
So what actually is RankEmbed?
The DOJ document gives RankEmbed a clear definition:
“The RankEmbed model itself is an AI-based deep learning system with strong natural language understanding. This allows the model to more efficiently identify the best documents to retrieve, even when a query is missing certain words.”
And here’s the truly surprising part:
“RankEmbed uses 1/100th of the training data of earlier ranking models, yet produces higher quality search results.”
1/100th of the data, better results.
What does this data point tell us? It tells us Google has switched from “winning with data volume” to “winning with model architecture.” RankEmbedBERT is built on BERT-style Transformer architecture (the same technology family behind ChatGPT), and it can extract rich semantic relationships from small amounts of high-quality data.
In other words, Google’s judgment of your content is increasingly about “understanding what you’re saying” rather than “counting how many times you got clicked.”
Neo’s take: This has massive implications for independent site content strategy. If Google is training smarter models on less data, then your content has to win on “semantic quality,” not on “keyword density” or “click numbers.” One genuinely deep article that solves a user’s problem can outperform ten keyword-stuffed filler posts.
The 2006 patent: clicks stopped being a direct ranking factor a long time ago
A lot of people don’t know that Google filed a click-related patent back in 2006, titled:
“Modifying search result ranking based on implicit user feedback”
The inventors did something particularly clever: they strictly separated “signal generation” from “ranking behavior” itself.
Here’s what the patent says:
“The ranking subsystem may include a ranking modifier engine that uses implicit user feedback to cause re-ranking of search results. User selections (click data) of search results can be tracked and converted into a ‘click fraction’ used to re-rank future search results.”
Pay attention to that “click fraction” — it’s not a single click; it’s an aggregated statistic.
Technically it’s called the LCIC score (Long Click divided by Clicks). Note that “Clicks” is plural — because the decision is based on the sum of many clicks (aggregation), not any single click.
How is this score calculated? Three steps:
1. Summation For a specific query-document pair, add up all weighted clicks.
2. Normalization Divide that sum by the total click count.
3. Statistical smoothing Apply a “smoothing factor” to the result so a single click on a rare query can’t skew the outcome — this is the anti-cheating mechanism.
The formula from the patent:
LCC_BASE = #WC(Q,D) / [#C(Q,D) + S0]
where:
WC(Q,D) = the weighted click sum for that query-URL pair
C(Q,D) = the total click count for that query-URL pair
S0 = the smoothing factor
The key implication of this formula:
If someone deliberately fakes clicks (the “rare query + single click” pattern), the smoothing factor S0 crushes the influence of those manipulative clicks to almost zero.
In other words, as far back as 2006, Google had already neutralized click fraud at the patent level. The “CTR optimization” tools being sold today are selling a defense Google figured out 20 years ago.
Neo’s take: Guys still paying for click-fraud tools — wake up. This business was killed at the patent level two decades ago. If your SEO budget is still burning on it, stop now. Spend that money on things that generate real long clicks (users who land on your site, stay, read, and convert), like: content depth, page experience, and brand trust.
So what is a “long click,” and why does it matter?
The patent breaks clicks into categories:
- Short Click: the user clicks in and bounces within seconds
- Medium Click: the user stays for a while
- Long Click: the user reads deeply and interacts
- Last Click: the click that ends the user’s search task
Google weights these clicks differently.
The simple version: a user clicks your site, then quickly returns to the results and clicks someone else — that’s a negative signal. A user clicks your site and stays for a long time, or even completes their search task — that’s a positive signal.
But again: these are all aggregated statistics, not individual clicks. You need consistent, sustained long clicks from real users for any of this to matter.
What does this actually mean for running an independent site?
After all that technical detail, let’s get to the practical question: how does any of this help me run my independent site?
Here are 5 concrete, actionable recommendations:
1. Stop any “click fraud” activity immediately Whether it’s manually faking clicks on Google Search or simulating clicks with bots. The 2006 patent tells you Google crushes manipulative data with smoothing factors down to near zero. It’s pure money burning.
2. Shift your SEO budget toward “long-click” performance Your goal isn’t getting users to “click in” — it’s getting users to click in and not want to leave. That means:
- Load speed (above-the-fold content within 3 seconds)
- Content structure (headings, table of contents, clear paragraphs)
- Visual comfort (readable font size, line height, colors)
- Content depth (solve real user problems, don’t pad word count)
3. Prioritize “semantic matching” over “keyword matching” RankEmbedBERT is a BERT-architecture deep learning model — it understands meaning. Writing “how to choose waterproof shoes for hiking” vs “hiking waterproof shoes selection guide” means roughly the same thing to it. Instead of stacking keywords, genuinely nail one topic.
4. Be wary of paid “AI visibility” and “CTR optimization” tools If a tool promises to “boost your CTR to move rankings,” remember: Google treats clicks as raw signals that must pass through noise removal, aggregation, training, signal generation, and ranking fusion. Any tool claiming to “directly manipulate CTR to influence rankings” either doesn’t understand the mechanics or is selling you a fantasy.
5. Build long-term brand traffic, not just keyword traffic Branded searches → direct clicks to your site → deep browsing → conversion. This path generates the highest-quality signals (including clicks) of all. That’s why big brands’ SEO always looks steadier — their branded traffic is continuously feeding Google’s systems “high-quality training data.”
Neo’s take: Every time the SEO world gets an “insider reveal,” the real winners are the people who’ve been doing the right things all along — publishing quality content, chasing real user experience, building brand trust. Every “mechanism exposé” just proves once again that there are no shortcuts.
Summary
This post is dense, so let me pull the key points together:
-
The DOJ’s September 2025 antitrust filing peeled back Google’s internal mechanics — it explicitly says clicks, human ratings, content, and queries are all “raw signals,” not direct ranking factors.
-
Raw signals pass through multiple processing layers before influencing rankings: noise removal → aggregation → AI model training → quality signal generation → ranking fusion. At least 5–6 stages in between.
-
Navboost isn’t a “rank-by-clicks” system — it’s a “measure popularity” system. Those are two different things.
-
The “70-day search log” has been badly misread — it’s used to train the RankEmbedBERT AI model, not fed directly into the ranking engine.
-
RankEmbed delivers better results with 1/100th of the data — Google has already shifted toward “semantic understanding” rather than “statistical counting.”
-
The 2006 patent already proved click fraud is useless — the S0 smoothing factor neutralized manipulative clicks 20 years ago.
-
What actually works is “long clicks” and “aggregated signals” — and those can only come from genuine user experience and content quality. No shortcuts.
I hope this analysis saves you the money you’d otherwise spend on “click tools,” and points your SEO budget where it actually counts. If you’re still figuring out your independent site SEO strategy, follow along — let’s master this game together.
See you in the next post.