Google thinks a casino page is the canonical version of your site — and the real culprit is an error screen you never saw


Hi everyone, this is Neo.

In mid-September, Roger Montti published a piece on Search Engine Journal that I ended up reading twice. A site owner on Reddit reported that his pages were slowly but steadily dropping out of Google’s index, and that after digging in, Google had picked a casino betting page — a site with zero connection to his business — as the canonical version of his pages.

His own words, near enough: their pages are about companies and suppliers, and the page Google treats as canonical is a casino betting page. He searched that casino site top to bottom and could not find a single page that came anywhere near his content.

Then John Mueller replied to the thread. And between his answer and a second Reddit user’s follow-up, the story points at something that matters to every site owner — and it has almost nothing to do with cross-domain canonicals.

The short version: your site may be serving Googlebot an error screen instead of a page.

Let me break it into three parts: why the popular explanation doesn’t hold up, what the real causal chain looks like, and what you can check in half an hour.

First, the two concepts

A cross-domain canonical is a rel="canonical" tag on a page of site A pointing at a URL on site B. It tells search engines “the original version of this content lives over there.” It’s the cross-site version of the canonical tag you already know, and Google has always described it as a strong hint, not an instruction.

It used to have two legitimate jobs:

  1. A fallback for domain migrations — when a 301 redirect genuinely wasn’t possible, you could use it to pass signals to the new domain. Today, that situation should basically never happen.
  2. Syndicated content — you publish an article on your own site and hand the same piece to a partner, using a cross-domain canonical to declare who has the original. Google has since changed its advice on this too.

Google’s current guidance on syndicated content is blunt: the republishing site should use robots meta tags to keep the copy out of the index — <meta name="Googlebot-News" content="noindex"> if you only want to keep it out of Google News, <meta name="Googlebot" content="noindex"> if you want it out of Search as well.

Which means there is no good reason to use a cross-domain canonical anymore. It’s a hint. A 301 redirect and a noindex directive are things Google is obligated to respect. Using a tool the other side can ignore, to solve a problem that already has an absolute solution, is just bad engineering.

Does “de-indexed by a cross-domain canonical” even happen?

This is where Montti’s analysis matters, and it boils down to one operational sentence:

For a cross-domain canonical to take effect, the tag has to be on your page, pointing at theirs.

Signals travel outward from you. They don’t get stolen. If the other site puts a canonical pointing at itself, that has nothing to do with you. Only your page declaring “the original is over there” moves anything.

So if you audit your own pages and find no canonical pointing off-domain, you’re left with two possibilities:

  • Your site was hacked. Someone injected a canonical tag into your template or database pointing somewhere else. Treat that as a security incident.
  • This was never a cross-domain canonical problem. You’re looking at two things that happened around the same time and assuming one caused the other.

Montti used an analogy I liked: what does an SEO coincidence look like? He pointed at two familiar ones. Plenty of people swear a disavow file works within weeks; in reality it plays out over months. And plenty of people label “we have several similar pages about one topic” as keyword cannibalization — when a site focused on a topic naturally has several similar pages. He compared it to blaming a stomach bug on the last thing you ate when it was really that shopping cart handle.

Neo’s take: this isn’t about mocking the person who posted. It’s about something more useful for the rest of us — most “penalties” in SEO are labels we apply ourselves. Stolen canonical, algorithmic penalty, competitor sabotage: these labels are satisfying because they point outward and require us to change nothing. The moment you accept that the output might be coming from your own site, your entire diagnostic direction flips.

The real culprit: the error shell served to Googlebot

The most valuable part of the thread was the follow-up from a second Reddit user, who ran into the identical situation and found the mechanism. His account, paraphrased:

Searching the third-party canonical URL in Google showed it indexed with the title “Application error: a client-side exception has occurred (see the browser console for more information)” — a generic JavaScript application error. And that is the exact same message our own site has occasionally displayed during temporary outages when the app fails to load properly. So we suspect Googlebot at some point crawled an error or fallback response instead of real page content, and Google treated multiple URLs showing that same error shell as duplicates. That would explain the cross-domain canonical selection and the de-indexing that followed.

All three possible outcomes land in the same place — Mueller confirmed as much afterwards:

Outcome What it means What it costs you
A. Your page is canonical but indexed with the server error message Google indexed you, but indexed the error copy Your page doesn’t surface for its real content
B. Your page is treated as a soft 404 Google decides the page has no content Same
C. The other page becomes canonical Your page is treated as the copy Same

Notice the absurdity of the third one: A and B are your own problem, and C is just one of the ways it shows up. That’s why “de-indexed by a cross-domain canonical” is a misleading description — it isn’t the cause, it’s a symptom.

Where that error message comes from

I went and looked up the string. Application error: a client-side exception has occurred (see the browser console for more information) is Next.js’s default client-side exception screen — the fallback React shows in the browser when your app throws.

That symptom has been tangled up with Search Console soft 404s for years in the developer community. There are threads in the Next.js repo from 2018 asking whether Next.js causes soft 404s, and discussions running through 2024–2026 from teams whose URLs return 200, render fine for users, and still get flagged as soft 404s — with the URL Inspection tool showing exactly that Application error screen for Googlebot.

One mechanism that keeps coming up is version skew:

  1. Googlebot’s first pass loads the JS chunks from the build you had live at that moment.
  2. You redeploy between passes. Chunk filenames change. The old URLs now 404.
  3. On its second pass, Googlebot requests the old chunks, the load fails, rendering breaks, and all that’s left is the error shell.
  4. Google classifies it as a soft 404 — or resolves it as outcome A, B or C above.

That’s what makes this class of bug so hard to catch. Open the page yourself and it’s always fine, because your browser pulls the current chunks. Only a crawler returning with a stale snapshot sees the broken version. And it isn’t permanently broken — it happens intermittently, which is why it always feels like “sometimes it works.”

There’s a second, older problem sitting underneath it: AI crawlers mostly don’t run JavaScript. Multiple third-party AI-readability audits make the same point — GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot take the HTML your server returns, and only Googlebot renders. For a site that renders client-side and occasionally serves an error shell, you’re not just losing Google indexing. You’re losing your entire eligibility for AI search.

Mueller’s full answer: all three outcomes are bad, so fix it before launch

Mueller first agreed the explanation was plausible, then recommended using Search Console’s live URL checker to see how Google actually renders the page.

His second reply is the practical one. The points, as I read them:

  • He flatly says the classification doesn’t really matter. Yes, the three outcomes are confusing, but the consequence is identical in all three: your page doesn’t show for its normal content. So arguing about whether it’s A or B is a waste of time.
  • The correct fix is catching this kind of error on your side, before Google finds it. Mueller described what he does on his own smaller sites: run a pile of automated tests before pushing live, and every time something goes wrong in production, have a code agent add a test for that scenario.
  • Add monitoring. Fetch your most critical pages on a schedule — hourly, or whatever fits — and check for problems, so you fix them before they harden into a search engine issue.

Neo’s take: this is really an argument about cost structure. You either spend a few hours on automated tests and page monitoring, or you spend months waiting for indexing to recover — and recovery moves in months, not weeks. We wrote about that in our September 4 piece. The math isn’t close.

A 30-minute audit, cheapest checks first

Step How to check Pass condition
1. Look for off-domain canonicals Search your page source and your Link: <...>; rel="canonical" response headers for any non-your-domain value No off-domain canonical (unless you’re deliberately migrating and know exactly why)
2. Check for a hack Search Console security issues; unexpected scripts, redirects, or canonical injections in templates Nothing injected
3. Probe the fetch layer curl a batch of key pages without JS; inspect the status and the body 200, and the body contains real content
4. See what Google sees Search Console URL Inspection → Test Live URL → check the crawled page screenshot and HTML Matches what your users get
5. Check build/version consistency Look for stale-chunk 404 risk (asset filenames changing per build, cache policy) Old build assets stay reachable briefly, or deploys switch atomically
6. Gate the deploy After each release, auto-fetch core pages and assert a key sentence is present Fail loud; never ship a broken render

Step 3 is the one teams skip and the one that actually catches this. Most teams check “does the page open.” The failure mode here is “does the page open correctly for a crawler.” Your browser is not the test subject. The crawler is.

If you’re already seeing suspicious soft 404s, or “Duplicate, Google chose a different canonical than the user,” work in this order: confirm the page reliably outputs clean content → then look at redirects and canonicals → only then request re-indexing. Fix first, request second. Otherwise you’re just helping Google confirm the page is broken a little faster.

One more thing about syndication

If you run a multilingual site, or you hand content to partners — both extremely common for cross-border brands — stop reaching for cross-domain canonicals. Per Google’s current advice:

  • 301 if you possibly can.
  • Republishers add a noindex (Googlebot-News if you only care about News, Googlebot if you want it out of Search too).
  • The original site keeps the full version. Don’t design a setup where both sides are trying to rank.

The beauty of this arrangement is that it doesn’t depend on a search engine’s goodwill. A canonical is a hint. A noindex and a 301 are enforcement.

Neo’s take: what this case really teaches

Three things stuck with me, and none of them are canonical mechanics.

One: the word “penalty” is hiding the real problem. We attribute every ranking drop to an algorithm, a competitor, or a platform, because those explanations demand nothing from us. In this case, the culprit was an error page our own stack wrote.

Two: AI search made technical hygiene more important, not less. When Googlebot was the only reader, client-side rendering problems were survivable. Today your content has to be readable by Googlebot, GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot — and most of them don’t execute JavaScript. An intermittent error shell that used to mean a handful of soft 404s now means your pages may never enter the retrieval pool at all.

Three: monitoring is the only certainty you can buy. Mueller’s “fetch your critical pages hourly” suggestion costs almost nothing and buys you one thing: finding the problem before Google does. No tool makes that call for you. You have to run it on a schedule.

Final word

If you do one thing today: check whether any page of yours carries an off-domain canonical, then curl three core pages and confirm the body is real content and not an application error screen.

Neither check costs money or requires a tool, and together they rule out a class of problems that is genuinely hard to see and slowest to recover from.

Most “cross-domain canonical penalties” aren’t cross-domain canonical problems. The real problem is that the version of your page the crawler gets is not the version you see.