How to find out whether AI can actually *see* your page — a 30-minute retrieval self-check with no tools required


Hi everyone, this is Neo.

Search Console gives you a Pages report that tells you what’s indexed and what isn’t. Bing Webmaster Tools does something similar.

AI search gives you nothing.

Which is why every site team ends up in the same loop: the boss says “ChatGPT never mentions us,” you say “maybe it hasn’t crawled us,” the boss asks “how would you know?”, and you have no answer — because there’s no panel that shows whether your page sits inside an AI engine’s retrieval reach.

On September 18, Chris Green published a short and genuinely practical piece on Search Engine Journal (a repost of his own post) that addresses exactly this: using a technique SEOs have relied on for decades to test the retrieval layer.

Here’s my operational version, plus two cross-checks I’d add.

The method: bait your URL with a verbatim passage

It comes from two familiar moves:

  • To check indexing, you use the site: operator.
  • To check whether content is indexed, you copy a long, distinctive chunk of text off the page and search it in quotation marks.

Green’s twist is to run the second one inside a chatbot. The prompt looks roughly like this:

Search for “paste your snippet here” and return any results which contain that exact text only.

The important caveat: use a chatbot with search enabled — ChatGPT’s web search, Gemini, Perplexity, Copilot. Pure chat mode will just improvise from training memory, which tells you nothing.

How to read the result:

  • It returns your URL → whatever search source it called contains that passage, correctly attributed to your URL. Technically, the page is retrievable.
  • It returns nothing → this is the valuable case, because now you have a troubleshooting list.

Green keeps stressing that this is not a replacement and not the truth. Interpret the output. Run it four to five times, because different engines call different search sources and results can differ run to run. If it’s ambiguous, change the snippet and change the engine.

When it returns nothing, work this order

The order is the dependency chain. If step one fails, everything below it is irrelevant:

Order Question How to check
1 Can it be discovered? Any internal links? In your sitemap? Green calls out a specific case: AI-generated pages deliberately orphaned so nothing links to them
2 Can it be fetched? Blocked by a WAF, firewall or robots.txt? Note that AI retrieval crawlers are a different set of names — see the next section
3 Can it be rendered/extracted? Does the body only exist after JS runs? Is the crawler getting an empty shell?
4 Can it be indexed? Any noindex? Is a canonical pointing somewhere else?
5 Is the content distinctive enough? Is your passage so generic that it can’t out-rank anyone else’s version?
6 Has it had time? It may simply not be discovered or indexed yet. That part is patience

On step 5, Green is refreshingly blunt: if your content isn’t distinctive enough that a verbatim passage can’t bring back your own page, it probably wasn’t going to be a valuable page anyway.

Add-on one: get the crawler names right first

Step 2 hides a high-frequency accident: plenty of teams check GPTBot and call it done.

AI-related crawlers are separate, and they do separate jobs:

User agent Owner Purpose
GPTBot OpenAI Training/index crawling
OAI-SearchBot OpenAI Retrieval crawling for ChatGPT search
ChatGPT-User OpenAI Live fetch when a user shares a URL
ClaudeBot / Claude-SearchBot Anthropic Claude training and retrieval
PerplexityBot Perplexity Perplexity search results
Google-Extended Google Training control for Gemini / Vertex AI

This is why “we blocked GPTBot in robots.txt, so why does ChatGPT still cite other people?” is such a common confusion — you blocked training; retrieval may be a different line in the file. The reverse happens too: a site blocks GPTBot on training-data grounds and accidentally severs its retrieval path. Plenty of cross-border stores have done exactly that.

Third-party audits put the numbers at different levels (one study found about 19% of audited sites completely inaccessible to AI agents; another says around 30% block AI crawlers without realizing it). The direction is consistent: the most common failure point is access, not content. Teams rewrite the article when the actual problem is one robots.txt line or one rendering choice.

Neo’s take: this is the third face of one problem, alongside our September 4 piece on whether to block AI crawlers and our September 2 piece on the technical signals audit. In technical work, sequence usually beats effort — confirm the door is open before debating how nice the room looks.

Add-on two: track three distinct states, never one

Every retrieval test returns one of three outcomes, and they mean completely different things:

State Meaning What to do
Absent The engine neither mentioned nor cited you Go back to the troubleshooting list; fix retrievability first
Mentioned (no link) The model named your brand but gave no link Most likely training-data memory, not this retrieval
Cited (with link) The answer links to your page or site Retrieval worked — now the GEO work starts

Blend them together and you produce a useless conclusion like “we’re doing fine in AI search.” Mentioned and cited are different things — we covered that in detail back on August 30.

Cross-check it: pin the finding down with curl and GSC

A chatbot’s answer is one observation, not a fact. So every time, verify the same pages with two harder instruments:

  1. curl without JS: curl -A "Mozilla/5.0" -s https://yourdomain.com/key-page | grep "a keyword from your passage". A hit means the body is in the server-returned HTML, which AI retrieval crawlers (which mostly don’t execute JS) can read. A miss means your body depends on JS rendering — the problem is in the render layer, not the content layer.
  2. Search Console URL Inspection: check that Google’s rendered HTML matches what users see. This also surfaces the “error shell” problem from our other piece today. If you already have Search Console’s generative AI report, look at it too — but remember it only covers AI Overviews and AI Mode, not ChatGPT or Perplexity.

Should you install the Chrome extension?

Green built one to save himself the grunt work (Exactly Matchy). It reads the rendered page, strips boilerplate like navigation, cookie banners and footers, then uses Chrome’s on-device model to pick 20–30 word passages containing specific names, numbers and unusual phrasing, and turns each into an exact-match prompt with one-click links into ChatGPT, Claude and Gemini.

A few details worth knowing: the on-device model only ranks and selects passages, it doesn’t rewrite them, so page text never goes to a third-party API — no keys, no cost. If Chrome’s AI isn’t available, it falls back to simpler scoring. The author also flags the risk: right now it’s only usable by forking the repo and loading it unpacked via chrome://extensions in developer mode. Install at your own risk.

Neo’s take: the extension solves the tedious half — picking snippets. But you don’t need it. Hand-pick a few sentences with concrete numbers or proper nouns (something like “across the 137 pages we tested…”), drop them into a spreadsheet, and run the test manually. That’s under 30 minutes. The extension just saves you that step every time — a month-two optimization.

The 30-minute checklist

Do it in this order once, then repeat monthly:

  1. Pick five pages: three core product/service pages, one long blog post, and one page you’re suspicious about. Don’t test five homepages.
  2. Hand-pick two distinctive passages per page (20–30 words, containing specific numbers, names or unusual phrasing).
  3. curl every page and confirm the body is in the raw HTML without JS.
  4. Read robots.txt against the crawler list above and confirm, line by line, which bots are intentionally not blocked.
  5. Run the exact-match test: three engines (ChatGPT / Gemini / Perplexity) × two passages per page × three samples each. Record absent / mentioned / cited.
  6. Keep the sheet: date, engine, page, passage ID, result, note (e.g. “returned after switching passages”). The sheet only becomes valuable at the third run — that’s when you start telling fluctuation apart from real change.

Final word

The value of this method isn’t what it proves. It’s that it converts “AI can’t see us” from a vague complaint into a layer you can isolate: not discovered, not fetchable, not renderable, or not indexable.

And the priority order for any site owner is this:

Retrievable > citable > traffic.

Most teams start from the last one — researching how to write more “AI-friendly” content while the front door has been shut the whole time. Spend 30 minutes confirming the door opens. Then talk about content.

If you do one thing today: pick one core page, lift a real passage off it, wrap it in quotes, and throw it at ChatGPT. See whether it hands your own page back to you.