AI Can Find Your Site. Can It Understand It? A 50-Site Audit Reveals What Most Sites Are Missing


Hi everyone, this is Neo.

AI visibility has dominated SEO conversations for the past two years. Teams have cleaned up code, chunked content, and made it easier for AI bots to crawl and ingest. If your brand has started showing up regularly in AI-generated answers or AI Overviews, you could be forgiven for thinking the battle for AI visibility is nearly won.

SALT.agency co-founder Reza Moaiandin’s latest audit pours cold water on that optimism. His team audited 50 major global websites across retail, SaaS, travel, publishing, and finance — and the verdict is sobering: most sites have made it easier for AI to find them, but almost none have made it possible for AI to truly understand them. And nearly two-thirds leave the question of which AI bots can access which content entirely to luck.

Here’s the full breakdown: the three-layer AI readiness model, the 27 audit elements, the real scores behind big-name sites, and the homework we as independent sellers can copy.

First, Fix a Misconception: AI Visibility ≠ Being Cited

Ask most marketing teams to define “AI visibility” and the answer comes back as mentions and citations — if someone asks an AI platform a question relevant to your category, you want a better-than-average chance of appearing in the response.

Reza argues that definition is old SEO thinking in a new wrapper — get ranked, get found, get clicked — transplanted into a very different form of search. But visibility in AI isn’t just about getting your brand and links in front of the right eyeballs. It’s also about how AI understands your information and interacts with your website.

His analogy: teaching a six-year-old to read isn’t the same as teaching them to understand. A kid can sound out the words fluently — cookie from the jar. But ask what they just read, and you may find they absorbed less than you thought. Once they master both reading and comprehension, they can act on it: follow a recipe, write a story, run down follow-up research.

Reading is one step toward comprehension. Comprehension is one step toward agency. The same progression applies to AI and your website.

The Three-Layer AI Readiness Framework

That’s the logic behind SALT.agency’s audit framework, built on three distinct layers:

Layer One: Retrievability — Can AI fetch and parse your content without stumbling?

This is the “reading” layer, the foundational one that overlaps most with conventional technical SEO — ARIA labeling, AI user-agent directives in robots.txt. Most teams already optimize here, because it maps directly to citations and brand mentions.

Layer Two: Attribution & Meaning — Can AI determine what your pages are about and who owns them?

This is the “comprehension” layer. How does AI know which number on the page is the product price, and which is the member-only discount? How does it determine who the brand or author really is, and whether they’re a trusted authority? How does it tell whether an article is about the car brand Jaguar, or one of the many Jaguar football clubs?

Higher comprehension confidence does two things at once: it increases the odds of your content being cited in relevant AI responses, and it reduces the risk of your brand being misrepresented.

Layer Three: Agency — Can AI agents access your site’s capabilities to carry out tasks?

This is the “action” layer — the difference between AI parroting your information back and AI transacting with your business on a user’s behalf.

The full framework runs 27 audit elements: 11 in Layer One, 3 in Layer Two, 13 in Layer Three. Each is classified by industry-standard maturity: Established, Emerging, or Frontier.

The 50-Site Audit: The Gaps Are Bigger Than You’d Think

Using an instrumented browser, SALT.agency captured live HTTP responses, rendered DOM, raw server HTML, and machine-discovery endpoints for all 50 sites on the same day (June 12, 2026), scoring 12 Established signals 0/1/2.

Layer One: mostly okay, but not a pass. Only three sites scored below 50%. Sounds fine — until you remember this layer overlaps most with conventional technical SEO, so much of it was already in place.

Finding one: 70% of homepages have JSON-LD structured data. 35 out of 50, with 32 scoring the full 2 points. But flip it around — nearly a third of major-brand homepages have no JSON-LD at all. Schema is how AI knows “this is a product, this is the price, this is the brand that sells it.” Without it, the LLM makes its best guess — and a wrong guess means wrong answers, or outright hallucinations.

Finding two: only 5 of 50 sites have implemented Cloudflare’s Content Signals Policy. These robots.txt directives spell out exactly what crawlers may do with your content across search indexing, live AI query responses, and model training. With it, your AI strategy stops being a crude “block all / allow all” binary and becomes something you can manage.

Layer Two: scores drop off a cliff. The gap between layers is the most uncomfortable part of the audit — most sites have done almost nothing to make themselves understandable.

Layer Three: near-total wipeout. Agentic browsers and agentic commerce only arrived in late 2025; of the 13 protocols tied to this layer, only two qualify as Established. Of the 48 sites where endpoint testing was possible, 46 scored zero. Even the best performers — airbnb.com and vercel.com — only implemented OAuth authorization server metadata (50% each) and skipped the protected resource metadata that would let AI interact with their sites securely.

Three Counterexamples Worth Studying

Counterexample one: Expedia — “open for business” sign, locked door.

Expedia published an llms.txt file that reads like a love letter to AI — canonical identity, neatly chunked copy about what the brand is and does. And yet Expedia.com scored just 33.3% overall, one of the lowest in the entire cohort: no JSON-LD, no sitemap declaration, only a quarter of content server-side rendered.

That’s the equivalent of putting an “open for business” sign in your window and forgetting to unlock the door. Reza’s point: a beautifully written llms.txt is only an intent signal — it doesn’t replace the rest of the technical work.

Counterexample two: Amazon — 29.2%, and it’s deliberate.

amazon.com scored just 29.2% overall, one of the three sites below 50% on Layer One. But Reza is explicit: I don’t think Amazon overlooked AI search. Like the BBC, Amazon’s robots.txt blocks almost all AI bots — this is design, not oversight. BBC, CNN, and The Guardian block most or all AI bots for the same understandable reason: their business models depend on people visiting their sites.

Counterexample three: eBay and Tripadvisor — the pickiest of the bunch. They’ve made selective, bot-by-bot decisions: allow the AI bots they want, block the ones they don’t. Airbnb and Cloudflare haven’t imposed blanket blocks but have written specific access rules for most major AI bots.

Meanwhile, 29 of the 50 sites (nearly two-thirds) appear to have made no deliberate decision about AI agent access one way or the other. Nothing blocked, nothing explicitly allowed, no rules for crawlers to follow. What AI bots can access, how they interpret it, and how they use it in responses is largely left to chance.

Let’s Be Clear: robots.txt Directives Are Not Law

Before you rush to edit your robots.txt or add an llms.txt, Reza flags some uncomfortable facts:

AI directives in robots.txt are preference statements, not hard rules — compliance is voluntary. GPTBot, ClaudeBot, Google-Extended, and Applebot-Extended broadly respect them; beyond that group, things get patchier, and some bots may crawl you regardless.

Content signals carry even less enforcement weight. Well-behaved bots can choose to follow them; there’s no technical mechanism forcing compliance.

llms.txt is, for now, still an intent signal. It’s an unratified standard with no agreed specification body — SALT’s framework classifies it as Frontier and excludes it from scoring. Still, 11 of the 50 sites have published one, which says something about the industry’s appetite for giving AI a “CliffsNotes” guide to their content.

Bottom line: none of these measures is guaranteed to work. But when has anything in SEO been guaranteed? Optimization has always been about strengthening as many signals as possible — never one silver bullet.

Neo’s Take: What Independent Sellers Can Copy

The most valuable thing about this audit is that it turns “AI technical optimization” from mysticism into a checklist. Here’s how I’d translate it into actions:

First, run your own three-layer self-check.

  • Layer One: does your robots.txt carry explicit directives for the major AI crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, PerplexityBot)? Are you on “block all” or “allow all” — or is there a middle ground?
  • Layer Two: is there JSON-LD on your homepage and core pages? Are key facts — entity, price, brand — machine-readable? Do you have an entity map?
  • Layer Three: this is still far out for most sellers, but know it exists. UCP and ACP are on the way; paying attention now is a head start.

Second, if you decide to block, block on purpose. The lowest-scoring sites in the audit — Amazon, BBC — are precisely the ones that made deliberate choices. The most dangerous position isn’t blocking AI; it’s having no strategy at all, neither blocked nor allowed. At minimum, you should know which crawlers are reading your content and whether you want them to.

Third, don’t fetishize llms.txt. I wrote about this when llms.txt first blew up: it’s valuable, but it’s a garnish — it doesn’t replace structured data or a sitemap. Expedia is the perfect cautionary tale. Fix the fundamentals before you chase the bonus points.

Fourth, Content Signals Policy is worth investigating. If you’re on Cloudflare, this mechanism lets you set different permissions for indexing, live AI answers, and model training — far finer-grained than “block all / allow all.” Only 5 of 50 major sites use it today. That’s a rare area where an independent site can get ahead of the big guys.

Fifth, make peace with one reality: whether AI reads your site is not up to you. It’s probably already happening — through default settings, legacy robots.txt files, and security policies written for a different era. But the content is still yours. You can still exert control over how AI accesses it, interprets it, and interacts with your brand.

The most uncomfortable line in Reza’s audit is this: mentions and citations were never the whole story — they’re only the first layer of real AI visibility. Your brand can be cited regularly and still lose customers to inaccuracies, hallucinations, or an AI that wasn’t confident enough in your information to recommend you.

The competition in AI search has moved from “being seen” to “being understood.” Is your website ready to be read — and believed — by machines?

I’m Neo, and I write about independent site SEO. Run this checklist against your own site and come back — I’d love to hear what you find.