That Red Cell In Your AI Visibility Report Might Not Be A Content Problem At All


Hi everyone, this is Neo.

If you’re running GEO (generative engine optimization) work for an independent site, you probably look at a table like this every month: twenty questions, and next to each one a marker for whether your brand showed up in the AI answer. Some cells are green. Some are red.

Then comes the real question: what does a red cell actually mean?

Not enough content depth? Not enough authority? Or did the model simply never learn your brand exists?

Most of the time, the answer depends on who has the loudest voice in the meeting. The content people say write more articles. The link people say earn authoritative citations. The brand people say buy PR. Those three prescriptions cost wildly different amounts of money, and all three are being justified by the same red cell.

On September 16, Pedro Dias published a piece on Search Engine Journal called “It Was There A Minute Ago.” He’s a former Google Search Quality team member now working as an independent consultant, and he’s one of the few people in this space willing to fire at his own industry. The article looks like a paper review. It’s actually making one point: the leap from “your brand wasn’t mentioned” to “your content is insufficient” isn’t missing a little evidence. It’s missing an entire layer of evidence.

Let me reorganize his argument, add my own read, and work out what it means if you run an independent store.

A red cell has at least four different causes, and they need different medicine

Let’s be blunt. When your brand fails to appear on a given prompt, an AI visibility counter can only tell you the outcome. It has no ability to tell you the cause. And there are at least four plausible causes:

Cause Plain-language version The prescription Where the money goes
Encoding failure The model doesn’t know you exist Content, crawlable structured information, authority signals Content team + technical SEO
Recall failure The model knows, but couldn’t retrieve it this time Different query angles, more consistent association signals (entities, co-occurrence) Entity SEO / brand building
Source conflict The retrieved content contains errors, or a competitor’s version is crowding you out Govern external sources, correct inaccurate descriptions Brand monitoring + PR
Query framing The question you tested was never going to name a brand anyway Test a set of questions closer to real buyer behavior No spend — fix the methodology

Four causes, four prescriptions. And a red/green table that only records “mentioned / not mentioned” cannot distinguish between any of them.

The real value of that SEJ piece is that it walks through three papers to show how hard this distinction actually is. How hard? Hard enough that even opening up the model doesn’t settle it.

Study one: the model had the right answer, and still caved to a wrong tool result

The first paper is MemToC. It tests a very realistic scenario: when a model’s own knowledge collides with what a tool (retrieval) hands back, which one wins?

The design is clean:

  1. No tools first — the model answers a set of factual questions from its own knowledge;
  2. Filter down to the cases where the model answered correctly;
  3. Then deliberately have the tool return a wrong answer and see whether the model changes its mind.

The results aren’t flattering: across four instruction-tuned models, correct-answer retention was only 6.5% to 17.1%.

Flip it around and the same models followed a correct tool result 86.0–93.1% of the time. So this isn’t skepticism about tools — it’s unconditional deference to them. The most telling condition is the third one: when both the model’s own answer and the tool output were wrong, the models still repeated the tool’s wrong answer in 78.4–86.0% of eligible cases.

One more detail worth filing away: researchers hand-annotated 120 responses where the tool returned something incorrect. Not one of them explicitly acknowledged the disagreement. The model doesn’t say “hold on, that contradicts what I just told you.” It just quietly delivers the wrong answer.

Neo’s take: the sample here is five open-weight models in the 7–9B range, and the authors themselves stress that those percentages cannot be applied to ChatGPT search or Google AI Overviews. So don’t use “6.5%” to scare anyone.

But the structural lesson matters for independent sites: what’s actually doing the talking inside an AI answer is usually that one passage of your content that got retrieved. What the model memorized carries surprisingly little weight next to what the retrieval layer hands it. That explains a complaint I hear constantly — “my brand has been in this industry for ten years, why doesn’t AI recognize us?” Because in this particular answer, the AI isn’t recognizing your ten years. It’s recognizing the three lines it just pulled.

Which also means: how other people write about you carries as much weight as how you write about yourself. Sometimes more. That’s the same reason I keep writing about conflicting brand information — it’s the same mechanism.

Study two: models encode 95% of the facts, but retrieval is the bottleneck

The second paper is “Empty Shelves or Lost Keys?” from Google Research and Technion, published at ICML 2026 (arXiv 2602.14080).

It asks a deceptively simple question: when a model gets a fact wrong, did it never learn it (empty shelves), or did it learn it but fail to retrieve it (lost keys)?

To separate the two, the team built a benchmark called WikiProfile — 2,150 facts, ten questions each, using strong-context reproduction to determine whether a fact is “encoded,” and requiring correct answers across four variants to count as “reliably retrievable.” After 13 models and 4 million responses:

  • Encoding is nearly saturated. GPT-5 and Gemini-3 pass the encoding probes for 95–98% of facts. Frontier models have basically stored the facts;
  • Retrieval is the bottleneck. A large share of errors previously blamed on missing knowledge are actually failures to access knowledge;
  • The failures are systematic, and disproportionately hit long-tail facts and reverse questions;
  • “Thinking” recovers a substantial share of failures, which suggests future gains may come less from scaling and more from helping models use what they already encode.

Neo’s take: there’s a translation error we need to guard against. When your brand is missing on a prompt, plenty of GEO vendors will tell you “the model doesn’t know you.” This paper shows those are two different illnesses.

Worse, look where the failures cluster: long-tail facts and reverse questions.

Now think about how buyers actually ask. They don’t ask “is Neo’s independent-site SEO consulting any good” (forward question, brand name included, easy hit). They ask “who are some reliable independent-site SEO consultants” — reverse question plus long tail. And those are exactly the two categories models are worst at.

So strong performance on brand-name prompts tells you nothing about how you perform on category prompts. Those two tests are not measuring the same thing.

Study three: even opening up the model doesn’t settle the diagnosis

The third paper, “From Parameters to Answers,” hit arXiv on September 10 (2609.11859), with first author Wenkang Wei at the University of Science and Technology of China.

This one goes straight into the model’s internals, but the idea translates. Studying country-continent questions, the researchers show that a model’s hidden state carries two distinct things:

  • Content — the actual knowledge that produces the answer;
  • Routing — the mechanism that controls where later layers go to find knowledge.

They then intervene on both, including a “redirect” that points the model at a wrong answer. The findings:

  • Intervening on routing generally changes the output in earlier layers, but becomes less effective in later layers, where the model increasingly holds onto the correct answer;
  • Intervening on answer-supporting content remains consequential throughout.

The authors are upfront that the conclusion depends on which signal you measure and how you change it, and that this provides no universal diagram of how every model retrieves facts.

Neo’s take: what this paper really undermines is the diagnostic vocabulary itself — specifically the phrase “recall failure.”

Because it shows that even with full access to the model and its internal activations, you still have to carefully separate what you can observe from what you can prove affected the answer. So on what basis does someone look at an external red/green table and put “this is a recall failure” on a slide?

Adding technical vocabulary doesn’t supply the missing experiment.

Neo’s take: read the three together and you get diagnostic inflation

These three papers measure three different things. The first tests responses to conflicting evidence. The second contrasts strongly-cued reproduction with reliable answering. The third intervenes on internal activations. As Dias warns, stacking them into one tidy funnel would produce a model that none of the papers actually tested.

But side by side, they point at an industry pattern: we have started writing prescriptions at scale without doing diagnoses.

I call it diagnostic inflation. The symptom (one red cell) is fixed, while the explanations keep getting more sophisticated — insufficient content, insufficient authority, recall failure, training-data coverage, entity ambiguity. Every extra explanation is another service someone can sell you.

Dias includes a self-deprecating example I think is brilliant. Two months ago he announced on LinkedIn that he was the “world’s most renowned AI visibility expert.” He made the title up, as a joke. As of now, searching worlds most renowned ai visibility expert still surfaces his original post cited in Google’s AI Overviews.

He draws out two things:

  1. The AI Overview even states in the answer that the title is a joke he gave himself — and still names him and cites the post;
  2. A counter that only records whether a brand name appeared would log that answer as a clean win.

His point: turning a red cell green isn’t the job. You have to read what the sentence actually says.

Neo’s take: this example is fun to play with if you do brand work — search a category keyword plus a self-appointed superlative and see whether AI takes the bait. But if you’re the one reporting up the chain, it’s a warning: visibility counters are misleading, and the ways they mislead are embarrassingly basic.

So how should you actually diagnose this? A checklist for independent sites

I don’t want to just complain. Here are six rules I now use when reviewing AI visibility, for myself and for clients. Take them.

1. Store the raw answer, not a boolean

Recording only “did the brand appear” wastes a sample. Save the full text of the AI answer, including who it cited, how it described you, and whether it described you as someone else entirely. One sample can answer far more than a true/false.

2. Track brand prompts and category prompts separately

These two have entirely different difficulty and failure mechanics (see study two). Averaging them into one table is like combining body temperature and blood pressure into a single number.

3. Sample the same prompt repeatedly and look at variance first

“We appeared 12 times this month, 15 last month” may be statistical noise. Ask the same prompt twenty times. If you see it bouncing randomly between 11 and 17, you have no change to explain. Confirm the difference exceeds routine variance before you theorize about causes.

4. Watch which sources get cited, not just your own site

Those three lines in the AI answer often come from someone else’s page. The question to ask is: whose content was cited, is it correct, and who wrote it?

If the cited source is a scraper site, an aggregator, or a stale page that describes your brand wrong, your problem isn’t on your own website. It’s in external source governance.

5. Keep interventions small and comparable

Don’t do “full-site rewrite plus ten links plus a PR round” all at once and then attribute the movement to one metric. Change one variable, then watch for two weeks. The lesson from the research is blunt: a correct answer can be displaced by a single wrong tool return, so this system has a lot of moving parts. Without controlling variables, you don’t get conclusions.

6. When you propose something, say what you ruled out

“Content is insufficient” is a conclusion, not a starting point. If you’re recommending more content budget, add one sentence: why do I believe this is a content problem rather than a source conflict or a recall failure?

Dias ends his piece with a question I think every GEO practitioner should keep for self-checks:

If you’re selling me a remedy for that red cell, what evidence tells you which problem I have?

When “insufficient content” is actually a defensible call

In fairness, content problems are real. I think the following genuinely support that diagnosis:

  • Consistent absence on category and general-knowledge prompts — not occasional dips, but weeks of it — while competitors appear consistently;
  • Your site genuinely has no usable content on the topic — nothing of yours shows up in what the AI retrieves;
  • The cited third-party sources describe you inaccurately, or omit you, and the error is locatable and fixable;
  • You re-tested with longer-tail questions closer to real buyer phrasing and got the same result.

Conversely, these are not enough to conclude your content is the problem:

  • You appeared a few times less than last month (could be variance);
  • You only tested brand-name prompts (that measures the brand, not the content);
  • Your content is being cited in AI answers but without your name attached (that’s a citation-versus-attribution problem);
  • Your only evidence is the red/green table itself.

Wrapping up

The point of this isn’t “stop tracking AI visibility.” The opposite — tracking has to happen. The papers themselves concede that with a clearly defined sample, counting is meaningful. You don’t need to understand a model’s internals to count whether it mentioned you.

What I want to say is something else: counting is observation, diagnosis is inference, and between them sits a whole chain of evidence.

Right now, a lot of people are selling inference at observation prices.

As an independent-site owner, you can’t demand that every vendor become a researcher. But you can demand one thing: when someone tells you “AI isn’t recommending your brand because your content isn’t good enough,” make them show you where that came from.

That’s not a high bar. Researchers writing papers are careful to state what they tested and what they didn’t. Why should commercial claims get an exemption?