INQUIRING LINE

Could counting how often an AI mentions your brand beat chasing its rank in any single answer?

Can visibility percentage across many runs measure AI brand prominence more reliably than ranking position?

This explores whether counting how often a brand shows up across many repeated AI answers (its visibility percentage) is a sturdier way to measure how prominent the brand is than tracking where it lands in any single ranked list.


This explores whether counting how often a brand appears across many repeated AI answers tells you more than where it sits in any single list. The corpus points toward yes, mainly because single-answer ranking position barely holds still. SparkToro's experiment collected 2,961 responses and found that AI recommendation lists repeat the same set of brands less than 1 time in 100, and repeat them in the same order only about 1 time in 1,000 How consistent are AI brand recommendation lists across repeated prompts?. If the order almost never repeats, then 'we ranked #2' describes one roll of the dice, not your brand. Inclusion is what survives repetition. Whether a brand appears in 40% of answers or 5% is a stable property you can track over time. Its position in any one answer isn't.

There's a second reason position matters less than it used to: people may not read AI answers as ranked lists at all. Eye-tracking work from RMIT and Microsoft found that when an AI Overview appears, attention to the top-ranked search result falls from 31% to 9%. The AI summary takes over the 'golden' spot, and readers trust it as much as they trusted ranked results Where do searchers look when AI Overviews appear?. If attention goes to the synthesized answer as a whole, being in it matters more than being first in it.

Recommender-system research adds a warning: position itself distorts measurement. YouTube's ranker needs a dedicated component to separate 'this item was good' from 'this item was shown higher' Why do ranking systems need to model selection bias explicitly?. Brand-tracking metrics built on ranking position inherit the same problem, since a slot can reflect where the item happened to be placed as much as real preference. Frequency across many runs spreads that noise out.

Visibility percentage has blind spots too. It depends on which model you ask. Frontier models differ in whether they favor their own company: Claude shows a small but consistent pro-Anthropic lean, while GPT shows one only when grading Do frontier AI models favor their own company?. A visibility score measured on one model is partly a fact about that model, so the reliable version is a percentage across models as well as across runs. There's also a useful parallel in the UK government's Consult tool. Its AI theme-mapping disagreed with human reviewers about as often as the reviewers disagreed with each other, yet those disagreements rarely changed which themes came out on top Does AI theme-mapping perform as well as human reviewers?. The general lesson is that individual outputs can be noisy while totals across many of them stay stable, which is exactly the logic behind visibility percentage.

The corpus has a gap here. It shows why ranking position is unreliable, but it doesn't directly test visibility percentage as a metric: how many runs you need, how sensitive it is to the wording of the prompt, or whether it predicts real outcomes like traffic or sales. The case for it rests on the instability evidence, not on validation of the metric itself.


Sources 5 notes

How consistent are AI brand recommendation lists across repeated prompts?

SparkToro's 2,961-response experiment found AI recommendation lists rarely repeat the same brands (less than 1 in 100) and almost never in the same order (about 1 in 1,000). The instability stems from AI's probabilistic design combined with natural variation in how people phrase similar questions.

Where do searchers look when AI Overviews appear?

Eye-tracking data shows AI Overviews receive significantly longer fixation times, reducing attention to the first-ranked result from 31% to 9%. Trust ratings between AI Overviews and ranked results remained equally high despite this attention shift.

Why do ranking systems need to model selection bias explicitly?

YouTube's multi-objective ranker uses MMoE for conflicting objectives and a shallow position tower to remove selection bias from training data. Without both mechanisms, models converge on degenerate equilibria that amplify their own past decisions.

Do frontier AI models favor their own company?

Claude models show consistent small pro-Anthropic bias across four evaluation tasks, while GPT models show bias only in agentic grading, and Gemini shows weak anti-Google bias. The differences warn against treating company favoritism as universal.

Does AI theme-mapping perform as well as human reviewers?

UK government's Consult tool achieved F1 0.76 against expert reviewers, compared to F1 0.81 between two human reviewers. Differences rarely affected which themes ranked top, suggesting AI performance was competitive despite inherent subjectivity in theme assignment.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.