SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

How consistent are AI brand recommendation lists across repeated prompts?

Can AI tools like ChatGPT and Claude provide reliable, repeatable brand rankings for tracking market visibility? Understanding this matters because companies may be paying for AI tracking products based on metrics that don't actually measure what they claim.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

SparkToro, working with Gumshoe.ai, recruited 600 volunteers to run 12 prompts each through ChatGPT, Claude, and Google's AI Overview (or AI Mode), producing 2,961 recorded responses, to test whether AI tools give consistent enough brand or product recommendation lists to be worth tracking as a "visibility" metric. Their finding: "there's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses," and for the order of items within a list, "it's more like 1 in 1,000 runs before you'd see two lists in the same order." Claude repeated the same set of brands "slightly more" often than ChatGPT or Google AI, but was still unlikely to repeat the order. Separately, across 142 human-written prompts about choosing headphones (994 total AI responses), brands like Bose, Sony, Sennheiser, and Apple appeared in 55-77% of answers; a City of Hope cancer hospital appeared in 69 of 71 ChatGPT answers (97%) for a West Coast cancer-care prompt; and an influencer named Adam Gallagher appeared in 36 of 73 Google AI responses on a men's-fashion prompt.

The post attributes the randomness to two sources of noise stacking on top of each other. First, the tools are "probability engines: they're designed to generate unique answers every time," not a fixed ranked index the way a search engine returns one. Second, real users almost never phrase prompts alike even when their intent matches — semantic similarity across the study's own 142 prompts on one topic averaged just 0.081, which the post compares to "Kung Pao Chicken and Peanut Butter: key ingredients had overlap, but other than being foods with peanuts in them, they're not especially close." Because sampling noise and prompt-wording noise vary independently, ranking position becomes close to meaningless, while appearance frequency across many runs — "visibility %" — still tracks something real, since it mirrors how much material on a topic exists for the model to draw from (a handful of Volvo dealerships in Los Angeles versus thousands of recent science-fiction novels).

This sits alongside Why did AI article share stop growing after 2025?, where a tracking percentage for AI-related content is also confounded by factors the source itself cannot isolate — SparkToro's piece goes further by naming its confound explicitly (model sampling plus prompt variance) and proposing a specific fix (measure visibility % across many runs, not position). It also parallels Do LinkedIn's AI hiring tools actually produce better hires?: both are company-published metrics for an AI system's output where the number being marketed (hiring speed, brand rank) is not the number that matters (hire quality, true market salience) — though SparkToro is unusual in naming the gap inside its own product category and telling buyers to "stop throwing money at AI tracking products that don't provide stats-backed, publicly-reviewable research."

The excerpt does not resolve its own conflict of interest: it runs on SparkToro's blog and leans on Gumshoe.ai's analysis tooling, so a party active in the AI-visibility-tracking market is also the one arguing most of that market is unreliable; the self-critical framing argues against a simple sales motive but does not remove it. The study also measured list membership and order only — the authors state they "didn't even try to collect data on how the AIs described each brand or how positive/negative sentiment" ran — so tone, phrasing, and any downstream effect on actual consumer choice are unmeasured. The data covers three US tools and a handful of topic spaces tested by US volunteers; the specific odds (1-in-100, 1-in-1,000) should not be assumed to hold for other models, languages, or topics, even though the general mechanism they describe plausibly does.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How do educators verify student capability when AI can produce indistinguishable work?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 88 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

SparkToro's experiment finds AI brand recommendation lists repeat less than 1 in 100 times, and in the same order about 1 in 1,000