Recommendation Intelligence Research™ · Study #36

The cited source changes. The winner almost never does.

Live Retrieval found that when ChatGPT's web search actually fires, it returns real citations, but which pages it cites is not fixed, ask the identical question twice and search can surface a different page each time. This study asks the next question directly: when the cited source changes between repeats of the same prompt, does the recommended brand change with it. Eight real product categories, ten independent repeats each, two full rounds, real competing brands instead of this series' usual synthetic four, because live search needs real content to find. Across the categories where the comparison could actually be run, the winner stayed the same brand about 90% of the time regardless of which exact domains search happened to cite that round, and a cluster-level test found no measurable link between which domain got cited and which brand won.

ZenodoCite this study: 10.5281/zenodo.22819482
90%Winner stability, 8 categories
100%Search invocation rate
+7.9 ptsCluster diff, not significant
8Real categories, both rounds
Where this comes from

Live Retrieval showed search is unpredictable. This asks if the pick is too.

Live Retrieval found two things: real web search does not fire reliably on a plain product question, and when it does fire and return citations, those citations are extractable, a real URL, a real title. Neither study in this series had asked the obvious next question directly, when the same prompt is repeated and search happens to cite a different page that time, does the recommended brand change along with it, or does the model land on the same pick regardless of which exact source it was handed. That is the question this study measures directly.

This also breaks from the rest of the series on purpose. Winner vs Loser, Authority, and Brand Familiarity all use the same 4 synthetic brands with hand-written facts and a forced two-name choice, because nothing there is actually being searched for, the model reacts only to what is in the system message. Live search needs real content to find, so this study uses real product categories with real, already-competing brands, and an open-ended “what's the best X” question instead of a forced choice between two named brands. A forced two-name choice hands the model both candidates up front, which Live Retrieval already showed reduces the model's own reason to search at all.

Model: gpt-4o, Responses API with web_search_preview
Categories: 8 real consumer product categories
Design: 10 repeats per category, 2 independent rounds
Total calls: 320 (160 search, 160 judge extraction)
Winner extraction: gpt-4o LLM judge on free text, no forced two-name choice
Statistics: per-category split by citation-domain overlap (Jaccard ≥ 0.5 vs < 0.5), cluster-permutation test across categories as primary evidence
Recency-cued prompts, by design. Live Retrieval found generic product prompts often never trigger search at all. Every prompt here includes an explicit “right now in 2026” cue, a small pilot confirmed 100% invocation across 4 candidate categories before the full 8-category design was committed to, and that 100% rate held across the entire real run.
One real example

The same question, ten times. Almost the same eight pages, every time.

Here is exactly what came back for one category, budget gaming mice, across all 10 repeats in round 1. The prompt never changed. Neither, almost, did the search results, and neither did the winner.

“What's the best budget gaming mouse right now in 2026?”, round 1, all 10 repeats
9 of 10 runsIdentical domain set
pickedtested.com, ryugear.in, search.rakuten.co.jp, coolbox.pe, ebay.co.uk, itechguides.com, overclockers.co.uk, rtings.com → Logitech
1 of 10 runsOne domain swapped
pickedtested.com, search.rakuten.co.jp, ttkgear.com, coolbox.pe, ebay.co.uk, itechguides.com, overclockers.co.uk, rtings.com → Logitech (unchanged)
Every one of the 45 possible pairs across those 10 runs shares at least 7 of 8 domains. Because no pair ever fell below the 50% overlap threshold, there was never a “different source” comparison group to test this category against, which is exactly why it could not enter the primary statistical test below. It is still real data, and it is the most extreme version of this study's own finding: both the source and the winner stayed almost completely fixed for the same real question asked over and over, in both independent rounds.
The finding

In most categories, the winner didn't move. Whether the source did or not.

Winner stability below is the share of repeats, both rounds averaged where both were measurable, that named the same modal brand as the winner. Six of eight categories landed at or effectively at 100%. Source stability (how often the exact same set of cited domains repeated) moved far more, from 2.2% up to 80.0% in the very same repeats, yet the winner mostly didn't follow it.

Winner stability by category, both rounds averaged
n=10 repeats per round · gpt-4o LLM judge on free-text response, 0 low-confidence extractions
Mechanical keyboard
100%
100% → 100%
Air fryer
100%
100% → 100%
Bluetooth speaker
100%
100% → 100%
Gaming mouse
100%
100% → 100%, source 80%→24%
Headphones · round 1 only
100%
round 2 not computable
Robot vacuum · round 2 only
100%
round 1 not computable
Electric toothbrush
90%
80% → 100%
Meal kit
50%
40% → 60%
Meal kits is the one real exception. Every other category held near-total winner stability in both rounds independently, headphones, robot vacuums, mechanical keyboards, air fryers, bluetooth speakers, and gaming mice all included, even while the citation domains behind those same repeats varied far more, from 6.7% to 80.0% exact-match rate. Meal kits is the only category where the winner itself moved meaningfully (40% round 1, 60% round 2) and the only category with a real positive cluster diff in both rounds independently, worth treating as an open question, not folded into the headline average.
Making sure it holds

Every category, round 1 vs. round 2, side by side

Round 2 reran the entire design from scratch on an independent seed, not a replay of round 1's calls. Four of the sixteen category-rounds could not enter the pairwise comparison at all, not because of missing or low-confidence data, but because every single repeat-pair in that category-round shared at least 50% of its cited domains, so there was never a “low overlap” group left to compare against. That is the same defensive behavior demonstrated directly for gaming mice above, applied automatically wherever it occurs.

CategoryWinner stability (R1 → R2)Source stability (R1 → R2)Cluster diff
Meal kit 40.0% → 60.0%only positive diff, both rounds 2.2% → 2.2% +39.3pp / +30.0pp
Electric toothbrush 80.0% → 100.0% 13.3% → 48.9% +25.8pp / +0.0pp
Mechanical keyboard 100.0% → 100.0% 48.9% → 20.0% +0.0pp / +0.0pp
Air fryer 100.0% → 100.0% 6.7% → 13.3% +0.0pp / +0.0pp
Bluetooth speaker 100.0% → 100.0% 26.7% → 46.7% +0.0pp / +0.0pp
Headphones 100.0% → n/a 33.3% → n/a +0.0pp / n/around 2 not computable
Robot vacuum n/a → 100.0% n/a → 46.7% n/a / +0.0ppround 1 not computable
Gaming mouse 100.0% → 100.0% 80.0% → 24.4% n/a / n/ano low-overlap group, both rounds
Primary evidence: cluster-permutation test across categories, independent unit = category, not individual call. Round 1, n=6 usable categories, mean diff +10.8pp, 95% CI [0.0, +24.0]pp, p=0.501. Round 2, n=6, mean diff +5.0pp, 95% CI [0.0, +15.0]pp, p=1.000. Combined, treating both rounds as 12 separate category-clusters, mean diff +7.9pp, 95% CI [0.0, +16.5]pp, p=0.248. None of the three clears significance at any conventional threshold. That supports the winner being largely independent of which exact source search happened to surface, not the reverse, with meal kits as the one category pulling the average up in both rounds and still not enough on its own to move the combined result past chance.
Supporting evidence

A judge that never hedged, and a defensive check that did its job

0 low-confidence Out of 160 judge calls, both rounds

Unlike the rest of this series, the winner here isn't fixed by a forced two-name choice, it's extracted by a gpt-4o judge reading the free-text response and naming whichever brand is presented first or called the top pick. The judge is instructed to flag “low” confidence when a response doesn't clearly rank one item first. Across all 160 judged responses in both rounds combined, none were flagged low-confidence, so every winner stability and source stability number on this page rests on a clear extraction, not a partial guess.

4 of 16 Category-rounds excluded, same reason each time

The analysis script's own defensive design returns no result rather than fabricate a diff when one comparison side is empty, built in from the start, not a later fix. That triggered for headphones (round 2), robot vacuums (round 1), and gaming mice (both rounds independently), each time because essentially every repeat in that category cited an overlapping enough set of domains that no meaningful “different source” group existed to compare against. Confirmed directly by hand for gaming mice: all 45 possible pairs across 10 repeats shared at least 50% of their cited domains in both rounds. Not a bug, the expected behavior when a category's search results are themselves maximally stable.

Why it matters

A brand's AI visibility isn't mostly about which page search happens to find

If the winner were driven mainly by whichever page search happens to surface that time, a brand could move its odds largely by getting picked up on whichever handful of review sites tend to get cited. This study instead found the model's own pick largely independent of the specific source: in 6 of 8 categories the winner barely moved across 20 real, live-search repeats spanning two independent rounds, even while the exact cited domains moved far more. That points influence toward what the model already associates with a brand, the subject of most of the rest of this series, rather than toward chasing whichever page search happens to cite in a given moment. Meal kits, the one category where source and winner moved together in both rounds, is the exception worth someone's attention, not the rule this study found.

What this doesn't prove

A new extraction step, and a small sample by this series' own standard

The winner here is extracted by an LLM judge reading free text, not fixed by a forced two-name choice like the rest of this series, a genuinely more fragile design even though this run had zero low-confidence flags. “Source” is defined here as the full set of cited domains per response, compared by Jaccard similarity between repeat pairs, one reasonable definition, not the only one a different study could choose. Four of sixteen category-rounds could not enter the primary test at all, so the combined result effectively rests on 12 category-rounds, not 16, smaller than most of this series' cluster counts and worth reading with that in mind. Eight consumer product categories is a small, narrow sample, it does not generalize to services, B2B, or categories with a single dominant market leader. Single model (gpt-4o) through the Responses API, not tested cross-model, and meal kits' outlier status is one data point from one study, not independently replicated elsewhere.

The page search cites mostly didn't decide the winner

Free AI Commerce Score™ in 10 seconds.

Across most of the categories this study could measure, the recommended brand held steady even as the exact cited source changed round to round. Worth knowing where your own store actually stands with AI models before assuming a single review-site placement is what moves the needle.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study is the direct follow-up to Live Retrieval, asking what happens to the recommended winner once search's own unpredictability is on the table.