Live Retrieval found that when ChatGPT's web search actually fires, it returns real citations, but which pages it cites is not fixed, ask the identical question twice and search can surface a different page each time. This study asks the next question directly: when the cited source changes between repeats of the same prompt, does the recommended brand change with it. Eight real product categories, ten independent repeats each, two full rounds, real competing brands instead of this series' usual synthetic four, because live search needs real content to find. Across the categories where the comparison could actually be run, the winner stayed the same brand about 90% of the time regardless of which exact domains search happened to cite that round, and a cluster-level test found no measurable link between which domain got cited and which brand won.
Live Retrieval found two things: real web search does not fire reliably on a plain product question, and when it does fire and return citations, those citations are extractable, a real URL, a real title. Neither study in this series had asked the obvious next question directly, when the same prompt is repeated and search happens to cite a different page that time, does the recommended brand change along with it, or does the model land on the same pick regardless of which exact source it was handed. That is the question this study measures directly.
This also breaks from the rest of the series on purpose. Winner vs Loser, Authority, and Brand Familiarity all use the same 4 synthetic brands with hand-written facts and a forced two-name choice, because nothing there is actually being searched for, the model reacts only to what is in the system message. Live search needs real content to find, so this study uses real product categories with real, already-competing brands, and an open-ended “what's the best X” question instead of a forced choice between two named brands. A forced two-name choice hands the model both candidates up front, which Live Retrieval already showed reduces the model's own reason to search at all.
Here is exactly what came back for one category, budget gaming mice, across all 10 repeats in round 1. The prompt never changed. Neither, almost, did the search results, and neither did the winner.
Winner stability below is the share of repeats, both rounds averaged where both were measurable, that named the same modal brand as the winner. Six of eight categories landed at or effectively at 100%. Source stability (how often the exact same set of cited domains repeated) moved far more, from 2.2% up to 80.0% in the very same repeats, yet the winner mostly didn't follow it.
Round 2 reran the entire design from scratch on an independent seed, not a replay of round 1's calls. Four of the sixteen category-rounds could not enter the pairwise comparison at all, not because of missing or low-confidence data, but because every single repeat-pair in that category-round shared at least 50% of its cited domains, so there was never a “low overlap” group left to compare against. That is the same defensive behavior demonstrated directly for gaming mice above, applied automatically wherever it occurs.
| Category | Winner stability (R1 → R2) | Source stability (R1 → R2) | Cluster diff |
|---|---|---|---|
| Meal kit | 40.0% → 60.0%only positive diff, both rounds | 2.2% → 2.2% | +39.3pp / +30.0pp |
| Electric toothbrush | 80.0% → 100.0% | 13.3% → 48.9% | +25.8pp / +0.0pp |
| Mechanical keyboard | 100.0% → 100.0% | 48.9% → 20.0% | +0.0pp / +0.0pp |
| Air fryer | 100.0% → 100.0% | 6.7% → 13.3% | +0.0pp / +0.0pp |
| Bluetooth speaker | 100.0% → 100.0% | 26.7% → 46.7% | +0.0pp / +0.0pp |
| Headphones | 100.0% → n/a | 33.3% → n/a | +0.0pp / n/around 2 not computable |
| Robot vacuum | n/a → 100.0% | n/a → 46.7% | n/a / +0.0ppround 1 not computable |
| Gaming mouse | 100.0% → 100.0% | 80.0% → 24.4% | n/a / n/ano low-overlap group, both rounds |
Unlike the rest of this series, the winner here isn't fixed by a forced two-name choice, it's extracted by a gpt-4o judge reading the free-text response and naming whichever brand is presented first or called the top pick. The judge is instructed to flag “low” confidence when a response doesn't clearly rank one item first. Across all 160 judged responses in both rounds combined, none were flagged low-confidence, so every winner stability and source stability number on this page rests on a clear extraction, not a partial guess.
The analysis script's own defensive design returns no result rather than fabricate a diff when one comparison side is empty, built in from the start, not a later fix. That triggered for headphones (round 2), robot vacuums (round 1), and gaming mice (both rounds independently), each time because essentially every repeat in that category cited an overlapping enough set of domains that no meaningful “different source” group existed to compare against. Confirmed directly by hand for gaming mice: all 45 possible pairs across 10 repeats shared at least 50% of their cited domains in both rounds. Not a bug, the expected behavior when a category's search results are themselves maximally stable.
If the winner were driven mainly by whichever page search happens to surface that time, a brand could move its odds largely by getting picked up on whichever handful of review sites tend to get cited. This study instead found the model's own pick largely independent of the specific source: in 6 of 8 categories the winner barely moved across 20 real, live-search repeats spanning two independent rounds, even while the exact cited domains moved far more. That points influence toward what the model already associates with a brand, the subject of most of the rest of this series, rather than toward chasing whichever page search happens to cite in a given moment. Meal kits, the one category where source and winner moved together in both rounds, is the exception worth someone's attention, not the rule this study found.
The winner here is extracted by an LLM judge reading free text, not fixed by a forced two-name choice like the rest of this series, a genuinely more fragile design even though this run had zero low-confidence flags. “Source” is defined here as the full set of cited domains per response, compared by Jaccard similarity between repeat pairs, one reasonable definition, not the only one a different study could choose. Four of sixteen category-rounds could not enter the primary test at all, so the combined result effectively rests on 12 category-rounds, not 16, smaller than most of this series' cluster counts and worth reading with that in mind. Eight consumer product categories is a small, narrow sample, it does not generalize to services, B2B, or categories with a single dominant market leader. Single model (gpt-4o) through the Responses API, not tested cross-model, and meal kits' outlier status is one data point from one study, not independently replicated elsewhere.
Across most of the categories this study could measure, the recommended brand held steady even as the exact cited source changed round to round. Worth knowing where your own store actually stands with AI models before assuming a single review-site placement is what moves the needle.
This study is the direct follow-up to Live Retrieval, asking what happens to the recommended winner once search's own unpredictability is on the table.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →