Recommendation Intelligence Research™ · Study #40

Reword the question three ways. The winner mostly stays the same.

Source Stability tested one thing: ask the identical question ten times, does the winning brand change. It mostly didn't. A LinkedIn reader pushed back with a different number, SparkToro's finding that two separate searches on the same topic return the same list less than 1% of the time, and argued that stability depends on wording, not on the model landing on a real, fixed pick. This study tests that objection directly: same eight product categories, but three differently worded questions per category, all carrying the same buying intent. Two independent rounds, 240 live search calls. The gap between how often the winner repeats on identical wording versus reworded wording was not statistically significant in either round on its own, and combined across both rounds it lands right at the edge of significance, p=0.063. Reformulating the question did not reliably flip the winner.

ZenodoCite this study: 10.5281/zenodo.22833586
+3.1ppCombined cluster diff, borderline (p=0.063)
8Real product categories
240Live search calls, 2 rounds
98.8%Search invocation rate
Where this comes from

A specific objection, tested with data, not argued back on the thread

Source Stability's headline number, winner stability near 90%, came from repeating the exact same prompt ten times per category. That is a test-retest measure: how much noise is in the model's own repetition, holding the question completely fixed. When that number was cited on LinkedIn, Vural Cifci pointed to SparkToro's finding that two independent searches on the same topic return the same list less than 1% of the time, and argued stability is highest exactly where an incumbent is already entrenched, and reads closer to zero for a challenger. Both points are fair. Test-retest and cross-intent variance are genuinely different kinds of variance, and they can both be true at once, which is exactly why this study measures the second one directly instead of just asserting it in a reply.

Recycles the same eight categories and the same recency-cued anchor ("right now in 2026") as Source Stability, so the two studies' numbers sit side by side. The only thing that changes here is the question's syntax: "what's the best X", "which X should I buy", "can you recommend X", same category, same price or time anchor, same buying intent, three different ways of asking for it.

Model: gpt-4o, Responses API with web_search_preview
Categories: 8, identical to Source Stability
Design: 3 phrasings per category, 5 repeats each, 2 independent rounds
Total calls: 480 (240 search, 240 judge extraction)
Winner extraction: gpt-4o LLM judge on free text, identical prompt to Source Stability
Statistics: per-category split by same-phrasing vs. different-phrasing pairs, cluster-permutation test across categories as primary evidence
Pre-registered before the run. The design, the two hypotheses (formulation matters vs. stability holds across formulations), and the statistical plan were written into STUDY-DESIGN.md before either round was collected, so this page reports what the pre-registered test found, not a test picked after seeing the data.
One real example

Same buying intent, three ways of asking. Sometimes the winner moves, sometimes it doesn't.

Electric toothbrushes is the category where rewording made the most difference, in round 1. Here is exactly what each phrasing returned, five repeats each.

1
“What's the best...”
“What's the best electric toothbrush right now in 2026?” → Philips, 5 of 5
2
“Which... should I buy”
“Which electric toothbrush should I buy currently, in 2026?” → SURI, 2 · Philips, 1 · Philips Sonicare, 1 · Oral-B, 1
3
“Can you recommend...”
“Can you recommend an electric toothbrush right now in 2026?” → Philips, 3 · Suri, 1 · SURI, 1
This is what a real +16.2pp category-level diff looks like up close. Phrasing 1 was maximally boring, Philips every single time. Phrasing 2 fragmented across four different brands. That variance is genuinely there in electric toothbrushes, round 1. It just didn't hold as a general pattern once round 2 reran the same design (round 2's diff for the same category was +2.7pp), and it wasn't the pattern most other categories showed at all, headphones, robot vacuums, air fryers, gaming mice, and bluetooth speakers all landed within 1.3pp of zero in both rounds.
The finding

Most categories showed almost no gap. A couple showed a real one.

For each category, every pair of calls within that category is split into two groups: pairs that used the same phrasing, and pairs that used two different phrasings. The number below is the difference in how often those two groups named the same winning brand, both rounds averaged.

Same-phrasing minus different-phrasing gap, by category, both rounds averaged
n=15 usable calls per category per round (3 phrasings × 5 repeats) · gpt-4o LLM judge, 3 of 240 low-confidence
Meal kit
10.0pp
0.0pp → 20.0pp
Electric toothbrush
9.5pp
16.2pp → 2.7pp
Mechanical keyboard
4.7pp
0.0pp → 9.3pp
Headphones
0.7pp
1.3pp → 0.0pp
Robot vacuum
0.0pp
0.0pp → 0.0pp
Air fryer
0.0pp
0.0pp → 0.0pp
Gaming mouse
0.0pp
0.0pp → 0.0pp
Bluetooth speaker
0.0pp
0.0pp → 0.0pp
Five of eight categories landed at or effectively at zero, in both rounds independently. Robot vacuums, air fryers, gaming mice, and bluetooth speakers showed no gap at all, either round. Headphones was close to zero both times. Meal kit and mechanical keyboard both moved from 0.0pp in round 1 to a real gap in round 2, the opposite direction of electric toothbrush, which moved from the largest gap in round 1 down to near zero in round 2. That inconsistency, not a clean repeated pattern, is exactly what "not significant in either round on its own" looks like at the category level.
Making sure it holds

Round 1 vs. round 2, and the cluster test that decides it

Round 2 reran the entire design from scratch on an independent seed. Per the house standard for this series (independent unit = category, not individual call or pair), the primary test cluster-bootstraps and permutation-tests the eight per-category diffs, separately for each round.

CategoryWinner stability (R1 → R2)Cross-phrasing diff (R1 → R2)Modal brand
Meal kit 100.0% → 80.0% +0.0pp → +20.0pplargest R2 diff hellofresh
Electric toothbrush 64.3% → 73.3% +16.2pp → +2.7pplargest R1 diff philips
Mechanical keyboard 53.3% → 53.3% +0.0pp → +9.3pp razer / keychron
Headphones 80.0% → 100.0% +1.3pp → +0.0pp sony
Robot vacuum 100.0% → 100.0% +0.0pp → +0.0pp dreame
Air fryer 100.0% → 100.0% +0.0pp → +0.0pp cosori
Gaming mouse 100.0% → 100.0% +0.0pp → +0.0pp logitech
Bluetooth speaker 100.0% → 100.0% +0.0pp → +0.0pp jbl
Primary evidence: cluster-permutation test across categories, independent unit = category. Round 1, n=8 categories, mean diff +2.2pp, 95% CI [0.0, +6.2]pp, p=0.499. Round 2, n=8, mean diff +4.0pp, 95% CI [0.0, +9.3]pp, p=0.245. Neither round clears significance on its own. Combined, treating both rounds as 16 separate category-clusters, mean diff +3.1pp, 95% CI [+0.4, +6.4]pp, p=0.063, closer to conventional significance than either round alone, but still on the wrong side of the usual p<0.05 line. Per this series' no-file-drawer standard, that means the honest read is: no significant effect of wording found in either independent round, with a suggestive but unconfirmed trend when both rounds are pooled.
Vural's other point

Does it matter more when nobody's clearly winning yet?

Vural's exchange raised a second, separate point: stability should be highest where an incumbent is already entrenched, and closer to meaningless for a challenger with no fixed position to defend. Splitting the eight categories by this study's own combined winner stability (both rounds pooled) into “incumbent” (a single brand named at least 90% of the time) and “open” (below that) gives a directional answer, not a confirmed one.

Incumbent categories · avg cross-phrasing diff +0.1pp
Combined winner stability ≥ 90%, both rounds pooled Robot vacuum (100%, dreame) · Air fryer (100%, cosori) · Gaming mouse (100%, logitech) · Bluetooth speaker (100%, jbl) · Headphones (90%, sony)
Open categories · avg cross-phrasing diff +8.0pp
Combined winner stability below 90%, both rounds pooled Meal kit (89.7%, hellofresh) · Electric toothbrush (69.0%, philips) · Mechanical keyboard (50.0%, keychron)
The direction fits Vural's point. The consistency doesn't, yet. Open categories averaged an 8.0pp gap against 0.1pp for incumbent categories, a real difference on its face. But look at the per-round detail in the table above: none of the three open categories replicated its own gap across both rounds, meal kit moved 0.0pp then 20.0pp, mechanical keyboard moved 0.0pp then 9.3pp, and electric toothbrush moved the opposite way, 16.2pp then 2.7pp. With only three open-category data points and each one bouncing round to round, this reads as a real pattern worth testing with more categories and more repeats, not a confirmed one.
Supporting evidence

Search fired almost every time, and the judge barely hedged

98.8% invoked 237 of 240 live search calls, both rounds

Every prompt carries the same recency cue that got 100% invocation in Source Stability and Live Retrieval, “right now in 2026” or “currently, in 2026”. Round 2 hit 120 of 120. Round 1 hit 117 of 120, three calls across meal kit, electric toothbrush, and gaming mouse answered without firing web search. High enough that search invocation itself is not a confound in the comparison above.

3 of 240 Low-confidence judge extractions, both rounds

The winner is extracted by the identical gpt-4o judge prompt used in Source Stability, reading free text and naming whichever brand is presented first or called the top pick. All 3 low-confidence flags landed in round 1; round 2 had zero. Every winner-stability and cross-phrasing number on this page rests on a clear extraction in all but those 3 of 240 calls, which were excluded rather than guessed at.

Why it matters

Wording alone is not a reliable lever on who wins

If cross-intent variance were the dominant force behind AI recommendations, a brand's odds should swing noticeably depending on how a buyer happens to phrase their question, a genuinely alarming idea for anyone trying to measure or influence recommendation share. This study instead found that in five of eight categories, wording made no detectable difference in either of two independent rounds, and in the cluster-level test that treats category as the unit of evidence, the overall effect of wording did not clear significance in either round on its own. That points toward what the earlier studies in this series have already found: what the model already associates with a brand (fame, ratings, authority claims, the fact it already occupies the answer) carries more weight than the specific syntax of the question asked. The one open thread worth someone testing further: whether that holds less firmly in categories with no clear incumbent yet.

What this doesn't prove

Three phrasings, not every possible phrasing

Three formulations per category is a specific, narrow slice of possible syntactic variation, not a claim that covers “any way of asking.” The judge logic and its known fragility (extracting a winner from free text, rather than a forced two-name choice) is recycled unchanged from Source Stability, on purpose, for comparability, which also means it inherits that study's limitations. Eight consumer product categories is a small sample by this series' own standard, and the incumbent-vs-open split above rests on just three open-category data points, too few to treat as confirmed rather than suggestive. Single model (gpt-4o) through the Responses API, not tested cross-model. The combined p=0.063 result sits close enough to conventional significance that a third independent round could plausibly tip it either way, this page reports two rounds, not three.

Wording alone didn't reliably move the winner

Free AI Commerce Score™ in 10 seconds.

Across most of the categories this study measured, the recommended brand held steady no matter how the buyer's question was phrased. Worth knowing where your own store actually stands with AI models before assuming a different search phrase would change the outcome.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study is the direct follow-up to Source Stability, testing the specific objection that stability depends on wording rather than on a genuinely fixed pick.