Recommendation Intelligence Research™ · Study #19

Hand it a rating, and it follows every single time.

How AI Decides flags Evaluation, what tips the model once two brands are genuinely close, as the biggest open gap in the whole decision path. So we built the test. First we found 16 real, close contests inside 50 shopping categories. Then, one fact at a time, we handed the model a price, a star rating, or a spec sheet for both brands and asked it to choose. A better rating flipped the verdict in 160 of 160 runs, zero exceptions. A better price barely moved it. A longer spec list landed in between.

100%Runs that followed the better rating
16Genuinely contested brand pairs
3Comparison facts tested
1,140Model calls across three phases
Where this comes from

Memory explains who's in the running. It doesn't explain who wins

The Model Predicts Itself showed that a brand's own recommendation history, not its store quality, not its fame, explains 61.4% of what the model recommends next. Web Search Changes 77% of Recommendations showed that access to live retrieval flips most of those picks entirely. Neither study answers a narrower, more practical question: once a buyer intent has two or three real, close candidates, live in the same conversation, what actually tips the final call? We built this study to answer it directly, by handing the model the kind of comparison facts a store page or a review widget would actually expose.

Three phases. First, 50 open "what's the best X" prompts across 10 categories, 10 runs each, no injected data, to find out which intents even have a real contest, not just a runaway winner. Second, for every genuinely contested pair, the exact same paired question with no comparison data at all, the baseline. Third, the same paired question again, but this time with one fact block about both brands added: price and availability, star rating and review count, or spec depth, matched to three factors already inside the AI Commerce Score™.

Model: gpt-4o, temperature 0.7
Phase 1, open baseline: 50 intents × 10 runs = 500 calls
Contested pairs found: 16 of 50 (top-2 within 4 of 10 wins)
Phase 2, paired baseline: 16 pairs × 10 runs = 160 calls
Phase 3, paired + one fact: 16 pairs × 3 signals × 10 runs = 480 calls
Total model calls: 1,140
A data issue we caught before it mattered. Two of the 18 initial contested pairs turned out to be the same brand named two different ways across runs, "Gravity" vs. "Gravity Blanket", "Anker" vs. its own "Anker Soundcore" sub-brand. We dropped both before running the paired phases, since comparing a brand against itself would have quietly inflated every flip-rate number downstream. That left 16 genuine two-brand contests.
The finding

Rating flips it. Every time

Across the three injected-fact conditions, the model followed the better number at very different rates. A star rating and review count won 160 of 160 runs, 100%, with zero exceptions across all 16 pairs. A longer, more specific spec list won 81.9% of runs. A better price and faster availability won only 60.6% of runs, barely above a coin flip.

Share of runs that followed the better number, by signal
480 Phase 3 runs across 16 contested pairs, one injected fact per run
Price + availability
60.6%
Spec depth
81.9%
Rating + review count
100%
Zero exceptions
The rating number isn't rounded down for effect. It really is 160 out of 160, and the per-pair breakdown is just as clean: every one of the 16 contested pairs split exactly 5 of 10 runs one way and 5 of 10 the other, matching precisely which brand had the better rating in that run. There is effectively zero variance between pairs, which is why the confidence interval on this number is 100% to 100%, not a typo, a genuinely deterministic pattern in this sample.
Which fact matters most

Rating beats price by a wide, real margin

Follows-the-numbers is the simplest read, but the more rigorous test compares each signal's flip rate, how often the injected fact changed the verdict versus the no-data paired baseline, against a measured noise floor and against each other. The paired baseline itself barely moved on its own: an 11.2% noise floor (95% CI 3.8–20.0%), far lower than the 46–47% floor we measured for open, single-brand recall prompts in the search study. Forced-choice framing alone makes the model far more consistent, before any comparison data is added at all.

Flip rate vs. the no-data paired baseline, by signal
Cluster bootstrap 95% CI over 16 intents · noise floor 11.2% (dashed reference)
Price · p=0.055, not quite significant
26.9%
Specs · p=0.0004, real
41.9%
Rating · p<0.0001, real
50.0%
Strongest
Rating and price are statistically different, not just numerically different. A pairwise permutation test, Bonferroni-adjusted across the three comparisons, finds rating beats price by 23.1 points, p=0.0012, real. Price vs. specs (p=0.203) and rating vs. specs (p=0.479) don't clear that same bar after correction, so the safest claim is: rating is the standout, and specs sits somewhere in the middle without a confirmed gap to either neighbor.
Why it matters

Evaluation isn't theater. It's just unevenly weighted

The null result we half expected, that the model ignores injected facts and defaults to whatever "memory" already favors, didn't happen. Handed real comparison data, the model used it, decisively for rating, moderately for specs, weakly for price. That means the three factors this study maps to inside the AI Commerce Score™ aren't equally worth chasing once a store is already in a close contest: Recommendation Confidence (reviews schema and sentiment) is the highest-leverage lever by a wide margin, User Intent Match (spec and description depth) is a real but smaller lever, and Commerce & Feed Accuracy (price and inventory sync) is the weakest of the three for tipping a close call, even though it's the easiest data for a store to expose.

The practical read for a merchant already inside a contested category: getting review data into a machine-readable form (schema markup, visible counts, current sentiment) is worth more, per unit of effort, than a price-matching push, at least for the specific moment when an AI system is choosing between you and a named competitor.

What this doesn't prove

Read the 100% carefully, not as a universal law

Stating the limits up front, the same way the rest of this series does. The injected facts were synthetic, generated to be realistic, not scraped from real product listings, so this measures what happens when a model is handed clean numbers, not what happens when it has to find messier real ones itself. The fact block sat immediately before the question in every run, which likely inflates how deterministically it gets used, a live retrieval pass pulling the same numbers from a real page might weight it differently. This is gpt-4o only, single-turn prompts only, and the 4-of-10 "contested" threshold from Phase 1 is a design choice, not a natural constant, though the main result held under the stricter 3-of-10 threshold as a robustness check.

None of that changes the core, cross-signal finding, that the three facts are not interchangeable and rating clearly outweighs price, but it does mean "100%" should be read as this sample, this model, this prompt shape, not as a permanent constant the way we'd caveat any other number in this series.

Supporting evidence

Two more signs the setup was clean

11.2% paired noise floor

Versus 46–47% for open, single-brand recall prompts in the earlier search study. A forced two-way choice is inherently far more stable than asking the model to generate one brand out of an unbounded list, before any comparison data is added at all.

2 of 18 were one brand, twice

Gravity vs. Gravity Blanket, Anker vs. its own Soundcore line. The model named its own top pick two different ways across runs. Caught and dropped before the paired phases, so they couldn't quietly inflate the flip-rate numbers.

Know which fact is worth fixing first

Free AI Commerce Score™ in 10 seconds.

If reviews carry the most weight in a close call, that's where a free scan starts.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This fills the Evaluation gap flagged on the decision map. Read the map for where it fits, or the memory study for the result this one builds on.