How AI Decides flags Evaluation, what tips the model once two brands are genuinely close, as the biggest open gap in the whole decision path. So we built the test. First we found 16 real, close contests inside 50 shopping categories. Then, one fact at a time, we handed the model a price, a star rating, or a spec sheet for both brands and asked it to choose. A better rating flipped the verdict in 160 of 160 runs, zero exceptions. A better price barely moved it. A longer spec list landed in between.
The Model Predicts Itself showed that a brand's own recommendation history, not its store quality, not its fame, explains 61.4% of what the model recommends next. Web Search Changes 77% of Recommendations showed that access to live retrieval flips most of those picks entirely. Neither study answers a narrower, more practical question: once a buyer intent has two or three real, close candidates, live in the same conversation, what actually tips the final call? We built this study to answer it directly, by handing the model the kind of comparison facts a store page or a review widget would actually expose.
Three phases. First, 50 open "what's the best X" prompts across 10 categories, 10 runs each, no injected data, to find out which intents even have a real contest, not just a runaway winner. Second, for every genuinely contested pair, the exact same paired question with no comparison data at all, the baseline. Third, the same paired question again, but this time with one fact block about both brands added: price and availability, star rating and review count, or spec depth, matched to three factors already inside the AI Commerce Score™.
Across the three injected-fact conditions, the model followed the better number at very different rates. A star rating and review count won 160 of 160 runs, 100%, with zero exceptions across all 16 pairs. A longer, more specific spec list won 81.9% of runs. A better price and faster availability won only 60.6% of runs, barely above a coin flip.
Follows-the-numbers is the simplest read, but the more rigorous test compares each signal's flip rate, how often the injected fact changed the verdict versus the no-data paired baseline, against a measured noise floor and against each other. The paired baseline itself barely moved on its own: an 11.2% noise floor (95% CI 3.8–20.0%), far lower than the 46–47% floor we measured for open, single-brand recall prompts in the search study. Forced-choice framing alone makes the model far more consistent, before any comparison data is added at all.
The null result we half expected, that the model ignores injected facts and defaults to whatever "memory" already favors, didn't happen. Handed real comparison data, the model used it, decisively for rating, moderately for specs, weakly for price. That means the three factors this study maps to inside the AI Commerce Score™ aren't equally worth chasing once a store is already in a close contest: Recommendation Confidence (reviews schema and sentiment) is the highest-leverage lever by a wide margin, User Intent Match (spec and description depth) is a real but smaller lever, and Commerce & Feed Accuracy (price and inventory sync) is the weakest of the three for tipping a close call, even though it's the easiest data for a store to expose.
The practical read for a merchant already inside a contested category: getting review data into a machine-readable form (schema markup, visible counts, current sentiment) is worth more, per unit of effort, than a price-matching push, at least for the specific moment when an AI system is choosing between you and a named competitor.
Stating the limits up front, the same way the rest of this series does. The injected facts were synthetic, generated to be realistic, not scraped from real product listings, so this measures what happens when a model is handed clean numbers, not what happens when it has to find messier real ones itself. The fact block sat immediately before the question in every run, which likely inflates how deterministically it gets used, a live retrieval pass pulling the same numbers from a real page might weight it differently. This is gpt-4o only, single-turn prompts only, and the 4-of-10 "contested" threshold from Phase 1 is a design choice, not a natural constant, though the main result held under the stricter 3-of-10 threshold as a robustness check.
None of that changes the core, cross-signal finding, that the three facts are not interchangeable and rating clearly outweighs price, but it does mean "100%" should be read as this sample, this model, this prompt shape, not as a permanent constant the way we'd caveat any other number in this series.
Versus 46–47% for open, single-brand recall prompts in the earlier search study. A forced two-way choice is inherently far more stable than asking the model to generate one brand out of an unbounded list, before any comparison data is added at all.
Gravity vs. Gravity Blanket, Anker vs. its own Soundcore line. The model named its own top pick two different ways across runs. Caught and dropped before the paired phases, so they couldn't quietly inflate the flip-rate numbers.
If reviews carry the most weight in a close call, that's where a free scan starts.
This fills the Evaluation gap flagged on the decision map. Read the map for where it fits, or the memory study for the result this one builds on.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →