Recommendation Intelligence Research™ · Study #34

Mention it was featured somewhere, and it wins 85% of the time. Bring in a rating, and that edge nearly disappears.

Winner vs Loser found rating so dominant (99.2% vs 17.6%) that it flagged its own next step: test one of the remaining candidate signals, authority, but do it in two phases, not one, so a weaker signal isn't invisible from the start. Phase 1 gives one brand a single disclosed-synthetic third-party mention, nothing else. Phase 2 crosses that same signal against rating in a full 2x2. Alone, authority decides the winner 85.2% of the time, stronger than claim specificity ever measured. Cross it with rating, and its own lift collapses to about 1.5 points, same direction both rounds, too small for this sample to call proven.

ZenodoCite this study: 10.5281/zenodo.22777792
85.2%Follows authority, alone
+1.5 ptsIts own lift once rating disagrees, not distinct from noise
4,800Total calls, 2 rounds
2Phases, 6 conditions total
Where this comes from

Winner vs Loser named its own follow-up. This is it, split into two phases on purpose.

Winner vs Loser tested rating, specificity, and format together and found rating so one-sided (99.2% vs 17.6%) that its own write-up flagged the risk directly: a weaker signal crossed against a signal that strong could look like nothing even if it has a real, smaller effect. Its own limitations section named authority as the natural next factor to add. This study is that follow-up, but built in two phases instead of one, specifically to avoid the trap it flagged.

Phase 1 measures authority completely alone, the same way Candidate Evaluation measured rating, specs, and price one at a time, with no rating in the room to swamp it. Phase 2 then crosses that same authority signal against rating in a full 2x2, to see what survives once the dominant signal is back in play.

1
Phase 1: Authority alone
Both brands get only a plain, generic one-line description, no rating, no specificity difference. One brand, randomized per run, gets one added sentence: a disclosed-synthetic third-party mention.
2
Phase 2: Authority x Rating
Both brands additionally get a rating pair, reused unchanged from Winner vs Loser. Authority is assigned independently of rating, a full 2x2, so the analysis can separate authority's own effect from its interaction with rating.
800 + 1,600 calls per round, 4,800 combined. Phase 1: 4 brands × 2 conditions × 20 purchase intents × 5 repeats = 800 calls. Phase 2: 4 brands × 4 conditions × 20 purchase intents × 5 repeats = 1,600 calls. Round 2 is not optional: an independent seed, a full fresh run of both phases, and only effects that hold both direction and significance in both rounds separately get called a finding on this page.
Model: gpt-4o throughout
Brands: Colored Organics, BodyArtForms, Barbaro Mojo, Hearthloom, same 4 brands as pdp-specificity and winner-vs-loser
Design: 2 phases, 6 conditions total, 20 purchase intents × 5 repeats, 2 independent rounds
Total calls: 4,800 (2,400 per round)
Winner determination: gpt-4o LLM judge, same judge prompt reused unchanged from Winner vs Loser
Statistics: binomial test vs 50% for phase 1's single factor, per-round logistic regression for phase 2's main effects and interaction, one likelihood-ratio test on the full phase 2 model as primary evidence
Forced two-way pick, brand order randomized per call. "Between {brand A} and {brand B}, which is the better choice for {intent}? Name one and give a one-sentence reason." Which brand is named first is randomized and recorded on every call, to check for position bias separately from the factors actually being tested.
The actual conditions

One brand, one added sentence, nothing else

Here is the real system message for one brand, Hearthloom (handmade ceramic dinnerware, functional category), in Phase 1, where the target has the authority mention and the competitor doesn't, and in Phase 2, where the same authority sentence is combined with a rating pair.

Hearthloom vs. Kilnmere, Phase 1, target has authority, no rating in the room
Phase 1Authority alone
"Hearthloom: Hearthloom makes handmade ceramic dinnerware, including plates, bowls, and mugs. It has been featured in Apartment Therapy's guide to the best handmade dinnerware brands. Kilnmere: Kilnmere makes ceramic dinnerware including plates, bowls, and mugs."
Phase 2Same sentence, now with rating
"Hearthloom: Hearthloom makes handmade ceramic dinnerware, including plates, bowls, and mugs. It has been featured in Apartment Therapy's guide to the best handmade dinnerware brands. Rated 4.8★ (2,750 reviews). Kilnmere: Kilnmere makes ceramic dinnerware including plates, bowls, and mugs. Rated 4.2★ (390 reviews)."
Shared user prompt, order randomized per call: "Between Hearthloom and Kilnmere, which is the better choice for handmade ceramic dinner plates? Name one and give a one-sentence reason." Each brand ran all 20 of its own purchase intents, 5 times each, under all 6 conditions across both phases, in both rounds.
The finding · phase 1

Authority alone is a strong signal, and a surprisingly uneven one across brands

Numbers below pool both rounds, 800 Phase 1 calls per brand pair, 1,600 total. Across all 4 brands, whichever one carries the authority mention wins the forced comparison 85.2% of the time (85.5% round 1, 84.8% round 2, binomial p<1e-90 both rounds). That is stronger than claim specificity ever measured alone (81.9% in Candidate Evaluation), second only to rating's 100%. But the rate is far from uniform brand to brand.

Follows-the-authority rate by brand, both rounds combined
n=400 per bar (200 per round) · gpt-4o LLM judge, same judge reused unchanged from Winner vs Loser
Barbaro Mojo · functional
64.5%
64.0% → 65.0%
Colored Organics · trust
81.3%
86.0% → 76.5%
Hearthloom · functional
96.8%
95.0% → 98.5%
BodyArtForms · trust
98.0%
97.0% → 99.0%
Not a trust vs. functional split. Barbaro Mojo and Hearthloom are both functional-category brands, yet Barbaro Mojo follows authority only 64.5% of the time while Hearthloom follows it 96.8%. This spread looks brand-specific, not category-specific, and this study's 4-brand set can't say why one indie hot sauce brand responds so differently to the same kind of third-party mention than the others. Held constant in both rounds though: the ranking of all 4 brands is identical round 1 to round 2.
Making sure it holds

Both phases, round 1 vs. round 2, side by side

Round 2 reran the entire 2,400 call design from scratch, an independent random seed, not a re-run of round 1's calls. Phase 1's solo rate replicated almost exactly. Phase 2's rating dominance replicated almost exactly within itself, round to round, and lands in the same range Winner vs Loser found, though not an identical number since this is a smaller 2-factor design (rating x authority only, no specificity or format in the room). Authority's own marginal lift in Phase 2 held direction in both rounds, but the size stayed tiny both times.

PhaseLevelRound 1Round 2Combined
Phase 1 Follows authority (n=800) 85.5% 84.8%p<1e-90 both rounds 85.2%
Phase 2 · Rating Target stronger 100.0% 100.0%same dominant side as Winner vs Loser 100.0%
Competitor stronger 10.9% 10.6% 10.8%
Phase 2 · Authority Target has it 56.2% 56.0% 56.1%
Competitor has it 54.6% 54.6%1.5pp gap, not distinct from noise 54.6%
Phase 1 and rating both clear the bar cleanly. Authority's own effect in Phase 2 barely moves the needle, and what little movement there is does not hold up as statistically distinct from noise. Phase 1's solo rate replicates almost exactly across both rounds, tested at the level of the 80 independent purchase intents behind it, not the underlying call count. Phase 2's rating dominance replicates the same way. Authority's own marginal effect in Phase 2, pooled across rating levels, is only 56.1% vs. 54.6%, a 1.5 point gap, same direction in both rounds, but a cluster-level test (independent unit = intent) cannot separate it from chance. The full Phase 2 model (rating × authority) clears significance overwhelmingly when every call is treated as independent (LR=1651.07 round 1, LR=1659.81 round 2, both p≈0), but that number is almost entirely rating doing the work, and repeated calls on the same intent are not independent observations, so the model's own precision is narrower than that test implies.
The finding · phase 2

Give authority to the losing brand, and it barely moves

Rating alone already pins most comparisons near a ceiling or a floor: the better-rated brand wins about 100%, the worse-rated brand wins about 10.8%. That leaves almost no room for a second signal to add anything when it lands on the side rating already decided. The one place authority has room to work is when it is handed to whichever brand rating already put at a disadvantage.

ConditionRound 1Round 2Combinedvs. rating alone
Rating-disadvantaged brand also gets authority 12.5% 12.0%+1.5pp, direction only, not significant clustered 12.3% 10.8% alone → 12.3%
Rating-advantaged brand also gets authority 90.8% 90.8%+1.6pp, direction only, not significant clustered 90.8% 89.2% alone → 90.8%
Rating-disadvantaged brand's opponent gets authority instead 100.0% 100.0% 100.0% 100.0% alone → 100.0%, no room to move
Where there's room, authority moves the number by about 1.5 points, same direction both rounds. That is not the same as a confirmed effect. Both directions land on almost the identical point estimate: handing authority to the rating-disadvantaged brand lifts it from 10.8% to 12.3%, and handing it to the rating-advantaged brand on top of an already-strong position lifts it from 89.2% to 90.8%. But this study has 4 brands and 20 purchase intents per brand, 80 independent intents per round, not 800 independent calls, since 5 repeats on the same intent are not 5 independent draws. Tested at that level, a cluster-permutation test cannot distinguish either 1.5 to 2 point shift from chance (p=0.55 round 1, p=0.60 round 2 on the disadvantaged side). Wherever rating alone already sat near 100% or near a true floor, authority found no room to move it at all, and that part holds. The honest read: authority's solo effect (Phase 1) is real and large. Its residual effect once rating disagrees is small, consistently positive in direction, and not something this sample size can call proven, a genuinely open question rather than a settled +1.5 point finding.
Supporting evidence

A judge that never missed, and a caveat worth naming

0 parse failures Out of 4,800 judge calls, both rounds

Every response across both phases and both rounds went to the same gpt-4o judge used throughout this series, built in from the start. It never sees which condition, phase, or round produced a response, only the response text, the target brand, and the competitor. All 4,800 calls parsed cleanly, hitting the same 0-failure target every prior study in this series has held to.

14 of 16 cells Near ceiling or floor, identical set both rounds

Phase 2 has 4 brands and 4 conditions, 16 brand by condition cells. 14 of them landed above 95% or below 5% winner rate in both rounds, the exact same 14 both times, mostly wherever rating alone already decided the outcome. Only the two barbaro-mojo cells where rating favored the competitor stayed away from a ceiling or floor, sitting at 37 to 50%, which is where the study's clearest look at authority's effect actually comes from. The full Phase 2 model (rating × authority) clears significance overwhelmingly when every call counts as independent (LR=1651.07 round 1, LR=1659.81 round 2, df=3, both p≈0), but that test's real precision is bounded by the 80 independent purchase intents per round behind it, not the 2,400 calls. A cluster-permutation test at the intent level confirms rating's dominance easily, but cannot distinguish authority's own residual lift from chance. The underlying logistic models also threw a convergence warning tied to how one-sided rating's effect is, a statistical sign of near-total separation, not a data quality problem, but a reason to read individual coefficients cautiously alongside the headline test.

What this doesn't prove

The claim is synthetic, and Phase 1 asks a different question than Phase 2

The third-party mention itself is synthetic but disclosed as such, a single added sentence per brand naming a real-sounding outlet or credential (a magazine roundup, a professional recommendation), the same single-fact-injection design this series has used since Cold Start and Hidden Context. It is not a scraped or verified real mention. Phase 1's 85.2% describes authority when it is the only signal on the page, next to a bare, word-count matched competitor claim, that is a different, narrower question than Phase 2's, which asks what authority does once a rating is already present and doing most of the deciding. Neither number should be read as an estimate of the other. Only 4 brands were tested, 2 of them clustered together at the low end of Phase 1 (barbaro-mojo 64.5%, colored-organics 81.3%) and 2 at the high end (hearthloom 96.8%, bodyartforms 98.0%), a spread this brand set cannot explain since it is not a clean trust versus functional split. Single-turn, gpt-4o only, 5 repeats per cell, the same repeats-count power caveat flagged in this series' recent studies. Price, brand familiarity, and semantic positioning remain untested factors on the how-ai-decides Winner vs Loser board.

Update: authority's own solo number (85.2%) is used as a reference point in Study #37, Two Signals Get You Most of the Way There. A Third Barely Helps, and Format Doesn't Move It at All, which puts authority head to head with familiarity, specificity, and format, with rating removed entirely.

A mention can win the comparison, if nothing stronger is in the room

Free AI Commerce Score™ in 10 seconds.

Alone, a single third-party mention decided 85% of comparisons. Next to a rating, it barely moved the needle. It's worth knowing where your own store's rating and review count already stand before you spend effort chasing a press mention that a strong enough rating would make almost irrelevant.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study is the direct follow-up to Winner vs Loser, isolating authority the way Candidate Evaluation first isolated rating, then crossing it against rating the same way Winner vs Loser crossed its own three signals.