Winner vs Loser found rating so dominant (99.2% vs 17.6%) that it flagged its own next step: test one of the remaining candidate signals, authority, but do it in two phases, not one, so a weaker signal isn't invisible from the start. Phase 1 gives one brand a single disclosed-synthetic third-party mention, nothing else. Phase 2 crosses that same signal against rating in a full 2x2. Alone, authority decides the winner 85.2% of the time, stronger than claim specificity ever measured. Cross it with rating, and its own lift collapses to about 1.5 points, same direction both rounds, too small for this sample to call proven.
Winner vs Loser tested rating, specificity, and format together and found rating so one-sided (99.2% vs 17.6%) that its own write-up flagged the risk directly: a weaker signal crossed against a signal that strong could look like nothing even if it has a real, smaller effect. Its own limitations section named authority as the natural next factor to add. This study is that follow-up, but built in two phases instead of one, specifically to avoid the trap it flagged.
Phase 1 measures authority completely alone, the same way Candidate Evaluation measured rating, specs, and price one at a time, with no rating in the room to swamp it. Phase 2 then crosses that same authority signal against rating in a full 2x2, to see what survives once the dominant signal is back in play.
Here is the real system message for one brand, Hearthloom (handmade ceramic dinnerware, functional category), in Phase 1, where the target has the authority mention and the competitor doesn't, and in Phase 2, where the same authority sentence is combined with a rating pair.
Numbers below pool both rounds, 800 Phase 1 calls per brand pair, 1,600 total. Across all 4 brands, whichever one carries the authority mention wins the forced comparison 85.2% of the time (85.5% round 1, 84.8% round 2, binomial p<1e-90 both rounds). That is stronger than claim specificity ever measured alone (81.9% in Candidate Evaluation), second only to rating's 100%. But the rate is far from uniform brand to brand.
Round 2 reran the entire 2,400 call design from scratch, an independent random seed, not a re-run of round 1's calls. Phase 1's solo rate replicated almost exactly. Phase 2's rating dominance replicated almost exactly within itself, round to round, and lands in the same range Winner vs Loser found, though not an identical number since this is a smaller 2-factor design (rating x authority only, no specificity or format in the room). Authority's own marginal lift in Phase 2 held direction in both rounds, but the size stayed tiny both times.
| Phase | Level | Round 1 | Round 2 | Combined |
|---|---|---|---|---|
| Phase 1 | Follows authority (n=800) | 85.5% | 84.8%p<1e-90 both rounds | 85.2% |
| Phase 2 · Rating | Target stronger | 100.0% | 100.0%same dominant side as Winner vs Loser | 100.0% |
| Competitor stronger | 10.9% | 10.6% | 10.8% | |
| Phase 2 · Authority | Target has it | 56.2% | 56.0% | 56.1% |
| Competitor has it | 54.6% | 54.6%1.5pp gap, not distinct from noise | 54.6% |
Rating alone already pins most comparisons near a ceiling or a floor: the better-rated brand wins about 100%, the worse-rated brand wins about 10.8%. That leaves almost no room for a second signal to add anything when it lands on the side rating already decided. The one place authority has room to work is when it is handed to whichever brand rating already put at a disadvantage.
| Condition | Round 1 | Round 2 | Combined | vs. rating alone |
|---|---|---|---|---|
| Rating-disadvantaged brand also gets authority | 12.5% | 12.0%+1.5pp, direction only, not significant clustered | 12.3% | 10.8% alone → 12.3% |
| Rating-advantaged brand also gets authority | 90.8% | 90.8%+1.6pp, direction only, not significant clustered | 90.8% | 89.2% alone → 90.8% |
| Rating-disadvantaged brand's opponent gets authority instead | 100.0% | 100.0% | 100.0% | 100.0% alone → 100.0%, no room to move |
Every response across both phases and both rounds went to the same gpt-4o judge used throughout this series, built in from the start. It never sees which condition, phase, or round produced a response, only the response text, the target brand, and the competitor. All 4,800 calls parsed cleanly, hitting the same 0-failure target every prior study in this series has held to.
Phase 2 has 4 brands and 4 conditions, 16 brand by condition cells. 14 of them landed above 95% or below 5% winner rate in both rounds, the exact same 14 both times, mostly wherever rating alone already decided the outcome. Only the two barbaro-mojo cells where rating favored the competitor stayed away from a ceiling or floor, sitting at 37 to 50%, which is where the study's clearest look at authority's effect actually comes from. The full Phase 2 model (rating × authority) clears significance overwhelmingly when every call counts as independent (LR=1651.07 round 1, LR=1659.81 round 2, df=3, both p≈0), but that test's real precision is bounded by the 80 independent purchase intents per round behind it, not the 2,400 calls. A cluster-permutation test at the intent level confirms rating's dominance easily, but cannot distinguish authority's own residual lift from chance. The underlying logistic models also threw a convergence warning tied to how one-sided rating's effect is, a statistical sign of near-total separation, not a data quality problem, but a reason to read individual coefficients cautiously alongside the headline test.
The third-party mention itself is synthetic but disclosed as such, a single added sentence per brand naming a real-sounding outlet or credential (a magazine roundup, a professional recommendation), the same single-fact-injection design this series has used since Cold Start and Hidden Context. It is not a scraped or verified real mention. Phase 1's 85.2% describes authority when it is the only signal on the page, next to a bare, word-count matched competitor claim, that is a different, narrower question than Phase 2's, which asks what authority does once a rating is already present and doing most of the deciding. Neither number should be read as an estimate of the other. Only 4 brands were tested, 2 of them clustered together at the low end of Phase 1 (barbaro-mojo 64.5%, colored-organics 81.3%) and 2 at the high end (hearthloom 96.8%, bodyartforms 98.0%), a spread this brand set cannot explain since it is not a clean trust versus functional split. Single-turn, gpt-4o only, 5 repeats per cell, the same repeats-count power caveat flagged in this series' recent studies. Price, brand familiarity, and semantic positioning remain untested factors on the how-ai-decides Winner vs Loser board.
Alone, a single third-party mention decided 85% of comparisons. Next to a rating, it barely moved the needle. It's worth knowing where your own store's rating and review count already stand before you spend effort chasing a press mention that a strong enough rating would make almost irrelevant.
This study is the direct follow-up to Winner vs Loser, isolating authority the way Candidate Evaluation first isolated rating, then crossing it against rating the same way Winner vs Loser crossed its own three signals.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →