Authority found that a specific third-party mention keeps a small, real 1.5 point lift even after rating enters the room. This is the same two-phase design applied to a related but plainer claim: no named outlet, no third party, just a bare, unsourced assertion that the brand is "one of the most recognized" names in its category. Phase 1 gives one brand a single added sentence, nothing else. Phase 2 crosses that same signal against rating in a full 2x2. Alone, the claim decides the winner 79.8% of the time, comparable to claim specificity. Cross it with rating, and its own marginal lift falls to 0.2 points, direction holds in both rounds, size is essentially zero, the most completely swamped signal measured in this series so far.
Authority tested a disclosed third-party mention, a magazine roundup or a professional recommendation, and found its point estimate held at a small, same-direction 1.5 points even once rating was back in the room, though not large enough for this sample to call it confirmed. This study asks the same question about a related but deliberately plainer signal: no named outlet, no attribution to anyone, just a brand asserting its own broad recognition in the same flat, unsourced style as the baseline claim itself. Distinct from claim-attribution too, which varied who a claim is attributed to, not what the claim says. Here only the content changes, a popularity assertion instead of a functional description.
Phase 1 measures familiarity completely alone, the same way Candidate Evaluation measured rating, specs, and price one at a time, with no rating in the room to swamp it. Phase 2 then crosses that same familiarity signal against rating in a full 2x2, to see what survives once the dominant signal is back in play.
Here is the real system message for one brand, Hearthloom (handmade ceramic dinnerware, functional category), in Phase 1, where the target has the familiarity claim and the competitor doesn't, and in Phase 2, where the same familiarity sentence is combined with a rating pair.
Numbers below pool both rounds, 800 Phase 1 calls per brand pair, 1,600 total. Across all 4 brands, whichever one carries the familiarity claim wins the forced comparison 79.8% of the time (80.6% round 1, 78.9% round 2, binomial p<1e-60 both rounds). That is comparable to claim specificity alone (81.9% in Candidate Evaluation) and a little below authority's own 85.2%. But one brand sits apart from the other three.
Round 2 reran the entire 2,400 call design from scratch, an independent random seed, not a re-run of round 1's calls. Phase 1's solo rate replicated closely. Phase 2's rating dominance replicated almost exactly within itself, round to round. Familiarity's own marginal lift in Phase 2 held direction in both rounds, but the size is the smallest of any signal measured against rating in this series, close enough to zero that round 2 alone rounds to a flat 0.0 point gap.
| Phase | Level | Round 1 | Round 2 | Combined |
|---|---|---|---|---|
| Phase 1 | Follows familiarity (n=800) | 80.6% | 78.9%p<1e-60 both rounds | 79.8% |
| Phase 2 · Rating | Target stronger | 99.9% | 99.6%same dominant side as Authority | 99.8% |
| Competitor stronger | 11.2% | 11.1% | 11.2% | |
| Phase 2 · Familiarity | Target has it | 55.8% | 55.4% | 55.6% |
| Competitor has it | 55.4% | 55.4%only 0.2pp gap, round 2 exactly 0.0pp | 55.4% |
Rating alone already pins most comparisons near a ceiling or a floor: the better-rated brand wins about 99.8%, the worse-rated brand wins about 11.2%. Authority's point estimate in the same spot was a point and a half; familiarity's is under half a point. Neither one clears a properly clustered significance test (independent unit = purchase intent, not individual call), so read both as small, same-direction, unproven signals rather than a confirmed effect that familiarity merely falls short of.
| Condition | Round 1 | Round 2 | Combined | vs. rating alone |
|---|---|---|---|---|
| Rating-disadvantaged brand also gets familiarity | 11.8% | 11.5%+0.4pp over baseline, both rounds | 11.6% | 11.2% alone → 11.6% |
| Rating-advantaged brand also gets familiarity | 99.8% | 99.2% | 99.5% | 99.8% alone → 99.5%, no additional room |
| Rating-disadvantaged brand's opponent gets familiarity instead | 100.0% | 100.0% | 100.0% | 100.0% alone → 100.0%, no room to move |
Every response across both phases and both rounds went to the same gpt-4o judge used throughout this series, built in from the start. It never sees which condition, phase, or round produced a response, only the response text, the target brand, and the competitor. All 4,800 calls parsed cleanly, hitting the same 0-failure target every prior study in this series has held to.
Phase 2 has 4 brands and 4 conditions, 16 brand by condition cells. 14 of them landed above 95% or below 5% winner rate in both rounds, the exact same 14 both times, mostly wherever rating alone already decided the outcome. Only the two barbaro-mojo cells where rating favored the competitor stayed away from a ceiling or floor, sitting at 43 to 47%, close to a coin flip, which is where the study's clearest look at familiarity's effect actually comes from. That is why the full Phase 2 model, not any single cell, is the primary evidence here (LR=1621.71 round 1, LR=1605.73 round 2, df=3, both p≈0). The underlying logistic models also threw a convergence warning tied to how one-sided rating's effect is, a statistical sign of near-total separation, not a data quality problem, but a reason to read individual coefficients cautiously alongside the headline test.
The familiarity claim itself is synthetic but disclosed as such, a single added sentence per brand asserting broad recognition with no name attached, the same single-fact-injection design this series has used since Cold Start and Hidden Context. It is not a scraped or verified real popularity signal, and it deliberately avoids naming any third party (that is Authority) or changing who a claim is attributed to (that is claim-attribution). Phase 1's 79.8% describes familiarity when it is the only signal on the page, next to a bare, word-count matched competitor claim, that is a different, narrower question than Phase 2's, which asks what familiarity does once a rating is already present and doing most of the deciding. Neither number should be read as an estimate of the other. Only 4 brands were tested, and one of them, Barbaro Mojo, responded to familiarity at barely above chance (53.5%, identical in both rounds) while the other three ranged from 81.8% to 96.8%, a spread this brand set cannot explain since it is not a clean trust versus functional split. Single-turn, gpt-4o only, 5 repeats per cell, the same repeats-count power caveat flagged in this series' recent studies. Price, semantic positioning, and structured information remain untested factors on the how-ai-decides Winner vs Loser board.
Alone, a bare claim of being widely known decided 80% of comparisons. Next to a rating, it added next to nothing. It's worth knowing where your own store's rating and review count already stand before you spend effort on a broad brand-awareness claim that a strong enough rating would make almost irrelevant.
This study is the direct sibling to Authority, isolating a plainer claim of familiarity the same way Authority isolated a specific third-party mention, then crossing it against rating the same way Winner vs Loser crossed its own three signals.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →