Recommendation Intelligence Research™ · Study #35

Claim the brand is widely known, and it wins 80% of the time. Bring in a rating, and the edge is nearly gone.

Authority found that a specific third-party mention keeps a small, real 1.5 point lift even after rating enters the room. This is the same two-phase design applied to a related but plainer claim: no named outlet, no third party, just a bare, unsourced assertion that the brand is "one of the most recognized" names in its category. Phase 1 gives one brand a single added sentence, nothing else. Phase 2 crosses that same signal against rating in a full 2x2. Alone, the claim decides the winner 79.8% of the time, comparable to claim specificity. Cross it with rating, and its own marginal lift falls to 0.2 points, direction holds in both rounds, size is essentially zero, the most completely swamped signal measured in this series so far.

ZenodoCite this study: 10.5281/zenodo.22778177
79.8%Follows familiarity, alone
+0.2 ptsIts own marginal lift once rating is added
4,800Total calls, 2 rounds
2Phases, 6 conditions total
Where this comes from

Authority named itself as one data point. This asks the same question about a plainer claim.

Authority tested a disclosed third-party mention, a magazine roundup or a professional recommendation, and found its point estimate held at a small, same-direction 1.5 points even once rating was back in the room, though not large enough for this sample to call it confirmed. This study asks the same question about a related but deliberately plainer signal: no named outlet, no attribution to anyone, just a brand asserting its own broad recognition in the same flat, unsourced style as the baseline claim itself. Distinct from claim-attribution too, which varied who a claim is attributed to, not what the claim says. Here only the content changes, a popularity assertion instead of a functional description.

Phase 1 measures familiarity completely alone, the same way Candidate Evaluation measured rating, specs, and price one at a time, with no rating in the room to swamp it. Phase 2 then crosses that same familiarity signal against rating in a full 2x2, to see what survives once the dominant signal is back in play.

1
Phase 1: Familiarity alone
Both brands get only a plain, generic one-line description, no rating, no specificity difference. One brand, randomized per run, gets one added sentence: an unsourced, self-asserted claim of broad recognition.
2
Phase 2: Familiarity x Rating
Both brands additionally get a rating pair, reused unchanged from Winner vs Loser and Authority. Familiarity is assigned independently of rating, a full 2x2, so the analysis can separate familiarity's own effect from its interaction with rating.
800 + 1,600 calls per round, 4,800 combined. Phase 1: 4 brands × 2 conditions × 20 purchase intents × 5 repeats = 800 calls. Phase 2: 4 brands × 4 conditions × 20 purchase intents × 5 repeats = 1,600 calls. Round 2 is not optional: an independent seed, a full fresh run of both phases, and only effects that hold both direction and significance in both rounds separately get called a finding on this page.
Model: gpt-4o throughout
Brands: Colored Organics, BodyArtForms, Barbaro Mojo, Hearthloom, same 4 brands as pdp-specificity, winner-vs-loser, and authority-signal
Design: 2 phases, 6 conditions total, 20 purchase intents × 5 repeats, 2 independent rounds
Total calls: 4,800 (2,400 per round)
Winner determination: gpt-4o LLM judge, same judge prompt reused unchanged from Winner vs Loser and Authority
Statistics: binomial test vs 50% for phase 1's single factor, per-round logistic regression for phase 2's main effects and interaction, one likelihood-ratio test on the full phase 2 model as primary evidence
Forced two-way pick, brand order randomized per call. "Between {brand A} and {brand B}, which is the better choice for {intent}? Name one and give a one-sentence reason." Which brand is named first is randomized and recorded on every call, to check for position bias separately from the factors actually being tested.
The actual conditions

One brand, one added sentence, nothing else

Here is the real system message for one brand, Hearthloom (handmade ceramic dinnerware, functional category), in Phase 1, where the target has the familiarity claim and the competitor doesn't, and in Phase 2, where the same familiarity sentence is combined with a rating pair.

Hearthloom vs. Kilnmere, Phase 1, target has familiarity, no rating in the room
Phase 1Familiarity alone
"Hearthloom: Hearthloom makes handmade ceramic dinnerware, including plates, bowls, and mugs. It is one of the most recognized and widely known handmade dinnerware brands among home cooks. Kilnmere: Kilnmere makes ceramic dinnerware including plates, bowls, and mugs."
Phase 2Same sentence, now with rating
"Hearthloom: Hearthloom makes handmade ceramic dinnerware, including plates, bowls, and mugs. It is one of the most recognized and widely known handmade dinnerware brands among home cooks. Rated 4.8★ (2,750 reviews). Kilnmere: Kilnmere makes ceramic dinnerware including plates, bowls, and mugs. Rated 4.2★ (390 reviews)."
Shared user prompt, order randomized per call: "Between Hearthloom and Kilnmere, which is the better choice for handmade ceramic dinner plates? Name one and give a one-sentence reason." Each brand ran all 20 of its own purchase intents, 5 times each, under all 6 conditions across both phases, in both rounds.
The finding · phase 1

Familiarity alone is a strong signal, and one brand barely responds to it at all

Numbers below pool both rounds, 800 Phase 1 calls per brand pair, 1,600 total. Across all 4 brands, whichever one carries the familiarity claim wins the forced comparison 79.8% of the time (80.6% round 1, 78.9% round 2, binomial p<1e-60 both rounds). That is comparable to claim specificity alone (81.9% in Candidate Evaluation) and a little below authority's own 85.2%. But one brand sits apart from the other three.

Follows-the-familiarity rate by brand, both rounds combined
n=400 per bar (200 per round) · gpt-4o LLM judge, same judge reused unchanged from Winner vs Loser and Authority
Barbaro Mojo · functional
53.5%
53.5% → 53.5%
Hearthloom · functional
81.8%
83.0% → 80.5%
Colored Organics · trust
87.0%
89.5% → 84.5%
BodyArtForms · trust
96.8%
96.5% → 97.0%
Barbaro Mojo barely moves. Every other brand in this series has responded to every signal tested well above the 50% baseline. Barbaro Mojo's 53.5% is the closest any brand-signal combination has come to pure chance anywhere in this series, and it lands at the identical value in both rounds, 53.5% and 53.5%. Hearthloom, the other functional-category brand, still follows familiarity 81.8% of the time, so this isn't a clean trust versus functional split, it looks specific to Barbaro Mojo. This study's 4-brand set can't say why.
Making sure it holds

Both phases, round 1 vs. round 2, side by side

Round 2 reran the entire 2,400 call design from scratch, an independent random seed, not a re-run of round 1's calls. Phase 1's solo rate replicated closely. Phase 2's rating dominance replicated almost exactly within itself, round to round. Familiarity's own marginal lift in Phase 2 held direction in both rounds, but the size is the smallest of any signal measured against rating in this series, close enough to zero that round 2 alone rounds to a flat 0.0 point gap.

PhaseLevelRound 1Round 2Combined
Phase 1 Follows familiarity (n=800) 80.6% 78.9%p<1e-60 both rounds 79.8%
Phase 2 · Rating Target stronger 99.9% 99.6%same dominant side as Authority 99.8%
Competitor stronger 11.2% 11.1% 11.2%
Phase 2 · Familiarity Target has it 55.8% 55.4% 55.6%
Competitor has it 55.4% 55.4%only 0.2pp gap, round 2 exactly 0.0pp 55.4%
Phase 1 and rating both clear the bar cleanly. Familiarity's own effect in Phase 2 is barely there. Phase 1's solo rate (p<1e-60 in both rounds independently) and Phase 2's rating dominance replicate closely. Familiarity's own marginal effect in Phase 2, pooled across rating levels, is 55.6% vs. 55.4% combined, a 0.2 point gap, smaller than authority's own 1.5 point gap and about an order of magnitude smaller than rating's 88 point swing. The full Phase 2 model (rating × familiarity) clears significance overwhelmingly (LR=1621.71 round 1, LR=1605.73 round 2, both p≈0), but as with authority, that test is almost entirely rating doing the work.
The finding · phase 2

Give familiarity to the losing brand, and it barely moves. Give it to the winning brand, and it moves even less.

Rating alone already pins most comparisons near a ceiling or a floor: the better-rated brand wins about 99.8%, the worse-rated brand wins about 11.2%. Authority's point estimate in the same spot was a point and a half; familiarity's is under half a point. Neither one clears a properly clustered significance test (independent unit = purchase intent, not individual call), so read both as small, same-direction, unproven signals rather than a confirmed effect that familiarity merely falls short of.

ConditionRound 1Round 2Combinedvs. rating alone
Rating-disadvantaged brand also gets familiarity 11.8% 11.5%+0.4pp over baseline, both rounds 11.6% 11.2% alone → 11.6%
Rating-advantaged brand also gets familiarity 99.8% 99.2% 99.5% 99.8% alone → 99.5%, no additional room
Rating-disadvantaged brand's opponent gets familiarity instead 100.0% 100.0% 100.0% 100.0% alone → 100.0%, no room to move
Where rating leaves any room, familiarity's point estimate moves by well under a point. Where rating already decided the outcome, it adds nothing. Handing familiarity to the rating-disadvantaged brand lifts it from 11.2% to 11.6%, a 0.4 point move, present in the same direction in both rounds, smaller than authority's equivalent +1.5 point estimate in the same spot. Handing familiarity to the already rating-advantaged brand does not add anything measurable on top, 99.8% alone versus 99.5% with familiarity added. A cluster-permutation test at the intent level (the independent unit here, not the individual call) cannot distinguish either study's disadvantaged-side lift from chance, so neither familiarity's smaller number nor authority's larger one should be read as a confirmed effect yet, just a consistent direction worth more data before calling it real.
Supporting evidence

A judge that never missed, and a caveat worth naming

0 parse failures Out of 4,800 judge calls, both rounds

Every response across both phases and both rounds went to the same gpt-4o judge used throughout this series, built in from the start. It never sees which condition, phase, or round produced a response, only the response text, the target brand, and the competitor. All 4,800 calls parsed cleanly, hitting the same 0-failure target every prior study in this series has held to.

14 of 16 cells Near ceiling or floor, identical set both rounds

Phase 2 has 4 brands and 4 conditions, 16 brand by condition cells. 14 of them landed above 95% or below 5% winner rate in both rounds, the exact same 14 both times, mostly wherever rating alone already decided the outcome. Only the two barbaro-mojo cells where rating favored the competitor stayed away from a ceiling or floor, sitting at 43 to 47%, close to a coin flip, which is where the study's clearest look at familiarity's effect actually comes from. That is why the full Phase 2 model, not any single cell, is the primary evidence here (LR=1621.71 round 1, LR=1605.73 round 2, df=3, both p≈0). The underlying logistic models also threw a convergence warning tied to how one-sided rating's effect is, a statistical sign of near-total separation, not a data quality problem, but a reason to read individual coefficients cautiously alongside the headline test.

What this doesn't prove

The claim is synthetic, and one brand's response is unexplained

The familiarity claim itself is synthetic but disclosed as such, a single added sentence per brand asserting broad recognition with no name attached, the same single-fact-injection design this series has used since Cold Start and Hidden Context. It is not a scraped or verified real popularity signal, and it deliberately avoids naming any third party (that is Authority) or changing who a claim is attributed to (that is claim-attribution). Phase 1's 79.8% describes familiarity when it is the only signal on the page, next to a bare, word-count matched competitor claim, that is a different, narrower question than Phase 2's, which asks what familiarity does once a rating is already present and doing most of the deciding. Neither number should be read as an estimate of the other. Only 4 brands were tested, and one of them, Barbaro Mojo, responded to familiarity at barely above chance (53.5%, identical in both rounds) while the other three ranged from 81.8% to 96.8%, a spread this brand set cannot explain since it is not a clean trust versus functional split. Single-turn, gpt-4o only, 5 repeats per cell, the same repeats-count power caveat flagged in this series' recent studies. Price, semantic positioning, and structured information remain untested factors on the how-ai-decides Winner vs Loser board.

Update: familiarity's own solo number (79.8%) is used as a reference point in Study #37, Two Signals Get You Most of the Way There. A Third Barely Helps, and Format Doesn't Move It at All, which puts familiarity head to head with authority, specificity, and format, with rating removed entirely.

A popularity claim can win the comparison, if nothing stronger is in the room

Free AI Commerce Score™ in 10 seconds.

Alone, a bare claim of being widely known decided 80% of comparisons. Next to a rating, it added next to nothing. It's worth knowing where your own store's rating and review count already stand before you spend effort on a broad brand-awareness claim that a strong enough rating would make almost irrelevant.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study is the direct sibling to Authority, isolating a plainer claim of familiarity the same way Authority isolated a specific third-party mention, then crossing it against rating the same way Winner vs Loser crossed its own three signals.