Recommendation Intelligence Research™ · Study #37

Two signals get you most of the way there. A third barely helps, and format doesn't move it at all.

Authority, familiarity, and specificity had each been measured alone or crossed against rating specifically. None had ever been put head to head against each other, or stacked in combination, with rating removed from the room entirely. This study runs 12 of the 16 possible combinations of authority, familiarity, specificity, and format, isolating the question of how these four non-rating signals actually rank and combine. Specificity comes out strongest, authority and familiarity trail close behind and are near-indistinguishable from each other, and format's own marginal effect rounds to zero, replicated across 2 independent rounds.

ZenodoCite this study: 10.5281/zenodo.22819610
75.5%Baseline, zero signals on
99.9%Full stack, all four signals
+3.25ppSpecificity's own marginal lift, strongest of the four
9,600Total calls, 2 rounds
Where this comes from

Four signals, measured one at a time. Never against each other, until now.

Winner vs Loser stacked rating, specificity, and format together and found rating so dominant it swamped the other two. Authority Signal and Brand Familiarity each isolated one third-party-style signal alone, then crossed it against rating specifically, and found the same pattern: strong alone, nearly invisible once rating entered the room. PDP Specificity found a signal that flips sign depending on category type. But no study in this series had ever put authority, familiarity, specificity, and format against each other directly, with rating removed from the comparison entirely.

This study fills that gap. Every signal manipulation is on the target side only, the competitor always gets the same fixed, plain claim. 12 of the 16 possible on/off combinations are run, the 4 single-signal-alone cells are skipped on purpose, reusing the already-published solo numbers from Authority Signal (85.2%), Brand Familiarity (79.8%), and PDP Specificity as reference points instead.

1
Baseline (0000)
Nothing on. Target gets only its plain, generic one-line description, prose format, no facts added. The zero-signal reference point.
2
6 pairs (AF, AS, AM, FS, FM, SM)
Exactly two of the four signals on. Tells us how each pair behaves together, reinforcing, redundant, or interfering.
3
4 leave-one-out triples (drop-A, drop-F, drop-S, drop-M)
Three signals on, one held back. Compared against the full stack, isolates each signal's own marginal contribution inside a realistic, information-dense listing.
4
Full stack (AFSM)
All four signals on at once. The ceiling.
4,800 calls per round, 9,600 combined. 4 brands × 12 conditions × 20 purchase intents × 5 repeats = 4,800 calls per round. Round 2 reran the entire design from scratch on an independent seed, not a re-run of round 1's calls. Only effects that hold direction and significance in both rounds separately get called a finding on this page.
Model: gpt-4o throughout
Brands: Colored Organics, BodyArtForms, Barbaro Mojo, Hearthloom, same 4 brands as the rest of this series
Design: 12 conditions, 20 purchase intents × 5 repeats, 2 independent rounds
Total calls: 9,600 (4,800 per round)
Winner determination: gpt-4o LLM judge, same judge prompt reused unchanged since Winner vs Loser
Statistics: likelihood-ratio test, main effects plus all 2-way interactions (the model this 12-of-16-cell design can actually identify) against an intercept-only null, per round
Forced two-way pick, brand order randomized per call. "Between {brand A} and {brand B}, which is the better choice for {intent}? Name one and give a one-sentence reason." Which brand is named first is randomized and recorded on every call, to check for position bias separately from the factors actually being tested.
The actual conditions

Same brand, zero signals vs. all four

Here is the real system message for Hearthloom (handmade ceramic dinnerware, functional category) at the two extremes of the design: baseline, where neither brand has anything but a plain one-line description, and the full stack, where the target carries all four signals, bulleted, against the same plain competitor.

Hearthloom vs. Kilnmere, baseline vs. full stack
Baseline (0000)Nothing on, prose
Hearthloom: Hearthloom makes handmade ceramic dinnerware, including plates, bowls, and mugs. Kilnmere: Kilnmere makes ceramic dinnerware including plates, bowls, and mugs.
Full stack (AFSM)All four on, structured
Hearthloom - Founded by two sisters, hand-thrown in small batches with lead-free, food-safe glazes. Dishwasher and microwave safe despite being handmade. Packaging is 100 percent recyclable. - It has been featured in Apartment Therapy's guide to the best handmade dinnerware brands. - It is one of the most recognized and widely known handmade dinnerware brands among home cooks. Kilnmere - Kilnmere makes ceramic dinnerware including plates, bowls, and mugs.
Shared user prompt, order randomized per call: "Between Hearthloom and Kilnmere, which is the better choice for handmade ceramic dinner plates? Name one and give a one-sentence reason." Each brand ran all 20 of its own purchase intents, 5 times each, under all 12 conditions, in both rounds.
The finding · ranking

Specificity leads. Authority and familiarity trail close behind, format is barely there.

Numbers below pool both rounds, 800 calls per condition (400 per round). Any pair or triple that includes specificity sits at or near the top. The baseline sits well below everything else, and the two pairs that combine format with only one other signal, and never specificity, sit at the bottom of the non-baseline conditions.

Winner rate by condition, both rounds combined, high to low
n=800 per bar (400 per round) · gpt-4o LLM judge, same judge reused unchanged from Winner vs Loser
AS
100.0%
authority + specificity
drop-M
100.0%
A+F+S, no format
AFSM
99.9%
full stack
FS
99.65%
familiarity + specificity
drop-F
99.0%
A+S+M, no familiarity
drop-A
98.0%
F+S+M, no authority
AF
96.7%
authority + familiarity
drop-S
96.65%
A+F+M, no specificity
SM
94.85%
specificity + format
FM
90.65%
familiarity + format, no specificity
AM
87.9%
authority + format, lowest pair
0000
75.5%
baseline, nothing on
The pattern holds in both rounds separately, not just combined. AS and drop-M tie for first in round 1 and again in round 2. AM is the lowest pair in both rounds (89.0% round 1, 86.8% round 2). Every condition containing specificity outranks every condition that combines format with only one other signal. The baseline sits 11 to 24 points below every other condition in both rounds, a real gap, not noise.
The finding · saturation

Once two signals are on, a third or fourth barely adds anything

The full stack (AFSM, 99.9%) sits only a few points above several two-signal pairs. Marginal contribution below is each drop-X condition subtracted from the full stack, how much each signal is worth once the other three are already present.

Marginal contribution to the full stack, both rounds combined
Full stack (AFSM) minus each leave-one-out condition · positive = that signal still helps once the other three are present
Specificity
+3.25pp
strongest, both rounds
Authority
+1.9pp
2.0pp round 1, 1.8pp round 2
Familiarity
+0.9pp
0.8pp round 1, 1.0pp round 2
Format
−0.1pp
-0.2pp then 0.0pp, both round to zero
Saturation, not addition. Specificity's solo main effect (on vs. off, pooled across all 12 conditions) is the largest of the four at +9.3 points (98.75% on vs. 89.45% off). Authority and familiarity are close behind each other, roughly +5.5 to +5.9 points, and swap which one is slightly ahead depending on whether you look at solo main effect or marginal contribution to the full stack, close enough that this study can't call a strict order between them. Format's on/off gap is the smallest by a wide margin, under 1 point (95.25% vs. 94.4%), and its marginal contribution to an already-strong stack rounds to zero in both rounds separately. Getting to two signals does most of the work. A third or fourth adds only a few more points, and which third or fourth matters far less than which two you started with.
Making sure it holds

Round 1 vs. round 2, side by side

Round 2 reran the entire 4,800-call design from scratch, an independent random seed, not a re-run of round 1's calls. The ranking is nearly identical condition by condition. The marginal-contribution ranking (specificity strongest, format weakest) holds in both rounds separately. Format's marginal contribution technically flips sign, -0.2pp to 0.0pp, but both numbers are within noise of zero, not a real reversal.

ConditionRound 1Round 2Combined
0000 (baseline)76.5%74.5%75.5%
AS100.0%100.0%tied for 1st, both rounds100.0%
drop-M100.0%100.0%tied for 1st, both rounds100.0%
AFSM (full stack)99.8%100.0%99.9%
AM89.0%86.8%lowest pair, both rounds87.9%
Marginal · Specificity+3.3pp+3.2pp+3.25pp
Marginal · Authority+2.0pp+1.8pp+1.9pp
Marginal · Familiarity+0.8pp+1.0pp+0.9pp
Marginal · Format-0.2pp0.0ppboth round to zero-0.1pp
Position bias, target named first95.0%94.8%94.9%
Position bias, target named second94.7%95.1%94.9%
0 judge parse failures across 9,600 calls, both rounds. No position bias worth naming, first vs. second named brand sits within half a point of each other in both rounds. The 12-condition ranking and the 4-signal marginal-contribution ranking both replicate cleanly. Only effect that does not survive as a real, distinguishable-from-zero finding: format's own marginal contribution, which is the correct, honest result for a signal that was never expected to move the needle much on its own.
Supporting evidence

A model spec that needed fixing, and a ceiling that compresses the range

LR=424.37 Round 2, p=6.07e-85, converged cleanly

The originally specified full 4-way interaction model (16 parameters) failed to converge on round 1's data: a saturated model that size isn't fully identified by a design that only runs 12 of 16 possible signal combinations, and two conditions landed at exactly 100% (0/400 losses), textbook quasi-complete separation. Fixed by reducing to main effects plus all 2-way interactions (11 parameters, matching what a 12-condition design can actually estimate), the same level the rest of this page's analysis already reports on. Round 1 still needed an L2-penalized fallback to get a stable fit (LR=371.75, p=3.29e-77). Round 2, run under the corrected model from the start, converged with plain MLE, no fallback needed (LR=424.37, df=10, p=6.07e-85). Both rounds land nowhere near the significance boundary either way.

38 of 48 cells Near ceiling or floor, both rounds

4 brands × 12 conditions = 48 brand-by-condition cells. 38 of them landed above 95% or below 5% winner rate in both rounds, almost the identical set both times. The baseline itself sits unusually high (75.5% combined, vs. a near-50% split that would be expected of two matched plain claims), a real structural feature of this design worth naming as a limitation, not a finding: it compresses how much room the remaining conditions have to spread out, and is likely why so many cells sit near a ceiling once even one or two signals are added on top of that already-elevated baseline.

What this doesn't prove

Four brands, a high baseline, and an incomplete factorial by design

All signal facts are synthetic (disclosed-synthetic authority mention, synthetic familiarity claim), the same convention as the rest of this series, not scraped from real press or real market-recognition data. Only 4 brands were tested, 2 trust-type and 2 functional-type, not a balanced design for testing whether category moderates any of this, that is a separate follow-up, not attempted here. The 12-condition subset deliberately skips the 4 single-signal-alone cells, reusing prior studies' published numbers as reference instead, so cross-study comparisons to those numbers are directional, not a strict apples-to-apples replication, their baseline conditions were not identical to this study's. The baseline itself (75.5%) sits well above a coin flip, a real structural asymmetry in this design worth flagging clearly, not explaining away. Single-turn, gpt-4o only, 5 repeats per cell, the same repeats-count power caveat flagged throughout this series. Price, brand fame beyond the familiarity claim tested here, and real (not synthetic) third-party mentions remain untested factors on the how-ai-decides board.

Update: format's marginal contribution to winning a forced comparison rounds to zero here, but that only tests selection. Continued in Study #38, We Turned the Same Facts Into Bullets. The Model Cited Fewer of Them., which shifts the outcome to citation fidelity and vocabulary reuse instead, and finds structured markup performs measurably worse than plain prose carrying the identical facts.

Two signals get you most of the way there. Which two matters more than how many

Free AI Commerce Score™ in 10 seconds.

Specificity paired with almost anything outranked every other pairing tested. Stacking a third or fourth signal on top barely moved the number further. Worth knowing whether your own store's listings are missing that first specific, concrete fact before spending effort on a fourth or fifth signal that won't add much once the first two are in place.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study pulls together reference numbers from three studies that each isolated one of these signals alone, and sits alongside the study that first found rating dominant enough to swamp everything else.