Authority, familiarity, and specificity had each been measured alone or crossed against rating specifically. None had ever been put head to head against each other, or stacked in combination, with rating removed from the room entirely. This study runs 12 of the 16 possible combinations of authority, familiarity, specificity, and format, isolating the question of how these four non-rating signals actually rank and combine. Specificity comes out strongest, authority and familiarity trail close behind and are near-indistinguishable from each other, and format's own marginal effect rounds to zero, replicated across 2 independent rounds.
Winner vs Loser stacked rating, specificity, and format together and found rating so dominant it swamped the other two. Authority Signal and Brand Familiarity each isolated one third-party-style signal alone, then crossed it against rating specifically, and found the same pattern: strong alone, nearly invisible once rating entered the room. PDP Specificity found a signal that flips sign depending on category type. But no study in this series had ever put authority, familiarity, specificity, and format against each other directly, with rating removed from the comparison entirely.
This study fills that gap. Every signal manipulation is on the target side only, the competitor always gets the same fixed, plain claim. 12 of the 16 possible on/off combinations are run, the 4 single-signal-alone cells are skipped on purpose, reusing the already-published solo numbers from Authority Signal (85.2%), Brand Familiarity (79.8%), and PDP Specificity as reference points instead.
Here is the real system message for Hearthloom (handmade ceramic dinnerware, functional category) at the two extremes of the design: baseline, where neither brand has anything but a plain one-line description, and the full stack, where the target carries all four signals, bulleted, against the same plain competitor.
Numbers below pool both rounds, 800 calls per condition (400 per round). Any pair or triple that includes specificity sits at or near the top. The baseline sits well below everything else, and the two pairs that combine format with only one other signal, and never specificity, sit at the bottom of the non-baseline conditions.
The full stack (AFSM, 99.9%) sits only a few points above several two-signal pairs. Marginal contribution below is each drop-X condition subtracted from the full stack, how much each signal is worth once the other three are already present.
Round 2 reran the entire 4,800-call design from scratch, an independent random seed, not a re-run of round 1's calls. The ranking is nearly identical condition by condition. The marginal-contribution ranking (specificity strongest, format weakest) holds in both rounds separately. Format's marginal contribution technically flips sign, -0.2pp to 0.0pp, but both numbers are within noise of zero, not a real reversal.
| Condition | Round 1 | Round 2 | Combined |
|---|---|---|---|
| 0000 (baseline) | 76.5% | 74.5% | 75.5% |
| AS | 100.0% | 100.0%tied for 1st, both rounds | 100.0% |
| drop-M | 100.0% | 100.0%tied for 1st, both rounds | 100.0% |
| AFSM (full stack) | 99.8% | 100.0% | 99.9% |
| AM | 89.0% | 86.8%lowest pair, both rounds | 87.9% |
| Marginal · Specificity | +3.3pp | +3.2pp | +3.25pp |
| Marginal · Authority | +2.0pp | +1.8pp | +1.9pp |
| Marginal · Familiarity | +0.8pp | +1.0pp | +0.9pp |
| Marginal · Format | -0.2pp | 0.0ppboth round to zero | -0.1pp |
| Position bias, target named first | 95.0% | 94.8% | 94.9% |
| Position bias, target named second | 94.7% | 95.1% | 94.9% |
The originally specified full 4-way interaction model (16 parameters) failed to converge on round 1's data: a saturated model that size isn't fully identified by a design that only runs 12 of 16 possible signal combinations, and two conditions landed at exactly 100% (0/400 losses), textbook quasi-complete separation. Fixed by reducing to main effects plus all 2-way interactions (11 parameters, matching what a 12-condition design can actually estimate), the same level the rest of this page's analysis already reports on. Round 1 still needed an L2-penalized fallback to get a stable fit (LR=371.75, p=3.29e-77). Round 2, run under the corrected model from the start, converged with plain MLE, no fallback needed (LR=424.37, df=10, p=6.07e-85). Both rounds land nowhere near the significance boundary either way.
4 brands × 12 conditions = 48 brand-by-condition cells. 38 of them landed above 95% or below 5% winner rate in both rounds, almost the identical set both times. The baseline itself sits unusually high (75.5% combined, vs. a near-50% split that would be expected of two matched plain claims), a real structural feature of this design worth naming as a limitation, not a finding: it compresses how much room the remaining conditions have to spread out, and is likely why so many cells sit near a ceiling once even one or two signals are added on top of that already-elevated baseline.
All signal facts are synthetic (disclosed-synthetic authority mention, synthetic familiarity claim), the same convention as the rest of this series, not scraped from real press or real market-recognition data. Only 4 brands were tested, 2 trust-type and 2 functional-type, not a balanced design for testing whether category moderates any of this, that is a separate follow-up, not attempted here. The 12-condition subset deliberately skips the 4 single-signal-alone cells, reusing prior studies' published numbers as reference instead, so cross-study comparisons to those numbers are directional, not a strict apples-to-apples replication, their baseline conditions were not identical to this study's. The baseline itself (75.5%) sits well above a coin flip, a real structural asymmetry in this design worth flagging clearly, not explaining away. Single-turn, gpt-4o only, 5 repeats per cell, the same repeats-count power caveat flagged throughout this series. Price, brand fame beyond the familiarity claim tested here, and real (not synthetic) third-party mentions remain untested factors on the how-ai-decides board.
Specificity paired with almost anything outranked every other pairing tested. Stacking a third or fourth signal on top barely moved the number further. Worth knowing whether your own store's listings are missing that first specific, concrete fact before spending effort on a fourth or fifth signal that won't add much once the first two are in place.
This study pulls together reference numbers from three studies that each isolated one of these signals alone, and sits alongside the study that first found rating dominant enough to swamp everything else.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →