Recommendation Intelligence Research™ · Study #39

Authority adds 9 points in trust categories. Familiarity adds 14.

PDP Specificity found that specific claims help functional categories and cost trust categories, the first time this series split a signal's own effect by category. Authority and Brand Familiarity were each measured earlier, but neither ever split its own result by category, both reported one pooled number. This study reruns both signals on the same 8 brands, 4 functional and 4 trust and safety sensitive, no rating in the room, and asks whether the same category split changes anything for social proof the way it did for specificity. It moves the opposite direction: authority and familiarity both move the recommended brand more in trust categories than in functional ones, not less.

ZenodoCite this study: 10.5281/zenodo.22832667
74%Trust follow rate, both signals
+9.1 ptsAuthority gap, trust vs. functional
+13.8 ptsFamiliarity gap, trust vs. functional
6,400Calls, 2 independent rounds
Where this comes from

Specificity's effect flipped by category. Did social proof's?

PDP Specificity was the only study in this series to test whether a signal's own effect changes by category, and it does: specific claims added 15.7 points in functional categories and cost nearly 5 points in trust categories. Authority Signal and Brand Familiarity were each tested on the same kind of category split conceptually, real vs. vague facts about a brand, but on their original 4 brands, 2 functional and 2 trust, they never actually reported the two category types separately. Both reported one pooled follow-the-signal number and stopped there.

This study closes that gap directly. Same no-rating-in-the-room design as the rest of the series, same forced two-way choice, same 20 purchase intents and 5 repeats per condition, but built from the start to report functional and trust separately, on 8 brands instead of 4, 2 recycled from the original Authority and Familiarity studies and 2 new ones matched in word count and claim length.

Model: gpt-4o, forced two-way choice
Brands: 8, 4 functional, 4 trust and safety sensitive
Conditions per brand: 4 (authority or familiarity x target has fact or competitor has it)
Design: 20 purchase intents, 5 repeats, 2 independent rounds
Total calls: 6,400 (3,200 per round), plus matched judge calls
Statistics: category_type x signal_level LR interaction test, one for authority, one for familiarity, each run separately in each round
No rating in the room, same as PDP Specificity and Winner vs Loser. A star rating overwhelms every other signal tested in this series, so it is left out here too, isolating what authority and familiarity do on their own, by category.
The 8 brands

Four categories where the decision is mostly about specs. Four where it's mostly about trust.

Same axis PDP Specificity used to define its own categories. Functional: the decision comes down to taste, fit, or whether it does the job. Trust: the decision comes down to what touches a baby's skin, what goes through a piercing, what goes on skin, or what a pet wears.

Functional
Decision is mostly about specs Barbaro Mojo · Cuban-style hot sauce, vs. Gindo's
Hearthloom · handmade ceramic dinnerware, vs. Kilnmere
Bellroy · slim leather wallets, vs. Herschel
Zigpoll · Shopify post-purchase survey tool, vs. SurveyMonkey
Trust and safety sensitive
Decision is mostly about trust Colored Organics · organic baby clothes, vs. Finn + Emma
BodyArtForms · body piercing jewelry, vs. Painful Pleasures
Wild One · dog gear, vs. Ruffwear
Primally Pure · natural deodorant, vs. Native Deodorant

4 brands (Barbaro Mojo, Hearthloom, Colored Organics, BodyArtForms) reuse byte-identical authority and familiarity facts from the original Authority Signal and Brand Familiarity studies, already word-count matched and verified across 5 prior studies. The other 4 (Bellroy, Zigpoll, Wild One, Primally Pure) are new, freshly written and matched to the same 13 to 17 word range as the originals.

The finding

Both signals move the winner more when trust is on the line

Follow-the-signal rate is the share of calls where the model picked whichever brand had the added authority or familiarity sentence that round, both rounds averaged. A flat 50% would mean the sentence made no difference at all.

Follow-the-signal rate by category, both rounds averaged
n=800 per category per signal per round · gpt-4o forced two-way choice, no rating in the room
Authority · trust categories
74.3%
R1 74.5% → R2 74.0%
Authority · functional categories
65.2%
R1 65.4% → R2 64.9%
Familiarity · trust categories
74.0%
R1 74.4% → R2 73.5%
Familiarity · functional categories
60.2%
R1 60.0% → R2 60.4%
Both interaction tests clear significance by a wide margin, in both rounds separately. Authority: category x signal_level LR = 461.93 in round 1, 433.11 in round 2, both p well below 1e-90. Familiarity: LR = 381.87 in round 1, 380.2 in round 2, both p well below 1e-80. Same direction, same rough size, in both independent rounds, which is what this series calls a confirmed finding rather than a round-1 fluke.
Compared to specificity

The exact opposite pattern from PDP Specificity

PDP Specificity found a hard fact about the product (dimensions, ingredients, compatibility) helps functional categories and actively costs trust categories, +15.7 points functional, -4.5 points trust. Authority and familiarity, both a form of social proof rather than a hard fact, move the opposite direction here: both help trust categories more than functional ones. Put together, the two studies suggest category doesn't just change how much a signal matters, it can flip which kind of signal matters more.

PDP Specificity · a hard fact
Functional minus trust +15.7pp functional
-4.5pp trust
Functional wins, trust loses
Authority & Familiarity · social proof
Trust minus functional +9.1pp trust (authority)
+13.8pp trust (familiarity)
Trust wins, functional gains less
Making sure it holds

Round 1 vs. round 2, signal by signal

Round 2 reran the full 3,200-call design on an independent seed, not a replay of round 1's calls. Direction, rough size, and statistical significance all held for both signals, separately.

SignalFunctional (R1 → R2)Trust (R1 → R2)LR interaction (R1 / R2)
Authority 65.4% → 64.9% 74.5% → 74.0% 461.93, p=8.5e-100 / 433.11, p=1.5e-93
Familiarity 60.0% → 60.4% 74.4% → 73.5% 381.87, p=1.9e-82 / 380.2, p=4.3e-82
Position bias checked and small. Across both rounds combined, the target brand won 69.2% of the time when named first (n=1,595) and 67.5% of the time when named second (n=1,605), a 1.7-point gap nowhere near large enough to explain a 9 to 14 point category effect.
Supporting evidence

A clean judge, and real heterogeneity underneath the average

0 parse failures Out of 6,400 judge calls, both rounds combined

Every one of the 6,400 forced two-way choices across both rounds was cleanly parsed into a winner by the LLM judge, no ambiguous or unparseable responses excluded or estimated around. Every number on this page rests on the full dataset, not a subset.

14 of 64 Brand x condition cells flagged ceiling or floor, same cells both rounds

The category-level averages hide real brand-to-brand spread. Two of the four functional brands, Bellroy and Zigpoll, sit close to 50% for both signals in both rounds, essentially no effect at the brand level. The other two, Hearthloom and Barbaro Mojo, show real effects, Hearthloom near 97 to 98%. Trust brands are more consistently elevated: BodyArtForms sits near 95 to 97%, Colored Organics near 81 to 85%, while Wild One and Primally Pure sit lower, around 52 to 68%. The category average is real and replicates, but it is an average over brands that vary a lot on their own.

Why it matters

The signal that helps most depends on what the category is actually about

A single, one-size-fits-all recommendation, "always lead with specifics" or "always lead with social proof", would be wrong for roughly half the categories in this dataset. When the category is mostly a specs decision, PDP Specificity found a concrete claim adds the most and social proof adds less. When the category is mostly a trust decision, this study found the reverse: authority and familiarity add 9 to 14 points more than they do in functional categories, while specificity actively costs points there. Which signal to lead with is itself a function of what kind of decision the buyer is making, not a fixed hierarchy that holds everywhere.

What this doesn't prove

A strong category average, a small and heterogeneous brand set underneath it

8 brands is a small sample to generalize a category-type effect from, even with 800 calls behind each category-type cell. The brand-level spread documented above is real: this is an average tendency across a small, deliberately mixed set of brands, not a claim that every functional brand responds low and every trust brand responds high. The “functional” and “trust” labels are the same 2-way axis PDP Specificity used, a simplification of a much larger space of possible categories, and this study only tested authority and familiarity, not specificity itself, on these 8 brands directly, so the specificity comparison above is between two different studies' brand sets, not a single unified test. Single model (gpt-4o), forced two-way choice only, no rating in the room by design, so this does not speak to what happens once a star rating is present, which the rest of this series has shown tends to dominate other signals when it is.

Know which signal actually fits your category

Free AI Commerce Score™ in 10 seconds.

Authority and familiarity carried more weight in trust categories, specificity carried more weight in functional ones. Worth knowing where your own store stands before deciding which one to lead with.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study closes the category-as-moderator question PDP Specificity opened, this time for social proof instead of a hard fact.