Recommendation Intelligence Research™ · Study #33

Give it the better rating, and it wins 91% of the time. Two signals we thought would matter next, barely do.

Every study in this series so far changed one thing at a time while the competitor sat there as a bare name with no information at all. This one puts three validated signals in the same room at once, rating, claim specificity, and message format, and for the first time gives the competitor real information too, its own rating, its own plain description. 4 brands, 8 conditions, 2 independent rounds, 6,400 calls. Rating decided the winner 90.8% of the time, almost identically in both rounds, regardless of what the other two factors said. Specificity still added a smaller, real boost of its own. Format did not hold up. Neither did the category pattern this series found last time.

ZenodoCite this study: 10.5281/zenodo.22777385
90.8%Follows the stronger rating
+5.8 ptsSpecificity's own gain
6,400Total calls, 2 rounds
8Conditions per brand, 3 factors at once
Where this comes from

Two studies found two strong signals, one at a time. This tests what happens once they compete.

Candidate Evaluation gave two real brands a rating and review count and found it decided the winner 100% of the time, no exceptions, when it was the only signal in the room. PDP Specificity found that rewriting the same real facts as concrete, specific claims instead of vague marketing language added 15.7 points in functional categories and cost close to 5 in trust sensitive ones, again with the competitor as a bare name. Both are real, replicated findings. Neither tells you what happens when a brand's rating, its claim, and how that claim is formatted are all in play at the same time, against a competitor that also has real information instead of nothing at all.

This is also the first study in the series where the competitor is not a blank slate. Both brands get a rating and review count, and both get a real description, so any advantage the target brand shows has to survive a competitor that is no longer empty-handed.

1
Rating
Who has the stronger rating and review count, target or competitor. Values are synthetic but realistic, held fixed per brand, not scraped from a real listing.
2
Specificity
The target's own claim, vague or specific, reusing the exact same claim text as pdp-specificity. The competitor's claim is a plain, generic one-liner, held fixed either way, so this stays a clean factor on the target alone.
3
Format
The same facts for both brands, shown either as one flowing paragraph per brand or as a short bulleted spec list per brand. The one signal in this trio never tested before.
2 x 2 x 2, full factorial, 8 conditions per brand. 4 brands × 8 conditions × 20 purchase intents × 5 repeats = 3,200 calls per round. Round 2 is not optional: an independent seed, a full fresh run, and only effects that hold both direction and significance in both rounds separately get called a finding on this page. Anything that only shows up when the two rounds are pooled together is reported as unconfirmed, not as a result.
Model: gpt-4o throughout
Brands: Colored Organics, BodyArtForms, Barbaro Mojo, Hearthloom, same 4 brands as pdp-specificity round 1
Design: 4 brands × 8 conditions × 20 purchase intents × 5 repeats, 2 independent rounds
Total calls: 6,400 (3,200 per round)
Winner determination: gpt-4o LLM judge, built in from the first run
Statistics: cluster-bootstrap confidence intervals over intents, per-round logistic regression for every factor and interaction, one likelihood-ratio test on the full model as primary evidence
Forced two-way pick, brand order randomized per call. "Between {brand A} and {brand B}, which is the better choice for {intent}? Name one and give a one-sentence reason." Which brand is named first is randomized and recorded on every call, to check for position bias separately from the three factors actually being tested.
The actual conditions

Both brands, real information, two formats

Here is the real system message for one brand, Barbaro Mojo (Cuban style hot sauce, functional category), in the condition where the target has the stronger rating and the specific claim, shown once as prose and once as structured.

Barbaro Mojo vs. Gindo's, target has the stronger rating, specific claim
ProseOne paragraph per brand
"Barbaro Mojo: Founded in 2020 in Miami by father and son Mario and Kevin Cruz, starting as homemade Christmas gifts. Made with sour orange, garlic, oregano, and cumin. Every bottle is gluten-free and vegan, with no gums, thickeners, or high-fructose corn syrup. Free shipping on orders over $32. Rated 4.9★ (3,400 reviews). Gindo's: Gindo's makes a variety of hot sauces and spicy condiments. Rated 4.3★ (410 reviews)."
StructuredBulleted spec list per brand
"Barbaro Mojo - Founded in 2020 in Miami by father and son Mario and Kevin Cruz... free shipping on orders over $32. - Rating: 4.9★ (3,400 reviews) Gindo's - Gindo's makes a variety of hot sauces and spicy condiments. - Rating: 4.3★ (410 reviews)"
Shared user prompt, order randomized per call: "Between Barbaro Mojo and Gindo's, which is the better choice for Cuban mojo marinade? Name one and give a one-sentence reason." Each brand ran all 20 of its own purchase intents, 5 times each, under all 8 conditions, in both rounds.
The finding

Rating wins the room, almost identically in both rounds

Numbers below pool both rounds, 3,200 calls per rating level. When the target brand has the stronger rating, it wins 99.2% of the time. When the competitor does, the target wins only 17.6% of the time (95% cluster-bootstrap CI 11.3% to 24.5%, resampled over purchase intents). That gap held almost exactly in both rounds separately: 99.2% vs. 17.2% in round 1, 99.2% vs. 18.0% in round 2. Across all 6,400 calls, the winner matches whichever brand has the stronger rating 90.8% of the time, regardless of what specificity or format said.

Winner rate by rating level, both rounds combined
n=3,200 per bar · gpt-4o LLM judge, built in from the start
Target has stronger rating
99.2%
Both rounds: 99.2%
Competitor has stronger rating
17.6%
17.2% → 18.0%
This is stronger evidence than Candidate Evaluation's 100%, not weaker. Candidate Evaluation tested rating alone, with the competitor as a bare name. Here, the competitor also has a real rating, a real claim, and a real format working for it, and rating still decides the winner in the overwhelming majority of calls. The 8.8 points of daylight between 100% and 99.2%, and the fact that the loser still wins 17.6% of the time when its own rating is behind, is where the other two factors are actually doing something, just not nearly as much as expected going in.
Making sure it holds

All 3 main effects, round 1 vs. round 2, side by side

Round 2 reran the entire 3,200 call design from scratch, an independent random seed, not a re-run of round 1's calls. Rating and specificity replicated cleanly. Format's direction repeated but never cleared significance in either round on its own.

FactorLevelRound 1Round 2Combined
Rating Target stronger 99.2% 99.2%p<0.0001 both rounds 99.2%
Competitor stronger 17.2% 18.0% 17.6%
Specificity Vague 55.2% 55.8% 55.5%
Specific 61.2% 61.5%p<0.001 both rounds 61.3%
Format Prose 59.2% 59.9% 59.5%
Structured 57.2% 57.4%p=0.25, p=0.15, neither significant 57.3%
Rating and specificity both clear this study's own bar. Format doesn't, yet. Rating (p<0.0001 in both rounds independently) and specificity (p=0.0006 round 1, p=0.001 round 2) hold direction and significance separately in both rounds, the pre-registered standard for calling something a finding. Format moved the same small direction both times, structured slightly behind prose, but was not significant alone in either round (p=0.25 and p=0.15). Pooling both rounds together nudges it toward significance (p=0.068 raw, p=0.014 controlling for word count), but that is exactly the kind of number this study's own rule says not to trust. H3 is reported as unconfirmed, not as a finding.
What didn't replicate · category

PDP Specificity's trust penalty doesn't show up here

PDP Specificity's headline finding was a sharp category interaction: specificity helped functional brands and hurt trust brands, one brand even reversed outright. This study reuses those same 4 brands, so it is a direct test of whether that pattern holds once the competitor also has a real rating on the table. It mostly doesn't.

BrandCategoryVagueSpecificChange
Barbaro Mojo Functional 71.5% 79.0% +7.5pp
Hearthloom Functional 50.2% 58.5% +8.3pp
Colored Organics Trust 50.0% 54.8% +4.8pp
BodyArtForms Trust 50.1% 53.1% +3.0pp
No reversal this time. Colored Organics went up, not down. In pdp-specificity, Colored Organics moved backward when its claim got more specific, 93.5% to 85.0%, p=0.006. Here, with a real competitor rating in the picture, it moved up instead, 50.0% to 54.8%. Functional brands still gained more than trust brands, +7.9 points pooled vs. +3.8, but a likelihood-ratio test for the category by specificity interaction, controlling for brand, clears significance pooled (χ²=4.02, df=1, p=0.045) only because pooling adds power. Tested the way this study's own rule requires, round by round, it does not replicate: χ²=3.45, p=0.063 in round 1, χ²=0.96, p=0.328 in round 2. One honest reading: once a competitor also carries real information, especially a rating, the trust penalty that showed up when the competitor was a blank slate has less room to operate.
What didn't replicate · format

Structured lists did not clear their own bar either

H3 predicted structured, bulleted facts would produce a small lift over prose, the same facts read more efficiently. The direction went the other way in both rounds, prose slightly ahead, but never by enough to call it real on its own terms.

The word count check actually strengthens the case for caution, not against it. Structured messages ran longer on average (58.9 words vs. 54.9 for prose, identical in both rounds since the underlying facts are fixed). Controlling for that word count difference in a logistic model makes format's own negative coefficient larger, not smaller (-0.093 uncontrolled, -0.128 controlling for word count), meaning length is not explaining away the small gap. But neither round alone reaches significance, so the honest read is that format's effect, if it exists at all, is too small for this sample size to confirm, not that structured formatting hurts. This is exactly the kind of question the statistical-power audit now underway across this series was designed to catch: some studies in this series ran 5 repeats per cell, a number that may simply be too few to resolve an effect this size, if it is real at all.
Supporting evidence

A judge that never missed, and a caveat worth naming

0 parse failures Out of 6,400 judge calls, both rounds

Every response across both rounds went to a real gpt-4o judge, built in from the start. It never sees which condition or round produced a response, only the response text, the target brand, and the competitor. All 6,400 calls parsed cleanly, hitting the same 0-failure target every prior study in this series has held to.

21 of 32 cells Near ceiling or floor in both rounds

With 4 brands and 8 conditions, that is 32 brand by condition cells. 21 of them landed above 95% or below 5% winner rate in both rounds, mostly wherever the target's rating already decided the outcome. That leaves comparatively little room in most individual cells to see the smaller specificity and format effects clearly, which is exactly the power limitation the study design flagged in advance for the 3-way interaction, and part of why the full-model likelihood-ratio test, not any single cell, is the primary evidence here (χ²=5,572.16, df=7, p<0.0001 pooled, essentially driven by rating). The underlying logistic models also threw a convergence warning tied to how one-sided the rating effect is, a statistical sign of near-total separation, not a data quality problem, but a reason to read individual coefficients cautiously alongside the headline test.

What this doesn't prove

3 of a longer list, and 2 hypotheses that didn't survive

Rating and review counts here are synthetic but realistic, held fixed per brand, not scraped from a real listing, the same disclosed limitation as Candidate Evaluation's injected facts. This study tested 3 factors out of a longer list on the how-ai-decides Winner vs Loser board (price, authority, brand familiarity, and semantic positioning were left out to keep this a full factorial instead of a much larger fractional one), and a follow-up round could add one of those as a 4th factor. The category by specificity interaction and the format main effect are reported here as unconfirmed, not as proven-null. This sample size and this brand set did not detect them reliably in both rounds separately, which is a different claim than proving they don't exist, especially once 21 of 32 cells were already sitting near a ceiling or floor. Single-turn, gpt-4o only, 4 brands split 2 and 2 on category type, same acknowledged limitations as every study in this series.

Update: authority and brand familiarity were picked up as follow-up factors. Continued in Study #34, Mention It Was Featured Somewhere, and It Wins 85% of the Time. Bring In a Rating, and It Nearly Disappears., and Study #35, Claim the Brand Is Widely Known, and It Wins 80% of the Time. Bring In a Rating, and the Edge Is Nearly Gone. Study #37, Two Signals Get You Most of the Way There. A Third Barely Helps, and Format Doesn't Move It at All, puts all four non-rating signals from this series head to head at once, with rating removed entirely. Study #38, We Turned the Same Facts Into Bullets. The Model Cited Fewer of Them., tests format a third time, on citation fidelity instead of selection, and finds structured markup performs measurably worse than plain prose.

Know which signal you're actually competing on

Free AI Commerce Score™ in 10 seconds.

When three signals compete at once, one of them decided 9 out of 10 comparisons. It's worth knowing whether your own rating and review count are strong enough to be doing that work for you, or against you, before you spend more effort polishing the claim underneath it.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study looks at what happens when three previously validated signals compete inside one comparison. It sits alongside the two studies that established each signal on its own, and its own limitations section named the direct follow-up below.