Every study in this series so far changed one thing at a time while the competitor sat there as a bare name with no information at all. This one puts three validated signals in the same room at once, rating, claim specificity, and message format, and for the first time gives the competitor real information too, its own rating, its own plain description. 4 brands, 8 conditions, 2 independent rounds, 6,400 calls. Rating decided the winner 90.8% of the time, almost identically in both rounds, regardless of what the other two factors said. Specificity still added a smaller, real boost of its own. Format did not hold up. Neither did the category pattern this series found last time.
Candidate Evaluation gave two real brands a rating and review count and found it decided the winner 100% of the time, no exceptions, when it was the only signal in the room. PDP Specificity found that rewriting the same real facts as concrete, specific claims instead of vague marketing language added 15.7 points in functional categories and cost close to 5 in trust sensitive ones, again with the competitor as a bare name. Both are real, replicated findings. Neither tells you what happens when a brand's rating, its claim, and how that claim is formatted are all in play at the same time, against a competitor that also has real information instead of nothing at all.
This is also the first study in the series where the competitor is not a blank slate. Both brands get a rating and review count, and both get a real description, so any advantage the target brand shows has to survive a competitor that is no longer empty-handed.
Here is the real system message for one brand, Barbaro Mojo (Cuban style hot sauce, functional category), in the condition where the target has the stronger rating and the specific claim, shown once as prose and once as structured.
Numbers below pool both rounds, 3,200 calls per rating level. When the target brand has the stronger rating, it wins 99.2% of the time. When the competitor does, the target wins only 17.6% of the time (95% cluster-bootstrap CI 11.3% to 24.5%, resampled over purchase intents). That gap held almost exactly in both rounds separately: 99.2% vs. 17.2% in round 1, 99.2% vs. 18.0% in round 2. Across all 6,400 calls, the winner matches whichever brand has the stronger rating 90.8% of the time, regardless of what specificity or format said.
Round 2 reran the entire 3,200 call design from scratch, an independent random seed, not a re-run of round 1's calls. Rating and specificity replicated cleanly. Format's direction repeated but never cleared significance in either round on its own.
| Factor | Level | Round 1 | Round 2 | Combined |
|---|---|---|---|---|
| Rating | Target stronger | 99.2% | 99.2%p<0.0001 both rounds | 99.2% |
| Competitor stronger | 17.2% | 18.0% | 17.6% | |
| Specificity | Vague | 55.2% | 55.8% | 55.5% |
| Specific | 61.2% | 61.5%p<0.001 both rounds | 61.3% | |
| Format | Prose | 59.2% | 59.9% | 59.5% |
| Structured | 57.2% | 57.4%p=0.25, p=0.15, neither significant | 57.3% |
PDP Specificity's headline finding was a sharp category interaction: specificity helped functional brands and hurt trust brands, one brand even reversed outright. This study reuses those same 4 brands, so it is a direct test of whether that pattern holds once the competitor also has a real rating on the table. It mostly doesn't.
| Brand | Category | Vague | Specific | Change |
|---|---|---|---|---|
| Barbaro Mojo | Functional | 71.5% | 79.0% | +7.5pp |
| Hearthloom | Functional | 50.2% | 58.5% | +8.3pp |
| Colored Organics | Trust | 50.0% | 54.8% | +4.8pp |
| BodyArtForms | Trust | 50.1% | 53.1% | +3.0pp |
H3 predicted structured, bulleted facts would produce a small lift over prose, the same facts read more efficiently. The direction went the other way in both rounds, prose slightly ahead, but never by enough to call it real on its own terms.
Every response across both rounds went to a real gpt-4o judge, built in from the start. It never sees which condition or round produced a response, only the response text, the target brand, and the competitor. All 6,400 calls parsed cleanly, hitting the same 0-failure target every prior study in this series has held to.
With 4 brands and 8 conditions, that is 32 brand by condition cells. 21 of them landed above 95% or below 5% winner rate in both rounds, mostly wherever the target's rating already decided the outcome. That leaves comparatively little room in most individual cells to see the smaller specificity and format effects clearly, which is exactly the power limitation the study design flagged in advance for the 3-way interaction, and part of why the full-model likelihood-ratio test, not any single cell, is the primary evidence here (χ²=5,572.16, df=7, p<0.0001 pooled, essentially driven by rating). The underlying logistic models also threw a convergence warning tied to how one-sided the rating effect is, a statistical sign of near-total separation, not a data quality problem, but a reason to read individual coefficients cautiously alongside the headline test.
Rating and review counts here are synthetic but realistic, held fixed per brand, not scraped from a real listing, the same disclosed limitation as Candidate Evaluation's injected facts. This study tested 3 factors out of a longer list on the how-ai-decides Winner vs Loser board (price, authority, brand familiarity, and semantic positioning were left out to keep this a full factorial instead of a much larger fractional one), and a follow-up round could add one of those as a 4th factor. The category by specificity interaction and the format main effect are reported here as unconfirmed, not as proven-null. This sample size and this brand set did not detect them reliably in both rounds separately, which is a different claim than proving they don't exist, especially once 21 of 32 cells were already sitting near a ceiling or floor. Single-turn, gpt-4o only, 4 brands split 2 and 2 on category type, same acknowledged limitations as every study in this series.
When three signals compete at once, one of them decided 9 out of 10 comparisons. It's worth knowing whether your own rating and review count are strong enough to be doing that work for you, or against you, before you spend more effort polishing the claim underneath it.
This study looks at what happens when three previously validated signals compete inside one comparison. It sits alongside the two studies that established each signal on its own, and its own limitations section named the direct follow-up below.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →