Recommendation Intelligence Research™ · Study #31

Say it yourself, and the model picks you 67.5% of the time. A third party saying the same thing only gets 51.1%.

This study started as a reply to a specific claim: that what matters isn't your own marketing, but who else is backing up your claims. So we tested it directly, twice. Same brand, same fact, same length, same model. Only who appears to be saying the fact changes: the brand's own voice, a plain third-person statement with no named source, or an apparent independent source like reviews and buyer discussions. A first pass found the brand's own voice winning the most, then a robustness check found the three wrappers weren't quite the same length, so we rebuilt the wording to close that gap and ran the whole thing again. The result held. Combined across both rounds, the brand's own voice still wins the most, in every single brand.

ZenodoCite this study: 10.5281/zenodo.22758452
67.5%Self-claim winner rate, combined
51.1%Third-party winner rate, combined
3,200Total calls, 2 rounds
4Brands, held in all 4
Where this comes from

A LinkedIn comment, turned into a measurable question

A commenter, J.M., replied to a post about AI search with a specific theory: "Not just on the website but across the ecosystem. Where else and who is backing your claims. Not your marketing but real claims in the structure the LLM looks for." That's a testable claim. It says attribution, who appears to be making a factual statement about a brand, should change whether an AI system picks that brand, separately from what the fact actually says.

Every earlier study in this series changed the underlying facts, the context around them, or the shape of the buyer's question. This one holds the fact completely fixed and only changes who appears to be saying it. The identical sentence gets tested under four different framings.

1
No fact
Control. No claim about the brand at all, just a generic comparison-shopping context. Replicates the basic finding from Cold Start and Fact Injection: does injecting any fact move selection at all.
2
Self-claim
The identical fact, framed as the brand's own marketing voice, first person, as if quoted from its own website.
3
Neutral
The identical fact, plain third-person statement, no source named at all. This is close to what most of the rest of this series has already been doing, and works as the baseline for the other two attribution frames.
4
Third party
The identical fact, framed as something independent reviews and buyer discussions have noted, not something the brand said about itself.
Why this is measured as a forced head-to-head, not open candidacy. Brand Legibility (Study #29) already showed that once a brand is named in context, it gets mentioned almost every time, regardless of framing. Measuring candidacy here would likely hit that same ceiling and tell us nothing about attribution specifically. So this study uses the same forced-comparison mechanism as Cold Start and Hidden Context: one fixed, established competitor per brand, across 20 purchase intents, and winner rate (does the target brand get the actual recommendation) is the primary number, not candidacy.
Model: gpt-4o throughout
Brands: 4, spanning organic baby clothes, body piercing jewelry, Cuban-style hot sauce, and ceramic dinnerware
Design: 4 brands × 4 conditions × 20 purchase intents × 5 repeats, run twice
Total calls: 3,200 (1,600 per round, 2 independent rounds)
Winner determination: gpt-4o LLM judge, built in from the first run, not added after the fact
Statistics: two-proportion z-tests, chi-square tests, logistic regression controlling for brand, round, and message length
The target brand's name never appears in the user's question. The fact, in whichever framing, lives entirely in the system message, the same mechanism used by Hidden Context and Fact Injection: information the model already has going in, not something the shopper just said. The user's question names only the one fixed competitor, as something the shopper is already considering, and asks for a clear pick. Any time the target brand wins, it won purely from what was sitting in context, never from the user handing the brand's name back to the model.
The actual prompts

The exact same fact, four different voices, tested twice

Here is the real text, not a summary, for one of the four brands (Colored Organics, organic baby clothes). Round 1's wrappers turned out to be different lengths, which is exactly what round 2 fixed. Both are shown below, word for word.

Round 1
No factControl
"The user has been comparing options in this category for a few days and wants a clear, confident recommendation."
Self-claim27 words
"On Colored Organics's own website, the brand describes itself this way: \"We use GOTS-certified organic cotton, with dyes that are free of azo compounds and heavy metals.\""
Neutral15 words
"Colored Organics uses GOTS-certified organic cotton, with dyes free of azo compounds and heavy metals."
Third party23 words
"Independent reviews and buyer discussions have noted that Colored Organics uses GOTS-certified organic cotton, with dyes free of azo compounds and heavy metals."
Round 2, wrappers rebuilt to near-identical length
No factControl, unchanged
"The user has been comparing options in this category for a few days and wants a clear, confident recommendation."
Self-claim22 words
"According to the brand's own marketing materials: Colored Organics uses GOTS-certified organic cotton, with dyes free of azo compounds and heavy metals."
Neutral21 words
"According to general publicly available information: Colored Organics uses GOTS-certified organic cotton, with dyes free of azo compounds and heavy metals."
Third party22 words
"According to independent reviews and buyer discussions: Colored Organics uses GOTS-certified organic cotton, with dyes free of azo compounds and heavy metals."
Shared user prompt, unchanged across both rounds and all four conditions, only the purchase intent changes: "What's the best organic baby onesies? I'm already considering Finn + Emma, but I want a clear pick, name one brand and give a one-sentence reason." Finn + Emma is Colored Organics's one fixed competitor for this study. Each brand ran all 20 of its own purchase intents, five times each, under all four conditions, in each round.
The finding

Attribution matters. Just not in the direction the theory predicted

Numbers below are the combined result of two independent rounds, 800 calls per condition, 3,200 calls total (see the next section for why there are two rounds and what changed between them). The order is not what the LinkedIn comment that started this study predicted. Third-party attribution beats a plain neutral statement, which is a real and useful finding on its own. But self-claim, the brand's own voice, beats third-party attribution too. Every pairwise gap here is statistically decisive: self-claim vs. third-party (z=6.67, p<0.0001), third-party vs. neutral (z=8.45, p<0.0001), and neutral vs. no fact (z=16.83, p<0.0001). A chi-square test across all four conditions, pooled, gives χ²=866.67 on 3 degrees of freedom.

Winner rate by attribution condition, combined across both rounds and all 4 brands
n=800 per condition · gpt-4o LLM judge, built in from the start
No fact
0.1%
Baseline
Neutral
30.4%
+30.3pp
Third party
51.1%
+51.0pp
Self-claim
67.5%
Highest
The LinkedIn theory said the opposite would happen. The claim was that independent, third-party-sounding validation should beat a brand's own marketing voice. In this mechanism, on these four brands, it's reversed: self-claim outperforms third-party attribution by 16.4 points, and third-party comfortably outperforms a source-free neutral statement by 20.7 points. Both gaps are statistically significant on their own in every single one of the 4 brands once both rounds are combined (see below).
Making sure it holds

We found a real weakness in round 1. Round 2 was built to break the finding, not confirm it.

Round 1 wrapped the identical fact in three different framings, but the wrapping text itself wasn't the same length. The self-claim wrapper ("On {brand}'s own website, the brand describes itself this way: ...") ran 9 to 12 words longer than the bare fact. The third-party wrapper ran 8 words longer. Neutral had no wrapper at all. That's a real hole: maybe the longer, more concrete-sounding wrapper won just because it was longer and more concrete-sounding, not because of who it claimed was speaking.

So we rebuilt the wrappers to be word-count matched, within a single word of each other ("According to the brand's own marketing materials: ...", "According to general publicly available information: ...", "According to independent reviews and buyer discussions: ..."), all wrapping the exact same third-person fact, and ran the entire 1,600-call design again from scratch, a fresh independent draw, not a re-run of the same prompts.

ConditionRound 1 (original wrappers)Round 2 (length-matched)Combined
No fact 0.2% 0.0% 0.1%
Neutral 23.0% 37.8% 30.4%
Third party 43.2% 59.0% 51.1%
Self-claim 65.2% 69.8%Highest, both rounds 67.5%
The order held. The exact size of the self-claim edge moved. Self-claim still beat third-party in round 2 (69.8% vs. 59.0%, z=3.17, p=0.0015), but the gap shrank from 22.0 points to 10.8 points, and per-brand it went from being clean in essentially every category to being clean in 2 of 4 individually, with the other 2 right-direction but not significant alone at n=100. Word count clearly explained some of round 1's margin. It did not explain the finding itself. A logistic regression on all 3,200 calls combined, controlling for brand, round, and system-message word count, still finds condition highly significant (likelihood-ratio test, χ²=572.93, df=3, p=7.4×10-124), and word count itself is not significant once condition is in the model (p=0.35). We are showing both rounds separately, not just the combined number, because that is the honest version of how this got tested.
The finding · by brand

Combine both rounds, and every brand clears significance on its own

This is the table that makes the finding hold up on its own, brand by brand, not just in the pooled total. With both rounds combined (n=200 per brand per condition, double the power of either round alone), self-claim beats third-party and third-party beats neutral, individually statistically significant, in all four brands.

BrandNo factNeutralThird partySelf-claim
Colored Organics 0% 45% 58% 84%p<0.001
Barbaro Mojo 0% 42% 66% 80%p=0.003
Hearthloom (fictional control) 0% 34% 72% 86%p<0.001
BodyArtForms 0% 0% 8% 21%p<0.001
BodyArtForms sits far below the other three brands across every condition, self-claim included (21% vs. 80 to 86%). Reading the actual responses explains why: body piercing jewelry is a category where gpt-4o already has strong opinions, naming brands like Neometal, Anatometal, and Industrial Strength on its own, unprompted. The other three categories (obscure organic baby clothing, a small Cuban hot sauce maker, and a fictional ceramic dinnerware brand) have far less competing brand knowledge already baked into the model, so an injected fact has more room to move the outcome. The p-values shown are for self-claim vs. third-party specifically, in that brand, on the combined 200-call sample; third-party vs. neutral is also significant in all four (p=0.009 to p<0.001, weakest in Colored Organics, strongest in Hearthloom).
A more cautious read

This wasn't really a two-brand race. Most responses picked neither.

The user prompt was open-ended ("what's the best X"), not a strict forced pick between exactly two names. Across all 3,200 calls in both rounds, the named competitor itself only won 37 times (1.2%). The rest of the non-target outcomes, 1,970 out of 3,200 (61.6%), went to some other real brand gpt-4o already knows, like Burt's Bees Baby for baby clothes or Neometal for piercing jewelry, brands that were never named in any prompt at all. So "winner rate" here is best read as: does the injected context, whichever framing it uses, pull the model's answer away from its own strong incumbent priors and toward the specific brand being described. That's still a meaningful, well-defined thing to measure, and the no-fact baseline (0.1%) shows how rarely it happens without any injected context at all. But it's a different claim than "brand A beats brand B head-to-head," and the page should be read that way.

Why it matters

Getting your own facts right beats chasing outside mentions

A lot of current AI-SEO advice, including the comment that started this study, pushes toward getting mentioned elsewhere: reviews, press, forums, third-party sites, on the theory that an AI system trusts an outside voice more than a brand's own words. Third-party attribution genuinely does help here, roughly a 20-point lift over a source-free statement, replicated across both rounds. But in this mechanism, on these four brands, tested twice with two different wordings, a brand's own clearly stated, specific description outperformed that same fact dressed up as independent proof. The practical read isn't "ignore third-party mentions." It's that getting a clear, specific, well-written statement of your own facts into the places a model actually draws from may matter more than the current advice assumes, and shouldn't be skipped in favor of chasing outside citations alone.

Supporting evidence

A real judge, and a confound we went looking for ourselves

0 parse failures Out of 3,200 judge calls, both rounds

Every response in both rounds went to a real gpt-4o judge, built in from the start per the Study #30 lesson. It never sees which condition or round produced a response, only the text, the target brand, and the competitor, so it cannot be biased toward any framing. All 3,200 calls parsed cleanly, and spot-checks confirmed it separates a real recommendation from a passing mention.

p=0.35 Word count's own effect, once condition is controlled

We looked for a way this finding could be wrong. Round 1 alone showed word count was itself predictive, exactly the confound round 2 was built to remove. Controlling for brand, round, and word count across all 3,200 calls, condition stays overwhelmingly significant (p=7.4×10-124) while word count's own effect is not (p=0.35). Length was worth checking. It isn't what's driving this.

What this doesn't prove

One model, one mechanism, honestly narrowed

This study used one model, gpt-4o, and one system-message injection mechanism, across two independent rounds. It has not been tested on other models or on live web search. Round 2's wrappers are still not a perfect linguistic match. "According to the brand's own marketing materials" is a slightly different grammatical construction than "according to independent reviews and buyer discussions," even at matched word count, and self-claim in the real world would typically read in first person ("we"), which round 2 deliberately removed to isolate attribution as the only variable, so round 2 trades some ecological realism for cleaner isolation. This round tested one fact per brand and three attribution frames; it did not test finer distinctions within third-party sourcing, like a review aggregator versus a news article versus word of mouth, which was flagged as an open question before this study ran. As the previous section covers, most responses in every condition picked neither the target brand nor its named competitor, so this measures how much a specific injected framing pulls selection toward the target, not a clean two-brand contest. One small data quality note for transparency: in round 1, 4 of 1,600 responses for Barbaro Mojo echoed a phrase from the injected fact ("Every Barbaro Mojo hot sauce...") without clearly presenting it as the pick, and the judge correctly declined to credit those as wins, which is the judge working as intended, not a bug, but is worth naming since it touches the raw text.

What does your own site say about you, and how?

Free AI Commerce Score™ in 10 seconds.

If how a fact is framed can be worth 22 points in a forced comparison, it's worth knowing what a model already has to work with about your brand today, and whether it reads like your own clear voice or like nothing at all.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study looks only at who appears to be making a claim. It sits alongside the series' other studies on what shapes candidacy and selection once a brand's facts are already fixed.