This study started as a reply to a specific claim: that what matters isn't your own marketing, but who else is backing up your claims. So we tested it directly, twice. Same brand, same fact, same length, same model. Only who appears to be saying the fact changes: the brand's own voice, a plain third-person statement with no named source, or an apparent independent source like reviews and buyer discussions. A first pass found the brand's own voice winning the most, then a robustness check found the three wrappers weren't quite the same length, so we rebuilt the wording to close that gap and ran the whole thing again. The result held. Combined across both rounds, the brand's own voice still wins the most, in every single brand.
A commenter, J.M., replied to a post about AI search with a specific theory: "Not just on the website but across the ecosystem. Where else and who is backing your claims. Not your marketing but real claims in the structure the LLM looks for." That's a testable claim. It says attribution, who appears to be making a factual statement about a brand, should change whether an AI system picks that brand, separately from what the fact actually says.
Every earlier study in this series changed the underlying facts, the context around them, or the shape of the buyer's question. This one holds the fact completely fixed and only changes who appears to be saying it. The identical sentence gets tested under four different framings.
Here is the real text, not a summary, for one of the four brands (Colored Organics, organic baby clothes). Round 1's wrappers turned out to be different lengths, which is exactly what round 2 fixed. Both are shown below, word for word.
Numbers below are the combined result of two independent rounds, 800 calls per condition, 3,200 calls total (see the next section for why there are two rounds and what changed between them). The order is not what the LinkedIn comment that started this study predicted. Third-party attribution beats a plain neutral statement, which is a real and useful finding on its own. But self-claim, the brand's own voice, beats third-party attribution too. Every pairwise gap here is statistically decisive: self-claim vs. third-party (z=6.67, p<0.0001), third-party vs. neutral (z=8.45, p<0.0001), and neutral vs. no fact (z=16.83, p<0.0001). A chi-square test across all four conditions, pooled, gives χ²=866.67 on 3 degrees of freedom.
Round 1 wrapped the identical fact in three different framings, but the wrapping text itself wasn't the same length. The self-claim wrapper ("On {brand}'s own website, the brand describes itself this way: ...") ran 9 to 12 words longer than the bare fact. The third-party wrapper ran 8 words longer. Neutral had no wrapper at all. That's a real hole: maybe the longer, more concrete-sounding wrapper won just because it was longer and more concrete-sounding, not because of who it claimed was speaking.
So we rebuilt the wrappers to be word-count matched, within a single word of each other ("According to the brand's own marketing materials: ...", "According to general publicly available information: ...", "According to independent reviews and buyer discussions: ..."), all wrapping the exact same third-person fact, and ran the entire 1,600-call design again from scratch, a fresh independent draw, not a re-run of the same prompts.
| Condition | Round 1 (original wrappers) | Round 2 (length-matched) | Combined |
|---|---|---|---|
| No fact | 0.2% | 0.0% | 0.1% |
| Neutral | 23.0% | 37.8% | 30.4% |
| Third party | 43.2% | 59.0% | 51.1% |
| Self-claim | 65.2% | 69.8%Highest, both rounds | 67.5% |
This is the table that makes the finding hold up on its own, brand by brand, not just in the pooled total. With both rounds combined (n=200 per brand per condition, double the power of either round alone), self-claim beats third-party and third-party beats neutral, individually statistically significant, in all four brands.
| Brand | No fact | Neutral | Third party | Self-claim |
|---|---|---|---|---|
| Colored Organics | 0% | 45% | 58% | 84%p<0.001 |
| Barbaro Mojo | 0% | 42% | 66% | 80%p=0.003 |
| Hearthloom (fictional control) | 0% | 34% | 72% | 86%p<0.001 |
| BodyArtForms | 0% | 0% | 8% | 21%p<0.001 |
The user prompt was open-ended ("what's the best X"), not a strict forced pick between exactly two names. Across all 3,200 calls in both rounds, the named competitor itself only won 37 times (1.2%). The rest of the non-target outcomes, 1,970 out of 3,200 (61.6%), went to some other real brand gpt-4o already knows, like Burt's Bees Baby for baby clothes or Neometal for piercing jewelry, brands that were never named in any prompt at all. So "winner rate" here is best read as: does the injected context, whichever framing it uses, pull the model's answer away from its own strong incumbent priors and toward the specific brand being described. That's still a meaningful, well-defined thing to measure, and the no-fact baseline (0.1%) shows how rarely it happens without any injected context at all. But it's a different claim than "brand A beats brand B head-to-head," and the page should be read that way.
A lot of current AI-SEO advice, including the comment that started this study, pushes toward getting mentioned elsewhere: reviews, press, forums, third-party sites, on the theory that an AI system trusts an outside voice more than a brand's own words. Third-party attribution genuinely does help here, roughly a 20-point lift over a source-free statement, replicated across both rounds. But in this mechanism, on these four brands, tested twice with two different wordings, a brand's own clearly stated, specific description outperformed that same fact dressed up as independent proof. The practical read isn't "ignore third-party mentions." It's that getting a clear, specific, well-written statement of your own facts into the places a model actually draws from may matter more than the current advice assumes, and shouldn't be skipped in favor of chasing outside citations alone.
Every response in both rounds went to a real gpt-4o judge, built in from the start per the Study #30 lesson. It never sees which condition or round produced a response, only the text, the target brand, and the competitor, so it cannot be biased toward any framing. All 3,200 calls parsed cleanly, and spot-checks confirmed it separates a real recommendation from a passing mention.
We looked for a way this finding could be wrong. Round 1 alone showed word count was itself predictive, exactly the confound round 2 was built to remove. Controlling for brand, round, and word count across all 3,200 calls, condition stays overwhelmingly significant (p=7.4×10-124) while word count's own effect is not (p=0.35). Length was worth checking. It isn't what's driving this.
This study used one model, gpt-4o, and one system-message injection mechanism, across two independent rounds. It has not been tested on other models or on live web search. Round 2's wrappers are still not a perfect linguistic match. "According to the brand's own marketing materials" is a slightly different grammatical construction than "according to independent reviews and buyer discussions," even at matched word count, and self-claim in the real world would typically read in first person ("we"), which round 2 deliberately removed to isolate attribution as the only variable, so round 2 trades some ecological realism for cleaner isolation. This round tested one fact per brand and three attribution frames; it did not test finer distinctions within third-party sourcing, like a review aggregator versus a news article versus word of mouth, which was flagged as an open question before this study ran. As the previous section covers, most responses in every condition picked neither the target brand nor its named competitor, so this measures how much a specific injected framing pulls selection toward the target, not a clean two-brand contest. One small data quality note for transparency: in round 1, 4 of 1,600 responses for Barbaro Mojo echoed a phrase from the injected fact ("Every Barbaro Mojo hot sauce...") without clearly presenting it as the pick, and the judge correctly declined to credit those as wins, which is the judge working as intended, not a bug, but is worth naming since it touches the raw text.
If how a fact is framed can be worth 22 points in a forced comparison, it's worth knowing what a model already has to work with about your brand today, and whether it reads like your own clear voice or like nothing at all.
This study looks only at who appears to be making a claim. It sits alongside the series' other studies on what shapes candidacy and selection once a brand's facts are already fixed.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →