This study started as a reply to a specific theory: the more concrete you are on your pricing and product pages, real certifications, sizes, materials, dollar amounts, the more an AI system will recommend you over generic marketing language. We tested it three times. Round one found concrete claims beating vague ones overall, but one brand quietly moved backward. Round two replicated the entire design from scratch and the same brand moved backward again. Round three added four new brands, chosen deliberately, two in categories where trust and safety are most of the sale, two in categories that are pure specs, to find out why. Across all eight brands combined, the pattern is clean: concreteness helps, reliably, in functional categories. In categories where trust is the real product, it stops helping, and can cost you instead.
A commenter, K.E., replied to a post about AI visibility with a specific argument: that what actually moves an AI recommendation is not just having a pricing or product page, but how concrete the language on it is. Generic marketing copy, the argument went, gives a model nothing to repeat back. Specific claims, real certifications, sizes, materials, dollar amounts, give it something. That is a testable claim about the shape of the language itself, separate from who appears to be saying it, which an earlier study in this series already measured in Claim Attribution.
Every earlier study in this series changed who says a fact, how much context a model already has, or how a buyer phrases their question. This one holds the underlying facts completely fixed and only changes how concrete the language describing them is: the same real facts about a brand, written once as generic marketing copy, and once as specific, PDP and pricing page style detail.
Here is the real text, not a summary, for two of the eight brands: Colored Organics, from the original round one and two brand set, and Primally Pure, one of the four new brands added in round three. Both show the identical wrapper used in every condition across all three rounds: "Here is what {brand} says about itself: {claim}".
Numbers below are the combined result of rounds one and two, run on the same four original brands, 800 calls per condition, 2,400 calls total (the next section covers why there are two rounds and what replicated). Rewriting a brand's real facts as specific, PDP style claims beat the exact same facts written as vague marketing language, and both comfortably beat having no claim at all. The vague vs. specific gap is statistically decisive (z=3.55, p=0.0004), as is no claims vs. vague (p<0.0001). A chi-square test across all three conditions, pooled, gives χ²=780.09 on 2 degrees of freedom (p<0.0001).
Round one wrapped the same real facts as vague and specific claims and found specificity ahead overall. Before trusting that, we reran the entire 1,200 call design from scratch, a new independent draw with a new random seed, not a re-run of the same prompts. The pooled result barely moved.
| Condition | Round 1 | Round 2 (fresh replication) | Combined |
|---|---|---|---|
| No claims | 0.0% | 0.0% | 0.0% |
| Vague | 54.2% | 54.0% | 54.1% |
| Specific | 62.8% | 63.0%Replicated | 62.9% |
This is the table that started the real question behind this study. With both rounds combined (n=200 per brand per condition, double the power of either round alone), three of the four original brands show specificity clearly helping. The fourth, Colored Organics, organic baby clothes, shows the opposite: the vague version outperformed the specific one.
| Brand | Category | Vague | Specific | Change |
|---|---|---|---|---|
| Colored Organics | Organic baby clothes | 93.5% | 85.0%p=0.006, reversed | −8.5pp |
| BodyArtForms | Body piercing jewelry | 9.0% | 10.0%p=0.73, flat | +1.0pp |
| Barbaro Mojo | Cuban style hot sauce | 50.5% | 68.0%p<0.001 | +17.5pp |
| Hearthloom | Ceramic dinnerware | 63.5% | 88.5%p<0.001 | +25.0pp |
Round 3 did not reuse the original four brands. It added four new ones, real brands with real, already verified facts from Atom Foundry's own recommendation reports, or freshly verified for this study, chosen deliberately to sit on either side of a category axis: is the category one where trust and safety are most of the sale, or one where the decision is really about specs.
Same design as rounds 1 and 2: no claims, vague, specific, one fixed competitor per brand (Ruffwear, Native Deodorant, Herschel, and SurveyMonkey), 20 purchase intents, 5 repeats, 1,200 fresh calls. Because these are all real, already established brands rather than obscure ones, the no claims baseline is much higher here, 24.5 percent, than it was for the original four, which is expected and covered in the more cautious read below.
This combines every brand from all three rounds, grouped by category type instead of by round, 600 calls per condition per category, 2,400 calls total. In purely functional categories, specificity adds 15.7 points and the gap is overwhelming (p<0.0001). In trust and safety sensitive categories, specificity does not help, and the direction reverses, though the pooled trust gap alone is not statistically significant (p=0.11). What is significant is the difference between the two patterns: a likelihood-ratio test for a category type by specificity interaction, controlling for brand, across all 2,400 calls, is decisive (χ²=39.26, df=1, p<0.0001).
Round 3's category split above is real and highly significant pooled, but not every individual brand had room to show it. Two of the four new brands were already recommended almost every time in both conditions, a ceiling effect that limits how much any claim, vague or specific, could move the number.
| Brand | Category type | Vague | Specific | Read |
|---|---|---|---|---|
| Wild One | Trust (dog gear) | 87.0% | 74.0%p=0.020, reversed | |
| Primally Pure | Trust (natural deodorant) | 99.0% | 100.0%p=0.32, ceiling | |
| Zigpoll | Functional (Shopify survey tool) | 85.0% | 96.0%p=0.008 | |
| Bellroy | Functional (slim leather wallets) | 100.0% | 98.0%p=0.16, ceiling |
Rounds 1 and 2 used obscure, low profile brands, so with no claim injected at all, the target brand was named 0 percent of the time in either round, 1,440 out of 1,440 no claims responses in the pooled set going to some other brand gpt-4o already knew. Round 3 deliberately used already established brands with real recommendation history, so its no claims baseline is 24.5 percent, the model already has some opinion about Wild One, Bellroy, Primally Pure, and Zigpoll before any claim is injected at all. That is expected, and it is also why round 3's per brand read is cleaner in some ways (less noise from an unfamiliar name) and noisier in others (2 of 4 brands already near the ceiling). The category by specificity split, tested on the pooled 8 brand set where both obscure and famous brands are represented in both categories, is the number that should travel, not any single round's own baseline.
The general AI-SEO advice to be concrete on pricing and product pages, real numbers over adjectives, holds up here, and holds up strongly: a 15.7 point gain in categories where the decision is genuinely about specs. But the same rewrite that helps a hot sauce brand or a Shopify app does not reliably help, and can cost, a brand whose actual product is trust: what goes on a newborn's skin, what goes through a piercing, whether a natural deodorant is really free of harsh chemicals. One reading is that piling on certifications and specifics in a trust category can start to read as trying to prove something, which is a different signal than a category where the buyer already assumes the specs are the whole pitch. The practical takeaway is not to stop being specific. It is to know which kind of category you are in before you decide how much proof to lead with.
Every response across all 3 rounds went to a real gpt-4o judge, built in from the start. It never sees which condition, round, or brand produced a response, only the response text, the target brand, and the competitor, so it cannot be biased toward vague or specific framing. All 3,600 calls parsed cleanly.
Specific claims ran 7 to 20 words longer than vague ones on the original 4 brands, a real design imperfection worth checking. Controlling for brand and word count across all 2,400 rounds 1 and 2 calls, specificity stays significant (p=0.0005) while word count's own effect is not (p=0.24). Round 3's new brands showed a genuine word count effect of their own (p=0.004), but specificity remained significant there too (p=0.007) once word count was controlled for. Length is a real thing to check. It is not what is driving this finding in either brand set.
This study used one model, gpt-4o, and one system message injection mechanism, across three independent rounds. It has not been tested on other models or on live web search. The trust versus functional axis was tested with 2 brands on each side, not an exhaustive survey of every possible category, and 2 of the 4 round 3 brands (Bellroy, Primally Pure) were already so close to 100 percent recommended in both conditions that there was little room left to detect a difference for those two specifically, even though their direction still matched the overall pattern. Most responses in every condition, in rounds 1 and 2 especially, picked some other brand entirely rather than the target or its named competitor, so this measures how much a given version of a brand's claims pulls selection toward that brand, not a clean two way contest. Round 3's brands also carried a real word count difference between vague and specific claims, and while specificity held up after controlling for it, that is a design detail worth fixing in any follow up round rather than something to gloss over.
If the same rewrite can be worth 15.7 points in one category and cost you 4.5 in another, it is worth knowing what a model already has to work with about your brand today, and whether your own pages read like proof or like trying too hard.
This study looks only at how concrete a brand's own language is. It sits alongside the series' other studies on who appears to make a claim and what shapes candidacy once a brand's facts are already fixed.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →