Recommendation Intelligence Research™ · Study #32

Specific claims add 15.7 points when a category is pure specs. They cost nearly 5 points when the category is about trust.

This study started as a reply to a specific theory: the more concrete you are on your pricing and product pages, real certifications, sizes, materials, dollar amounts, the more an AI system will recommend you over generic marketing language. We tested it three times. Round one found concrete claims beating vague ones overall, but one brand quietly moved backward. Round two replicated the entire design from scratch and the same brand moved backward again. Round three added four new brands, chosen deliberately, two in categories where trust and safety are most of the sale, two in categories that are pure specs, to find out why. Across all eight brands combined, the pattern is clean: concreteness helps, reliably, in functional categories. In categories where trust is the real product, it stops helping, and can cost you instead.

ZenodoCite this study: 10.5281/zenodo.22776560
+15.7 ptsSpecific claim gain, functional
−4.5 ptsSpecific claim cost, trust
3,600Total calls, 3 rounds
8Brands, split by category
Where this comes from

A LinkedIn comment, turned into a testable claim

A commenter, K.E., replied to a post about AI visibility with a specific argument: that what actually moves an AI recommendation is not just having a pricing or product page, but how concrete the language on it is. Generic marketing copy, the argument went, gives a model nothing to repeat back. Specific claims, real certifications, sizes, materials, dollar amounts, give it something. That is a testable claim about the shape of the language itself, separate from who appears to be saying it, which an earlier study in this series already measured in Claim Attribution.

Every earlier study in this series changed who says a fact, how much context a model already has, or how a buyer phrases their question. This one holds the underlying facts completely fixed and only changes how concrete the language describing them is: the same real facts about a brand, written once as generic marketing copy, and once as specific, PDP and pricing page style detail.

1
No claims
Control. No claim about the brand at all, just a generic comparison shopping context. Establishes the baseline used across the series.
2
Vague
The brand's real facts, written as generic marketing language, no numbers, certifications, or specifics named.
3
Specific
The exact same underlying real facts, rewritten as concrete, PDP and pricing page style claims: sizes, certifications, dollar amounts, founding dates, and other named specifics.
Why this is measured as a forced head to head, not open candidacy. Brand Legibility (Study #29) already showed that once a brand is named in context, it gets mentioned almost every time, regardless of how that context is written. Measuring candidacy here would likely hit that same ceiling and tell us nothing about specificity. So this study uses the same forced comparison mechanism as Claim Attribution and Cold Start: one fixed, established competitor per brand, across 20 purchase intents, and winner rate (does the target brand get the actual recommendation) is the primary number, not candidacy.
Model: gpt-4o throughout
Brands: 8 total, spanning organic baby clothes, body piercing jewelry, Cuban style hot sauce, ceramic dinnerware, dog gear, natural deodorant, slim leather wallets, and a Shopify survey tool
Design: 4 brands × 3 conditions × 20 purchase intents × 5 repeats, run 3 times
Total calls: 3,600 (1,200 per round, 2 exact replications on the original 4 brands, 1 fresh round on 4 new brands)
Winner determination: gpt-4o LLM judge, built in from the first run, not added after the fact
Statistics: two-proportion z-tests, chi-square tests, logistic regression controlling for brand and word count, likelihood-ratio interaction tests
The target brand's name never appears in the user's question. The brand's facts, in whichever version, live entirely in the system message, the same mechanism used by Hidden Context and Fact Injection: information the model already has going in, not something the shopper just said. The user's question names only the one fixed competitor, as something the shopper is already considering, and asks for a clear pick. Any time the target brand wins, it won purely from what was sitting in context, never from the user handing the brand's name back to the model.
The actual prompts

The same real facts, written two different ways

Here is the real text, not a summary, for two of the eight brands: Colored Organics, from the original round one and two brand set, and Primally Pure, one of the four new brands added in round three. Both show the identical wrapper used in every condition across all three rounds: "Here is what {brand} says about itself: {claim}".

Colored Organics, organic baby clothes (round 1 and 2 brand set)
No claimsControl, 19 words
"The user has been comparing options in this category for a few days and wants a clear, confident recommendation."
Vague37 words
"Here is what Colored Organics says about itself: Premium organic materials your baby will love. Thoughtfully made with safety and sustainability in mind, for the whole family. Proudly small and values driven, giving back because it matters."
Specific47 words
"Here is what Colored Organics says about itself: GOTS-certified organic cotton, sized newborn to 24 months. Woman owned, founded by a mother. Dyes are non-toxic, free of azo compounds and heavy metals, and snaps are nickel-free. A portion of profits goes to a children's charity every month."
Primally Pure, natural deodorant (one of 4 new round 3 brands, trust category)
Vague29 words
"Here is what Primally Pure says about itself: Clean, natural ingredients you can trust on your skin. Made with care, free from the harsh chemicals found in conventional products."
Specific44 words
"Here is what Primally Pure says about itself: Founded in 2015 on a family farm in Southern California, starting with a homemade natural deodorant. Made with grass-fed tallow, organic coconut oil, beeswax, and arrowroot powder, no synthetic ingredients, no outside investors, still independently run."
Shared user prompt, unchanged across all 3 rounds and all 3 conditions, only the purchase intent changes: "What's the best organic baby onesies? I'm already considering Finn + Emma, but I want a clear pick, name one brand and give a one-sentence reason." Each brand ran all 20 of its own purchase intents, 5 times each, under all 3 conditions, in its round or rounds.
The finding

Concrete beats vague, pooled across the first 4 brands

Numbers below are the combined result of rounds one and two, run on the same four original brands, 800 calls per condition, 2,400 calls total (the next section covers why there are two rounds and what replicated). Rewriting a brand's real facts as specific, PDP style claims beat the exact same facts written as vague marketing language, and both comfortably beat having no claim at all. The vague vs. specific gap is statistically decisive (z=3.55, p=0.0004), as is no claims vs. vague (p<0.0001). A chi-square test across all three conditions, pooled, gives χ²=780.09 on 2 degrees of freedom (p<0.0001).

Winner rate by claim specificity, rounds 1 and 2 combined, original 4 brands
n=800 per condition · gpt-4o LLM judge, built in from the start
No claims
0.0%
Baseline
Vague
54.1%
+54.1pp
Specific
62.9%
Highest
Concreteness itself is a real, separate lever, not just a byproduct of having a claim at all. Going from no claim to a vague one already produces most of the movement, from 0 percent to 54.1 percent. But going from vague to specific, the exact same underlying facts, just written with real numbers and names, adds another 8.8 points on top. A logistic regression on all 2,400 calls, controlling for brand and system message word count, still finds specificity significant (p=0.0005), and word count itself is not significant once specificity is in the model (p=0.24). This is the pooled result. It is not the whole story, as the next two sections show.
Making sure it holds

Round 2 was a fresh, independent replication. The pooled number held. One brand's behavior did not.

Round one wrapped the same real facts as vague and specific claims and found specificity ahead overall. Before trusting that, we reran the entire 1,200 call design from scratch, a new independent draw with a new random seed, not a re-run of the same prompts. The pooled result barely moved.

ConditionRound 1Round 2 (fresh replication)Combined
No claims 0.0% 0.0% 0.0%
Vague 54.2% 54.0% 54.1%
Specific 62.8% 63.0%Replicated 62.9%
The pooled effect replicated almost exactly. The per-brand breakdown is where it got interesting. Specific claims were 7 to 20 words longer than vague ones in every brand, a real design imperfection worth checking. Controlling for brand and word count, specificity stayed significant (p=0.0005) and word count did not (p=0.24), so length is not what is driving the pooled result. But when we broke the same combined data down brand by brand, one brand, Colored Organics, was moving in the opposite direction from the other three, consistently, in both rounds. A likelihood-ratio test for a brand by specificity interaction, across all 2,400 calls, found real heterogeneity (χ²=35.87, df=3, p<0.0001). That is not something four uncorrected per-brand tests could tell you on their own. It is what sent us looking for a pattern in which brand goes backward, which is what round 3 was built to test.
The finding · by brand, rounds 1 and 2

One brand went backward. The other three went up.

This is the table that started the real question behind this study. With both rounds combined (n=200 per brand per condition, double the power of either round alone), three of the four original brands show specificity clearly helping. The fourth, Colored Organics, organic baby clothes, shows the opposite: the vague version outperformed the specific one.

BrandCategoryVagueSpecificChange
Colored Organics Organic baby clothes 93.5% 85.0%p=0.006, reversed −8.5pp
BodyArtForms Body piercing jewelry 9.0% 10.0%p=0.73, flat +1.0pp
Barbaro Mojo Cuban style hot sauce 50.5% 68.0%p<0.001 +17.5pp
Hearthloom Ceramic dinnerware 63.5% 88.5%p<0.001 +25.0pp
Look at what Colored Organics, BodyArtForms, Barbaro Mojo, and Hearthloom actually sell. The two brands where specificity helped most, a hot sauce maker and a ceramic dinnerware brand, are purely functional categories: taste, food safety, dishwasher compatibility. The one brand where it reversed, organic baby clothing, is about as trust and safety sensitive as a category gets, a parent choosing what touches a newborn's skin. BodyArtForms, body piercing jewelry, is trust sensitive too, and shows no significant movement either way. That pattern, concreteness helping in functional categories and not helping, or reversing, in trust sensitive ones, was a hypothesis after round 2. Round 3 was built specifically to test it on brands we had never touched before.
Why it reverses

4 new brands, chosen 2 and 2, on purpose

Round 3 did not reuse the original four brands. It added four new ones, real brands with real, already verified facts from Atom Foundry's own recommendation reports, or freshly verified for this study, chosen deliberately to sit on either side of a category axis: is the category one where trust and safety are most of the sale, or one where the decision is really about specs.

Trust and safety sensitive
New brand facts, real and verified Wild One · dog gear
Primally Pure · natural deodorant
Purely functional
New brand facts, real and verified Bellroy · slim leather wallets
Zigpoll · Shopify survey tool

Same design as rounds 1 and 2: no claims, vague, specific, one fixed competitor per brand (Ruffwear, Native Deodorant, Herschel, and SurveyMonkey), 20 purchase intents, 5 repeats, 1,200 fresh calls. Because these are all real, already established brands rather than obscure ones, the no claims baseline is much higher here, 24.5 percent, than it was for the original four, which is expected and covered in the more cautious read below.

The finding · by category, all 8 brands

Pool all 8 brands, and the split is clean

This combines every brand from all three rounds, grouped by category type instead of by round, 600 calls per condition per category, 2,400 calls total. In purely functional categories, specificity adds 15.7 points and the gap is overwhelming (p<0.0001). In trust and safety sensitive categories, specificity does not help, and the direction reverses, though the pooled trust gap alone is not statistically significant (p=0.11). What is significant is the difference between the two patterns: a likelihood-ratio test for a category type by specificity interaction, controlling for brand, across all 2,400 calls, is decisive (χ²=39.26, df=1, p<0.0001).

Winner rate, vague vs. specific, by category type, all 8 brands combined
n=600 per bar · gpt-4o LLM judge
Functional, vague
68.8%
Baseline
Functional, specific
84.5%
+15.7pp
Trust, vague
65.2%
Baseline
Trust, specific
60.7%
−4.5pp
This is the cleanest version of the finding, because it pools everything. It is not just round 3's four new brands. It is all eight brands from all three rounds, sorted into the same two buckets. Functional categories (Cuban style hot sauce, ceramic dinnerware, slim leather wallets, a Shopify survey tool): vague 68.8 percent, specific 84.5 percent, gain of 15.7 points, p<0.0001. Trust and safety sensitive categories (organic baby clothes, body piercing jewelry, dog gear, natural deodorant): vague 65.2 percent, specific 60.7 percent, a drop of 4.5 points that alone does not clear significance at this sample size, but which is the opposite direction from every functional brand, and the category by specificity interaction is one of the strongest effects in this entire study.
Round 3 · by brand

2 of the 4 new brands moved cleanly. 2 sat too close to the ceiling to tell.

Round 3's category split above is real and highly significant pooled, but not every individual brand had room to show it. Two of the four new brands were already recommended almost every time in both conditions, a ceiling effect that limits how much any claim, vague or specific, could move the number.

BrandCategory typeVagueSpecificRead
Wild One Trust (dog gear) 87.0% 74.0%p=0.020, reversed
Primally Pure Trust (natural deodorant) 99.0% 100.0%p=0.32, ceiling
Zigpoll Functional (Shopify survey tool) 85.0% 96.0%p=0.008
Bellroy Functional (slim leather wallets) 100.0% 98.0%p=0.16, ceiling
Both already-famous brands per category sat near the ceiling, in both conditions. Bellroy and Primally Pure both already get recommended 99 to 100 percent of the time regardless of how their facts are worded, so there is almost no room left for a claim to move the number, in either direction. Where there was real room to move, Wild One (trust) reversed significantly, matching Colored Organics from round 1 and 2, and Zigpoll (functional) improved significantly, matching Barbaro Mojo and Hearthloom. Across all 8 brands measured in this study, zero of the four trust category brands showed specificity significantly helping, and three of the four functional category brands did.
A more cautious read

Round 3's brands were already famous. That changes the baseline, not the pattern.

Rounds 1 and 2 used obscure, low profile brands, so with no claim injected at all, the target brand was named 0 percent of the time in either round, 1,440 out of 1,440 no claims responses in the pooled set going to some other brand gpt-4o already knew. Round 3 deliberately used already established brands with real recommendation history, so its no claims baseline is 24.5 percent, the model already has some opinion about Wild One, Bellroy, Primally Pure, and Zigpoll before any claim is injected at all. That is expected, and it is also why round 3's per brand read is cleaner in some ways (less noise from an unfamiliar name) and noisier in others (2 of 4 brands already near the ceiling). The category by specificity split, tested on the pooled 8 brand set where both obscure and famous brands are represented in both categories, is the number that should travel, not any single round's own baseline.

Why it matters

Be specific, unless the sale is really about trust

The general AI-SEO advice to be concrete on pricing and product pages, real numbers over adjectives, holds up here, and holds up strongly: a 15.7 point gain in categories where the decision is genuinely about specs. But the same rewrite that helps a hot sauce brand or a Shopify app does not reliably help, and can cost, a brand whose actual product is trust: what goes on a newborn's skin, what goes through a piercing, whether a natural deodorant is really free of harsh chemicals. One reading is that piling on certifications and specifics in a trust category can start to read as trying to prove something, which is a different signal than a category where the buyer already assumes the specs are the whole pitch. The practical takeaway is not to stop being specific. It is to know which kind of category you are in before you decide how much proof to lead with.

Supporting evidence

A real judge, and a confound we checked twice

0 parse failures Out of 3,600 judge calls, all 3 rounds

Every response across all 3 rounds went to a real gpt-4o judge, built in from the start. It never sees which condition, round, or brand produced a response, only the response text, the target brand, and the competitor, so it cannot be biased toward vague or specific framing. All 3,600 calls parsed cleanly.

p=0.24 Word count's own effect, once specificity is controlled

Specific claims ran 7 to 20 words longer than vague ones on the original 4 brands, a real design imperfection worth checking. Controlling for brand and word count across all 2,400 rounds 1 and 2 calls, specificity stays significant (p=0.0005) while word count's own effect is not (p=0.24). Round 3's new brands showed a genuine word count effect of their own (p=0.004), but specificity remained significant there too (p=0.007) once word count was controlled for. Length is a real thing to check. It is not what is driving this finding in either brand set.

What this doesn't prove

One model, one axis, honestly narrowed

This study used one model, gpt-4o, and one system message injection mechanism, across three independent rounds. It has not been tested on other models or on live web search. The trust versus functional axis was tested with 2 brands on each side, not an exhaustive survey of every possible category, and 2 of the 4 round 3 brands (Bellroy, Primally Pure) were already so close to 100 percent recommended in both conditions that there was little room left to detect a difference for those two specifically, even though their direction still matched the overall pattern. Most responses in every condition, in rounds 1 and 2 especially, picked some other brand entirely rather than the target or its named competitor, so this measures how much a given version of a brand's claims pulls selection toward that brand, not a clean two way contest. Round 3's brands also carried a real word count difference between vague and specific claims, and while specificity held up after controlling for it, that is a design detail worth fixing in any follow up round rather than something to gloss over.

Update: specificity was picked up as a factor to test head to head against authority, familiarity, and format in Study #37, Two Signals Get You Most of the Way There. A Third Barely Helps, and Format Doesn't Move It at All, with rating removed entirely. Specificity came out as the strongest single marginal contributor of the four.

Know which category your brand is in, before you write a word

Free AI Commerce Score™ in 10 seconds.

If the same rewrite can be worth 15.7 points in one category and cost you 4.5 in another, it is worth knowing what a model already has to work with about your brand today, and whether your own pages read like proof or like trying too hard.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study looks only at how concrete a brand's own language is. It sits alongside the series' other studies on who appears to make a claim and what shapes candidacy once a brand's facts are already fixed.