Recommendation Intelligence Research™ · Study #30

Ask like you're Googling it. The AI recommends the brand 15.6 points less.

Every earlier study in this series changed the facts about a brand, or the context around it. This one changes neither. Same brand, same facts, same named competitors, same model. Only the way the buyer asks the question changes. We ran two separate tests: a 15-cell grid that varies the length and grammar of the question, and 7 writer-style personas, from a Google-style keyword typer to a formal older buyer. Both tests point to the same thing. A question that reads like a search-bar keyword loses to a full, natural sentence by the same 15.6 points, in both tests, on their own.

ZenodoCite this study: 10.5281/zenodo.22758351
15.6ppGap from keyword-style phrasing
704Total calls
15Length & type combinations
7Personas tested
Where this comes from

Same brand, same facts. Only the question's shape changes

An earlier study, Brand Legibility, found that once a brand is named, it almost always gets mentioned. The real contest is winner rate: who gets presented as the top pick against named competitors. This study keeps every fact fixed and changes only the buyer's own question. We split the test into two separate layers, so a result in one layer can't be mistaken for a side effect of the other.

A1
Length × sentence type
5 length levels (bare keyword, qualified phrase, "best X" phrase, full question, paragraph with context) × 3 grammatical types (declarative, imperative, question) = 15 cells. Demographically neutral, pure sentence mechanics.
B1
Neutral baseline
Plain, full-sentence question, no stylistic markers. The persona-layer control.
B2
Google-style
Bare search-bar keywords, no grammar, no punctuation.
B3
Voice assistant
Long, run-on, unpunctuated, the way a speech-to-text transcript reads.
B4
Young woman (Gen Z)
Informal, lowercase, light emoji and "!" use.
B5
Young man (Gen Z / millennial)
Short, informal, slang.
B6
Older woman
Formal, polite, full sentences.
B7
Older man
Formal, direct, full sentences, shorter than the older-woman style.
The writing styles for personas B4 through B7 are not guesses. They follow published research on how different age groups write online, covering things like period use, exclamation marks, and emoji. Full sources and the specific numbers are listed next to the persona results further down this page.
Model: gpt-4o throughout
Brands: 4, spanning organic baby clothes, body piercing jewelry, Cuban-style hot sauce, and ceramic dinnerware
Grid calls: 480 (15 cells × 4 brands × 8 repeats)
Persona calls: 224 (7 personas × 8 templates × 4 brands)
Winner determination: gpt-4o LLM judge, not a position heuristic
Statistics: chi-square tests, logistic regression controlling for brand
The brand's name never appears in the user's question. Only the system message carries the brand's real facts and its real named competitors. These stay identical across every cell for a given brand. The buyer's prompt only names the category, like "organic baby clothes" or "Cuban-style hot sauce." So any change in who wins comes from how the question is asked, not from what it says about the brand.
The actual prompts

What "tested from every angle" actually looked like

Here is the real text, not just summary numbers. Below is the full 15-cell length and sentence-type grid for one of the four brands we tested (Barbaro Mojo, Cuban-style hot sauce), plus one example prompt per persona. Every brand ran all 15 grid cells and all 56 persona prompts. What you see below is the exact wording used, word for word.

Length levelDeclarativeImperativeQuestion
L1 · bare "Hot sauce." "Hot sauce." "Hot sauce?"
L2 · qualified "Need Cuban-style hot sauce." "Find me Cuban-style hot sauce." "Cuban-style hot sauce?"
L3 · "best X" phrase "Want the best Cuban-style hot sauce." "Recommend the best Cuban-style hot sauce." "What's the best Cuban-style hot sauce?"Weakest cell
L4 · full question "I need a good brand of Cuban-style hot sauce." "Recommend a good brand of Cuban-style hot sauce." "What's a good brand of Cuban-style hot sauce?"Ceiling, 100%
L5 · paragraph "I'm looking for Cuban-style hot sauce. I grew up on real Cuban food and I'm tired of hot sauces that just taste like generic vinegar heat, so I want something that actually tastes like real mojo criollo, not a knockoff." "Please recommend Cuban-style hot sauce. I grew up on real Cuban food and I'm tired of hot sauces that just taste like generic vinegar heat, so I want something that actually tastes like real mojo criollo, not a knockoff." "What's a good option for Cuban-style hot sauce? I grew up on real Cuban food and I'm tired of hot sauces that just taste like generic vinegar heat, so I want something that actually tastes like real mojo criollo, not a knockoff."
Neutral baselineControl
"What's a good Cuban-style hot sauce?"
Google-styleWeakest persona
"best Cuban-style hot sauce"
Voice assistantUnpunctuated transcript
"hey can you find me a good Cuban-style hot sauce I need one pretty soon"
Young womanGen Z, ceiling 100%
"someone please recommend Cuban-style hot sauce i'm begging 🙏"
Young manGen Z / millennial
"yo whats a solid brand for Cuban-style hot sauce"
Older womanFormal, ceiling 100%
"Good afternoon, I was wondering if you could recommend a reliable brand for Cuban-style hot sauce."
Older manFormal, ceiling 100%
"I am looking for a dependable brand for Cuban-style hot sauce. What would you recommend?"
The finding · length

The short "best X" phrase loses. Not the shortest prompt. The middle one.

If length worked the way most prompting advice assumes, the shortest or the longest prompt would win. That is not what happened. A one-word query does fine. A full paragraph with context does fine too. The weakest cell in the whole grid is the short "best X" style phrase. This pattern held up across all three scoring methods we used, including the final, trusted LLM-judge pass (χ²=32.29, p=0.000002).

Winner rate by length level, pooled across brands · LLM judge, final scoring
Same brand, same facts · only the length of the question changes
L1 · bare keyword ("Hot sauce.")
99%
Near ceiling
L2 · qualified phrase
97.9%
Near ceiling
L3 · "best X" phrase
84.4%
Weakest
L4 · full natural question
100%
Ceiling
L5 · paragraph with context
91.7%
Strong
This is the strongest finding on the page. Under the first mechanical scoring rule, the rule-based re-score, and the final real LLM judge, L3 (the "best X" style phrase) came in last every time. L4 (a full, natural-sounding question) sat at or near 100% every time too. No other finding on this page held up this cleanly across all three scoring methods.
The finding · sentence type

Declarative edges out question and command. Only modestly, and only after a real correction.

This finding is worth explaining honestly, because our first read of it was wrong. A simple scoring rule (checking if the brand's first mention fell in the first 15% of the response) first showed declarative statements beating questions by 21 points. When we read a sample of the real responses by hand, we found the bug. Many replies open with a lead-in sentence before a numbered list. So a brand that was genuinely "1." on that list still got marked a loser, because its first mention came after that lead-in sentence. A rule-based fix only matched the original method 85.4% of the time, and it flipped the finding's direction. We did not trust either version. So every response was sent to a real gpt-4o judge and asked one question: was this brand the top recommendation? The numbers below come from that judge.

Winner rate by sentence type, pooled across brands · LLM judge, final scoring
χ²=9.11, p=0.011 · real, but a much smaller gap than first reported
Declarative ("I need X.")
98.8%
Highest
Question ("What's a good X?")
93.8%
Middle
Imperative ("Recommend X.")
91.2%
Lowest
Old method: 95.6% vs. 86.9% vs. 74.4% (declarative highest, 21-point spread). Rule-based re-score: 87.5% vs. 82.5% vs. 76.2% (question highest, direction flipped). Final LLM judge: 98.8% vs. 93.8% vs. 91.2% (declarative highest again, 7.6-point spread). The direction from the first method held up, but the size did not. The real effect is significant, but about a third the size we first reported. We are showing the correction, not just the final number, because that is the honest version of how we got here.
The finding · writer persona

Google-keyword style loses. Age and gender, on their own, don't

We tested seven writer-style personas. Each one used 8 different template phrasings, so no result depends on a single sentence. All seven ran the same way across all four brands.

Where the persona writing styles come from. The styles used for the younger and older personas are not guesses. They are based on real, published research on how different age groups write online. Older writers tend to read a period at the end of a sentence as clear and professional. One survey found 71% of adults over 50 read it that way. Gen Z writers more often skip the period and use "!" for enthusiasm instead. Emoji use also skews younger: 68% of Gen Z uses emoji regularly, including at work, compared to 36% of adults over 50. Sources: Stockton University, Talk Business (2025), and Thurlow (2002), "Generation Txt?"

These prompts were written by us, following these known patterns. They are not a claim about how every person in any group writes. The finding below is about how the model responds to a writing style, not about verified demographic behavior.

Winner rate by persona, pooled across brands · LLM judge, final scoring
χ²=17.38, p=0.008 · 7 personas × 8 templates × 4 brands
Older man
100%
Ceiling
Voice assistant
100%
Ceiling
Young woman
100%
Ceiling
Older woman
100%
Ceiling
Neutral baseline
93.8%
Control
Young man
90.6%
Below control
Google-style
84.4%
Weakest
Four of the seven personas sit at a 100% ceiling, including both older personas and one of the two younger ones. The one persona that does worse across every brand is Google-style bare-keyword typing. That is the same style of phrasing that came out weakest in the length grid above, found completely on its own, through a different set of prompts and a separate analysis. That is the real headline here. It is not about who is asking. It is about whether the question reads like a sentence.
A more cautious read

Persona-brand "affinity" is weaker than the first pass suggested

Our first look at the data hinted at a real connection between persona and brand: one brand spiked for specific personas in a way that looked like a genuine style-brand match. Once every response went through the same LLM-judge correction described above, that pattern mostly went away. Under the trusted scoring, the persona-by-brand results are close to 100% almost everywhere, with two exceptions. One brand (body piercing jewelry) scores lower across most personas. Google-style phrasing scores lower across all four brands. A logistic regression that controls for brand confirms both the length effect and the persona effect hold up on their own (both p<0.01). The honest read is two separate weak spots that add up, not one strong style-and-brand interaction.

Why it matters

The divide isn't who's asking. It's whether it reads like a sentence

The most reassuring part of this page might be what did not show up. There is no consistent age or gender bias among the six human-styled personas, once the scoring was done properly. A formal older buyer and an informal Gen Z buyer, asking for the exact same thing, land in roughly the same place, both near the ceiling. What actually moved the result, in two separate tests that have nothing to do with each other, was the same pattern: a bare, keyword-style phrasing that reads like a search bar instead of a real question. If there is a practical takeaway here, it is about how people write to AI assistants, not about who they are. A full sentence beats a keyword string by roughly the same amount, no matter who is typing it.

Supporting evidence

Three scoring methods, one trusted judge

90.1% / 90.5% Judge agreement with the old heuristic

The final gpt-4o judge agreed with the original method 90.1% of the time on the length grid and 90.5% on the personas. That is higher than the 85.4% agreement of our own rule-based re-score attempt. Even that in-between fix was not perfect. Only the real judge is treated as the trusted source on this page.

p<0.01 Length and persona effects, controlling for brand

A logistic regression that controls for brand confirms both effects hold up on their own. The length effect stays significant (likelihood-ratio test, p<0.00000001), and so does the persona effect (p=0.003). Neither one is just a side effect of one brand's overall strength.

What this doesn't prove

One scoring correction disclosed, a smaller one still open

The sentence-type finding changed direction once, then changed size again, across three rounds of scoring on the same data. We showed all three rounds here, not just the final number. Candidacy (whether the brand gets mentioned at all) sits at 97 to 100% across nearly every cell and persona. That matches the rest of this series. The real signal in this study is entirely in winner rate, not candidacy. We used four brands, reused from Brand Legibility because their facts were already verified. That is a narrow category spread: organic baby clothes, body piercing jewelry, Cuban-style hot sauce, and ceramic dinnerware. The persona layer tests writing style, not stated identity. A prompt that directly states the asker's age or gender, like "As a 65-year-old man...", is a separate follow-up we have not run yet. This study used one model, one system-message setup, and one round of 704 calls.

How does your own site read to a model?

Free AI Commerce Score™ in 10 seconds.

If a keyword-style query can cost a brand 15.6 points against the exact same facts, it's worth knowing what a model already has to work with when someone asks about you, however they happen to phrase it.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This study looks only at the shape of the buyer's question. It sits alongside the series' other studies on what shapes candidacy and selection once a brand's facts are already fixed.