Recommendation Intelligence Research™ · Study #38

We turned the same facts into bullets. The model cited fewer of them.

Format has lost twice already in this series on the question of who wins a forced comparison: rating swamped it in Winner vs Loser, and its own marginal contribution rounded to zero in Signal Hierarchy. Neither study asked what "structured markup helps AI understand your content" usually means in practice, whether structuring the same facts changes how much of them the model actually repeats back. This study holds three real facts about a brand exactly constant and only changes whether they arrive as a flowing paragraph or a bulleted list. Structured format did not increase citation rate or vocabulary reuse. It reduced both, by a small but statistically significant margin, replicated independently in two separate rounds.

ZenodoCite this study: 10.5281/zenodo.22819917
−5.2ppCitation rate, structured vs. prose
−3.9ppVocabulary lift, structured vs. prose
2Independent rounds, same direction
1,600Total calls, both rounds
Where this comes from

Format has been tested for selection. Never for citation

Winner vs Loser found rating decided a forced two-brand comparison 90.8% of the time regardless of format. Signal Hierarchy tested authority, familiarity, specificity, and format against each other directly and found format's own marginal contribution to the winning brand rounded to zero in both independent rounds (-0.2pp, then 0.0pp). Both studies asked the same kind of question: does structuring the page change who wins a comparison. Neither ever asked whether structuring the same facts changes how much of them the model actually repeats back, specifically or in its own wording, when someone later asks it an open buyer question. Selection and citation are not the same outcome, and this series had only ever measured the first one.

This study is deliberately single-brand and open-ended, not a forced two-name comparison, borrowing the system-message injection mechanism from Fact Injection and Hidden Context instead, since the question here is what the model volunteers, not which of two options it names. No competitor appears anywhere in the design.

Model: gpt-4o throughout
Brands: Colored Organics, BodyArtForms, Barbaro Mojo, Hearthloom
Design: 2 conditions × 20 purchase intents × 5 repeats × 2 rounds
Total calls: 1,600 (800 per round)
Judge calls: 4,800 (2,400 per round, one per injected fact)
Statistics: clustered paired t-test, brand × intent, n=80 per round
What "structured" means here. The same three facts, always all three, held fixed per brand: a concrete specific claim, a disclosed-synthetic authority mention, and a familiarity claim, the same "full stack" content used as the AFSM condition in Signal Hierarchy. Prose joins them into one paragraph. Structured presents them as a bulleted list, one fact per line, the exact same bullet-assembly logic used in Signal Hierarchy's own target blurb.
The test

Same three facts, only the separators change

Here is the real system message for Hearthloom (handmade ceramic dinnerware) at both extremes of the design. The facts are word-for-word identical. Only the format changes.

Prose condition
Additional context retrieved for this query: Hearthloom Founded by two sisters, hand-thrown in small batches with lead-free, food-safe glazes. Dishwasher and microwave safe despite being handmade. Packaging is 100 percent recyclable. It has been featured in Apartment Therapy's guide to the best handmade dinnerware brands. It is one of the most recognized and widely known handmade dinnerware brands among home cooks.
Structured condition
Additional context retrieved for this query: Hearthloom - Founded by two sisters, hand-thrown in small batches with lead-free, food-safe glazes. Dishwasher and microwave safe despite being handmade. Packaging is 100 percent recyclable. - It has been featured in Apartment Therapy's guide to the best handmade dinnerware brands. - It is one of the most recognized and widely known handmade dinnerware brands among home cooks.

User message, no brand name in it, only in the injected context: "I'm looking for handmade ceramic dinner plates. What would you recommend and why?" Each brand ran all 20 of its own real purchase intents, 5 times each, under both conditions, in both rounds.

The finding

Prose beat structured on both outcomes, both rounds

Citation rate is the share of the 3 injected facts an LLM judge found cited or clearly paraphrased in the response. Vocabulary lift is the share of 5 pre-registered distinctive terms per brand (certifications, names, specific phrases) that show up verbatim. Both measures moved in the same direction, structured below prose, in both independently-run rounds.

MetricConditionRound 1Round 2
Citation rate Prose 65.3% 63.8%
Structured 59.5% 59.2%
Difference (structured − prose) −5.83pp, p=0.0003 −4.67pp, p=0.0011
Vocabulary lift Prose 67.7% 68.9%
Structured 64.5% 64.4%
Difference (structured − prose) −3.25pp, p=0.0044 −4.50pp, p<0.0001
Word count Prose 144.6 145.6
Structured 142.5 140.0
Brand mentioned Prose / Structured 99.8% / 100.0% 99.8% / 100.0%
Clustered paired t-test, brand × intent (80 clusters per round, each averaged over 5 repeats), both directions and both p-values hold independently in round 1 and round 2. Per the series' no-file-drawer rule, that makes this a real finding, not a round-1 fluke: structured markup produced a lower citation rate and lower vocabulary reuse than plain prose carrying the identical facts, in both outcomes, in both rounds.
Making sure it holds

Not a length confound, not a mention-rate confound

Two obvious alternative explanations were checked directly. First, response length: structured responses were not longer than prose, they were if anything slightly shorter (141.3 words average vs. 145.1 for prose, both rounds combined), so the gap cannot be explained by prose simply giving the model more room to cite things. Second, brand mention: both conditions sat at or above 99.8% brand-mention rate in both rounds, essentially at ceiling, so the citation-rate gap is not a byproduct of structured responses failing to name the brand at all, it shows up specifically in whether the model repeats the injected facts once the brand is already named.

Why it matters

"Structure it for AI" is not a free upgrade

A common piece of AI-visibility advice is to reformat product and brand content into bullets and lists because that is supposedly what gets cited. This series has now tested that claim three separate times on three different outcomes, and it has lost every time: format's marginal contribution to winning a forced comparison rounded to zero in Signal Hierarchy, rating swamped it entirely in Winner vs Loser, and now, on the specific outcome the "structure it for AI" advice is actually about, whether facts get cited and specific vocabulary gets reused, structured format performs measurably worse than plain prose carrying the identical information. The effect here is small, five points on citation rate, four on vocabulary lift, not a reason to panic about existing bulleted content. But it is a real, twice-replicated result in the opposite direction from what the popular advice assumes, and citation counts and format-driven visibility claims common elsewhere online deserve the same scrutiny this series has been applying to authority, familiarity, and specificity.

What this doesn't prove

Four brands, one injection mechanism, synthetic facts

All injected facts are synthetic (disclosed-synthetic authority mention, synthetic familiarity claim), the same convention as the rest of this series, not scraped from real press or real market-recognition data. Only 4 brands were tested, the same acknowledged category-coverage limitation as every study in this series. Single-turn, gpt-4o only, a single injection mechanism, system-message "retrieved context" framing, not tested on live web search or on a model actually calling a retrieval tool. The 5 vocabulary terms per brand are a judgment call about what counts as distinctive, fixed before any data collection, but a judgment call nonetheless, not an exhaustive or algorithmically derived list. Repeats=5 per cell follows this series' standard, not a power calculation specific to this study, the same open item flagged against the series-level statistical-power audit.

Supporting evidence

Two checks that the direction is real, not noise

2 of 2 Outcomes replicate, both rounds

Both citation rate and vocabulary lift favored prose over structured in round 1 and again independently in round 2, with p-values under 0.005 in three of the four round-level tests and under 0.0001 in one.

141 vs 145 Words, structured vs. prose

Structured responses were not longer than prose responses, both rounds combined. The citation-rate gap cannot be explained by prose simply having more room to work with.

Is your product content actually getting cited?

Free AI Commerce Score™ in 10 seconds.

If reformatting for "AI readability" can make citation slightly worse, not better, it is worth knowing what your own store's content is actually giving AI to work with before restructuring anything.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This is the third and most direct test of the "format" signal in this series, and the first to measure citation fidelity instead of selection.