Format has lost twice already in this series on the question of who wins a forced comparison: rating swamped it in Winner vs Loser, and its own marginal contribution rounded to zero in Signal Hierarchy. Neither study asked what "structured markup helps AI understand your content" usually means in practice, whether structuring the same facts changes how much of them the model actually repeats back. This study holds three real facts about a brand exactly constant and only changes whether they arrive as a flowing paragraph or a bulleted list. Structured format did not increase citation rate or vocabulary reuse. It reduced both, by a small but statistically significant margin, replicated independently in two separate rounds.
Winner vs Loser found rating decided a forced two-brand comparison 90.8% of the time regardless of format. Signal Hierarchy tested authority, familiarity, specificity, and format against each other directly and found format's own marginal contribution to the winning brand rounded to zero in both independent rounds (-0.2pp, then 0.0pp). Both studies asked the same kind of question: does structuring the page change who wins a comparison. Neither ever asked whether structuring the same facts changes how much of them the model actually repeats back, specifically or in its own wording, when someone later asks it an open buyer question. Selection and citation are not the same outcome, and this series had only ever measured the first one.
This study is deliberately single-brand and open-ended, not a forced two-name comparison, borrowing the system-message injection mechanism from Fact Injection and Hidden Context instead, since the question here is what the model volunteers, not which of two options it names. No competitor appears anywhere in the design.
Here is the real system message for Hearthloom (handmade ceramic dinnerware) at both extremes of the design. The facts are word-for-word identical. Only the format changes.
User message, no brand name in it, only in the injected context: "I'm looking for handmade ceramic dinner plates. What would you recommend and why?" Each brand ran all 20 of its own real purchase intents, 5 times each, under both conditions, in both rounds.
Citation rate is the share of the 3 injected facts an LLM judge found cited or clearly paraphrased in the response. Vocabulary lift is the share of 5 pre-registered distinctive terms per brand (certifications, names, specific phrases) that show up verbatim. Both measures moved in the same direction, structured below prose, in both independently-run rounds.
| Metric | Condition | Round 1 | Round 2 |
|---|---|---|---|
| Citation rate | Prose | 65.3% | 63.8% |
| Structured | 59.5% | 59.2% | |
| Difference (structured − prose) | −5.83pp, p=0.0003 | −4.67pp, p=0.0011 | |
| Vocabulary lift | Prose | 67.7% | 68.9% |
| Structured | 64.5% | 64.4% | |
| Difference (structured − prose) | −3.25pp, p=0.0044 | −4.50pp, p<0.0001 | |
| Word count | Prose | 144.6 | 145.6 |
| Structured | 142.5 | 140.0 | |
| Brand mentioned | Prose / Structured | 99.8% / 100.0% | 99.8% / 100.0% |
Two obvious alternative explanations were checked directly. First, response length: structured responses were not longer than prose, they were if anything slightly shorter (141.3 words average vs. 145.1 for prose, both rounds combined), so the gap cannot be explained by prose simply giving the model more room to cite things. Second, brand mention: both conditions sat at or above 99.8% brand-mention rate in both rounds, essentially at ceiling, so the citation-rate gap is not a byproduct of structured responses failing to name the brand at all, it shows up specifically in whether the model repeats the injected facts once the brand is already named.
A common piece of AI-visibility advice is to reformat product and brand content into bullets and lists because that is supposedly what gets cited. This series has now tested that claim three separate times on three different outcomes, and it has lost every time: format's marginal contribution to winning a forced comparison rounded to zero in Signal Hierarchy, rating swamped it entirely in Winner vs Loser, and now, on the specific outcome the "structure it for AI" advice is actually about, whether facts get cited and specific vocabulary gets reused, structured format performs measurably worse than plain prose carrying the identical information. The effect here is small, five points on citation rate, four on vocabulary lift, not a reason to panic about existing bulleted content. But it is a real, twice-replicated result in the opposite direction from what the popular advice assumes, and citation counts and format-driven visibility claims common elsewhere online deserve the same scrutiny this series has been applying to authority, familiarity, and specificity.
All injected facts are synthetic (disclosed-synthetic authority mention, synthetic familiarity claim), the same convention as the rest of this series, not scraped from real press or real market-recognition data. Only 4 brands were tested, the same acknowledged category-coverage limitation as every study in this series. Single-turn, gpt-4o only, a single injection mechanism, system-message "retrieved context" framing, not tested on live web search or on a model actually calling a retrieval tool. The 5 vocabulary terms per brand are a judgment call about what counts as distinctive, fixed before any data collection, but a judgment call nonetheless, not an exhaustive or algorithmically derived list. Repeats=5 per cell follows this series' standard, not a power calculation specific to this study, the same open item flagged against the series-level statistical-power audit.
Both citation rate and vocabulary lift favored prose over structured in round 1 and again independently in round 2, with p-values under 0.005 in three of the four round-level tests and under 0.0001 in one.
Structured responses were not longer than prose responses, both rounds combined. The citation-rate gap cannot be explained by prose simply having more room to work with.
If reformatting for "AI readability" can make citation slightly worse, not better, it is worth knowing what your own store's content is actually giving AI to work with before restructuring anything.
This is the third and most direct test of the "format" signal in this series, and the first to measure citation fidelity instead of selection.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →