Here's the idea we started with: if a brand's own facts are written clearly, the model should recommend it more often. If the same facts are jumbled or buried, it should recommend it less. We tested this five different ways: on 9 well-known brands, on a harder set of prompts, with real competitors named in the same prompt, on 3 small real brands almost nobody has heard of, and on one brand we made up from nothing. Almost every time, the result was the same: once a brand's facts are placed in front of the model right before a matching question, the model recommends that brand, no matter how clearly or badly those facts are written. Only one brand broke this pattern, in three separate tests, and it broke in the opposite direction we expected.
A reader asked us a simple question: if a model understands a brand's identity more clearly, does it recommend that brand more often? To test this, we took 9 well-known brands and reused facts we had already fact-checked in an earlier study. We did not invent any new facts. For each brand, we wrote the same real facts three different ways. In the clear version, the brand's category and audience come first, and other details come after. In the ambiguous version, the same facts are listed with no order or hierarchy. In the misaligned version, a minor detail, like a certification or a price point, comes first, and the fact that actually defines the brand's category comes last. Nothing about what's true changes. Only the order changes.
We ran two tests per brand per version. The first, Understanding, checked whether the model could correctly state the brand's category and target customer just from reading the framing. The second, Candidacy and Selection, checked whether the brand got mentioned at all, and whether it got named first, when we asked a real, already-published buyer question. The Understanding test came back clean: the model got the category right 93 to 96% of the time, and the audience right 96 to 100% of the time, across all three framing versions. It reads clear, ambiguous, and misaligned writing about equally well. It was the Candidacy test that turned into a four-round investigation.
The first real run was 270 calls, across 9 brands and 3 framing versions. Almost nothing moved. Eight of the nine brands sat at a flat 100% candidacy rate in every single version. Clear, ambiguous, and misaligned framing made no difference at all for those eight. Only one brand moved: 40% candidacy under clear framing, versus 90% under both ambiguous and misaligned, the opposite direction we expected. Nearly all of the pattern in the pooled numbers below, where ambiguous and misaligned appear to beat clear, comes from that one brand alone.
The obvious first explanation was the buyer question itself. Each brand's original question was picked because it already had a high, near-ceiling mention rate on its own. So for 8 of the 9 brands, we swapped in a different, harder buyer question, one that on its own, with no brand facts attached, had only a 25% to 60% mention rate. (We kept the ninth brand on its original question, since it had already shown a real effect.) If framing needed more room to matter, this should have created it.
It didn't. All 8 brands went right back to 100% candidacy under ambiguous and misaligned framing, and 92.2% even under clear. How hard the question was on its own turned out not to matter, once the brand's own facts sat in the prompt right before that question, the model recommended the brand almost automatically, regardless of how "hard" that same question tested by itself. The ninth brand, run again on its original question, showed the same reversed pattern as before: 30% under clear framing versus 100% under both ambiguous and misaligned.
The next explanation: maybe candidacy was never a real choice. The prompt only ever described one brand, so of course the model mentioned it, there was nothing else to pick. To fix this, we named each brand's own real, already-known competitors in the same prompt. Those competitor names stayed identical across all 3 framing versions for a given brand, only the framing of the main brand's own facts changed. Now candidacy had to be a real choice among named options.
Five of the nine brands still landed at a flat 100% candidacy and 100% selection in all 3 versions, even against named rivals. Three more stayed at 100% candidacy but showed small, inconsistent wobble in selection, on a sample of 10 calls per cell, that's noise, not a real signal. One brand, tested this third separate way, produced its largest and cleanest effect yet.
One explanation was left. The 9 well-known brands show up constantly in "best brand" roundups and reviews. Maybe the model already knows enough about these specific brands that our framing barely registers next to what it already knows. So we ran the same named-rival test on 3 real, small, little-known brands. We checked online first and confirmed none of them show up in any "best brand" roundup, unlike every one of the original 9.
These 3 brands covered organic baby clothing, body-piercing jewelry, and Cuban-style hot sauce, 90 real calls total, real facts, real named rivals. The result: 100% candidacy and 100% selection, for all 3 brands, in all 3 framing versions. That rules out the fame explanation completely. A brand almost nobody has heard of behaved exactly like a household name. So we ran one more test, the most decisive one: a brand that does not exist. We invented a ceramics brand from scratch, with made-up facts and made-up rival names. It has zero possible history for the model to know about, by definition. It also landed at 100% candidacy and 100% selection, in all 3 versions.
We checked online first and confirmed none of the 3 brands show up in any "best brand" roundup. All 3 still hit 100% candidacy and 100% selection, in every framing version. Prior fame was not the reason for the ceiling.
We made this brand up entirely, so it has no possible history for the model to know about. Same test, same 100% candidacy and selection. This settles it: the ceiling comes from how the prompt is built, not from the brand.
Once a brand's facts sit in front of the model, right before a matching question, the model recommends that brand almost automatically. It doesn't matter if the brand is a household name, a small company almost nobody has heard of, or one we invented an hour before the test. Across 12 of the 13 brands we tested, using 4 different setups, how clearly that brand's identity was written made almost no measurable difference. So this study ended up answering a different question than the one we started with. It's not really about whether clear writing helps a brand win. It's about how low the bar for getting recommended actually is, once a brand's facts reach the model at all. We've seen hints of this same pattern in two earlier studies on this site, and this confirms it again from a new angle.
The one exception matters precisely because it's an exception. That one brand's real facts create a genuine conflict with its test question: a moderately premium product, tested against a question written around a tight budget. When that kind of conflict exists, how clearly the facts are written stops being irrelevant, and starts deciding the outcome, in the opposite direction plain intuition would predict. The cleaner, more organized version of the brand's premium story got it ruled out more often, not less. For most brands, most of the time, simply getting your facts in front of the model is the whole game. But if your brand's real position cuts against what a buyer is actually asking for, how you organize those facts can decide which way it breaks.
This is not a claim that framing never matters anywhere. It's a claim that framing didn't matter for 12 of the 13 brands we tested, under this one specific setup: a brand's facts placed in a system prompt, right before a matching-category user question. A different setup, like mentioning the brand mid-conversation instead, or a prompt that compares several brands at once instead of an open buyer question, could behave differently. We haven't tested those. Round 4's group of little-known and made-up brands is intentionally small, 3 real brands plus 1 made-up one. It was built to rule out one specific explanation (prior brand fame), not to estimate a rate across a wide sample of brands.
The rival-accuracy numbers we reported in Rounds 1 and 2 (whether the rivals the model names on its own match a brand's real rivals) stop being meaningful once Round 3 hands the model real rival names directly, so we only report that number for the first two rounds. We used one model (gpt-4o) throughout; we did not test other models. The one-brand exception, while confirmed three separate ways, is still a single brand's story. The price-versus-budget explanation behind it is grounded in reading the actual model responses, but confirming it holds more broadly needs new brands with a similar real conflict, not more tests of the same one brand.
If simply getting your facts in front of the model is most of what matters, then what's already in your product pages, reviews, and structured data is doing more work than how polished your pitch sounds. A free scan shows what the model currently has to work with.
This study reuses facts verified in Possession vs Deployment and Accuracy + Depth, and extends the same "what actually gets a brand recommended" question this series keeps returning to.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →