The last study in this series showed that a model can know six facts about a brand and use, on average, just one, most of what's stored never gets said. This is the causal follow-up: what happens if you stop waiting for the model to remember, and instead put one real, verified fact directly in front of it, as if a retrieval step had just succeeded. Across 4 brands that barely get recommended today, doing that raised the mention rate by an average of 77.9 percentage points. But the size of that lift, and how much of it comes from the model actually citing the fact versus just being nudged to say the name, differs brand by brand in a way that matters.
Possession vs Deployment measured a passive gap: facts a model already has in memory, and how often those facts show up unprompted. This study manipulates the same mechanism directly. Four brands were chosen specifically because they barely get recommended today, real baseline recommend rates from 0% to 16.5% in the published Recommendation Reports. For each, one real, verified, distinctive fact was identified and injected as a system message reading "Additional context retrieved for this query: {brand}: {fact}", simulating the moment a retrieval step has just succeeded, then the model was asked the brand's own already-published real buyer questions, unmodified, and the response was checked for whether the brand appears at all, and, on a capped sample of the cells where it does, whether the specific injected fact is actually cited or paraphrased, not just the brand name dropped in.
The baseline side of the comparison is entirely reused, the same already-published 20-prompt buyer-question data behind each brand's Recommendation Report. Only the injected condition (10 repeats per prompt) and the fact-usage scoring are new.
Three of the four brands went from barely-mentioned to near-universal: 1.25% to 97.5%, 5.75% to 97%, and 16.5% to 96%. The fourth, Brand D, the one with a true 0% baseline, moved much less, to 44.5%. That's still a real lift, but it's the smallest in the cohort by a wide margin, and the likely reason is visible in the response text itself: Brand D's injected fact is less directly relevant to the specific buyer questions it was tested against than the other three brands' facts, so the model is less willing to work it into an on-topic answer. A retrieved fact only helps as much as it actually fits the question being asked.
A bigger question than the lift itself: when the brand gets mentioned, is the model actually using the fact it was handed, or just nudged to say the name? On a capped sample of mentioned cells, a judge call checked whether the specific injected fact (or a clear paraphrase) shows up in the response. The lift-to-fact-usage ratio, mention-rate lift divided by fact-usage rate, is the diagnostic: a ratio near 1 means the brand only gets mentioned about as often as the fact gets cited, a tight coupling. A ratio well above 1 means mentions are outrunning fact citations, the brand is getting named more from the framing alone than from anything specific being said about it.
Possession vs Deployment showed the bottleneck isn't whether a model knows something about a brand, it's whether that knowledge reaches the specific sentence being generated. This study shows that bottleneck is fixable, at least in the narrow sense tested here: put the fact where the model can see it at the moment of the answer, and mention rate moves dramatically, even for brands that are functionally invisible today. That's the practical version of AIVO's Linkage Gap versus Reasoning Gap distinction, a Linkage Gap is fixable by getting information into context; a Reasoning Gap, where the model has the information and still doesn't act on it, is not. Every brand here showed lift outpacing fact-usage, the signature of a Linkage Gap being closed by simple presence in context, not a structural reasoning failure being overcome.
Stating the limits up front. This study injects the fact directly, it does not measure how often a real retrieval system would actually surface this exact fact for this exact query, that's a separate, harder question this doesn't answer. The judge that scores fact-usage runs on the same model that generated the responses, a self-grading risk. The mitigation is a manual spot-check: 15 judge calls checked by hand against the raw response text. 14 of 15 matched. The one disagreement was a likely false negative, a response that closely paraphrased the injected fact but wasn't marked as using it, which if anything means the fact-usage rates reported here are a slight underestimate, not an inflated one.
Four brands, deliberately selected as underperformers, not a representative cross-brand sample, this is a floor-case study by design. 10 repeats per prompt in the injected condition versus 20 in the reused baseline, the baseline's CI is already tight from the larger published sample. One category mix, consumer ecommerce, one model. Cross-platform and cross-model versions of this question are a separate, larger, and currently API-key-blocked follow-up.
15 judge calls checked by hand against the raw response text, spanning all 4 brands. 14 matched exactly. The one disagreement was a conservative false negative, a close paraphrase marked as not using the fact, biasing the reported fact-usage rates slightly downward, not upward.
Every brand tested showed mention-rate lift outpacing fact-usage rate, a consistent Linkage Gap signature across the whole cohort, not a result driven by one outlier brand.
If putting the right fact in front of a model can move mention rate this much, the fastest way to find out which facts about your brand are missing at the moment it matters is a free scan.
This is the causal half of the possession-deployment question, and sits next to the within-conversation displacement result measured one study earlier.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →