We asked gpt-4o directly what it knows about 9 brands. It answered confidently every time, certifications, price positioning, specific product features, the kind of detail a brand would want a shopper to hear. Then we checked whether any of that showed up in the moment that actually matters: a real buyer question, no brand named, where the model brought that brand up on its own. Across 1,392 fact checks, 83.6% of the facts a model just claimed to know never appeared. Knowing about a brand and using what it knows are two different things, and most of what gets stored never gets said.
Every prior study in this series measures one call type: what does the model recommend, under one condition or another. This one measures two call types against each other. First, a possession probe: in a fresh conversation with no purchase context, gpt-4o was asked to list the most distinctive, checkable facts it knows about a brand, its products, its pricing, what sets it apart from named competitors. Run three times per brand and merged down to the claims that showed up, in substance, across at least two of the three runs, so a one-off phrasing doesn't get counted as a stable fact.
Second, deployment: did any of those facts actually show up when the model made a real recommendation? This is where the design gets cheap. The deployment side already existed, 400 real gpt-4o buyer-question responses per brand, collected for the published Recommendation Reports on Bellroy, Caraway, and the rest of the Wave-1 cohort, full response text and a brand-mentioned flag already saved for every one of 3,600 observations across 9 brands. Instead of re-collecting that at AIVO-scale cost, every cell where the brand was actually mentioned was scored against the brand's own possessed-fact list: does this specific fact show up, even paraphrased, in this specific response. A tenth brand, Topicals, was excluded outright, it was mentioned in 0 of its 400 observations, so there is nothing to check deployment against, a complete Candidacy failure rather than a Linkage Gap case.
Across the full cohort, 83.6% of possession-deployment pairs never appear (95% CI 78.3-88.6%, n=1,392). The deployment rate, the flip side of that number, is 16.4%: roughly one fact in six that the model just told us it knew actually made it into a real recommendation where the brand was mentioned. The gap holds up brand by brand, though the size of it varies a lot.
A natural guess: maybe the brands that get recommended most often are also the ones whose facts get used most reliably, winning and fact-deployment could be the same underlying mechanism seen twice. Correlating each brand's deployment rate here against its published recommend rate from the Recommendation Reports series gives r = -0.27 (p = 0.47, n = 9), no relationship, and if anything a weak negative one. At nine brands this is underpowered by design, not a confirmed null, but there is no visible sign that winning more and deploying more of your own facts are the same thing.
That is a genuinely useful negative result. It means a brand cannot assume that simply winning more often will also fix its fact-deployment rate, and a brand that is losing cannot assume its facts are the reason, the mechanisms look separable rather than the same lever measured twice.
This series has already shown, in Candidate Evaluation, that when a specific comparison fact is put directly in front of the model at the moment it is deciding, a better rating flips the outcome 100% of the time. That is the deployment side working exactly as expected, once a fact is placed at the decision point, it gets used. This study shows the other half: a fact the model already has, sitting in memory, unprompted, is used only about one time in six when the moment to use it actually arrives. The bottleneck was never whether the model knows something about a brand. Every brand here returned specific, largely accurate facts the instant it was asked. The bottleneck is whether that fact makes it from memory into the specific sentence being generated at the moment of recommendation.
This lands in a broadly similar range to an unrelated industry measurement of the same phenomenon published elsewhere in 2026 (a comparable, non-overlapping brand cohort reported a gap in the low-to-high 70s), which is worth noting as a rough sanity check on the scale of the effect, not as a claim of shared methodology.
Stating the limits up front. The same model, gpt-4o, runs the possession probe, produces the deployment response, and judges whether a fact shows up in that response. That is a real limitation, not a hypothetical one, a model could in principle be systematically lenient or strict with itself in ways a human or a different model would not be. The mitigation here is a manual spot-check: a random sample of judge calls was checked by hand against the underlying response text before publishing this number, and every one checked matched, a fact marked used was genuinely present, paraphrased or not, and a fact marked absent genuinely was not in the text. That is reassuring, not conclusive, on a sample this small.
"Six distinctive facts" is a fixed target, not a natural unit, a looser or tighter extraction prompt would shift the denominator and the rate along with it. The deployment-side data was collected one to three weeks before the possession probe, reused rather than paired in time, though gpt-4o's behavior is not expected to have shifted materially in that window. Nine brands, one category mix, consumer ecommerce, one model. Cross-platform and cross-model versions of this question are a separate, larger, and currently API-key-blocked follow-up.
A random sample of judge calls was checked by hand against the real response text they scored. All three matched exactly, a fact marked used was genuinely reflected in the text, a fact marked absent genuinely was not. Small sample, but zero disagreements before publishing.
Topicals was mentioned in 0 of its 400 buyer-question observations, so there was nothing to check deployment against. Rather than drop it silently, it stays out of the cohort by name, a total Candidacy failure is a different mechanism than a Linkage Gap.
If the model already knows more about your brand than it says, the fastest way to find out which facts are missing at the moment it matters is a free scan, not a longer knowledge-base article.
This answers a Memory question flagged on the decision map, and sits next to the injected-fact result in Candidate Evaluation.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →