The Fame Study found that Wikipedia pageviews, the single most obvious proxy for brand fame, explain only 1.2% of the variance in how often a model spontaneously recommends a brand with no search, no injected facts, closed book. That leaves 98.8% of "why does the model already prefer this brand" unexplained. This study widens the net: does a broader public footprint, cross-lingual reach on Wikidata, how long the brand's domain has existed, how much recent media has covered it, do any better? Across 95 brands and four free public-data signals, the answer is barely. Every signal on its own is statistically indistinguishable from noise. Combined, all four explain 11.2%, about five times a single Wikipedia number, and still a small fraction of the 61.4% a model's own past behavior explains about itself.
How AI Decides names Memory as Stage 1 of the decision path: what exactly enters a model's memory, how strongly a brand is represented there, and how that internal representation shapes a recommendation later. The Fame Study took a first pass at that question using the most obvious public proxy for fame, Wikipedia pageviews, and found it explains almost nothing, 1.2% of the variance in how often a brand gets recommended with no search and no injected facts. That result left the door open: maybe Wikipedia alone is just too narrow a lens. A brand can be well known without a large English Wikipedia article.
So this study widens the footprint to four free, keyless public signals: Wikipedia pageviews again for direct comparison, Wikidata sitelinks (how many language editions cover the brand, a broader reach signal than English Wikipedia alone), domain age (how long the brand's own web presence has existed), and GDELT media mention volume (how much recent press has covered it). It reuses the same closed-book Phase 1 baseline as Cold Start and Candidate Evaluation, no search, no injected facts, so the recommendation-frequency side of the correlation is identical in method to the rest of the series. Nothing is injected or manipulated here; this design is observational, correlating real brands' real public footprint against a real closed-book baseline.
Wikipedia pageviews explain 2.3% of recommendation frequency here (p=0.135). Wikidata sitelinks explain 0.9% (p=0.365). Domain age explains 2.2% (p=0.147). Media mention volume, the closest of the four, explains 5.2% (p=0.095). A label-shuffle permutation test (10,000 reshuffles) checked each one against chance pairing, and none clears the conventional p<0.05 bar. That means none of these four signals, taken alone, can be reliably distinguished from noise with this sample.
A multiple regression using all four signals together, over the 50 brands with complete data on every signal, explains 11.2% of the variance (95% CI 3.8-37.5%), roughly five times any single signal alone. That is a genuine improvement over one number, but it is worth seeing next to the two benchmarks this series has already established: the original Fame Study's single-signal 1.2%, and the-model-predicts-itself's 61.4%, how well the model's own past behavior predicts its future behavior. Even the best-case combination of everything measurable here lands closer to the first number than the second.
Widening from one easily-measurable public signal to four barely moved the needle. That reinforces, and sharpens, the throughline of this whole series: recommendation frequency is a property of the model's own memory, not of how well documented a brand is out in the world. It is not just that Wikipedia fame fails to explain it, cross-lingual reach, domain age, and recent media volume fail too, individually and even stacked together.
The practical read for a brand with no recommendation history yet: don't expect a bigger Wikipedia footprint, more press coverage, an older domain, or more language editions to move a model's spontaneous, closed-book recommendation on their own, none of these four levers, even combined, gets anywhere close to what the model's own prior behavior explains about itself. That narrows, rather than answers, "where does the memory come from," and it is the direct motivation for the shift toward causal, controlled experiments on the Founder Lab noted on the research roadmap: at some point, measuring more correlates stops being the fastest way to find out.
Stating the limits up front. Unlike Candidate Evaluation or Cold Start, nothing is injected or manipulated here. A correlation between a footprint signal and recommendation frequency does not establish that the signal causes the model to remember the brand, both could share a common cause, genuinely being a large, old, well-covered brand. Domain age is looked up at a guessed {brand}.com address, so wrong-domain or no-domain brands are dropped from that signal only, not imputed. GDELT mentions are capped at 250 records per query and were unavailable for 41 of 95 brands in this run, so it distinguishes obscure from some coverage, not a lot from enormous.
Wikipedia, Wikidata, and GDELT all skew English-language and Western-media, so a brand strong in a non-English market could show a falsely low footprint here. This is single-turn, gpt-4o only, one snapshot in time, scoped deliberately to one model for this run, the cross-model version (GPT vs Claude vs Gemini) is parked as a follow-up study. And brand-frequency estimates get noisier for brands with few wins, which is exactly what the robustness check below tests directly.
Robustness check: an earlier, stricter pass required each brand to win at least twice before counting it (75 of 95 brands qualified) instead of once. It found the same shape: Wikipedia 0.5%, Wikidata sitelinks 0.4%, domain age 2.7%, media mentions 3.5%, combined model 6.1%. Every number in that pass was smaller, not larger, than the fuller 95-brand run reported above, meaning the headline numbers here are, if anything, the more generous reading, not an inflated one. The conclusion does not change under either cut: individually negligible, combined still modest, both far under 61.4%.
This study's Wikipedia-pageviews-alone R² lands close to the original Fame Study's 1.2% baseline, well within the bootstrap confidence interval (0.0-11.7%). That's the internal sanity check the design calls for before trusting the other three signals' numbers, and it passes.
Re-running the analysis with a stricter brand-inclusion threshold (75 brands instead of 95) still finds no individual signal reaching significance, and a smaller combined model, 6.1% instead of 11.2%, still a fraction of the 61.4% self-consistency benchmark either way.
If public footprint alone barely moves a model's spontaneous recommendation, a free scan is a faster place to find real, actionable gaps than chasing press or Wikipedia edits.
This answers the Memory question flagged on the decision map. Read the map for where it fits, or the original Fame Study for the single-signal baseline this one extends.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →