We took the most-recommended brands across five e-commerce categories and asked what separates the brands AI names from the ones it ignores. We first answered: fame. That answer was wrong, and this page now says so. Re-measured properly, neither store quality nor public fame predicts recommendation. What does still hold is that the recommendations are stable, which makes the emptiness stranger, not less.
The headline of this study was wrong. We reported that public fame explains 24.9% of how often AI recommends a brand, twelve times more than store quality. That figure was measured on 87 brands, and those 87 were not a random sample, they were the brands that a broken brand-to-store join happened to match.
We rebuilt the join and re-ran the identical regression on 872 brands. Fame explains 1.2% (adjusted 0.7%), which sits on top of the permuted null. Store quality explains 0.7%. The twelve-to-one ratio is gone: both are indistinguishable from noise.
What survives: the stability finding, which never depended on that join. What dies: fame as an explanation. We are leaving the mistake documented in place rather than quietly editing it out.
In the first part of this series we captured 20,000 AI product recommendations across beauty, supplements, coffee, pets, and home and living, and matched every one to the measured AI-readiness of the real store behind the brand. Across all five categories the relationship between how AI-ready a store is and how often AI recommends it was statistically indistinguishable from zero.
That answered the first question and raised a harder one. If a clean, well-structured, machine-readable store is not what earns the recommendation, then what is? This study is the attempt to answer it with the same discipline: real numbers, no invented deltas, and an honest account of what we could and could not measure.
What we originally published: we took the 200 most-recommended brands, measured public fame signals that do not depend on the store at all, such as Wikipedia readership, language editions, article length, and brand nameability, ran a multiple regression against recommendation frequency, and found fame explained 24.9% against store quality's 2.1%. Roughly twelve to one.
Why it was wrong. That regression ran on 87 brands, not 200, because only those had both a fame signal and a matched store. We treated that as an apples-to-apples subset. It was not. Those 87 were selected by a brand-to-store join built in June, and when we rebuilt the join in July we found it had matched only a fraction of the brands, and not at random. The sample was an artifact of a bug.
What the same regression gives on the corrected data. Same four predictors, same outcome, same method. 872 brands instead of 87:
Bar height is zoomed to the 0 to 1.5 percent range so the comparison is visible; all three values are genuinely this small. Fame lands exactly on the noise floor.
Share of recommendation frequency explained (R²), re-measured on 872 brands, July 2026. When we shuffle the outcome at random and re-run the identical regression, noise alone produces an R² of 1.2% at the 95th percentile, exactly what public fame produces.
Regression can feel abstract, so here is the same truth in plain terms. We sorted the 200 brands by how often they are recommended and compared the top 50 against the bottom 50. The top group is recommended about six times more often. You would expect their stores to be far better. They are not.
Left panel: a one-point gap in store quality, basically identical. Right panel: the top group is twice as likely to have a Wikipedia article.
It would be comforting to think the unexplained majority is just noise, that AI picks brands more or less at random and there is nothing to understand. We tested that directly. Every shopping question was run twenty times, and we measured how often the answer changed.
It almost never does. The same question puts the same brand in the top spot between 78 and 91 percent of the time, and across twenty runs only two or three brands ever reach first place. The brand-level variation across runs is effectively zero.
Share of runs that return the same brand in first place, by category.
Store quality does not matter. A cleaner, more machine-readable store does essentially nothing for how often AI recommends you. This held, and got stronger.
Neither does our fame measure. Sits exactly on the noise line. We reported 24.9% from a sample a bug selected. That claim is withdrawn.
And it is still stable. The same brands win again and again. We just no longer have any idea what they win on.
What went wrong, stated plainly. The original regression ran on the 87 brands that had both a fame signal and a matched store. We called that apples-to-apples. It was not: the match came from a brand-to-store map built in June that turned out to have matched only a minority of brands, and not at random. Rebuilding the map produced 872 usable brands, and on those the fame effect vanished into the noise. We also reported raw R² where adjusted R² was the honest figure, and we never ran a permuted null, so we had no idea what noise alone produces at n=87 with four predictors. It produces a lot. Two of those three errors are the same error: we let a number through without asking what it would look like if nothing were there.
Limitations that were always here. Wikipedia is a rough proxy for fame, not a perfect one. It captures encyclopedic prominence, while the fame that actually drives recommendation is more likely commercial presence across reviews, listicles, and forums. Two signals we believe matter, advertising spend and reliable web or Reddit mention volume, are not freely or credibly obtainable, so they are excluded by design rather than estimated. The top-50 versus bottom-50 comparison still rests on 30 and 18 brands from the old broken map and has not been re-measured; it is flagged in place rather than removed. The individual weights inside the fame regression should not be read on their own, because the fame signals overlap; the valid figure is the combined R². Every number here comes from a real query, never an assumed delta, which is exactly why the 24.9% survived as long as it did. It was a real number from a real run. It was just measured on a sample a bug chose.
The market measures whether AI can see you. We measure why AI chooses you, or why it chooses someone else instead. Get a Recommendation Intelligence read on your category.
This study is one of seven that isolate a single variable behind the flagship finding that store readiness does not predict recommendation.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →