Recommendation Intelligence Research™ · Study #2

Neither store quality nor fame explains AI recommendations

We took the most-recommended brands across five e-commerce categories and asked what separates the brands AI names from the ones it ignores. We first answered: fame. That answer was wrong, and this page now says so. Re-measured properly, neither store quality nor public fame predicts recommendation. What does still hold is that the recommendations are stable, which makes the emptiness stranger, not less.

872Brands re-measured
1.2%Explained by fame
0.7%Explained by store quality
78-91%Still stable
Correction · 17 July 2026

The headline of this study was wrong. We reported that public fame explains 24.9% of how often AI recommends a brand, twelve times more than store quality. That figure was measured on 87 brands, and those 87 were not a random sample, they were the brands that a broken brand-to-store join happened to match.

We rebuilt the join and re-ran the identical regression on 872 brands. Fame explains 1.2% (adjusted 0.7%), which sits on top of the permuted null. Store quality explains 0.7%. The twelve-to-one ratio is gone: both are indistinguishable from noise.

What survives: the stability finding, which never depended on that join. What dies: fame as an explanation. We are leaving the mistake documented in place rather than quietly editing it out.

Where this comes from

First we proved store quality does not drive recommendation

In the first part of this series we captured 20,000 AI product recommendations across beauty, supplements, coffee, pets, and home and living, and matched every one to the measured AI-readiness of the real store behind the brand. Across all five categories the relationship between how AI-ready a store is and how often AI recommends it was statistically indistinguishable from zero.

That answered the first question and raised a harder one. If a clean, well-structured, machine-readable store is not what earns the recommendation, then what is? This study is the attempt to answer it with the same discipline: real numbers, no invented deltas, and an honest account of what we could and could not measure.

The headline finding, retracted

Fame does not outweigh store quality. Neither one works.

What we originally published: we took the 200 most-recommended brands, measured public fame signals that do not depend on the store at all, such as Wikipedia readership, language editions, article length, and brand nameability, ran a multiple regression against recommendation frequency, and found fame explained 24.9% against store quality's 2.1%. Roughly twelve to one.

Why it was wrong. That regression ran on 87 brands, not 200, because only those had both a fame signal and a matched store. We treated that as an apples-to-apples subset. It was not. Those 87 were selected by a brand-to-store join built in June, and when we rebuilt the join in July we found it had matched only a fraction of the brands, and not at random. The sample was an artifact of a bug.

What the same regression gives on the corrected data. Same four predictors, same outcome, same method. 872 brands instead of 87:

What explains how often a brand is recommended?
Noise floor · 1.2%
0.7%
1.2%
1.2%
Store quality
Public fame
Pure noise

Bar height is zoomed to the 0 to 1.5 percent range so the comparison is visible; all three values are genuinely this small. Fame lands exactly on the noise floor.

Share of recommendation frequency explained (R²), re-measured on 872 brands, July 2026. When we shuffle the outcome at random and re-run the identical regression, noise alone produces an R² of 1.2% at the 95th percentile, exactly what public fame produces.

One precision that matters. This kills the claim that our fame measurement predicts recommendation. It does not prove fame is irrelevant. Wikipedia was always a rough proxy, and we said so in the original limitations: it captures encyclopedic prominence, not commercial presence across reviews, listicles and forums. So the honest statement is that we have no evidence fame drives recommendation, not that it doesn't. The difference matters, and we are not going to blur it to save the headline.
The same finding, in plain brands

The most-recommended and least-recommended brands have nearly identical stores

Regression can feel abstract, so here is the same truth in plain terms. We sorted the 200 brands by how often they are recommended and compared the top 50 against the bottom 50. The top group is recommended about six times more often. You would expect their stores to be far better. They are not.

Top 50 vs bottom 50 most-recommended brands
Store quality score (0 to 100)
≈ same
50.8
49.8
Top 50
Bottom 50
Has a Wikipedia article
2× more likely
56%
28%
Top 50
Bottom 50

Left panel: a one-point gap in store quality, basically identical. Right panel: the top group is twice as likely to have a Wikipedia article.

Treat this chart with suspicion too. The store-quality half rests on 30 brands in the top group and 18 in the bottom, and those came from the same broken join as the retracted 24.9%. The one-point gap may well be real, and it points the same way as the corrected regression (store quality doesn't separate winners from losers). But we have not re-measured it, so we are flagging it rather than leaning on it.
The part most people get wrong

These recommendations are not random. They are stable

It would be comforting to think the unexplained majority is just noise, that AI picks brands more or less at random and there is nothing to understand. We tested that directly. Every shopping question was run twenty times, and we measured how often the answer changed.

It almost never does. The same question puts the same brand in the top spot between 78 and 91 percent of the time, and across twenty runs only two or three brands ever reach first place. The brand-level variation across runs is effectively zero.

How consistent is the top recommendation?
Home & Living
78%
Beauty
84%
Pets
85%
Coffee
86%
Supplements
91%

Share of runs that return the same brand in first place, by category.

This part survives the correction. Stability was measured directly from the recommendation runs and never touched the brand-to-store join, so the bug that killed the fame figure does not reach it. And it makes the corrected picture stranger, not weaker: the same brands win again and again, reliably, and neither their store nor their public prominence explains why.
The conclusion

Three measurements, and no answer

0.7%

Store quality does not matter. A cleaner, more machine-readable store does essentially nothing for how often AI recommends you. This held, and got stronger.

1.2%

Neither does our fame measure. Sits exactly on the noise line. We reported 24.9% from a sample a bug selected. That claim is withdrawn.

78-91%

And it is still stable. The same brands win again and again. We just no longer have any idea what they win on.

What this means for a brand: if you are not already inside the winning set, you are not invisible by accident or by bad luck on a given day. You are invisible consistently, and after correcting our own work, we have to state plainly that we cannot tell you why. Not store quality. Not any public fame signal we can measure. Anyone selling you a route into AI recommendations is selling something they have not measured. The only thing that can be measured today is where you actually stand.
Method & honesty

How we measured it, and what we could not

Brands analyzed: 872 (corrected). Originally 87, a broken join
Recommendation data: 20 runs per intent, gpt-4o-mini
Fame signals: Wikipedia views, language editions, article length, name length
Entity matching: Wikidata, commercial entities only, humans excluded
Models used: multiple regression, adjusted R², permuted null
Stability: top-1 consistency across 20 runs per intent

What went wrong, stated plainly. The original regression ran on the 87 brands that had both a fame signal and a matched store. We called that apples-to-apples. It was not: the match came from a brand-to-store map built in June that turned out to have matched only a minority of brands, and not at random. Rebuilding the map produced 872 usable brands, and on those the fame effect vanished into the noise. We also reported raw R² where adjusted R² was the honest figure, and we never ran a permuted null, so we had no idea what noise alone produces at n=87 with four predictors. It produces a lot. Two of those three errors are the same error: we let a number through without asking what it would look like if nothing were there.

Limitations that were always here. Wikipedia is a rough proxy for fame, not a perfect one. It captures encyclopedic prominence, while the fame that actually drives recommendation is more likely commercial presence across reviews, listicles, and forums. Two signals we believe matter, advertising spend and reliable web or Reddit mention volume, are not freely or credibly obtainable, so they are excluded by design rather than estimated. The top-50 versus bottom-50 comparison still rests on 30 and 18 brands from the old broken map and has not been re-measured; it is flagged in place rather than removed. The individual weights inside the fame regression should not be read on their own, because the fame signals overlap; the valid figure is the combined R². Every number here comes from a real query, never an assumed delta, which is exactly why the 24.9% survived as long as it did. It was a real number from a real run. It was just measured on a sample a bug chose.

Find out where your brand stands

Free AI Commerce Score™ in 10 seconds.

The market measures whether AI can see you. We measure why AI chooses you, or why it chooses someone else instead. Get a Recommendation Intelligence read on your category.

Free · No signup · Results in 10 seconds
Explore the series

More mechanism studies

This study is one of seven that isolate a single variable behind the flagship finding that store readiness does not predict recommendation.