Recommendation Intelligence Research™ · Candidacy vs Selection

Every store we tested was already in the game. We never checked who doesn't make it.

Every correlation in this series so far, including the ones that found nothing, was measured only among brands the model already recommends. That is a range-restricted sample. We had never asked whether a better store gets recommended at all. So we tested the full population: 599 stores that appeared in at least one recommendation against 60,325 that never appeared in any. One factor separates the two groups. It still predicts nothing about who wins once you are in it.

60,924
Stores, full population
599
Ever recommended
55.4 vs 56.0
Total score, recommended vs never
+39%
Intent gap, the one real separator
Where this comes from

The catch: we only ever looked at brands already in the room

Every prior test that correlated store quality with recommendation frequency ran the comparison among brands the model already recommends. That answers whether a better store, once it is being recommended, gets recommended more often. It never asks the more basic question: does a better store get recommended in the first place, or does it never even enter consideration?

Those are different questions with different samples. Restricting to brands already recommended can hide a real effect that only shows up at the gate, before frequency ever becomes relevant. So we dropped the restriction and compared the full population of scanned stores: everyone who ever appeared in a recommendation against everyone who never did.

The test

599 recommended. 60,325 never.

We split every scanned store into two groups based on a single question: did this store appear in at least one AI recommendation, ever, or not. Then we compared the AI Commerce Score and each of its components across the full population, no restriction to brands already winning recommendations.

Total stores scanned: 60,924
Recommended at least once: 599
Never recommended: 60,325
Compared: total score and 7 score components
Intent robustness check: 9 of 9 niches
Frequency correlation: r = 0.112, R squared = 1.2%
What we are not reporting. An earlier pass also compared raw review counts and average price between the two groups. We are leaving both out. A later data-quality check found review count populated for only 9.4% of the 66,085 scanned stores and average price for 66.7%, which means those two comparisons were mostly comparing empty fields, not real values. Only the score components below are computed for every scanned store regardless of raw-field coverage, so only those are reported here.
The result

Recommended stores are not better. They are marginally worse.

Across the full population, recommended stores scored 55.4 on average against 56.0 for stores that were never recommended. That is a real number now, not a restricted-sample artifact. The 0.7% correlation this series already reported was measured among recommended brands only. This test removes that restriction entirely, across 60,924 stores, and the gap still does not appear. If anything it points the wrong way.

Score components, recommended vs never recommended
Full population, 599 recommended stores against 60,325 that were never recommended
Recommended (599) Never recommended (60,325)
intent 5.13 3.70 visual 7.43 7.22 schema 3.83 3.53 technical 11.27 12.13 trust 10.42 11.06 price 8.25 8.94 brand 3.07 3.14 Recommended stores score higher on intent, visual, and schema, and lower on technical, trust, and price. Same total either way.
Recommended stores trade technical, trust, and price for intent. The total comes out the same. Intent is the only gap large enough to matter, roughly 39% higher for recommended stores. The rest move by single-digit percentages.
Testing the one real gap

Intent holds up under every check we ran

A single gap this size could easily be an artifact of category mix, circular scoring, or which brands happen to own their own store. We tested for all three, plus whether the gap actually predicts anything once a store is in the recommended set.

9 of 9
Niches show the gap, range +0.78 to +1.59 points. Not a category artifact.
0.20 to 0.29
Intent's correlation with the store's other score factors. A free-floating brand-recognition signal would not correlate with anything.
1.40 vs 1.53
Gap size in brand-owned stores versus independent ones, nearly identical. Not a brand-selection artifact.
R² = 1.2%
Once a store is in the recommended set, intent does not predict how often it gets picked, r = 0.112. Same magnitude as the fame study, on the null.
Why it matters

Store quality gets you candidacy, not selection

One store factor genuinely separates the brands that ever get recommended from the brands that never do, and it survives niche breakdown, circularity checks, and a brand-ownership split. That is real. But the same factor, measured only among brands already in the recommended set, predicts almost nothing about which of them wins more often. Getting into the game and winning it are governed by different things.

One confound is worth naming directly: 3.70% of Magento stores appear in the recommended group against 0.79% of Shopify stores, a 4.7x difference. That looks dramatic until you notice Magento stores in this dataset skew older and larger, the platform itself is not doing the work. It is a marker for something else, not a lever anyone can pull.

What correlation can never settle: direction. Brands with sharply positioned stores may be recommended because their positioning is clear, or their positioning may be clear because they are already successful enough to invest in it. Both stories produce the same correlation. Nothing in this data, or any observational data, can tell them apart.
Does your store even make the list?
Get your free AI Commerce Score™ in 10 seconds and see where your intent signal stands against the 599 stores that made it into AI recommendations.
Get My Free AI Score