Every correlation in this series so far, including the ones that found nothing, was measured only among brands the model already recommends. That is a range-restricted sample. We had never asked whether a better store gets recommended at all. So we tested the full population: 599 stores that appeared in at least one recommendation against 60,325 that never appeared in any. One factor separates the two groups. It still predicts nothing about who wins once you are in it.
Every prior test that correlated store quality with recommendation frequency ran the comparison among brands the model already recommends. That answers whether a better store, once it is being recommended, gets recommended more often. It never asks the more basic question: does a better store get recommended in the first place, or does it never even enter consideration?
Those are different questions with different samples. Restricting to brands already recommended can hide a real effect that only shows up at the gate, before frequency ever becomes relevant. So we dropped the restriction and compared the full population of scanned stores: everyone who ever appeared in a recommendation against everyone who never did.
We split every scanned store into two groups based on a single question: did this store appear in at least one AI recommendation, ever, or not. Then we compared the AI Commerce Score and each of its components across the full population, no restriction to brands already winning recommendations.
Across the full population, recommended stores scored 55.4 on average against 56.0 for stores that were never recommended. That is a real number now, not a restricted-sample artifact. The 0.7% correlation this series already reported was measured among recommended brands only. This test removes that restriction entirely, across 60,924 stores, and the gap still does not appear. If anything it points the wrong way.
A single gap this size could easily be an artifact of category mix, circular scoring, or which brands happen to own their own store. We tested for all three, plus whether the gap actually predicts anything once a store is in the recommended set.
One store factor genuinely separates the brands that ever get recommended from the brands that never do, and it survives niche breakdown, circularity checks, and a brand-ownership split. That is real. But the same factor, measured only among brands already in the recommended set, predicts almost nothing about which of them wins more often. Getting into the game and winning it are governed by different things.
One confound is worth naming directly: 3.70% of Magento stores appear in the recommended group against 0.79% of Shopify stores, a 4.7x difference. That looks dramatic until you notice Magento stores in this dataset skew older and larger, the platform itself is not doing the work. It is a marker for something else, not a lever anyone can pull.