Recommendation Intelligence Research™ · The Strongest Signal We Have

Nothing about your brand predicts recommendation. The model's own past behavior does.

Store quality explains 0.7% of who gets recommended. Public fame explains 1.2%. Web traces explain 0.2%. Intent, the one factor that separates recommended stores from the rest, explains 1.2% of frequency once you are already in. Four independent signals, four numbers within a rounding error of zero. Then we asked a different question: does what the model recommended last month predict what it recommends this month? That number is 61.4%, and it still has not moved after 15 days.

1,082
Brand-intent pairs, month to month
61.4%
R squared, position predicts position
0.2%
Weakest external signal, web traces
0 pts
Drift after 15 days
Where this comes from

We ran out of external things to test

By this point in the series we had tested store quality, public fame, and how intent separates recommended stores from stores that never get picked. All three landed close to zero: 0.7%, 1.2%, and 1.2%. One layer was still untouched: web traces, meaning how visible a brand is across the open web, not just its own store.

We measured it using distinct domains mentioning the brand, whether the brand's own site turns up at all, and mentions on review sites, comparing our most reliably recommended brands, the 301-brand core that survives every model and search setting we test, against everything else.

Distinct domains, core: 8.1
Distinct domains, rest: 7.9
Own site present, core: 58%
Own site present, rest: 44%
Review site mentions, core: 2.7
Correlation to frequency: r = 0.046, R squared = 0.2%
A handicap we are disclosing. The search API we used for this measurement rejects quoted phrase queries on its free tier, so every search ran unquoted, which is a looser and noisier match than we would prefer. Both groups took the same handicap, so the comparison between them still holds, but the absolute counts should be read as approximate.
Four signals, four zeros

Every external layer we have tested lands in the same place

Put the four correlations next to each other and the pattern is not subtle. Store quality, fame, web presence, and intent all sit within a point of zero. Then look at what happens when you stop asking about the brand entirely and ask about the model instead.

R squared against recommendation frequency, four external signals and one internal one
Each bar is how much of recommendation frequency that signal explains on its own
Store quality 0.7% Public fame 1.2% Web traces 0.2% Intent, within set 1.2% Position, month to month 61.4% Four independent brand signals land within one point of zero. Only the model's own past output clears sixty percent.
The test

Ask the model what it did last time

We took June recommendation data and July recommendation data, two genuinely separate periods pulled from different tables, spanning a model change in between, and matched them on 1,082 brand-intent pairs built from the same 50 intents. Then we correlated June's position and frequency against July's, using Pearson correlation on the raw values.

Brand-intent pairs: 1,082
Intents: same 50, both periods
Periods compared: June vs July
Method: Pearson correlation
Circularity: different periods, different tables
Model: changed between periods
Why this is not circular. This is not the same measurement correlated with itself. June and July come from different collection runs, stored in different tables, spanning a change to the underlying model. If June still predicts July under those conditions, the thing being measured is more stable than the pipeline that measured it.
The finding

The strongest number in this entire series

June's position predicts July's position at r = 0.784, R squared = 61.4%. June's frequency predicts July's frequency at r = 0.737, R squared = 54.4%. And June's position predicts July's frequency at r = negative 0.565, R squared = 31.9%, negative because a better position number, lower is better, lines up with a higher frequency, which is exactly what you would expect if position and frequency are two views of the same underlying stability.

How well June predicts July
Pearson R squared, 1,082 brand-intent pairs, same 50 intents both periods
Position → Position 61.4% r = 0.784 Frequency → Frequency 54.4% r = 0.737 Position → Frequency 31.9% r = -0.565 Across a month, a model change, and separate datasets. Not the same measurement correlated with itself.
Do not mix this with the 53.6% ceiling. That earlier number was Spearman, and Spearman's version of this same self-prediction test comes out at 91.3%, not 61.4%. So 61.4% is roughly two thirds of what is achievable by this measure, not further ahead of the ceiling than seems possible. Pearson and Spearman answer related but different questions and should never be reported as if they were the same statistic.
The companion finding

And it does not move for 15 days

A signal that predicts itself a month out is only useful if it also holds steady in between. We already knew that comparing the same sweep against itself produces about 44% turnover from ordinary noise, nothing to do with real change. We extended that same comparison out to 1 day apart, 14 days apart, and 15 days apart, on 234,283 scans across 57,242 domains.

Percent of top picks that changed, by how far apart the two scans were
Baseline is the same sweep compared with itself, three repeats shown as small points
100% 75% 50% 25% 0% 43.3% 44.6% 44.7% 44.4% within one sweep 1 day apart 14 days apart 15 days apart On a full 0 to 100 percent scale, four measurements spread across 15 days sit on top of one another.
Nothing moves in 15 days. The gap between the noise floor and every later measurement rounds to zero. Waiting longer between scans does not reveal drift because there is no drift to reveal in this window.
Why it matters

This kills monitoring as a product, and tells you what to sell instead

Put the two findings together. Recommendation is not explained by anything measurable about the brand in the world, not its store, not its fame, not its web presence. It is explained, at 61.4%, by what the model already did last time. And that state does not drift for at least 15 days. That is not a weak signal buried in noise. It is a stable, self-consistent, internal property of the model, and external brand attributes barely touch it.

The practical consequence is blunt: there is nothing to monitor day to day, because nothing changes day to day. A single, well-timed scan tells you what will still be true weeks later. Continuous tracking would be selling reassurance about a number that was never going to move.

Supporting evidence

Two more signs this is structural, not noisy

28.5 to 30.3
Top 3 brands' combined share of recommendations across 10 categories, coffee to food and snacks. A range of 1.8 points. The same concentration law holds everywhere we look.
53.4 vs 55.0
Total score, our most stable core brands versus everyone else recommended. The core has worse stores, not better, echoing the same reversal we found across the full population.
Your store might already be locked in, one way or the other
Get your free AI Commerce Score™ in 10 seconds. Since the signal that matters most is not your store attributes, this tells you where you actually stand.
Get My Free AI Score