Recommendation Intelligence Research™ · Study #18

Two months later, the model still agrees with itself.

Our earlier study found that a model's own recommendation last month explains 61.4% of what it recommends this month, more than store quality, brand fame, or web presence combined, measured over a 15-day window. We wanted to know if that holds over something longer and messier: six independently triggered runs, spread across two full months, not one clean before-and-after comparison. We tracked the same 50 shopping intents across 10 categories, from June 30 to August 26 2026, same model, same temperature, every time. In 86% of intents, the model named the exact same #1 brand in all six sweeps. For the seven that didn't fully agree, we doubled the sample size to find out whether we were looking at real change or ordinary noise.

86%Intents with an unchanged #1, every sweep
50Intents tracked, 10 categories
6Independent sweeps, June 30 – Aug 26
30,170Recommendations analyzed
Where this comes from

The first study measured 15 days. We wanted two months

The Model Predicts Itself showed that June's recommendation position predicts July's at R² 61.4%, and that the ranking does not move across a 15-day window. Both of those measurements came from controlled, single comparisons: two snapshots pulled and lined up against each other. This study asks the same underlying question a rougher way, closer to how the system actually gets used: repeated, independently triggered calls to the same model, spread across two months, with no single "before and after" pair to lean on.

We picked 50 shopping intents across 10 categories, from action cameras to ashwagandha supplements, and asked gpt-4o-mini to name its top 10 picks for each, ten times per intent per sweep, temperature held at 0.7 throughout. We ran that sweep six separate times over eight weeks. Then we asked one blunt question, harsher than a correlation: is the #1 pick the exact same brand, every single time?

Model: gpt-4o-mini, temperature 0.7
Intents: 50, across 10 categories
Sweeps: 6, June 30 – Aug 26 2026
Span: 58 days
Recommendations analyzed: 30,170
Concordance rule: same #1 brand, every sweep
A data issue we found and fixed. Two of the six sweeps briefly shared a single identifier with a same-day repeat run, which looked like conflated data until we traced every row back to its true execution timestamp and split them apart. Once separated, the two independent runs on those days agreed on the #1 pick for the large majority of intents, an additional replication within the same day, not just across weeks. The one disagreement we found this way, discussed below, turned out to be genuinely useful rather than a data quality problem.
The finding

Locked in 43 of 50 intents

Across six sweeps and eight weeks, 43 of 50 intents, 86%, returned the exact same #1 brand every single time. GoPro won "best action camera" in five of six sweeps at 10 out of 10 runs, and in the sixth at 1 out of 10, still first place. Lululemon, Levi's, Vital Proteins, Optimum Nutrition, and PetSafe never lost their #1 spot once, in any sweep, across two months.

Intents by outcome
50 tracked intents, June 30 – August 26 2026
86%
43 of 50
14%
7 of 50
Locked, same #1 every sweep
Contested, #1 changed at least once
This is a stricter test than 61.4%. The original number was a continuous correlation, position tracking position across a 1 to 10 scale. This one asks a binary, harsher question: is it the exact same brand at #1, every single time, for two months? A softer, correlation-based version of this same dataset would likely score higher still. Measured this bluntly, it holds up at 86%.
The sample-size check

Two of three "ties" turned out to be nothing

Seven intents did not return the same #1 in every sweep. Three of them looked genuinely close: a near-tie in one 10-run sweep, not the kind of runaway win we saw everywhere else. Rather than call that "drift," we doubled the sample size, 20 runs instead of 10, and re-ran all three the same day to see whether the tie was real or just noise at a small sample.

Position-1 share at 20 runs, three previously borderline intents
Same day, same model, same temperature · run to settle what a 10-run sample could not
Plus-size activewear · Athleta
60%
Plus-size activewear · Nike
25%
Eye cream, dark circles · Olay
50%
Eye cream, dark circles · Clinique
25%
Ceramic dinnerware · Fiesta
50%
Still contested
Ceramic dinnerware · Corelle
40%
Two false alarms, one real split. At 20 runs, plus-size activewear resolved cleanly to Athleta and eye cream for dark circles resolved cleanly to Olay, both had simply hit an unlucky 10-run sample on an earlier sweep. Ceramic dinnerware did not resolve. Fiesta and Corelle split 50/40 even at double the sample size, the one intent in this dataset where the model itself does not appear to have settled on a favorite.
Why it matters

Most categories are decided. A few are still up for grabs

Put the two findings together and the practical takeaway splits in two. For the 86% of intents that are locked, daily monitoring buys you very little, the same conclusion the original study reached: there is nothing to catch because nothing is moving. But the 14% that are genuinely contested are exactly the opposite case. Those are the categories where the model has not settled, where a well-timed piece of content or a clearer product page has somewhere real to land, because the #1 spot is not already spoken for.

The practical filter is simple: before spending on visibility work for a category, check whether it is locked or contested first. Fighting for a spot GoPro or Lululemon has held for two straight months is a different bet than fighting for a spot that is already splitting 50/40.

Supporting evidence

Two more signs this is structural

1.6 vs 7.4

Average position of category specialists like Lululemon, GoPro, and Optimum Nutrition, versus generalist retailers like Target and Amazon, which appear across more categories than anyone, 17 to 21, but average position 7.4 and rarely win #1.

2 of 3 were noise

Borderline #1 finishes that looked like they might be shifting resolved into clear, stable winners once we doubled the sample size. Only ceramic dinnerware held up as a genuine, ongoing split.

A note on what #1 sometimes means

Not every winner is a company

The most consistent #1 across the ashwagandha-related intents was not a company at all. KSM-66 is a patented ashwagandha extract licensed by dozens of supplement brands, not a business you can buy from directly. For intents like that, there is no single competitor to outrank, because the winning name is an ingredient spec, not a retailer.

We flag it here rather than folding it into the 86%, since it changes what "winning the #1 spot" would even mean for a brand trying to compete in that category. There is nothing to out-market, no store to outrank; the fix is a formulation choice, not a content one.

Find out if your category is locked or contested

Free AI Commerce Score™ in 10 seconds.

If your category is already decided, the fight is somewhere else. If it isn't, this is where you'd want to know.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This follows directly from the strongest number in the whole series. Read the original for the full method and the four null results it stands against.