Recommendation Intelligence Research™ · The Confabulation Test

The model explains every pick. It never explains it the same way twice.

If you cannot find a signal that predicts recommendation, the next question is obvious. Just ask the model why it picked what it picked. We did, 29,633 times. 90% of the answers were completely unique text. The most repeated answer showed up 12 times. There is no rule in there to find. Only a new sentence, every time.

29,633
Reasons checked
26,812
Completely unique
90%
Uniqueness rate
12
Max times any reason repeated
Where this comes from

The obvious question we had not asked

Every measurable predictor in this series has come up empty. Store quality explains 0.7% of recommendation. Public fame explains 1.2%. Knowing where a brand's website lives explains nothing about whether the model bothers to recommend it. At some point the obvious move is to stop guessing at predictors and just read what the model itself says.

Every recommendation in our dataset comes with a short reason attached, a line the model writes to justify its own pick. If that line reflects a real rule the model is applying, the same rule should produce recognizably similar language across similar picks. If it does not, the reason is not a rule. It is something written after the fact to make the pick look justified.

The test

Read every reason. Count what repeats.

We took the full reason field from the recommendation dataset, all 29,633 entries, and checked how many were exact duplicates of another entry in the set. If the model applies a consistent internal rule, that rule should show up as the same handful of phrasings recurring across thousands of picks.

Reasons checked: 29,633
Distinct reasons: 26,812
Uniqueness rate: 90%
Most repeated reason: 12 occurrences
Scoring: exact-text match
Source: reason field, recommendation dataset
What this method cannot see. We counted exact-text duplicates, not paraphrases. Two reasons that say the same thing in different words still count as distinct here. If anything, that makes 90% a floor, not a ceiling. The true rate of the model saying something genuinely new could be even higher than what we measured.
The finding

No repertoire. A fresh sentence every time.

Of 29,633 reasons, 26,812 were completely unique. That is 90% of every reason the model ever gave. The single most repeated reason in the entire dataset showed up 12 times, out of nearly thirty thousand chances to repeat itself.

How often does the model reuse a reason?
29,633 reason strings, checked for exact-text repetition
90% unique text 26,812 reasons, 90% appear exactly once in the dataset 2,821 reasons, 10% repeat text used somewhere else Most repeated reason: 12 times out of 29,633 total reasons, that is 0.04% of the dataset

A model with a real, applied rule for why it picks what it picks would keep reaching for the same handful of justifications. Instead, out of 29,633 opportunities to repeat itself, it managed to do so meaningfully only a handful of times, and even then, never more than 12 times for any single phrasing. There is no repertoire to find. There is a new sentence, written on demand, every single time.

Why it matters

The reason is written after the pick, not before it

This is what confabulation looks like in a language model. The pick happens first, produced by whatever weighting of memory, training signal, and pattern completion actually drives the recommendation. The reason is generated afterward, as a plausible-sounding sentence that fits the pick it already made. It is not a report on the mechanism. It is a caption written for it.

That has a direct, practical consequence for anyone trying to reverse engineer why AI recommends one brand over another: asking the model why does not work. You can ask it, and it will always answer, fluently and confidently, with something new every time. But a fluent answer to a why question is not evidence that the model knows why. It is evidence that the model is good at writing justifications on request, which is a different skill entirely.

What this means: the reason column cannot be used to explain, predict, or optimize for recommendation. It is not the rule that produced the pick. It is text generated to sound like one. Anyone building a store description strategy around the language the model uses to justify recommendations is optimizing for a caption, not a cause.
Stop guessing at what AI is looking for. Measure it.
Get your free AI Commerce Score™ in 10 seconds and see exactly where your store stands on the signals that actually move recommendation, not the reasons the model writes after the fact.
Get My Free AI Score