We asked one AI model the same 50 buying questions twice. The only thing we changed was whether it could browse the web. 77% of the brands it recommended changed. This is a controlled look at the layer almost nobody measures: the difference between what AI recommends from memory and what it recommends from live retrieval.
A few weeks ago we published research showing that store quality explains almost nothing about which brands AI recommends. Public prominence explained far more. Most of the outcome stayed unexplained.
A data engineer named Rami read it and pushed back with a sharp idea. Maybe a chunk of that unexplained variance wasn't unexplainable at all. Maybe it was the retrieval layer: whatever the model's internal search decides to surface quietly decides the recommendation. He proposed a clean way to isolate it: run the same prompts with browsing on and off, and measure how much the answers overlap.
The design is deliberately boring. Change one thing, hold everything else still.
We also ran the whole thing against gpt-4o-mini, to separate the effect of search from the effect of using a bigger model. That turned out to matter more than we expected, and not in the way we first reported.
On the same model, the overlap between search-on and search-off recommendations was 23%. In other words, 77% of the recommended brands changed (95% CI 74 to 80%) when we turned web search on. Same model, same questions, same number of runs. One toggle.
Same model, gpt-4o. Same 50 buying prompts. 10 runs per prompt, per condition. The only variable is whether browsing was on or off.
That interval comes from a cluster bootstrap resampling the 50 intents rather than the individual brand-intent pairs, because brands cluster inside intents and the pairs are not independent. A naive binomial would have given a tighter interval, 75 to 79%. The wider one is the honest one.
With search off, the model recommends from memory. It names the brands it saw most often during training, which skews heavily toward the famous ones. With search on, it throws most of those out and recommends whatever its retrieval surfaces in the moment. These are not small adjustments to a stable list. They are two largely different lists.
The obvious objection to the headline: maybe gpt-4o simply recommends different brands than gpt-4o-mini, and we were measuring a model difference rather than a search effect.
So we ran gpt-4o with search off as a control. That comparison, gpt-4o with search versus gpt-4o without, is clean: same model, same run, one variable. It is where the 77% comes from, and it stands.
An earlier version of this piece claimed that changing the model moved only about 6% of recommendations, and concluded that which model you ask barely matters.
That 6% was never measured. We derived it, by subtracting two overlap figures from two different comparisons and treating the gap as the model effect. It is not the same quantity. We published arithmetic and called it a result.
Measured properly, gpt-4o-mini without search against gpt-4o without search, 66.9% of the recommended brands changed. Not 6%. Swapping the model rewrites roughly two thirds of the answer, close to what turning on search does.
The mini runs and the gpt-4o control runs were collected twelve days apart, so that figure mixes the model change with whatever drifted in between. We could not separate them from that data.
We have since re-run both conditions in the same window, which removes the ambiguity: the clean model effect is 68.9% (95% CI 66 to 71). The search comparison never had this problem, because both conditions ran together.
A better finding than the one we started with. Switch the model and roughly two thirds of the recommendations change. Turn on browsing and roughly three quarters change.
There is no such thing as what AI recommends. There is only what a particular model, in a particular configuration, at a particular moment, recommends.
Anyone selling you a position in AI recommendations should be asked which AI, in which mode, and on what day.
Every number above compares two conditions. But we never asked the obvious question: how different are two answers when nothing changes at all?
These models are non-deterministic. The same question, asked twice in the same second, does not return the same answer.
So we split a single sweep in half. Runs one to five against runs six to ten. Same model, same configuration, same moment. The only variable is the dice.
47% of the recommended brands were different. Nothing changed. Just asking twice.
Our first instinct was to subtract. If the floor is 47% and search changes 77%, then search is worth 30 points.
That is wrong, and it is the same mistake as the 6% earlier in this piece: a difference between two noisy measurements, dressed up as a quantity. It also assumes there is one floor, when in fact the floor is different for every comparison.
Pool the runs from both conditions. Reshuffle which run belongs to which label. Recompute the change rate. Repeat ten thousand times.
That gives a null distribution: what the change rate looks like when the condition label means nothing. The observed value either sits inside that distribution or it does not. No subtraction.
| Comparison | Observed | Null | p | Verdict |
|---|---|---|---|---|
| Turn on web search | 76.9% | 39.1% | <0.001 | Real |
| Swap the model | 68.9% | 54.5% | <0.001 | Real |
| Wait three days | 45.3% | 45.6% | 0.73 | Nothing |
10,000 label-shuffled resamples per comparison. p is the share of shuffles that matched or exceeded the observed change rate.
Search and model survive. Drift does not. We re-ran an identical condition three days later and the result sits exactly on top of the null. Over three days, with the same model in the same mode, the hierarchy does not move. It only looked like it moved because the model is noisy.
Notice the null is different every time. Pooling a small model with a large one gives a noisy null (54.5%). Pooling browsing-off with browsing-on gives a lower one (39.1%), because browsing is more repeatable. This is exactly why subtracting a single number from everything was never going to work.
Measure a brand's recommendation position with a single query and roughly half of what you see is a coin flip. Not a trend. Not an improvement. Not something you caused.
The only way to see signal is to ask many times and compare distributions against a null. Most tools ask once.
With browsing off it is 47%. With browsing on it is 27%. Retrieval nearly halves the randomness: when the model has something to read, it stops guessing. The smaller model is noisier still, at 57%.
And it is tempting to assume you can just run more and make the noise go away. So we ran a sweep at twenty runs per condition instead of ten, and measured the floor at every split from one run per side up to ten.
Bar width is zoomed to the 44 to 50 percent range so the flattening is visible; the percentages themselves are the real measured noise floor at each sample size.
It falls for the first three, then stops dead at 46.1% and does not move again. Doubling the sampling from five runs to ten changes nothing. Not a tenth of a point.
So the noise is not a budget problem. You cannot buy determinism back by running more. The randomness is a property of the model itself. There is a stable core of brands and a tail long enough that every query pulls different edges of it, no matter how many times you ask.
The effect was not uniform. Browsing rewrote recommendations far more in some categories than others. Pets: 88% of recommendations changed. Fitness: 61%. The pattern: the more fragmented and long-tail the market, the more browsing overrides memory. In categories dominated by a few household names, the model's memory and its search mostly agree. In categories full of small brands, they don't, and search wins.
Share of recommendations that changed when web search was turned on, by category. Bar width is relative to the highest category, Pets at 88%.
Pets, interestingly, has been the most extreme category in every study we've run. It was the sharpest in our earlier quality-versus-recommendation work too. Something about that market makes AI recommendations unusually unstable. We don't fully understand it yet.
A data engineer raised a fair objection: how do we know that's the category and not the prompts?
Broad questions like best dog food should lean on retrieval harder than narrow ones like best cat litter for odor control. If our 50 prompts skew broad in some categories and narrow in others, what looks like market structure could just be prompt style.
So we tagged all 50 prompts and split the results.
Pooled, broad prompts changed 74.2% and narrow prompts changed 80.0%. Narrow changed more, the opposite of the prediction. We could have stopped there and declared the objection handled.
That would have been wrong.
Inside each individual category, prompt style has no consistent direction at all: broad changes more in five categories, narrow in four. The pooled gap was composition, not causation. The categories that happen to hold most of our narrow prompts are also the categories that change the most, so narrow prompts appear to change more purely because of where they live.
This is Simpson's paradox, and it made the pooled number worse than useless: it looked like an answer.
Hold specificity constant and look at the category effect inside each stratum.
Fitness is lowest in both. Coffee and Pets are highest in both. The category ranking survives when prompt style is controlled for. Across the ten categories, the correlation between a category's share of broad prompts and its change rate is minus 0.56. If the confound were real, that would need to be strongly positive.
We would rather say this than have someone find it.
So the honest claim is that we looked for the confound and found the opposite, and the category effect survives stratification. Not that no confound exists.
Rami's original point had a second half we haven't cracked. He suggested a lot of the unexplained variance hides in training-data provenance: how often the model saw a brand, and in what sentiment context. We can see which brands get named. We can't yet see whether a model absorbed a brand warmly or coldly during training. That remains the biggest black box in the whole thing.
Getting recommended by AI is treated as one goal. It isn't. Being in the model's memory is a slow, prominence-driven game that rewards brands the model saw a lot during training. Ranking in what the model retrieves is a live game, closer to search, decided at query time by the retrieval layer.
Most AI-visibility tools measure a third thing entirely: whether the AI can see you at all. But visibility is not recommendation. A recommendation made from memory and one made from retrieval can be almost completely different answers to the same question. If you only optimize for one of these, you're leaving the other on the table, and depending on the user's settings, it might be the one that actually decides the sale.
All numbers come from a fixed set of 50 prompts, 10 runs each, per condition, measured directly. No estimates. Overlap is measured on the set of distinct recommended brands per intent. The clean search-effect figure uses gpt-4o in both conditions. One caveat worth stating: browsing implementations change over time, so treat the exact percentages as a snapshot of this model at this moment, not a physical constant. The direction and the size of the effect are the point.
Memory or retrieval, the first question is whether AI can read your store at all.
This study is one of seven that isolate a single variable behind the flagship finding that store readiness does not predict recommendation.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →