Recommendation Intelligence Research™ · Study #20

We invented a brand with zero history. Reviews got it picked anyway.

How AI Decides names Cold Start, how a brand with no memory ever gets considered, as one of the open gaps this series hasn't measured yet. So we built the hardest version of the test. We took 36 category leaders the model reaches for by default, ten times out of ten in most cases, and pitted each against a brand we made up, with no training-data footprint at all. With zero evidence, the invented brand won 0 times out of 360 runs. Give it a review score, and it wins 53.1% of the time. Press mentions and sales-volume claims did almost nothing.

0%Chosen with zero evidence, out of 360 runs
53.1%Chosen once given a review score
36Entrenched category leaders challenged
1,940Model calls across three phases
Where this comes from

Every study so far measured brands the model already knows

Hand It a Rating showed what tips the verdict once two known brands are genuinely close: a review score flips it 160 of 160 times. But every brand in that test already had a memory footprint, the model had recommended both before. That leaves the harder, more consequential question completely open: how does a brand with no memory at all, nothing the model has ever seen, get considered in the first place?

We built the hardest realistic version of that test. First, we found 36 intents with a genuinely entrenched incumbent, a brand the model names as the top pick in at least 7 of 10 open runs, not a close contest. Then, for each category, we invented one brand name with no training-data footprint at all, and paired it against that entrenched leader. First with nothing but the two names. Then again, with one piece of evidence about the newcomer at a time: a review score, a press mention, or a sales-volume claim.

Model: gpt-4o, temperature 0.7
Phase 1, open baseline: 50 intents × 10 runs = 500 calls
Entrenched leaders found: 36 of 50 (top brand ≥7 of 10 wins)
Phase 2, paired, zero evidence: 36 intents × 10 runs = 360 calls
Phase 3, paired + one evidence type: 36 intents × 3 types × 10 runs = 1,080 calls
Total model calls: 1,940
The invented brands, by category. One name per category, never used before this study: Barkwell (pets), Luminire (beauty), VitaCrest (supplements), Roastframe (coffee), Fibrelane (fashion), Ironmeadow (fitness), Freshcadence (food), Hearthloom (home), Calmroot (wellness), Circuitnest (electronics). Names were invented for this study and not checked against a trademark database; any resemblance to a real company is unintentional.
The finding

Zero evidence, zero wins. Every time

Across 360 zero-evidence runs, the invented brand was never once chosen over the entrenched incumbent, 0 out of 360, not close to zero, exactly zero. That is the cleanest confirmation this series has produced that the model's consideration set is a closed, memory-gated list. Then we handed the model one fact about the newcomer at a time. A review score won 53.1% of runs. A press mention won 1.9%. A sales-volume claim won 0.3%, both statistically indistinguishable from the zero-evidence floor.

Share of runs the invented brand won, by evidence type
1,080 Phase 3 runs across 36 entrenched intents, one injected fact per run, vs a 0.0% zero-evidence floor
Sales volume claim
0.3%
Press mention
1.9%
Review score
53.1%
Only one that works
The 53.1% is a real, uneven effect, not a flat artifact. Broken out across the 36 entrenched intents, the review-driven win rate ranges from 0% to 100%, with most intents landing somewhere in between rather than clustering at the extremes. It tracks by category: lowest in beauty and electronics (28-32%, categories with strong established expert-brand loyalty like dermatologist-recommended skincare or GoPro and Anker), highest in food and pets (70-87%). Sampled reasoning text cites the exact injected review numbers directly, not a fabricated justification.
Which evidence actually opens the door

Reviews aren't just the best lever. They're the only one

A permutation test comparing each evidence type against the zero-evidence baseline tells the same story with statistics: reviews clear significance by a wide margin (p<0.0001), while press mentions (p=0.24) and sales volume (p=1.0) do not clear it at all, meaning neither is reliably distinguishable from doing nothing. The gap between reviews and the other two is itself large and significant, reviews beat press by 51.1 points and volume by 52.8 points, both p<0.0001 after correction. Press and volume are not different from each other, both effectively sit at the floor.

New-brand win rate vs the zero-evidence floor, by evidence type
Cluster bootstrap 95% CI over 36 intents · zero-evidence floor 0.0% (95% CI 0.0-0.0%)
Volume · p=1.000, not significant
0.3%
Press · p=0.241, not significant
1.9%
Reviews · p<0.0001, real
53.1%
Strongest, by far
Why it matters

For a brand with no memory, reviews aren't a factor. They're the gate

Hand It a Rating already showed that reviews are the strongest single lever once a brand is already a known candidate. This study shows something stronger: reviews are powerful enough to manufacture candidacy from nothing at all, for a brand the model has never seen and cannot verify exists. Press coverage and sales-volume claims, the kind of proof a new brand's PR strategy usually chases first, did essentially nothing in this design. That is a specific, testable, and slightly uncomfortable claim for anyone investing in press over review infrastructure as an AI-visibility strategy.

The practical read for a brand with no recommendation history yet: getting review data into a machine-readable, visible form, the same Recommendation Confidence factor in the AI Commerce Score™ that mattered most in the prior study, appears to be the single highest-leverage lever available before a brand has any memory footprint to lean on at all.

What this doesn't prove

This tests injected evidence, not live search

Stating the limits up front. This study could not wire a live web-search tool into a personal API script, so it does not test whether turning on real browsing lets a genuinely new, real, unranked brand surface, the faithful version of the original search-on/search-off design from earlier in this series. It tests a narrower, more controlled question instead: given the model is handed one specific discoverable fact directly, how much that fact moves the outcome. A live-search version of this test, against a real small brand instead of an invented one, is a natural next extension, not part of this run.

The invented brand names were not checked against a trademark database, the "entrenched" threshold (top brand ≥7 of 10) is a design choice, and this is single-turn, gpt-4o only, the same disclosed limitations as the rest of the series.

Robustness check: re-running the same analysis at a stricter threshold (top brand ≥8 of 10, 27 of the 36 intents) holds up. Reviews: 47.8% (95% CI 36.3-59.6%). Press: 0.7%. Volume: 0.0%. The conclusion does not change under the stricter cut.

Supporting evidence

Two more signs the setup was clean

0 of 360 zero-evidence wins

Not a low rate rounded down, a literal zero across every one of the 36 entrenched intents. The strongest, cleanest confirmation yet that the model's consideration set is a closed, memory-gated list absent any evidence to the contrary.

27 of 36 held under a stricter cut

Re-running the analysis on only the most entrenched intents (top brand 8 or more of 10 wins) still gives reviews 47.8%, press 0.7%, volume 0.0%. Same conclusion, tighter sample.

Know which fact is worth fixing first

Free AI Commerce Score™ in 10 seconds.

If reviews are the gate for brands with no memory yet, that's where a free scan starts.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This fills the Cold Start gap flagged on the decision map. Read the map for where it fits, or the evaluation study for the result this one builds on.