We took 10 brands already measured in the Recommendation Reports series and ran a new kind of conversation: an explicit, favorable introduction of the brand, then three follow-up questions that never name it again, what else is out there, what stands out, and finally, pick just one. Across 200 four-turn conversations, the brand survived to the final answer 49.5% of the time, cohort-wide. That headline number hides almost everything interesting. Three brands survived every single run. Two never survived once. And one brand with a perfectly respectable one-in-four baseline recommend rate got swapped out for the same competitor in 18 of 20 conversations.
Every prior study in this series that touches conversation dynamics has looked at stability between separate sessions, ask the same question a week apart, does the model still agree with itself. This one looks inside a single conversation. Turn one is an explicit, favorable introduction: the brand is named and described positively, the kind of opening line a shopper who already likes the brand might use. Turns two and three never name the brand again, an open "what else should I consider" question, then a narrowing "what stands out to you" question. Turn four forces a single final pick. A cheap case-insensitive substring check flags whether the brand comes back up unprompted at T2 or T3, and a dedicated judge call classifies the T4 answer as SURVIVED, DISPLACED, or AMBIGUOUS, with AMBIGUOUS counted conservatively as not-survived in the headline numbers.
The correlation check against each brand's real-world performance reuses data rather than re-collecting it: the baseline recommend rate for each brand comes straight from its published Recommendation Report, hundreds of real buyer-question observations already collected and scored. Nothing new was gathered for that half of the comparison, only the T1-T4 conversations themselves are new calls. Brand identities are anonymized throughout this page, labeled Brand A through Brand J, to protect the individual businesses studied; the underlying rates and relationships are unchanged.
T1 recognition is 100% by construction, the brand was just named. By T2, with the brand unprompted, mention rate falls to 32.5%. By T3, it climbs back to 59%, likely because "what stands out to you" is a narrower question than "what else is out there," pulling the model back toward stronger candidates. T4's forced pick lands at 49.5% survival, cohort-wide (95% CI 25.5-74.5%, n=200), well above a similarly designed industry test published elsewhere in 2026, which reported 12.7% survival on its own headline number. The gap is large, and the most likely reason is the T1 design difference disclosed above: this test opens with an explicitly favorable introduction, most naturalistic conversations don't.
The dip-then-rebound shape is the interesting part, not just the endpoint. T2's open question loses the brand fastest, T3's narrower framing brings some of it back, and T4's forced choice settles in between.
The cohort-wide average hides a split cohort. Three brands, labeled A, B, and C here, survived all 20 of their conversations. Two, labeled I and J, survived none, consistent with their near-zero baseline recommend rates, they rarely get considered in the first place, so an artificially favorable T1 introduction doesn't have much to protect. The rest sit in between, and one of them, Brand H, is the most interesting case in the cohort, more on that below. Brand identities are anonymized here to protect the individual businesses studied.
Correlating each brand's T4 survival rate against its published baseline recommend rate gives r = 0.68 (p = 0.048, n = 9), Brand I excluded since its near-zero baseline makes the ratio undefined. That's on the edge of significance and moves in the opposite direction from the null result this series found for possession-vs-deployment, a stronger real-world track record does seem to buy some protection against mid-conversation displacement. But the relationship is loose, not tight, and the biggest miss in the cohort is the most telling data point.
The lock-in study elsewhere in this series showed that once a model settles on a recommendation, it tends to repeat that same answer weeks later, a between-session kind of stability. This study measures the other axis: within a single conversation, does an early, favorable mention survive to the end. The answer is that it depends heavily on which brand, and specifically on whether there's a strong category-default competitor waiting in the wings. Brands A, B, and C held their ground through every follow-up. Brand H, despite a respectable baseline, did not, because the model has a go-to answer for its category that isn't Brand H, and three follow-up questions were enough to surface it. Getting mentioned first is not the same as staying mentioned.
Stating the limits up front. T1 is an explicit, favorable introduction of the brand, a deliberate best-case starting condition, not a naturalistic first message. That is almost certainly why the 49.5% survival rate here reads so much higher than a similarly designed industry test's 12.7%, the two studies are not measuring the same starting point, and the comparison should be read as context on the shape of the effect, not a contradiction of that result. The same model, gpt-4o, runs the conversation and judges the T4 verdict, a real self-grading risk. The mitigation is a manual spot-check: 17 judge calls checked by hand against the raw T4 text, spanning the full range from 100%-survival brands to 0%-survival brands to the Brand H outlier and its one AMBIGUOUS case, and all 17 matched. That's reassuring on a sample this size, not conclusive.
The H3 correlation is n=9, genuinely underpowered, p=0.048 is a real but fragile signal, not a settled fact. One conversation was scored AMBIGUOUS rather than forced into a binary; it's counted conservatively as not-survived in every headline number, which if anything understates survival slightly. Ten brands, one category mix, consumer ecommerce, one model. Cross-platform and cross-model versions of this question are a separate, larger, and currently API-key-blocked follow-up. Brand identities are anonymized throughout this page and do not affect any of the reported rates or statistics.
17 judge verdicts checked by hand against the raw T4 conversation text, spanning full-survival brands, zero-survival brands, and the Brand H outlier including its one AMBIGUOUS case. All 17 matched the underlying text exactly.
Only one of 200 T4 answers genuinely hedged between two brands rather than picking one. It's counted as not-survived in every headline number here, the conservative choice, so the real survival rate is if anything a hair higher than reported.
Getting mentioned first is only half the picture. The fastest way to see where your brand actually stands once a real conversation keeps going is a free scan, not a guess.
This extends the Stability question from between-session lock-in to within-conversation displacement, and sits next to the possession-deployment gap measured in the previous study.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →