Recommendation Intelligence Research™ · Study #27

The model is almost never wrong about your brand. It just doesn't say much.

We took the closed-book brand claims already collected for Possession vs Deployment, same 9 Wave-1 brands, same gpt-4o possession probe, zero new API calls, and asked two different questions of it. First: when the model volunteers a specific, checkable fact about a brand with no prompting and no search, is it actually true? Second: how much ground does it cover, out of everything real that could be said? 33 of 34 checkable claims held up against independent web verification. But averaged across 9 brands and 8 real attribute categories, unprompted coverage was only 61.1%, and it was not random which parts got skipped.

ZenodoCite this study: 10.5281/zenodo.22754001
97.1%Checkable claims that held up (BAS)
61.1%Real ground covered, unprompted (CDI)
9Brands tested
34Checkable claims scored
Where this comes from

Two questions, over the same claims

Marcos Viladomiu commented on our flagship report that being mentioned is not the whole picture, a brand can show up in a recommendation and still get described inaccurately, or so thinly that the mention barely counts as knowledge. That splits into two separate, measurable questions. Brand Accuracy Score (BAS): when the model states something about a brand with no prompt, no search, and no context beyond the brand name, is it true. Content Depth Index (CDI): out of everything real that could be said, how much ground does it actually cover.

This reuses the possession side of Possession vs Deployment entirely, the same 40 closed-book claims across 9 Wave-1 brands, 3 gpt-4o runs merged to stable claims, zero new API calls. What is new here is the verification: every claim was independently checked against the brand's own site, independent press, retailers, or review sources through live web search, not scored by the same model that produced it. 6 of the 40 claims are purely stylistic description ("minimalist design," "vibrant colors") with nothing to fact-check, so they are excluded from BAS and scored only for CDI, which measures topic breadth, not truth.

Model producing claims: gpt-4o, reused, zero new calls
Verifier: independent live web search, not self-graded
Claims collected: 40, 9 brands
Checkable (factual) claims: 34, scored for BAS
Descriptive claims: 6, excluded from BAS, kept for CDI
CDI taxonomy: 8 categories, our own construction
Why reuse instead of re-collect. The claims already existed from the possession side of Possession vs Deployment. The only new work here is verification, checking each one against a real, independent source rather than trusting the model's own confidence. That is also why this design avoids the self-grading risk flagged in the earlier study: the judge here is a human researcher reading real sources, not the same model marking its own homework.
The finding, part one

Accuracy is basically solved

33 of 34 checkable claims held up (97.1%), everything from B Corp certification to specific warranty lengths to named product features, verified against the brand's own materials or independent coverage. One did not. Bellroy's closed-book claim states its pricing runs "lower than Peak Design and Nomatic." Real prices say otherwise: Bellroy's Slim Sleeve runs $85-135, while Nomatic's own flagship wallet is priced at $19.99, not lower, several times higher. Peak Design's comparable Passport Wallet sits close to Bellroy's range, so the miss is specifically about Nomatic, not a blanket confabulation.

Accuracy by brand (BAS)
Share of checkable claims independently confirmed · 3-4 claims scored per brand
Bellroy · n=4 claims
75%
1 claim contradicted
Boll & Branch · n=4 claims
100%
Branch · n=4 claims
100%
Caraway · n=3 claims
100%
Onyx Coffee Lab · n=4 claims
100%
Peak Design · n=5 claims
100%
Rumpl · n=3 claims
100%
Wild One · n=4 claims
100%
Zigpoll · n=3 claims
100%
One of these claims was flagged as likely confabulated in an earlier, less rigorous pass, and it was the wrong one. Onyx Coffee Lab's claim about solar-powered facilities and carbon-neutral shipping was initially marked suspicious. Re-verification found strong, independent corroboration: Onyx's own site states it invested in a solar energy system for its roastery in 2019, Arkansas Business trade coverage quotes the founders saying the facility runs entirely on solar, and carbon-neutral messaging appears consistently since around 2021. That claim is corrected to Confirmed here. The actual miss was Bellroy's pricing comparison, caught only once every claim was checked against a real source instead of judged from memory.
The finding, part two

Depth is where the real gap sits

Being accurate turns out to be the easy part. The harder question is breadth: out of 8 real attribute categories a shopper might care about, materials, pricing, sustainability, design, warranty, certifications, business model, and named competitor comparison, how many does the model volunteer without being asked. Averaged across all 9 brands, the answer is 61.1% (44 of a possible 72 category-brand pairs), and it ranges from 25% for Zigpoll to 87.5% for Bellroy.

Depth by brand (CDI)
Distinct attribute categories touched unprompted, out of 8
Bellroy · 7 of 8
87.5%
Boll & Branch · 6 of 8
75%
Wild One · 6 of 8
75%
Branch · 6 of 8
75%
Peak Design · 5 of 8
62.5%
Rumpl · 5 of 8
62.5%
Onyx Coffee Lab · 4 of 8
50%
Caraway · 3 of 8
37.5%
Zigpoll · 2 of 8
25%
Narrowest coverage

The category breakdown explains why some brands score higher than others, and it is not random.

Which attribute categories get volunteered, across all 9 brands
Share of brands where the model touched this category unprompted, no search, no hint
Sustainability / Ethics8 of 9
Materials7 of 9
Pricing / Positioning7 of 9
Design / Functionality7 of 9
Competitor Comparison5 of 9
Business Model / Distribution4 of 9
Warranty / Returns3 of 9
Certifications / Awards3 of 9
Sustainability shows up almost every time. Warranty and certifications almost never do, even when they are real and well documented. Peak Design's lifetime warranty is one of the most publicized facts about the brand, and it did come up. But Caraway, Onyx Coffee Lab, and Rumpl all have real, checkable warranty or certification facts of their own that never surfaced in the closed-book description. The model is not making things up here, it is just quieter on some real topics than others, and sustainability language appears to be the one category it reaches for almost by default.
Why it matters

Presence was never the hard part

This series has repeatedly shown that presence, showing up in a recommendation at all, is the real bottleneck (Candidacy vs Selection, Cold Start). Marcos's original point was that presence alone does not settle the question, a brand could be present, accurate, and still thin. This study shows accuracy is not where that risk lives, 97.1% of checkable claims held up, and the one miss was a specific comparative number, not a fabricated fact. Depth is where the real variance is: a brand's own warranty, its own certifications, its own supply chain claims can be entirely real and still never make it into an unprompted description, simply because sustainability and design language crowd out the rest. If a brand wants a fuller picture volunteered, the fix is not correcting the model, the model is already mostly right, it is making the underrepresented categories, warranty, certifications, business model, easier for the model to find and repeat.

What this doesn't prove

Our taxonomy, not a universal one

The 8-category attribute taxonomy (materials, pricing, sustainability, design, warranty, certifications, business model, competitor comparison) is our own construction for this study, built to be broad and consistent across very different product categories, from cookware to office furniture to a survey widget. A different, equally reasonable taxonomy could shift CDI up or down without the underlying facts changing. Verification was done by a single researcher through live web search rather than multiple independent judges or a formal inter-rater process, and web search itself can miss facts that exist only in a brand's private materials or that are too recent to be indexed. Comparative pricing claims, like the Bellroy miss, are judged against spot-checked retail prices at a point in time, not a continuously tracked price history, so a claim that was once true could look wrong here simply because prices moved. Nine brands is not enough to draw a confident BAS-to-CDI correlation, a brand's accuracy and its depth do not appear to move together in this cohort, but that is not a statistically powered claim at this sample size.

Purely descriptive language ("minimalist," "vibrant") is excluded from BAS because it is not falsifiable, but it still counts toward CDI as touching a category, since the question there is topic breadth, not truth. That is a deliberate design choice, not an attempt to inflate either number, and it is stated here so the two metrics are read the way they were built to be read.

Supporting evidence

Two checks that the verification held up to scrutiny

1 of 40 Flagged, and it changed under scrutiny

An earlier, less rigorous pass flagged Onyx Coffee Lab's sustainability claim as likely confabulated. A second, source-by-source check reversed that call, the claim held up, and surfaced a real, different miss instead (Bellroy's pricing comparison). The correction is documented above rather than quietly overwritten.

8 vs 3 Sustainability vs Warranty coverage

8 of 9 brands got unprompted sustainability language. Only 3 of 9 got their warranty or certifications mentioned, despite several having real, checkable ones. The gap held up brand by brand, not just in the average, which is why this reads as a category effect and not noise.

Know which real facts about your brand go unsaid

Free AI Commerce Score™ in 10 seconds.

If warranty, certifications, and business-model facts are the categories the model volunteers least, the fastest way to see what it is leaving out about your brand is a free scan, not a guess.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This extends a LinkedIn question about our flagship report, and sits next to the possession data it reuses.