We took the closed-book brand claims already collected for Possession vs Deployment, same 9 Wave-1 brands, same gpt-4o possession probe, zero new API calls, and asked two different questions of it. First: when the model volunteers a specific, checkable fact about a brand with no prompting and no search, is it actually true? Second: how much ground does it cover, out of everything real that could be said? 33 of 34 checkable claims held up against independent web verification. But averaged across 9 brands and 8 real attribute categories, unprompted coverage was only 61.1%, and it was not random which parts got skipped.
Marcos Viladomiu commented on our flagship report that being mentioned is not the whole picture, a brand can show up in a recommendation and still get described inaccurately, or so thinly that the mention barely counts as knowledge. That splits into two separate, measurable questions. Brand Accuracy Score (BAS): when the model states something about a brand with no prompt, no search, and no context beyond the brand name, is it true. Content Depth Index (CDI): out of everything real that could be said, how much ground does it actually cover.
This reuses the possession side of Possession vs Deployment entirely, the same 40 closed-book claims across 9 Wave-1 brands, 3 gpt-4o runs merged to stable claims, zero new API calls. What is new here is the verification: every claim was independently checked against the brand's own site, independent press, retailers, or review sources through live web search, not scored by the same model that produced it. 6 of the 40 claims are purely stylistic description ("minimalist design," "vibrant colors") with nothing to fact-check, so they are excluded from BAS and scored only for CDI, which measures topic breadth, not truth.
33 of 34 checkable claims held up (97.1%), everything from B Corp certification to specific warranty lengths to named product features, verified against the brand's own materials or independent coverage. One did not. Bellroy's closed-book claim states its pricing runs "lower than Peak Design and Nomatic." Real prices say otherwise: Bellroy's Slim Sleeve runs $85-135, while Nomatic's own flagship wallet is priced at $19.99, not lower, several times higher. Peak Design's comparable Passport Wallet sits close to Bellroy's range, so the miss is specifically about Nomatic, not a blanket confabulation.
Being accurate turns out to be the easy part. The harder question is breadth: out of 8 real attribute categories a shopper might care about, materials, pricing, sustainability, design, warranty, certifications, business model, and named competitor comparison, how many does the model volunteer without being asked. Averaged across all 9 brands, the answer is 61.1% (44 of a possible 72 category-brand pairs), and it ranges from 25% for Zigpoll to 87.5% for Bellroy.
The category breakdown explains why some brands score higher than others, and it is not random.
This series has repeatedly shown that presence, showing up in a recommendation at all, is the real bottleneck (Candidacy vs Selection, Cold Start). Marcos's original point was that presence alone does not settle the question, a brand could be present, accurate, and still thin. This study shows accuracy is not where that risk lives, 97.1% of checkable claims held up, and the one miss was a specific comparative number, not a fabricated fact. Depth is where the real variance is: a brand's own warranty, its own certifications, its own supply chain claims can be entirely real and still never make it into an unprompted description, simply because sustainability and design language crowd out the rest. If a brand wants a fuller picture volunteered, the fix is not correcting the model, the model is already mostly right, it is making the underrepresented categories, warranty, certifications, business model, easier for the model to find and repeat.
The 8-category attribute taxonomy (materials, pricing, sustainability, design, warranty, certifications, business model, competitor comparison) is our own construction for this study, built to be broad and consistent across very different product categories, from cookware to office furniture to a survey widget. A different, equally reasonable taxonomy could shift CDI up or down without the underlying facts changing. Verification was done by a single researcher through live web search rather than multiple independent judges or a formal inter-rater process, and web search itself can miss facts that exist only in a brand's private materials or that are too recent to be indexed. Comparative pricing claims, like the Bellroy miss, are judged against spot-checked retail prices at a point in time, not a continuously tracked price history, so a claim that was once true could look wrong here simply because prices moved. Nine brands is not enough to draw a confident BAS-to-CDI correlation, a brand's accuracy and its depth do not appear to move together in this cohort, but that is not a statistically powered claim at this sample size.
Purely descriptive language ("minimalist," "vibrant") is excluded from BAS because it is not falsifiable, but it still counts toward CDI as touching a category, since the question there is topic breadth, not truth. That is a deliberate design choice, not an attempt to inflate either number, and it is stated here so the two metrics are read the way they were built to be read.
An earlier, less rigorous pass flagged Onyx Coffee Lab's sustainability claim as likely confabulated. A second, source-by-source check reversed that call, the claim held up, and surfaced a real, different miss instead (Bellroy's pricing comparison). The correction is documented above rather than quietly overwritten.
8 of 9 brands got unprompted sustainability language. Only 3 of 9 got their warranty or certifications mentioned, despite several having real, checkable ones. The gap held up brand by brand, not just in the average, which is why this reads as a category effect and not noise.
If warranty, certifications, and business-model facts are the categories the model volunteers least, the fastest way to see what it is leaving out about your brand is a free scan, not a guess.
This extends a LinkedIn question about our flagship report, and sits next to the possession data it reuses.
Illustrative example · single-site signal for atomfoundry.dev.
View full signals →