Research

What we found after scanning thousands of real stores

We run automated scans on Shopify and DTC stores to understand how AI shopping agents actually read, evaluate, and recommend them. The data is live, continuously updated, and ours alone.

66,090Stores analyzed
55/100 Average AI Commerce Score™
24% AI Invisible Risk
7% High Confidence Stores
Score distribution across scanned stores in July 2026
7%High Confidence
7%Moderate Confidence
24%Low Confidence
62%AI Invisible Risk
AI Confidence Pyramid™
62%
Stores below 50 AI Commerce Score™
31%
Stores between 50 and 84 AI Commerce Score™
4,626 stores
Reached 85+ AI Commerce Score™
Average AI Commerce Score™ by category
56
Fashion
Risk: 20% · Best: 93/100
Fashion
57
Beauty
Risk: 17% · Best: 93/100
Beauty
57
Health and Wellness
Risk: 19% · Best: 94/100
Health
57
Food and Drink
Risk: 20% · Best: 87/100
Food
59
Home and Living
Risk: 15% · Best: 92/100
Home
57
Sports and Outdoor
Risk: 17% · Best: 89/100
Sports
54
Pets
Risk: 24% · Best: 84/100
Pets
57
Tech and Gadgets
Risk: 18% · Best: 93/100
Tech
56
Kids and Baby
Risk: 19% · Best: 83/100
Kids
59
Jewelry
Risk: 13% · Best: 85/100
Jewelry
AI Commerce Bar™
8Scoring factors · 0 to 49 AI Invisible Risk · 50 to 69 Low Confidence · 70 to 84 Moderate Confidence · 85 to 100 High Confidence · Live data · updated Jul 22, 2026
Research Library

Understanding AI commerce through measurement.

Explore public research built from real AI responses, large-scale store analysis and reproducible methodology.

Mechanism
Web Search Rewrites 77% of AI Product Recommendations
Turning on browsing changes 77% of recommended brands, same model, same prompts.
77% of brands changed
Read report
Mechanism: Corrected
The Fame Study, Corrected
Public fame explains 1.2% of recommendation, re-measured on 872 brands, sitting on the noise floor.
1.2% explained
Read report
Mechanism
AI Knows Your Website. It Still Won't Recommend You.
The model names the correct domain 75.9% of the time. Knowing where you are is not the same as reading your store.
75.9% domain accuracy
Read report
Mechanism
29,633 Reasons. 26,812 Unique. The Model Confabulates.
The model writes a fresh justification for every pick. Asking it why does not work.
90% unique reasons
Read report
Mechanism
Search Changes the Vocabulary, Not Just the Brands
With search on, the model names ingredients and specifications. With it off, impressions.
21x more specific
Read report
Mechanism
Candidacy vs Selection
Full population, 60,924 stores. Intent gets you into the recommended set. It wins you nothing once you're there.
R² 1.2%
Read report
Mechanism
Nothing About Your Brand Predicts Recommendation. The Model's Own Past Behavior Does.
June's position predicts July's position, the strongest signal in this entire series, and it does not drift for 15 days.
61.4% self-predictive
Read report
Mechanism: Follow-up
Two Months Later, the Model Still Agrees With Itself
50 intents, six independent sweeps across two months. 43 kept the exact same #1 brand every time.
86% locked in
Read report
Mechanism
Hand It a Rating, and It Follows Every Single Time
16 genuinely contested brand pairs. A better star rating flipped the verdict 160 of 160 runs. Price barely moved it.
100% follow rating
Read report
Mechanism
We Invented a Brand With Zero History. Reviews Got It Picked Anyway
Against an entrenched category leader, a brand with no memory won 0 of 360 runs with no evidence, 53.1% once given a review score.
0% to 53.1% with reviews
Read report
Mechanism
We Widened the Fame Signal Four Ways. It Barely Moved
Wikipedia, Wikidata, domain age, and media mentions combined explain 11.2% of recommendation frequency, still far under the 61.4% self-consistency benchmark.
11.2% combined, vs 61.4%
Read report
Key Findings

What the data actually shows

The model remembers itself.
The model remembers itself.
Previous recommendations predict future recommendations better than store quality, brand fame or web presence.
View Study
Web Search rewrites recommendations.
Web Search rewrites recommendations.
Enabling web search changed 77% of recommended brands, even with the same prompts and model.
View Study
Store quality doesn't predict recommendations.
Store quality doesn't predict recommendations.
High-quality stores are not automatically recommended by AI shopping systems.
View Study
AI Invisible Risk is real.
AI Invisible Risk is real.
Many stores remain effectively invisible to AI recommendations despite being fully operational.
View Study
Recommendation rankings stay stable.
Recommendation rankings stay stable.
Recommendation rankings showed virtually no measurable drift across a 15-day period.
View Study
Recognition isn't recommendation.
Recognition isn't recommendation.
AI can recognize and understand a brand without ever recommending it.
View Study
Memory beats web presence.
Memory beats web presence.
The model's own history explains recommendations better than external brand signals.
View Study
Big brands aren't guaranteed winners.
Big brands aren't guaranteed winners.
Well-known brands often outperform in visibility, but not necessarily in AI readiness.
View Study
Visibility isn't enough.
Visibility isn't enough.
Being discoverable by AI doesn't guarantee being selected as a recommendation.
View Study
Trust is cumulative.
Trust is cumulative.
Recommendation confidence emerges from many small trust signals rather than a single optimization.
View Study
Intent changes the outcome.
Intent changes the outcome.
The same brand can perform very differently depending on what the shopper is asking.
View Study
Industries behave differently.
Industries behave differently.
Each e-commerce category follows its own recommendation dynamics under the same methodology.
View Study
AI commerce is measurable.
AI commerce is measurable.
Recommendation behavior can be studied through controlled experiments and repeatable measurements.
View Study
Monitoring isn't always necessary.
Monitoring isn't always necessary.
Stable recommendation rankings suggest that continuous daily monitoring often adds little value.
View Study
Every recommendation has a pattern.
Every recommendation has a pattern.
AI recommendations are not random. They follow measurable, repeatable behavior across large datasets.
View Study
Two months later, it still agrees with itself.
Two months later, it still agrees with itself.
86% of 50 tracked intents kept the exact same #1 brand across six independent sweeps, June to August.
View Study
Hand it a rating, and it follows every single time.
Hand it a rating, and it follows every single time.
Across 16 contested brand pairs, a better star rating flipped the verdict 160 of 160 runs. A better price barely moved it.
View Study
A brand with no memory won 0 of 360 runs. Reviews changed that.
A brand with no memory won 0 of 360 runs. Reviews changed that.
Against an entrenched category leader, an invented brand never won with no evidence. Give it a review score, and it wins 53.1% of the time.
View Study
We widened the fame signal four ways. It barely moved.
We widened the fame signal four ways. It barely moved.
Wikipedia, Wikidata, domain age, and media mentions, none individually significant. Combined, 11.2%, still far under the model's own 61.4% self-consistency.
View Study
Research Topics

What we are studying

These are the areas where we are actively collecting data and publishing findings.

AI Visibility
Why do some stores never appear in AI-generated recommendations?
Active research
Recommendation Intelligence
What determines which brands AI chooses to recommend, and which it ignores?
Active research
Machine-Readable Trust
Which trust signals can AI actually read, verify and use during decision-making?
Active research
Agentic Commerce
How AI shopping agents discover, evaluate and purchase products across the web.
Active research
AI Commerce Benchmarks
Measuring how industries compare in AI readiness across categories and markets.
Active research
Commerce Accuracy
How product data quality affects AI understanding, confidence and recommendations.
Active research
AI Decision Science
Understanding how AI models rank, compare and choose between competing brands.
Active research
Recommendation Stability
Measuring how AI recommendations change over time, across models and retrieval systems.
Active research
Our Methodology

The AI Commerce Score Engine™

Every AI Commerce Score is calculated from eight weighted evaluation layers that measure how AI shopping agents understand, trust and recommend e-commerce stores.

Semantic Visuals & Image Clarity
15%
AI Structured Signals
15%
Core Technical & Interpretability
15%
AI Trust & Transaction Confidence
15%
Commerce & Feed Accuracy
15%
User Intent Match
10%
Recommendation Confidence
10%
External Authority Signals
5%
Explore more

Where this connects

The research sits on top of everything else we build. Here is how it links together.

See where your store stands

Free AI Commerce Score across all 8 factors.

Find out exactly what AI agents see when they evaluate your store.

Free · No signup · Results in 10 seconds
Research infrastructure · Founder Lab

Atom Foundry operates its own live Shopify store, Founder Lab, as a permanent controlled experiment. Most measurement of AI commerce is observational: you watch stores you don't control and infer what matters. Founder Lab lets us do the opposite. We own every variable, so we can change one thing at a time, a product description, a schema field, a trust signal, and watch, in real time, how AI systems respond.

What we measure there goes beyond a single score. We track how each change moves the store's AI Commerce Score™, but also which AI agents crawl the store and when, which pages they read, whether they retrieve it in live shopping answers, and how its recommendation profile shifts over time. Read the live results in the Founder Lab Log.