Recommendation Intelligence Research™ · Study #22

The model hedges most when it's most sure.

Candidate Evaluation found that a better rating decides a contested pick almost every time. This study asks the next question: when the model explains that pick, does its language, hedged ("seems," "may be") versus assertive ("clearly," "significantly"), actually track how confident it really is? Seven independent methods test this on the same 16 fixed scenarios: a hedge/booster lexicon, a blind second-model judge, real token logprobs, self-reported confidence, repetition stability, five ways of phrasing the same question, and replication across GPT-4o, Claude Sonnet and Gemini. The answer is consistent across all seven: verbal hedging correlates r=0.03 with real internal confidence, and for the pick type that draws the most hedged language, real confidence is the highest of any signal tested.

r=0.03Hedge language vs. real confidence
11.2%Hedge rate on the highest-confidence pick type
100%Picks agreeing across GPT-4o, Claude, Gemini
7Independent methods, one result
Where this comes from

Candidate Evaluation answered what. This answers how sure

How AI Decides names Confidence as Stage 9 of the decision path, open research at the time: "how certain is AI about the recommendation, and how much would the answer change with slightly different phrasing?" Candidate Evaluation had already shown that a better star rating decides a head-to-head comparison almost every time. That leaves an obvious question unanswered: when the model writes out its reasoning for a pick, does the language it reaches for, hedged phrasing like "seems" or "may be" versus assertive phrasing like "clearly" or "significantly", actually reflect how sure it is internally? Or is hedging just a writing habit, disconnected from what is really going on underneath?

To answer that cleanly, every method here reuses the exact same 16 fixed contested-pair scenarios from Candidate Evaluation's phase 3 run: the same brand_a/brand_b pairs, the same injected fact_a/fact_b, the same rating-signal setup where one brand has a strictly better rating. Nothing new is generated. Seven different measurement techniques are simply pointed at the identical underlying decisions, so every method's numbers are directly comparable row for row.

Fixed scenarios reused throughout: 16, identical facts every method
Total calls across all 7 methods: ~2,900
Primary model: gpt-4o, temperature 0.7 (methods 1-6)
Cross-model replication: + Claude Sonnet 4.5, Gemini 2.5 Flash (method 7)
Ground truth for real confidence: OpenAI token logprobs, constrained A/B response
Independent check on the lexicon labels: blind second-model (Claude) judge
Why reuse instead of regenerate. Every prior study in this series that changed the underlying facts also changed the decision, making cross-study comparison noisy. Holding the facts completely fixed and only changing the measurement technique isolates one question: does the language track confidence, independent of what the "right" pick even is.
Seven ways to ask the same question

Lexicon, judge, logprobs, self-report, repetition, phrasing, cross-model

Each method attacks a different weak point in the one before it. Together they triangulate on a single answer instead of resting on one technique's blind spots.

Method 1
Hedge lexicon
Regex word-boundary match against a hedge/booster word list, run across every prior study's free-text reasons.
11.2% hedge rate, rating signal
Method 2
Blind LLM judge
Claude rates 50 anonymized texts hedged/neutral/assertive with no access to the lexicon labels.
66% agreement with lexicon
Method 3
Token logprobs
Constrained A/B response, temperature 0, max_tokens=1. The actual probability gap the model assigns its own pick.
Real, unfakeable confidence
Method 4
Self-report
Asked directly for a 0-100 confidence number alongside its pick, same prompts as method 3.
r=0.345 vs. real logprobs
Method 5
Repetition stability
Does the model actually flip its pick across 10 repeated runs, and does that behavioral instability track hedge language?
r=-0.09 to +0.11
Method 6
Phrasing perturbation
Same fixed facts, 5 different ways of asking (baseline, casual, formal, reversed, imperative).
0/16 picks changed
Method 7
Cross-model
Identical prompts sent to GPT-4o, Claude Sonnet 4.5 and Gemini. Does the pattern hold outside one lab's house style?
16/16 pick agreement
All seven, together
One consistent answer
Hedged vs. assertive language is not a reliable window into real model confidence, on any of the seven tests.
The core finding

Hedging is highest exactly where real confidence is highest

Candidate Evaluation tested three signal types: price, rating, and specs. The lexicon classifier (method 1) found the rating signal draws hedged language far more often than the other two, 11.2% of responses versus 2.5% for price and 3.1% for specs, a statistically significant gap (p=0.03 vs. price). The obvious reading is that the model is genuinely less sure when a rating is doing the deciding. Token logprobs (method 3), the actual probability gap between the chosen brand and the alternative, say the opposite: the rating signal has the highest real confidence of the three, not the lowest.

Hedge rate by signal type (method 1, lexicon)
Cluster bootstrap 95% CI over 16 intents, 10,000 resamples
Price · p=1.0 (vs. rating)
2.5%
Specs · p=0.05 (vs. rating)
3.1%
Rating · p=0.03 (vs. price)
11.2%
Most hedged language
Real confidence by signal type (method 3, token logprobs)
Confidence gap = P(chosen) − P(alternative), scaled ×100 for display · n=160 per signal
Price · mean self-report 80.0
82.3
Specs · mean self-report 80.8
88.2
Rating · mean self-report 89.2
99.4
Highest real confidence
The signal that draws the most hedged language is the one the model is, in reality, most sure about. Row-level correlation across all 480 matched calls confirms this isn't a fluke of the group averages: hedge label correlates r=0.035 with the real logprobs confidence gap, and r=0.032 with the model's own self-reported confidence number. Both are statistically indistinguishable from zero. Even self-reported confidence, the model's own stated number, only weakly tracks its real internal signal (r=0.345), so the model is not fully able to introspect its own certainty either, and the words it chooses track that internal signal even less.
Before trusting any of these numbers

Does hedging even mean what we think it means?

A hedge/booster word list is a blunt instrument, so before leaning on it, a second model (Claude) blind-rated 50 anonymized reasoning texts with no access to the lexicon's labels or word lists, just the raw text. The two only agreed on the exact three-way label (hedged / neutral / assertive) 66% of the time, and on the coarser hedged-vs-not distinction 76% of the time. Most of the disagreement traces to one specific failure mode: sentences that mix a hedge word and a booster word in the same sentence, which the lexicon scores as a false "neutral" tie, but a human or model reader would usually call assertive overall.

Separately, method 5 asks whether hedge language tracks something else entirely: real behavioral instability, whether the model actually picks a different brand across 10 repeated runs of the identical prompt. Across four separate datasets (50-brand and 16-intent recall runs, with and without injected facts), the correlation between hedge rate and how often the top pick actually changes ranges from −0.09 to +0.11, essentially zero in every case.

Read every hedge-rate percentage in this study as directional, not exact โ€” the 66% inter-rater agreement is the honest error bar sitting underneath the lexicon method throughout.
Ask it differently, does the hedge follow

The decision is stable. The wording isn't

The same 16 fixed facts were asked five different ways: a plain baseline, a casual tone, a formal tone, the brands listed in reversed order, and an imperative instruction. Across all 800 calls, 0 of 16 intents changed their winning brand no matter which phrasing was used, the underlying decision is completely robust to how the question is asked. The hedge rate, however, moved by more than double depending on phrasing alone, nothing about the facts changed.

Hedge rate by phrasing (method 6), same facts, same pick every time
5 phrasings × 16 intents × 10 runs = 800 calls, 0 pick flips
Casual
6.9%
Formal
7.5%
Imperative
8.1%
Reversed order
11.2%
Baseline
15.6%
Most hedged phrasing
Does it hold outside one lab

GPT-4o, Claude and Gemini agree on the pick, not the tone

The same 16 fixed scenarios were sent, unchanged, to GPT-4o, Claude Sonnet 4.5 and Gemini. Across 480 calls, every one of the 16 intents landed on the same majority pick across all three models, and within each model the pick barely moved across its own 10 runs (0% flip rate for GPT-4o and Claude, 1.2% for Gemini). That is strong evidence the injected facts themselves, not any single lab's quirks, are what drives the decision.

Self-reported confidence, though, has a distinct house style per provider: GPT-4o averages 87.9, Claude the most conservative at 78.8, Gemini the most confident at 91.4 with the widest spread (30 to 100). But the language pattern in each model's stated reason is nearly identical: assertive/boosted language dominates everywhere, 87.5% of GPT-4o's reasons, 93.8% of Claude's, 81.9% of Gemini's, with true hedge language rare across the board (0.6% to 1.9%). The confident-sounding tone isn't a GPT house style, it's close to universal.

Pick agreement across all 3 providers: 16/16 intents
Within-model flip rate: GPT-4o 0%, Claude 0%, Gemini 1.2%
Self-reported confidence, mean: GPT-4o 87.9, Claude 78.8, Gemini 91.4
Assertive/boosted language in reasons: 82-94% across all 3 models
One caveat on this leg. Gemini ran on Gemini 2.5 Flash, not the Pro-tier model originally planned, after the Pro tier's free-tier quota (250 requests/day) was exhausted mid-run. So this is GPT-4o and Claude Sonnet, a mid-tier model, against a fast-tier Gemini, not three matched flagship models. The agreement result is, if anything, a more conservative finding given the tier mismatch.
Why it matters

Don't read the wording. Read the facts it was given

Seven independent tests, four different confidence proxies, three different providers, and five different phrasings all land on the same place: how confident or hedged a model's language sounds is not a reliable signal of how confident it actually is. For a brand watching how AI describes it, "the AI clearly recommends us" and "the AI seems to lean toward us" are, on this evidence, not meaningfully different signals of real model conviction, both could be sitting on an internal confidence anywhere in the same range.

What did stay stable throughout: the actual pick. 0 of 16 decisions changed across five phrasings, and 16 of 16 agreed across three separate model providers. That is the practical takeaway for the rest of this series: the lever that moves a recommendation is the fact injected into the comparison (a rating, a price, a spec), not the tone the model happens to write in afterward, and not which provider is asked. Chasing "sound more confident" copy is chasing the wrong variable.

What this doesn't prove

This is one signal class, tested seven ways

Stating the limits up front. All seven methods hold the same variable fixed: a rating-signal, contested-pair scenario where one brand has a strictly better rating (reused from Candidate Evaluation's phase 3). Methods 5 through 7 do not repeat this design for the price or specs signal types, so the cross-model and phrasing-robustness findings are demonstrated for the rating signal specifically, not confirmed identically for every signal type in this series.

The lexicon classifier (method 1) that most of the headline percentages rest on only reaches 66% three-way agreement with an independent model judge, treat any single hedge-rate percentage as directional. Token logprobs, the one genuinely unfakeable confidence measure, are only available through OpenAI's API, Anthropic does not expose them and Gemini's availability is inconsistent across API versions, so methods 3's "real confidence" comparison is GPT-4o only; Claude and Gemini are represented only through self-report, which method 4 already shows is an imperfect proxy (r=0.345 vs. real logprobs). And Gemini's leg of method 7 ran on a Flash-tier model rather than the originally planned Pro-tier model, for the free-tier quota reason noted above.

Everything here is correlational and observational within a controlled prompt design, not a causal test on a real store. It measures how a model talks and how sure it internally is about a fixed, artificial comparison; it does not measure whether changing a real brand's language in the wild would change how AI recommends it. That causal question is the direct motivation for the Founder Lab experiments noted on the research roadmap.

Supporting evidence

Three more signs the pattern is real, not noise

r=0.345 Self-report vs. real confidence

Even the model's own stated 0-100 confidence number only weakly tracks its real internal signal (token logprobs). It doesn't fully know its own mind, so the words it chooses on top of that number track real confidence even less.

66% Independent judge agreement

A second model (Claude), blind to the lexicon's labels, agreed on the exact hedge/neutral/assertive rating only 66% of the time on 50 sampled texts. Built into every hedge-rate number in this study as the honest error bar.

0/16 → 16/16 Stable within, stable across

0 of 16 picks flipped across 5 phrasings of the same question, and all 16 stayed identical across GPT-4o, Claude Sonnet and Gemini. The decision is the stable part of this system; the wording around it isn't.

Stop reading tone. Start reading facts

Free AI Commerce Score™ in 10 seconds.

If a model's confident-sounding language doesn't track its real confidence, a free scan is a faster way to find what actually moves a recommendation than parsing how AI describes you.

Free · No signup · Results in 10 seconds
Keep reading

The rest of the research series

This answers the Confidence question flagged on the decision map. Read the map for where it fits, or Candidate Evaluation for the original rating-signal decisions this study reuses.