Line of inquiry
Inquiring lines›Why are language models fragile de…›Why does model confidence diverge…›this line of inquiry
How do confident outputs distort user judgment of accuracy?
A broader line of inquiry — a family of 30 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 30
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do users track model confidence instead of actual accuracy?
- How should designers measure and explain semantic uncertainty to users?
- Why is confidence a dangerous proxy for accuracy in human-AI interaction?
- What happens when confident language masks uncertainty in AI outputs?
- Can organized response format trick users into overestimating AI reliability?
- How do confident system outputs weaken user skepticism about their reliability?
- Why do users trust overconfident AI outputs across different languages?
- Do verbal uncertainty estimates calibrate better than confidence scores for personalization?
- Does high model confidence increase the risk of human overreliance?
- How does uncertainty estimation drive computational resource allocation in models?
- Can humans maintain scrutiny capacity when routed only to uncertain decisions?
- Does user preference for confirmation override model capability for disagreement?
- Does confidence-based weighting in deliberation substitute for competence-based expertise?
- Do confidence signals mislead patients differently in medical versus other domains?
- Can dialogue systems abstain from responding when uncertainty is too high?
- Does adding multiple interpretations to ambiguous situations respect language more than resolving them?
- Why do users prefer AI responses that actually harm their decision-making?
- Does the same uncertainty-driven logic appear in other conversation systems?
- How do belief distributions help systems recover from speech recognition errors?
- How do linguistic norms for expressing certainty vary across languages and models?
- What happens when DSM categories are treated as ground truth in AI?
- How should dialogue systems represent and update uncertainty from noisy ASR input?
- What skills do users need to work effectively with stochastic outputs?
- How much does domain expertise actually improve human forecasting under uncertainty?
- Can teachers trained under uncertainty constraints distill better generalizing students?
- What makes emotion scores more stable than human preference labels?
- Can language models match competitive crowd forecasters on real future events?
- What information is lost when majority labels discard minority interpretations?
- How do moment-to-moment ToM fluctuations shape AI response quality?
- Why do novices accept AI output without validation in vibe coding workflows?