Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How does AI reshape human understa…›this line of inquiry
How do clinicians calibrate trust in AI medical recommendations?
A broader line of inquiry — a family of 59 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 59
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Why do users over-trust AI in some domains but under-trust it in medicine?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- Does medical AI accuracy depend more on knowledge or reasoning ability?
- How much do edited AI responses versus raw outputs affect clinician ratings?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- Why do clinicians fail to act on correct AI suggestions in real care?
- How does reliance on AI recommendations erode professional judgment over time?
- Does AI change clinician cognition or just increase reliance on predictions?
- Do confidence signals mislead patients differently in medical versus other domains?
- Why do medical diagnoses require human judgment even with AI assistance?
- Would clinicians' ratings change if authorship was visible from the start?
- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- Why do radiologists fail to benefit from AI decision support?
- Can annotation and explanation labels reduce automation bias in clinical settings?
- Can algorithmic aversion explain clinicians' skepticism of AI recommendations?
- How does expert annotation instability affect medical AI benchmarking?
- What prospective trials are needed to validate AI diagnostic claims?
- Have similar doubts about calculators or diagnostic aids actually disappeared over time?
- Can medical diagnosis depend less on knowledge and more on orchestration?
- Do expert physicians also prefer AI-written medical text when it is unlabeled?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Do physicians follow incorrect advice more when they trust its source?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- How does blinded rating of diagnoses compare to real clinical outcomes?
- How does the personal nature of medical decisions affect trust in AI?
- Do radiologists' beliefs about AI-assisted performance match their actual outcomes?
- Does showing AI confidence scores reduce radiologist over-reliance on wrong suggestions?
- Why do experts resist AI recommendations that contradict their own judgments?
- How much do physician scores improve when assisted by the same model?
- Does medical domain competency require knowledge injection or better prompting?
- How does trusting wrong AI advice change what medical action people decide to take?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Can clearer accountability structures reduce patient resistance to AI providers?
- Why does labeling advice as AI from a doctor change how people trust it?
- Why did clinicians guess authorship at chance level despite strong preferences?
- What evidence would prove medical AI actually works in clinics?
- Do patients show the same bias toward expert-labeled medical advice?
- What safeguards help radiologists maintain independent judgment when using AI assistance?
- How much does prompt selection bias favor medical models over base models?
- How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?
- What role does interface design play in clinician adoption of AI tools?
- Which medical domains depend most on up-to-date information retrieval?
- What makes reasoning auditable in medical AI decision support?
- What makes clinical theory grounding more effective than pattern matching alone?
- What role does cost estimation play in steering diagnostic test ordering?
- How well do curated benchmark cases represent real clinical deployment?
- What proportion of patients cannot complete AI interviews due to technology barriers?
- Can rubric-graded response quality predict real-world clinical workflow success?
- How do existing ICD-11 diagnostic categories already capture AI-related presentations?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Does AMIE's advantage hold when patients interact through speech or video instead of text?
- Why does medical knowledge require continuous access to current sources?
- Can trainees improve formulation skills by practicing against simulated patients?
- Does this colonoscopy finding apply to other medical specialties using AI?
- How would AI therapists compound the overestimation problem with patients?
- Are newspaper advice columns representative of broader professional psychological guidance?
- Can people tell which medical advice is accurate based only on how it reads?