Line of inquiry
Inquiring lines›How do language models learn and r…›How do language models learn and w…›this line of inquiry
Should models ask for clarification when facing ambiguous or under-specified information?
A broader line of inquiry — a family of 51 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 51
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models learn to ask clarifying questions instead of making assumptions?
- Can models learn to ask clarifying questions instead of answering prematurely?
- Should LLMs query users back when presented with under-specified scenarios?
- Can language models ask clarifying questions when sentences are ambiguous?
- Do models fail to identify what information they need without guidance?
- Can LLMs learn to ask clarifying questions instead of guessing?
- Can language systems learn when to ask for clarification instead of choosing one reading?
- What explains the gap between perplexity performance and actual reasoning capability?
- Do models naturally learn to ask clarifying questions without explicit supervision?
- Can models learn to identify what information is missing from questions?
- Why do specific clarifying questions outperform generic requests for clarity?
- Can correct model outputs prove that semantic meaning rather than surface patterns drove the response?
- Why might expressed satisfaction with explanations diverge from actual cognitive clarity?
- What makes a clarifying question aligned with user interests versus structurally sound?
- Can models identify information gaps without just guessing or refusing to answer?
- What distinguishes genuine understanding from correct output without coherent principles?
- Can contamination-free evaluation distinguish between memorization and genuine prediction ability?
- What makes some clarifying questions more useful than others?
- How do behavioral differentiation and paraphrase stability trade against accuracy?
- How do we measure genuine reasoning inside a language model?
- Can question quality be trained separately from the decision to ask?
- What training approach enables models to proactively request clarification?
- What is the difference between a truthful answer and an honest one?
- Does disambiguation on the input side differ from the output-side preview approach?
- Can models detect false presuppositions when they actually possess the knowledge?
- Can models identify what information they are missing in underspecified tasks?
- How can we measure whether an agent reasons correctly rather than just sounds plausible?
- How does linguistic calibration differ from token probability calibration?
- Does adding multiple interpretations to ambiguous situations respect language more than resolving them?
- How does ambiguity detection connect to models' ability to ask clarifying questions?
- Can models identify what information they are missing in underspecified problems?
- What structural changes enable agents to ask clarifying questions?
- How do human annotators disagree systematically on ambiguous examples?
- Can models distinguish between ambiguous and incomplete information inputs?
- Why do current speech benchmarks fail to measure reasoning over audio?
- Which types of clarifying questions actually help users versus wasting their time?
- What separates pattern matching from genuine language understanding?
- Why are ground truth labels missing from unlabeled domain evaluations?
- When and what should a model actually decide to delegate?
- What makes specific-facet questions outperform generic need-rephrasing requests?
- Can measuring semantic entropy help us detect unreliable generations?
- Do models learn different sophistry strategies for QA versus code generation?
- Why do specific clarifying questions outperform rephrased versions of user needs?
- Why does fairness depend on context and who you ask?
- Can clarification questions alone match the stability gains of previews plus questions?
- Why does document perplexity stay low while question-answering accuracy drops?
- Can simple diagnostic tests predict language model performance in production complexity?
- What classifier accuracy is needed to assign memory roles reliably at retrieval time?
- What makes factual verification difficult in inter-model debate?
- Why do missing survey responses need careful handling in risk assessment?
- Why does homework adherence remain low despite advances in language model capability?