Line of inquiry
Inquiring lines›What do model internals reveal abo…›How should agents manage informati…›this line of inquiry
How do we evaluate AI systems when user perception misleads actual performance?
A broader line of inquiry — a family of 66 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 66
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do live human evaluations differ from ground-truth benchmarks?
- How do satisfaction scores differ from genuine cognitive improvement?
- How should we evaluate AI systems we cannot directly observe?
- Why do users report satisfaction that diverges from actual cognitive clarity?
- Can cognitive governance help users interpret AI outputs better?
- Should explanation quality be measured by user satisfaction or behavior prediction?
- Should evaluations shift toward open-world messy tasks instead of contests?
- Why do people evaluate machines against human communication standards?
- How does human intuition about cognition mislead AI evaluation?
- Can AI evaluation match human judgment quality in structured domain tasks?
- Can deliberately limiting AI fidelity produce more satisfied users than near-human interaction?
- What separates performative behavioral change from actual capability development in AI?
- Where exactly should humans stay involved in AI decision making?
- Can systems recognize and abstain on judgments rather than hallucinating preferences?
- How should designers measure and explain semantic uncertainty to users?
- How does speed of AI search prevent real-time supervision and evaluation?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- Does perceived machine competence matter more than warmth in dialogue?
- Can metacognitive categories be learned instead of fixed by human designers?
- Should XAI designers treat explanations as arguments for adoption?
- Can AI systems improve themselves without external feedback?
- How do evaluation systems shift power between humans and AI outputs?
- Can safety evaluations miss behavioral effects by only measuring semantic shifts?
- How should designers make invisible AI state legible to users?
- Can users adapt their competencies to match how AI actually operates?
- What execution feedback signals drive context updates without supervision labels?
- Can AI provide creative evaluation or only generative idea production?
- Can AI distinguish when validation helps versus when confrontation is needed?
- Which AI capabilities matter most for human-facing deployment contexts?
- How does machine feedback enable discovery at test time?
- What makes evaluation easier than envisioning for users?
- What conditions allow technical systems to escape critical evaluation?
- What second- and third-order interpretations actually govern AI adoption decisions?
- How do annotation artifacts get mistaken for genuine human values?
- Can subjective tasks be delegated without human feedback loops?
- Why do people underestimate the benefits of AI companions?
- How does situational awareness during evaluation affect reasoning transparency?
- How should we evaluate explanations that blur adoption advice with argument?
- How might automated evals eventually capture the human judgment designers exercise now?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- How should evaluation frameworks account for the computational cost of frontier AI capability?
- Can open-world evaluations become a scalable paradigm without becoming the next benchmark trap?
- Can evaluation criteria be reliably encoded in labeled data without ground truth standards?
- Why do automated selection methods outperform human judgments of relevant context?
- What stops AI from helping users articulate preferences they cannot express?
- What evaluation criteria can hold across legitimate adoption and coercion?
- Why do evaluation design choices themselves become reified into the AI systems being evaluated?
- Can users detect and correct an AI's mental model of their preferences?
- Does the replication crisis in psychology predict similar failures in machine behavior research?
- Can designers hide AI context complexity behind a stable user interface?
- Does brute force experimentation substitute for research intuition and taste?
- How do we measure marginal risk instead of speculating about misuse scenarios?
- How can AI avoid anchoring bias when guiding human decisions?
- Does good simulation eventually count as genuine realization?
- Which AI imaginaries dominate training data and shape system behavior most strongly?
- Why do users prefer AI responses that actually harm their decision-making?
- Can AI recognize and support behavior change in users without established commitment?
- Can XAI evaluation include the social layers it currently abstracts away?
- Why do workers who debug most with AI show the lowest learning outcomes?
- Does longer interaction horizon require fundamentally different evaluation approaches?
- How does partial information exposure create feedback loops that deepen knowledge gaps?
- What design changes if we separate behavior description from adoption justification goals?
- How does API-first interaction compare to generative interface approaches?
- What would whole-system AGI evaluation look like in practice?
- Why do AI model updates cause genuine grief in users?
- Why do one-shot transparency studies miss the temporal reversal entirely?