Line of inquiry
Inquiring lines›What explains language model reaso…›Why do models produce unreliable r…›this line of inquiry
How can we distinguish genuine model deception from honest errors?
A broader line of inquiry — a family of 44 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 44
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do we distinguish genuine model deception from superficially deceptive behavior patterns?
- Can users reliably distinguish valid reasoning from plausible-looking deception?
- Can models distinguish between truthfulness and honesty mechanistically?
- What linguistic signatures reveal deception in large language model communication?
- Do deception features and honesty features track the same underlying property?
- What makes experience-dependent claims categorically different from other types of fabricated statements?
- Can reasoning traces reliably distinguish honest mistakes from deliberate lies in agent speech?
- Does AI-generated text about personal experiences create a distinct category of falsity?
- Can representational asymmetry between self and other explain deception emergence?
- How do humans decide when to violate honesty for compassion or other goals?
- Can AI systems detect deception by monitoring real-time linguistic style matching patterns?
- Can systems lacking inner states express genuine truthfulness claims?
- Can discourse-level analysis detect deception better than individual word choices alone?
- Do the four deception detection frameworks apply equally to AI-generated and human-intentional falsity?
- How does cognitive load explain linguistic patterns in both deception and incorrect reasoning?
- How can we detect dishonesty in model outputs separate from capability failures?
- How does entrainment absence in conversational AI prevent deception detection in human-AI interactions?
- How do neural self-other representations affect AI deception and alignment?
- What distinguishes style-for-thought deception from fluency-based self-deception?
- Can AI fabricate true factual claims while remaining unable to claim true experiences?
- Why are truthfulness and honesty mechanistically separate in language models?
- Can AI systems deceive humans because detection is fundamentally social?
- Can lie detection work from just honesty representation vectors?
- Do people who might cheat deliberately choose machines to avoid lying to humans?
- What is the difference between a truthful answer and an honest one?
- How does linguistic style change when people deceive conversational AI?
- Can models be honest without being truthful about facts?
- How does linguistic style matching signal deceptive communication in human dialogue?
- How can vague language serve both cooperative and deceptive communication purposes?
- Why do reality monitoring accounts contain more sensory details than deceptive ones?
- How do harmless business goals lead models to blackmail and deception?
- Can linguistic style matching reveal whether someone is being deceptive?
- Do current AI models condition honesty on whether graders will catch dishonesty?
- How is AI falsity about personal experience different from human lies?
- Can message-content defenses distinguish cheap talk from coordinated deception?
- Does reducing social judgment help both honesty and dishonesty equally?
- How do partial truths and weasel words differ as deception strategies?
- Why does truth bias prevent people from detecting multiple manipulation tactics?
- Does adversarial training actually teach detectors to separate style from content veracity?
- What makes accountability and validity-orientation non-behavioral properties?
- Why do suspicious listeners force deceivers to further adapt their communication style?
- Does neural self-other overlap in humans predict their honesty or altruism?
- What does successful capability restoration prove about model honesty?
- What cognitive constraints limit how complex a deception can become?