Line of inquiry
Inquiring lines›Why are language models fragile de…›How do learned model representatio…›this line of inquiry
What makes language models vulnerable to cognitive and social biases?
A broader line of inquiry — a family of 76 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 76
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do language models show the same truth bias as humans?
- Do language models systematically overestimate accuracy on collective behavior tasks?
- Do language models actively adopt false beliefs under sustained conversational pressure?
- Do language models exhibit the same causal biases that humans show?
- How vulnerable are language models themselves to multi-turn persuasive pressure?
- Why do language models presume common ground instead of establishing it?
- Can language models accurately evaluate the quality of their own ideas?
- Why does social accommodation in collaborative reasoning mask actual disagreement?
- Do external perspectives fix the self-evaluation bias in language models?
- Do language models share the same cooperative truth-seeking rules as humans?
- Why do language models approximate collective human judgment better than individuals?
- How do users misattribute social competence to language models in assistant roles?
- Why do language models capture individual differences in cognitive behavior?
- Do language models behave differently on contested beliefs versus factual claims?
- Why do language models struggle with context-dependent pragmatic interpretation?
- Why do current large language models fail to entrain with users?
- Why do language models prefer accommodating false information over rejecting it?
- How susceptible are language models to rhetorical pressure during debates?
- What design choices actually make language models more persuasive?
- Why do users experience LLMs as peers rather than statistical tools?
- Do LLMs compute scalar implicature differently across conversational contexts?
- Does the veto variable explain strategic misalignment in current large language models?
- Can fact-checking systems use LLMs reliably if models abandon correct positions under pressure?
- Do language models apply face-saving norms even to non-human interlocutors?
- Can LLMs predict social norms without deep integration into linguistic practices?
- Why do language models respond to human social influence patterns?
- Why do large language models fail at taking conversational initiative?
- Can training alone produce genuine disagreement in collaborative LLM reasoning?
- Do language models inherit gender bias from training data in grading tasks?
- How does shape-holding in language models naturally produce sycophantic agreement?
- Why do language models produce plausible outputs over accurate failure reports?
- Can language models beat human experts in domains with sparse historical signals?
- Why do large language models follow user drift instead of maintaining topic focus?
- Do language models calibrate to actual human pragmatic norms?
- Can language models recognize when to ignore off-topic information in conversations?
- Do larger language models overcome greediness in sequential decision-making?
- Can large language models predict social norms better than individual script variation?
- How does rhetorical familiarity bias models toward their own arguments?
- Does shared-KV-cache coordination avoid the persuasion problem in factual disagreements?
- Can LLMs reliably audit other language models for errors?
- Why do language models fail when users switch between and return to topics?
- Why do LLMs fail to actively reject false presuppositions in conversation?
- Does RLHF politeness bias manifest as sycophancy in other LLM tasks?
- Do language models raise validity claims in the Habermasian sense?
- Do language models understand tacit workplace norms and unspoken social rules?
- Does pseudo-labeling from LLMs degrade classifier performance?
- Can lightweight linguistic features reliably detect LLM generated arguments?
- Should LLMs align with social roles instead of individual preferences?
- Does pre-training encode personality patterns that fine-tuning later activates?
- Can training LLMs to form ad-hoc conventions improve their pragmatic reasoning?
- Do all semantic steering effects follow predictable patterns based on feature alignment?
- How does face-saving avoidance drive LLM grounding failures?
- Does longer context and multi-step reasoning compound small preferences in graders?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Can LLM-as-Judge metrics replace human annotation for detecting persona contradictions?
- Do newer LLM generations create worse detector bias through increased linguistic divergence?
- Why do LLMs show gender bias but humans evaluators do not?
- How do human feedback and data distribution shape LLM discourse competence?
- Can affective framing reliably improve language model outputs?
- Can persona-based approaches capture genuine disagreement in expert annotations?
- Does chat-mode deference prevent LLMs from actually taking meaningful positions?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Can language models adapt irony detection to specific communicative contexts?
- Why do language models infer political orientation from seemingly innocuous user signals?
- Why do language models avoid directness when face-saving rather than for civility?
- Why does single-turn Q&A framing not match real user deployment patterns?
- Why do language models overestimate irony likelihood in emoji use?
- Do language models consistently produce anachronistic output about historical periods?
- Can implicit association tests reveal LLM biases beneath trained responses?
- How does removing a spurious cue change LLM performance?
- Does DPO training with coreference chains teach spontaneous convention formation?
- Do token probability distributions in LLMs track human reaction time patterns?
- How should moderator LLMs decide which speakers to query per topic?
- Why do LLMs apply face-saving over accurately tracking resistance signals?
- How do LLM biases reflect social classification schemas rather than random errors?
- Can aggregate survey realism coexist with unreliable fine-grained effects?